Teyrex Logo

Paul H. · September 29, 2026

How to Calculate the ROI of AI Automation: A Small Business Worksheet

Most small businesses measure AI automation by the model bill. That is the wrong line. A worked example with current pricing shows where the money actually goes and when to kill the project.

Bar chart of the monthly cost of one AI email automation: staff review $750, maintenance $300, error rework $225, n8n cloud $25, and model tokens $7, which is 0.5% of the total

The short answer: monthly ROI = (hours saved x loaded hourly cost) minus (tokens + tooling + human review + error rework + maintenance), and payback = build cost divided by that monthly number. For a typical small business automation, the model tokens are under 1% of the real cost. The numbers that decide whether it pays are how long a human still spends checking each output and how often the AI gets it wrong. If you only track the API bill, you are measuring the one line that almost never matters.

  • Measure the manual baseline first. Time 20 real items with a stopwatch before anyone builds anything.
  • Count review time as a cost. An AI draft that takes a person 1 minute to check still costs 1 minute.
  • Do not optimize tokens first. At small business volumes they are usually the smallest line in the budget.
  • Kill line: if payback is longer than 12 months, or review takes more than half the original manual time, stop or redesign.

What goes into the ROI of an AI automation?

Five cost lines, measured against one benefit line. The benefit is the staff time the automation removes, priced at the loaded hourly cost (wages plus taxes and overhead, not the take-home rate). The costs are model tokens, the automation platform (n8n, Zapier, Make, or a server you run), the time a person still spends reviewing outputs, the time spent redoing the outputs the AI got wrong, and ongoing maintenance when prompts drift or an upstream API changes. There is also a one-off build cost, which is what payback is measured against.

On our Return on Tokens page we define ROT as value of output minus token cost, divided by token cost. That formula is the right lens for teams spending serious money on tokens. For a 10 to 50 person business, the token line is so small that the full cost stack below matters more.

A worked example: AI email triage for a 12-person service business

This is an illustrative model, not a client case, built with the kind of numbers we see when scoping these projects. The business gets 1,500 inbound emails a month. Today a person reads each one, routes it, and drafts a reply, which takes about 4 minutes. At a loaded cost of $30 per hour, that is 100 hours and $3,000 per month.

The automation: an n8n workflow sends each email to a model, which classifies it and drafts a reply for a person to approve. Assume 2,500 input tokens per email (instructions, the email, some customer context) and 400 output tokens. Here is the monthly cost, with model prices from Anthropic's pricing page and platform prices from n8n's pricing page, both as of 2026-09-29:

  • Model tokens: about $7. Claude Haiku 4.5 costs $1 per million input tokens and $5 per million output tokens. 3.75M input tokens is $3.75, 0.6M output tokens is $3.00.
  • Automation platform: about $25. n8n Cloud Starter is €20 per month billed annually and includes 2,500 workflow executions, enough for 1,500 emails. Self-hosting the Community Edition is free but moves the cost into server and upkeep time.
  • Human review: $750. A person still reads and approves every draft. At 1 minute each, that is 25 hours.
  • Error rework: $225. If 5% of drafts are wrong and take 6 minutes to fix (the original 4 plus 2 for cleanup), that is 7.5 hours.
  • Maintenance: $300. Roughly 3 hours a month of developer time at $100 per hour, mostly prompt fixes and broken nodes after upstream API changes.

Total: $1,307 per month against $3,000 of manual work, a net saving of $1,693 per month. If the build costs $5,000 (roughly where our n8n projects start for simple workflows), payback is about 3 months and the first-year net is about $15,300.

Bar chart of the monthly cost of one AI email automation: staff review $750, maintenance $300, error rework $225, n8n cloud $25, and model tokens $7, which is 0.5% of the total

Does the choice of AI model change the ROI?

Barely, at small business volumes. Running the same 1,500 emails on Claude Sonnet 5 ($2 input, $10 output per million tokens) costs about $13.50 a month. Running them on Claude Opus 5.5 ($4 and $20) costs about $27. Anthropic notes that its newer models use a tokenizer that produces roughly 30% more tokens for the same text, so budget up to about $35 for Opus. Compare that with review time: cutting review by 30 seconds per email saves 12.5 hours, or $375 a month. A better model is worth paying for when it makes drafts accurate enough that review gets faster. It is not worth debating for the token cost alone.

Which number should you watch after launch?

Review time per item. Everything else in the model is either small or fairly stable. Review time is large, it is easy to underestimate in a demo, and it creeps up quietly when staff stop trusting the drafts and start rewriting them. Here is the same workflow with only review time changed:

Bar chart of months to pay back a 5,000 dollar build by review time per email: 0.5 minutes pays back in 2.4 months, 1 minute in 3.0 months, 2 minutes in 5.3 months, and 3 minutes in 25.9 months

At 3 minutes of review per email, the automation saves $193 a month and takes over two years to pay back. Add a 15% error rate to that and it loses money every month. Nothing about the tokens changed. The project quietly turned from a win into a cost, and a dashboard that only tracks the API bill would never show it.

How do you measure the baseline before building?

With a stopwatch and a spreadsheet, before anyone writes a prompt. Pick 20 real items from last week, time the person who normally does the work, and note which items were unusual. Multiply the median time by monthly volume and loaded hourly cost. Then, during a two-week pilot, time review on the AI output the same way and log every draft that needed real rework. Those two numbers replace guesses in every line of the model above.

While you are measuring, look for the parts that do not need a model at all. Routing an email because it comes from a known customer domain or contains an order number is a rule, and a rule runs in plain workflow logic for free, every time, with no error rate. Anthropic's own guidance on building effective agents makes the same point: start with the simplest solution and add complexity only when it earns its place. Spend tokens on the judgment calls and let code handle the rest.

When should you kill an AI automation?

  • Payback on the build is longer than 12 months at measured (not demo) review time.
  • Review takes more than half of the original manual time per item. At that point staff are doing the work twice.
  • The error rate rises month over month and nobody owns fixing it.
  • The only person who understands the workflow has left, and maintenance has become guesswork.

Killing a project that does not pay is a good outcome. It frees the budget for the one that will. If you want to run your own numbers, the Return on Tokens Calculator models value against token spend, and the AI project cost estimator gives a rough build cost to divide into your monthly saving. If the math works and you want it built, talk to us about n8n automation; we will scope it against your measured baseline, not a demo.