AI Agent Running Costs: What One Agent Really Costs per Month

Axel Grubba
Axel Grubba
Sep 28, 2026
AI Agent Running Costs: What One Agent Really Costs per Month
Explore this topic with AI:
ChatGPTPerplexityGoogle

Last updated: September 2026. All prices are public list prices checked on September 28, 2026.

The model bill is rarely the biggest part of AI agent running costs. Your review time is. In the worked example below, one business agent doing 300 jobs a month costs $73.71 in model tokens and $167.70 in your time, at one minute of checking per job.

That does not make tokens irrelevant. The same agent can cost $37 or $387 a month in tokens depending on two settings most people never look at. This guide is for business owners and operators who want a real number before they switch an agent on, not a range from $15 to $15,000.

  • An agent run is many model calls, not one. Each step re-sends the whole conversation, so input tokens dominate and grow faster than the step count
  • Prompt caching cut the token bill by 62% in our example, more than switching to a cheaper model tier did
  • Reasoning tokens are billed as output, and after caching they were a third of the per-run cost
  • Tool calls and sandbox time are small at typical volumes: $7.20 a month combined here
  • Human review is the largest line until you move from checking every run to checking a sample

What AI Agent Running Costs Actually Include

An AI agent is a model working in a loop: it reads the task, calls a tool, reads the result, decides the next step, and repeats until it is done. Every loop is a separate paid model call. That is the first thing that separates agent costs from chat costs.

Anthropic's engineering team reported that "agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens as chats." The same write-up found token usage alone explained 80% of the performance variance in their research evaluations, with tool calls and model choice as the other two factors. Spending tokens is how agents get better answers, which is exactly why the bill needs watching.

Here is the full list of what you pay for:

Cost layer What drives it Typical size
Input tokens Instructions, tool definitions, history and tool results, re-sent every step The biggest token line
Output tokens Visible replies plus hidden reasoning Smaller volume, 5x the unit price
Tool calls Paid tools such as web search, per call Cents per run
Sandbox or compute time A workspace where the agent runs code or a browser Cents per hour
Integrations The apps and automation plans the agent connects to Often a subscription you already pay
Retries and loops Runs that fail, repeat, or wander A tail that can exceed the average
Human review Your minutes checking what the agent did Usually the largest line

Token Prices in September 2026

Every major provider now prices fresh input, cached input, and output separately, and most add a price for writing to the cache. The numbers below come straight from the Claude API pricing page, the OpenAI API pricing page, and the Gemini API pricing page, all per million tokens.

Claude API pricing page listing per-million-token input, output, and prompt caching prices for Claude Fable 5.1, Opus 5.5, Sonnet 5.5, and Haiku 4.5

Model Input Cache write Cached input Output
Claude Opus 5.5 $4.00 $5.00 $0.20 $20.00
Claude Sonnet 5.5 $2.00 $2.50 $0.20 $10.00
GPT-6 Sol $2.00 $2.50 $0.20 $10.00
Gemini 3.8 Flash $0.75 Storage billed hourly $0.075 $3.75
Claude Haiku 4.5 $1.00 $1.25 $0.10 $5.00

Three details on those pages change the math more than the headline rates do.

Cached input costs a tenth of fresh input or less. Anthropic prices a cache read at 0.1x the base input price, and 0.05x on Opus 5.5. A cache write costs 1.25x for the five-minute cache, so caching pays off after a single reuse. OpenAI's GPT-6 models follow the same pattern.

Reasoning tokens are output tokens. OpenAI's reasoning guide says reasoning tokens "are billed as output tokens" even though you never see them, and Google labels its output price "including thinking tokens." A model that thinks for 700 tokens before a 300-token answer bills you for 1,000.

Tools add tokens before they do anything. Anthropic lists a 286-token tool-use system prompt on Opus 5.5, before your own tool definitions. Its browser toolset adds about 6,600 input tokens to every request that declares it.

OpenAI API pricing page showing standard per-million-token rates for GPT-6 Astra, Sol, and Luna, including cached input and cache write prices

A Worked Example: One Agent, 300 Runs a Month

To get a real number, you need a real job. Take a common small-business lane: an agent that handles a customer request end to end. It reads the message, looks up the order, checks the policy, searches the web once or twice when needed, drafts a reply, and logs what it did.

The assumptions:

  • 12 model calls per run. About what a job with five tool calls needs, plus a step or two to finish
  • 12,000 tokens of fixed context per call: instructions, business context, and tool definitions
  • 1,800 tokens added per step: 1,500 of tool results and 300 of visible model output
  • 1,000 output tokens per step, of which 700 are reasoning
  • 2 web searches, 3 minutes of sandbox time, and 1 minute of your review per run
  • 300 runs a month, roughly 10 a day

The per-run token math

The first call sends 12,000 tokens. The twelfth call sends the fixed context plus eleven steps of history: 12,000 + 11 × 1,800 = 31,800 tokens. Add up all twelve and one run sends 262,800 input tokens to produce 12,000 output tokens. The model rereads its own history far more than it writes.

At mid-tier rates without caching:

  • Input: 262,800 × $2.00 per million = $0.53
  • Output: 12,000 × $10.00 per million = $0.12
  • Total: $0.65 per run

With caching, each call reads everything it already sent from cache and only writes what is new. Across the run that is 231,000 cached tokens and 31,800 newly written ones:

  • Cache writes: 31,800 × $2.50 per million = $0.08
  • Cache reads: 231,000 × $0.20 per million = $0.05
  • Output: 12,000 × $10.00 per million = $0.12
  • Total: $0.25 per run, 62% less

Notice what happened to the mix. Before caching, input was 81% of the bill. After caching, output is nearly half of it, and the 8,400 reasoning tokens alone cost $0.08, about a third of the run.

Model choice versus caching

Run the same arithmetic across the price table and the ranking is not what most people expect.

Bar chart of monthly model token cost for 300 agent runs: $387 frontier uncached, $194 mid-tier uncached, $134 frontier cached, $74 mid-tier cached, and $37 small model cached

A frontier model with caching ($134) costs less than a mid-tier model without it ($194). Before you downgrade the model to save money, check whether caching is on. Prompt caching is automatic on some providers and needs a flag on others, and a single changing value near the top of your prompt, such as a timestamp, quietly breaks it on every call.

The full monthly bill

Now add the layers that are not tokens. Web search is $10 per 1,000 searches on the Claude API and on OpenAI. Sandbox time uses Anthropic's published agent session rate of $0.08 per running hour. Integration glue uses Zapier's Professional plan at $19.99 a month billed annually. Review time uses $33.54 an hour, the May 2025 US mean wage of $69,770 a year across all occupations, divided by 2,080 working hours.

Line item Calculation Monthly
Model tokens 300 runs × $0.2457 $73.71
Loops and retries 5% of runs (15) wander to 30 steps, $0.43 extra each $6.46
Web searches 600 × $0.01 $6.00
Sandbox time 15 hours × $0.08 $1.20
Integration plan One automation subscription $19.99
Machine total $107.36
Human review 5 hours × $33.54 $167.70
All-in $275.06

Bar chart of the $275 monthly cost of one AI agent: $167.70 review time, $73.71 model tokens, $19.99 integration plan, $6.46 loops, $6.00 web searches, $1.20 sandbox time

That is $0.92 per completed job. Whether it is cheap depends entirely on what the job was worth before. If a person spent six minutes on each request, the same 300 jobs took 30 hours, or about $1,006 at the same wage.

How AI Agent Running Costs Scale

Costs do not all grow the same way, and the difference matters when you plan for more volume.

Runs scale linearly. Double the runs, double the token bill. There is little volume discount at small-business scale, although batch processing cuts token prices by 50% for work that can wait.

Steps scale faster than linearly. Each extra step re-sends everything before it, so cost grows roughly with the square of the run length. In our example, a run that wanders to 30 steps is 2.5 times longer but costs $0.68 instead of $0.25 with caching, 2.75 times more. Without caching it costs $2.59 instead of $0.65, four times more. Long runs are where budgets break, and caching is also what softens that curve.

Tool results scale with what you let in. One fetched web page is about 2,500 tokens and a research PDF about 125,000, according to Anthropic's figures. An agent that reads whole documents instead of the relevant section pays for that text on every remaining step of the run.

Review scales with your trust, not with volume. Checking every run keeps review as the largest line forever. Checking one run in ten, after the agent has earned it, brings the all-in figure from $275.06 to $124.13 a month. That is the biggest single saving available, and it has nothing to do with models.

How to Estimate Your Own Agent Costs

You do not need a spreadsheet. You need five numbers, and the fourth is the one people skip.

  1. Runs per month. Count how often the job actually happens today, not how often you wish it did
  2. Steps per run. Roughly two model calls per tool call, plus one to finish
  3. Tokens per step. Fixed context plus the size of a typical tool result. Most providers show usage per request, so run the job ten times and read the real number
  4. Review minutes per run. Time yourself checking five results. That number drives the total more than the model does
  5. Your failure rate. The share of runs that loop, retry, or need redoing, priced at their longer length

Then: monthly cost ≈ runs × (tokens per run × blended price + paid tool calls) + runs × review minutes × your hourly rate. Use the cached price for repeated context and the full price for everything new, and you will usually land close to the real bill.

How to Cap AI Agent Costs

Estimates drift. Caps do not. These are the controls that actually hold:

  • Set a spend limit on the provider account. Most provider consoles offer a monthly limit, a budget alert, or both. Set one before the first scheduled run, not after the first surprising invoice
  • Cap steps per run. A maximum of 20 or 25 steps turns a runaway loop into a failed run you can inspect, instead of an open-ended bill
  • Cap reasoning effort and output length. Most reasoning models expose an effort or budget setting. Hard jobs deserve it, routine lookups do not
  • Keep the top of the prompt stable. Anything that changes on every call, placed early, breaks caching for everything after it
  • Trim tool results before the model sees them. Return the three fields the agent needs, not the whole record or the whole page
  • Route by difficulty. A small model for sorting and lookups, a stronger one for the step that needs judgment
  • Move from full review to sampling once the agent has a clean track record on that job, and keep full review for anything that sends money or messages customers. That is the same draft-first rule we recommend when choosing which jobs to hand an agent first

What Nobody Tells You

  • The same text can cost more on a newer model. Anthropic notes that its newer tokenizer "produces approximately 30% more tokens for the same text." Compare real usage, not just the price per token
  • Promotional prices end. Gemini 3.8 Flash is $0.75 per million input tokens through December 31, 2026, then $1.50 from January 1, 2027. Budget on the price you will pay next year
  • Cached prices differ by model within one provider. Opus 5.5 reads cache at 5% of base input while most models read at 10%, so the cheapest setup is not always the cheapest model
  • Failures bill too. A run that errors on step nine has already paid for eight steps. Retrying from scratch pays for them again, and switching to a backup model mid-run starts over with a cold cache
  • The quiet months are cheap and the busy month is not. Agent cost follows your business volume, which is exactly when a variable bill is least welcome. Set the cap for your best month, not your average one

Where Crevio Fits

Crevio pricing page showing the free Starter plan, Pro at $20 a month with 1,000 credits, and Business at $50 a month with 2,500 credits

Everything above assumes you run the agent yourself and pay each provider separately. Crevio takes a different approach: it is an AI business builder, and all of the AI work it does for you, from building your website and writing product pages to generating images, sending marketing emails, and researching the web, draws from one credit balance. Bigger jobs use more credits than small ones, and the platform tracks the tokens, caching, and tool calls underneath so you do not have to.

Pro is $20 a month with 1,000 credits and Business is $50 a month with 2,500, and both let you pick a larger monthly allowance as the business grows. Starter is free with a small allowance for trying it out. Monthly credits reset each billing cycle, while bonus credits from signup and referrals do not expire. On a paid plan you can top up at any time or switch on automatic top-ups, and if credits do run out, your website and checkout keep running. Only new AI work pauses.

Being straight about the trade-off: a credit balance gives you one predictable number, but it hides the per-token detail this article walks through. If you want to control model choice and caching settings yourself, running your own agent stack gives you that. And no pricing model removes the review line. That minute per run is still yours until you decide the agent has earned less supervision. Transaction fees also apply on every plan: 5% on Starter, 2.5% on Pro, and 1% on Business.

AI Agent Running Costs FAQ

How much does it cost to run an AI agent per month?

For a single business agent handling about 300 jobs a month, expect roughly $40 to $400 in model tokens depending on the model and whether prompt caching is on, plus a few dollars for tools and compute. Our worked example came to $107.36 in machine costs and $275.06 including one minute of human review per job.

What is the biggest hidden cost of an AI agent?

Human review time. Checking every result for one minute at the US mean wage cost more than twice the entire token bill in our example. The second biggest is re-sent context: each step of a run pays again for everything the agent has already read, which is why long runs cost disproportionately more.

Do reasoning tokens cost extra?

They are billed as output tokens at the model's output price, even though you never see them. OpenAI and Google both state this on their documentation and pricing pages. In our example, reasoning made up about a third of the cost of each cached run.

How much does prompt caching save on agent costs?

In our example it cut the token bill by 62%, from $0.65 to $0.25 per run, because an agent re-sends the same instructions and history on every step. Cached reads cost a tenth of fresh input on most models. Caching only works if the start of the prompt stays identical between calls.

Is a subscription cheaper than paying per token for an agent?

It depends on your volume and how predictable it is. Flat plans cap your downside but can cost more when usage is light, while metered API pricing tracks usage exactly. For ChatGPT and Claude plans specifically, we worked out the break-even between a subscription and the API: about 80 runs a month for a $20 plan. We compared the platform side of this in our Squad alternatives breakdown, which weighs metered against flat pricing.

Tokens are the cost you can read on an invoice. Your attention is the cost you have to count yourself, and it is usually the bigger one.

What will you sell today?

Describe what you want to sell — Crevio builds, launches, and grows it. Products, payments, and marketing, all on autopilot.

Start for free