How to use TokenOps
The field manual. Ten minutes to read, then the tool takes ninety seconds in a room. Open TokenOps.
TLDR
- Answer three questions (what are you building, at what scale, can data leave), land on a grounded preset, verify the flagged assumptions, read the answer out loud.
- The big dollar figure is the ceiling a hardware quote must come UNDER to beat tokens. Type a real quote for an instant verdict.
- Everything is a range or a score with its rules showing. Nothing here is a quote.
1. What this is and when to open it
TokenOps answers the question every AI conversation eventually hits: what will this agent workload cost, and where should it run? It estimates token demand from how an agent is built, prices that demand across five providers at public list rates, derives the budget a hardware quote must come under, and scores nine placement routes with every rule visible.
Open it when you hear any of these in a meeting:
- "What would this AI thing actually cost us per month?"
- "Should we just buy GPUs instead of paying per token?"
- "Can we run this ourselves? Our data cannot leave."
- "Which platform should we build the agents on?"
It is for the architect or seller running the conversation, in front of the Customer or preparing for one. It is a conversation tool. It is never a quote, and it says so on every surface.
The spine of the whole tool is one equation, worth memorizing for the hostile-architect moment: monthly cost = users x adoption x runs per day x days x calls per run x (input tokens x input rate + output tokens x output rate), cache adjusted. Every screen in the tool is that equation with receipts attached.
The five providers priced at public list rates: Anthropic, OpenAI, Google Gemini, Azure OpenAI, and AWS Bedrock, three model tiers each, every cell editable. Quick-formula workloads (the RAG, agents, and coding estimators) are priced too, at the worker rate with an editable input/output split, so demand never shows tokens that cost zero dollars.
2. The front door: presets, always
The tool opens by asking what you are building. Eight patterns, ranked by how common they actually are in 2025-2026 production surveys, plus three example Customers you can walk in as. Your answers route to a starting point with every assumption stated and flagged for verification. From the landing you choose the depth:
- In a meeting: a four-step wizard, about twelve inputs, two minutes to a defensible answer. Use this with a Customer watching.
- Deep sizing (the tool calls this full view Architect Mode): every section, every formula, every assumption on one scrolling page. Use this at your desk, before the proposal, or when an engineer wants to argue. Arguing is encouraged; the math is on screen.
The three example Customers (Calloway Reed LLP, Harborline Mutual, Northgale Communications) are fake companies carrying real-magnitude numbers from published deployments, each with a variable-by-variable explanation of what every number means and what it drives. They double as the best way to learn the tool.
3. Meeting Mode, step by step
Step 1, the use case. Name the scenario, pick the Customer size, and answer the budget signal honestly (Unknown, Weak, Some budget, or Committed). Budget matters more than it looks: only Unknown trips the guard, but if it stays Unknown the tool will refuse to recommend a route and tell you to go do discovery instead. That refusal is a feature.
Step 2, scale. Users, runs per user per day, adoption percent. Adoption is the one everybody overestimates. Fifty percent adoption of the licensed population is a realistic pilot; type what you believe, not what the sponsor hopes. Days per month shows whenever a user count is present, and hides if you clear Users.
Step 3, the policy gate. Can data leave the environment: Yes, With controls, or No. This single answer moves more scoring weight than anything else in the tool. Answer No and the private routes surge while public providers take a visible penalty. Follow-up questions reveal themselves based on your answer.
Step 4, workload shape. Pick the agent topology (single assistant up to swarm), the prompt cache hit rate, and the retry rate. Cache matters because providers bill a cached input re-read at roughly a tenth of the normal input rate, so a stable system prompt is real money. Topology is the silent cost multiplier: a planner-worker-judge shape makes about eight model calls per run where a single assistant makes one. The tool shows that math rather than hiding it.
Then the answer page: recommendation, budget ceiling, provider costs, whiteboard card, and discovery questions, all live. Every number updates if you go back and change an input.
4. Deep sizing: the sections that matter most
Architect Mode shows sixteen sections, five of which appear only when the modern agent workload is switched on. All of them work; four of them decide the answer.
- Agent topology. Calls per run by role: planners, workers, judges, retries, replans. If you only tune one section after a preset, tune this one to match how the Customer's agent will actually be built.
- Token anatomy per role. Input and output tokens per call for each role, plus the cached input percent. Output tokens cost five to six times input tokens at every shipped provider rate, so the output columns are where money hides.
- Data and policy gate. Fifteen policy questions. Seven weighted conditions among them build the private policy score, and the score with its arithmetic renders right below the section.
- Economics and the ceiling. The savings threshold (default: owned must beat tokens by 40 percent), the amortization window, and the quote slot.
The other sections (RAG, tools, memory, context snowball, sizing, storage, network) add realism when the conversation goes there. Every one renders its formulas inline, always expanded.
5. Reading the recommendation card
- Route scores, 0 to 100. Nine routes, each scored by visible weighted rules: direct provider, cloud model service, Airia, Kamiwaza, Build Technology Group, HPE Private Cloud AI, rented GPU validation, owned hardware, and hybrid. The bars are comparable at a glance; the "rules that fired" list below them is your talking track.
- Two viable routes. When the top two land within ten points, the tool refuses to fake certainty and presents both with the tradeoff stated. Read that line to the Customer verbatim; it starts the right conversation.
- Do not size yet. If the policy gate, usage volume, or budget signal is missing, this beats every route and lists exactly what discovery is owed. Do not fight it. It is right.
- Missing data, primary risk, primary lever. Every recommendation names what would change it: no quote entered, no measured benchmark, estimated usage. Confidence shows High, Medium, or Low with the averaging math printed under it.
6. The hardware budget ceiling, the number that wins meetings
TokenOps never prices hardware. It inverts the question: given what the token route costs per month, and requiring ownership to beat it by the threshold, the recommended configuration must come in under a specific dollar figure. That figure is the big number on the economics card.
Use it exactly like this: when a Customer gets a hardware quote, type it into the quote slot. The tool answers UNDER or OVER, by how much, instantly. Under means ownership genuinely beats tokens by your margin. Over means negotiate or stay on tokens. The break-even chart below it shows the crossover volume, and both lines use the same billed-token math, so the chart always agrees with the sentence.
7. Levers, rates, and sliders
- Optimization levers. Each lever (raise cache hits, route workers to the mini tier, cap retries, summarize tool results) recomputes the entire model with the change applied and reports real dollars. A lever that cannot be computed honestly says so instead of guessing.
- Provider rates. Every price cell is editable. Type the Customer's negotiated rate and the row marks itself user supplied. Rows older than 60 days flag themselves STALE.
- Scoring weights. Every weight behind the routes, including the policy points and the co-recommend margin, is a drag bar with a reviewed default. Drag one in front of a skeptical architect and watch the routes reorder live. Reset restores the reviewed defaults.
8. Three meeting plays
- The token bill shock. Customer's OpenAI bill tripled. Meeting Mode, their real usage numbers, then straight to the provider comparison and the optimization levers. The levers usually find 30 to 60 percent before anyone mentions hardware.
- The GPU itch. Customer wants to buy hardware because it feels cheaper. Enter their workload, show the ceiling, and hand them the number their quote must beat. Either the quote clears it (great, proceed with evidence) or it does not (you just saved them from a very expensive rack).
- The locked-down environment. Policy gate: data cannot leave. Watch the private routes surge and the do-not-size guard demand real usage data. The discovery questions card becomes your next-meeting agenda.
9. Saving, sharing, exporting
- Work autosaves to your browser. Named scenarios save locally; nothing ever leaves the page on its own.
- Share links are created only when you click the button, and warn if a Customer name would be included. The sanitized variant strips identity fields.
- Exports: a Customer-friendly markdown summary, the full detailed-math report with every substitution, a print report with the formula appendix, and JSON scenario save and load.
10. What it will never do
- It will never present rough math as a quote.
- It will never price hardware; it tells you what hardware has to cost.
- It will never hide a formula, an assumption, or a rule that fired.
- It will never transmit what you type; the only external request on the whole site is a cookieless page counter, and the test harness fails the build if any other appears.
Methodology and receipts: docs/tokenops.md. Also see the Nutanix Conversation Sizer manual.