Token Usage
Tokens are the units AI models use to process text. Otto displays token details when OpenRouter reports them, but billing is based on verified usage, not a local token-rate formula.
What Tokens Describe
Input Tokens
Input tokens are context sent to the model, such as:
- Ticket title and description
- Code context and diffs
- Conversation history
- Agent system instructions
- Organization and project context
- Tool results from earlier steps
Output Tokens
Output tokens are generated by the model, such as:
- Clarifying questions
- Implementation plans
- Code changes
- PR summaries
- Review replies
Some models may also report cache, reasoning, or other token categories. Those fields are shown as provider-reported details when available.
Billing Relationship
Token details help explain usage, but they are not the source of truth for charges.
Important billing details:
- Published catalog prices and token estimates are estimates only; they are not billed.
- Usage can remain pending while it is unavailable or still being reconciled.
- Unknown usage stays pending until verified; Otto does not convert unknown usage into fabricated or zero-cost charges.
- Fractional cents are preserved and carried across settlements.
Typical Token Drivers
| Driver | Why it matters |
|---|---|
| Large tickets | More initial context is sent to the model. |
| Large diffs | Review and implementation workflows include more code context. |
| Long conversations | Prior turns may be needed for continuity. |
| Tool-heavy work | Tool results can be added to subsequent model calls. |
| Reasoning models | Some models report additional reasoning tokens or metadata. |
Reducing Usage
Write Concise Tickets
❌ Verbose:
"I would like you to please add a feature that allows users to upload their profile pictures..."
✅ Concise:
"Add user avatar upload: JPG/PNG, max 5MB, store in S3"
Keep Context Focused
Include the conventions and files that matter for the task. Avoid unrelated architecture documents, old chat logs, or broad repository dumps unless they are required.
Choose a Suitable Model
Otto's model catalog exposes compatible OpenRouter models and validates required capabilities such as tools, image input, or reasoning support before inference. Lower-cost models can be appropriate for simple clarification or summarization; complex implementation and review tasks may need stronger models.
Break Up Large Work
Smaller, focused tickets usually require less context per run and make human review easier.
Monitoring Usage
Run and billing views can show:
- Requested and returned model/provider
- Input, output, cache, and reasoning token details when available
- Billing label: pending, charged, waived, or internal
- Workflow, run, and request lineage for audit
