Token Usage

Tokens are the units AI models use to process text. Otto displays token details when OpenRouter reports them, but billing is based on verified usage, not a local token-rate formula.

What Tokens Describe

Input Tokens

Input tokens are context sent to the model, such as:

  • Ticket title and description
  • Code context and diffs
  • Conversation history
  • Agent system instructions
  • Organization and project context
  • Tool results from earlier steps

Output Tokens

Output tokens are generated by the model, such as:

  • Clarifying questions
  • Implementation plans
  • Code changes
  • PR summaries
  • Review replies

Some models may also report cache, reasoning, or other token categories. Those fields are shown as provider-reported details when available.

Billing Relationship

Token details help explain usage, but they are not the source of truth for charges.

Important billing details:

  • Published catalog prices and token estimates are estimates only; they are not billed.
  • Usage can remain pending while it is unavailable or still being reconciled.
  • Unknown usage stays pending until verified; Otto does not convert unknown usage into fabricated or zero-cost charges.
  • Fractional cents are preserved and carried across settlements.

Typical Token Drivers

DriverWhy it matters
Large ticketsMore initial context is sent to the model.
Large diffsReview and implementation workflows include more code context.
Long conversationsPrior turns may be needed for continuity.
Tool-heavy workTool results can be added to subsequent model calls.
Reasoning modelsSome models report additional reasoning tokens or metadata.

Reducing Usage

Write Concise Tickets

❌ Verbose:
"I would like you to please add a feature that allows users to upload their profile pictures..."

✅ Concise:
"Add user avatar upload: JPG/PNG, max 5MB, store in S3"

Keep Context Focused

Include the conventions and files that matter for the task. Avoid unrelated architecture documents, old chat logs, or broad repository dumps unless they are required.

Choose a Suitable Model

Otto's model catalog exposes compatible OpenRouter models and validates required capabilities such as tools, image input, or reasoning support before inference. Lower-cost models can be appropriate for simple clarification or summarization; complex implementation and review tasks may need stronger models.

Break Up Large Work

Smaller, focused tickets usually require less context per run and make human review easier.

Monitoring Usage

Run and billing views can show:

  • Requested and returned model/provider
  • Input, output, cache, and reasoning token details when available
  • Billing label: pending, charged, waived, or internal
  • Workflow, run, and request lineage for audit