Skip to main content

Review Token Usage

Use Monitor > Tokens to review token usage and optimization opportunities across MCP servers, toolkits, skills, and tool responses.

The page has two modes:

  • Analyze shows usage metrics and context breakdowns.
  • Optimize shows recommended actions and savings estimates.

Open Tokens

  1. Open Monitor > Tokens.
  2. Set the Start Date and End Date.
  3. Use Advanced Filters when you need to narrow results by user, AI agent, MCP server, or tool.
  4. Select Analyze, Optimize, or Devices depending on the task.
Willow Token Optimization Analyze view showing summary metrics, token usage panels, biggest spenders, used models, and messages per conversation

The Analyze view starts with four summary metrics:

MetricWhat it counts
Est. Conversation SpendEstimated cost of the conversations in the window, with the token count beside it.
Response TokensTokens returned by tools.
Init Context TokensTokens loaded before any tool runs, with the session count beside it.
Avg Tokens/CallMean tokens per tool call.

Below them, Tool Token Usage charts request, response, and raw response tokens, where the gap between raw and response is what response mapping saved. Toolkit Context Savings compares actual toolkit init tokens against the projected cost of exposing every MCP tool. Biggest spenders, used models, and messages per conversation appear further down when there is data for them.

The Optimize view groups suggested work into actions with estimated savings. Which recommendations appear depends on your own usage; the common ones are:

RecommendationWhy it helps
Build toolkits for high-context MCP serversMCP servers expose every tool definition to the agent. Curating a toolkit with only the tools a workflow needs reduces fixed context on every conversation.
Reduce skill sizeSkills are injected into every conversation. Moving long examples or documentation into references lowers the baseline session cost.
Willow Tokens Optimize view showing recommended actions and estimated total savings

Use the recommendation cards as the starting point. Then open the related detail page to identify the exact server, toolkit, skill, or tool response to change.

Build toolkits recommendation

The Build toolkits for high-context MCP servers card points you toward MCP servers whose full tool list is adding too much context. Its next step is usually MCP Context, followed by creating or tightening a toolkit.

Recommendation card for building toolkits for high-context MCP servers

Reduce skill size recommendation

The Reduce skill size card points you toward oversized skills. Its next step is Skill Token Analysis, where you can review the skill and extract long content into references.

Recommendation card for reducing an oversized skill

Analyze cards

Further down, the Analyze view summarizes the major context and response token sources.

CardWhat it means
MCP ContextTool descriptions and input schemas loaded from MCP servers.
Toolkit ContextContext exposed by curated toolkits.
Skills ContextSkill content loaded into conversations.
Tool ResponsesTool outputs and response-mapping opportunities.

Cards show totals, health indicators, and a View all action. Open the full view when a card shows a high count, a high token total, or an item marked for optimization.

MCP Context card

The MCP Context card opens MCP Context. Use it when one MCP server exposes many tools or has a large tool-description and input-schema footprint.

MCP Context card with Slack, Everything MCP, and Context7 context token rows

Toolkit Context card

The Toolkit Context card opens Toolkit Context. Use it to compare curated tool subsets and find toolkits that can be split or tightened.

Toolkit Context card with toolkit context token rows and View all action

Skills Context card

The Skills Context card opens Skill Token Analysis. Use it when a skill is large enough to raise the baseline cost of every conversation where it is loaded.

Skills Context card showing humanizer as a high-token skill and my skill as a low-token skill

Tool Responses card

The Tool Responses card opens Tool Response Optimization. Use it to find tool outputs that need response mappings or output-format changes.

Tool Responses card showing no tool call data in the selected period

Devices view

The Devices tab reports AI agent token usage collected from managed devices, rather than traffic through the gateway. Use it to see what agents cost on developer machines, including activity the gateway never sees.

Token Optimization Devices tab with All AI Agents and All Models filters, cards for Est. Device Spend, Messages, Sessions, and Devices Reporting, and Tokens by Model Over Time and Model Usage panels

Two filters at the top narrow the view by AI agent and by model. Below them, four summary cards show estimated device spend, messages and the tool calls behind them, sessions, and the number of devices reporting with the user count.

PanelWhat it shows
Tokens by Model Over TimeDaily input plus output tokens per model, parsed from local agent transcripts on managed devices.
Model UsageEstimated cost by model, covering input, output, and cache-read tokens, priced per model.
Usage by UserDevice-collected AI agent usage attributed to organization users.
Usage by DeviceEvery managed device that reported AI agent token usage.

Where the data comes from

This view is populated by AI Discovery scan agents, not by the gateway. Until scan agents report AI agent activity, the panels stay empty with the message that data appears once scan agents with the device token usage feature report activity.

If the Devices tab is empty while the Analyze tab has data, that is expected: Analyze reflects gateway traffic, and Devices reflects what scan agents observe on managed machines. Confirm scan agents are deployed and reporting before troubleshooting further.

Empty states

Some areas can show no data when Willow has not recorded enough token-tracked activity in the selected period. If that happens:

  • widen the date range
  • remove advanced filters
  • run tool calls through a connected AI client
  • confirm Token Usage Analytics is enabled
  • for the Devices tab, confirm scan agents are deployed and reporting