After every successful request, the system reads the usage data from the model response: prompt token (message/instructions sent) and completion token (model answer). Both are summed into the total token, then recorded to your account usage and per-model history.
If the upstream doesn't send complete prompt token data, the system uses a minimum estimate from the message content so usage is still recorded. For image requests, the image token estimate from the model configuration is also taken into account. Cache tokens are recorded as usage information, but the main quota always uses the total tokens billed by the model.
The simple formula: (prompt token + completion token) × model multiplier. The result is rounded and deducted from your quota. When total usage reaches the quota, the API key becomes exceeded and the next requests are rejected until more quota is added.
Say your starting quota is 1,000,000 tokens. You send a request with 800 prompt tokens and the model answers with 1,200 completion tokens.
Model 1x
(800 + 1.200) × 1 = 2,000 tokens used.
Quota left: 1,000,000 − 2,000 = 998,000 tokens.
Model 1.5x
(800 + 1.200) × 1.5 = 3,000 tokens used.
Quota left: 1,000,000 − 3,000 = 997,000 tokens.