Live models
auto×1auto-debug×20visionclaude-fable-5-b×20visionclaude-opus-5×12visionclaude-opus-5-b×12visionclaude-sonnet-5×10claude-sonnet-5-b×10visiondeepseek-v4-flash×2deepseek-v4-flash-0731×2deepseek-v4-flash-vision-exp×2.5visiondeepseek-v4-mod×3deepseek-v4-pro×1.15deepseek-v4-pro-0813×1.8deepseek-v4.1-flash×2.56visiondeepseek-v4.1-mod×3.2visionglm-5.1×1glm-5.2×1.25glm-5.2-mod×3glm-5.3×2glm-5.3-flash×2visionglm-5.3-flash-mod×3.5visionglm-5.3-flashx×2.5visionglm-5.3-flashx-mod×4visionglm-5.3-mod×3.2gpt-5.6×5visiongpt-5.6-luna×5visiongpt-5.6-luna-b×5visiongpt-5.6-sol×15visiongpt-5.6-sol-b×15visiongpt-5.6-sol-xhigh×15visiongpt-5.6-terra×10visiongpt-5.6-terra-b×10visionhy3×1hy4×1.4kimi-k2.6×1kimi-k2.7-code×1kimi-k2.7-code-highspeed×1.5kimi-k3×2visionkimi-k3-mod×3.5visionmimo-v2.5-pro×1.4mimo-v2.6-flash×2mimo-v2.6-pro×2visionminimax-m3×1.4auto×1auto-debug×20visionclaude-fable-5-b×20visionclaude-opus-5×12visionclaude-opus-5-b×12visionclaude-sonnet-5×10claude-sonnet-5-b×10visiondeepseek-v4-flash×2deepseek-v4-flash-0731×2deepseek-v4-flash-vision-exp×2.5visiondeepseek-v4-mod×3deepseek-v4-pro×1.15deepseek-v4-pro-0813×1.8deepseek-v4.1-flash×2.56visiondeepseek-v4.1-mod×3.2visionglm-5.1×1glm-5.2×1.25glm-5.2-mod×3glm-5.3×2glm-5.3-flash×2visionglm-5.3-flash-mod×3.5visionglm-5.3-flashx×2.5visionglm-5.3-flashx-mod×4visionglm-5.3-mod×3.2gpt-5.6×5visiongpt-5.6-luna×5visiongpt-5.6-luna-b×5visiongpt-5.6-sol×15visiongpt-5.6-sol-b×15visiongpt-5.6-sol-xhigh×15visiongpt-5.6-terra×10visiongpt-5.6-terra-b×10visionhy3×1hy4×1.4kimi-k2.6×1kimi-k2.7-code×1kimi-k2.7-code-highspeed×1.5kimi-k3×2visionkimi-k3-mod×3.5visionmimo-v2.5-pro×1.4mimo-v2.6-flash×2mimo-v2.6-pro×2visionminimax-m3×1.4

Frequently Asked Questions

How multipliers, grades, tokens, and quota work on Tokenly.

What is a Model Multiplier?

A Model Multiplier is a token-consumption multiplier applied to each model. By default, every model uses a 1x multiplier, meaning tokens consumed are counted 1:1 against your quota.

If a model is set to a 1.5x multiplier, every token consumed is counted 1.5 times against your quota. Example: the upstream counts 1,000 tokens, so your quota decreases by 1,500 tokens.

The multiplier is shown in the Available Models list in the Customer tab, for example glm-5.2-debug (1.5x). A model without parentheses stays at 1x.

What does Model Grade mean?

Grade helps you pick a model based on service quality and performance. Grade does not judge a model's base capability absolutely; experience may change depending on service conditions.

  • A

    Top quality.

    From well-known services with good quality and performance.

  • B

    Limited performance.

    From new services or services with very small requests-per-minute limits, so performance may be suboptimal.

  • C

    Budget choice.

    Very low service performance, but available on bandelbanget.xyz at a cheaper price.

Easiest analogy to understand Grade A & B — Grade A is like buying from big stores such as hypermarts and Indomaret. Grade B is like buying from a Bangladeshi wholesale store or a small warung. Grade B might still be good quality, but it stays grade B because the source is not a big store.

For those concerned about privacy and security, I recommend using grade A models. Safer against cloaking and other risks.

For grade B, bandelbanget.xyz cannot guarantee, but can compensate with +tokens if the grade B quality is very poor.

How are tokens and quota calculated?

After every successful request, the system reads the usage data from the model response: prompt token (message/instructions sent) and completion token (model answer). Both are summed into the total token, then recorded to your account usage and per-model history.

If the upstream doesn't send complete prompt token data, the system uses a minimum estimate from the message content so usage is still recorded. For image requests, the image token estimate from the model configuration is also taken into account. Cache tokens are recorded as usage information, but the main quota always uses the total tokens billed by the model.

The simple formula: (prompt token + completion token) × model multiplier. The result is rounded and deducted from your quota. When total usage reaches the quota, the API key becomes exceeded and the next requests are rejected until more quota is added.

Easy simulation

Say your starting quota is 1,000,000 tokens. You send a request with 800 prompt tokens and the model answers with 1,200 completion tokens.

Model 1x

(800 + 1.200) × 1 = 2,000 tokens used.

Quota left: 1,000,000 − 2,000 = 998,000 tokens.

Model 1.5x

(800 + 1.200) × 1.5 = 3,000 tokens used.

Quota left: 1,000,000 − 3,000 = 997,000 tokens.

Still have questions?

Our support team is available on Telegram. Reach out anytime for help with keys, quota, or billing.