01000001 01010000 01001001QWEN-3.8-MAXDEDICATEDQWEN-3.7-PLUS01001110 01000110 01011001QWEN-3.8-27BFRONTIERDEEPSEEK-V4-FLASHGLM-5.3-FLASHNO-FOREX-MARKUP01000001 01010000 01001001QWEN-3.8-MAXDEDICATEDQWEN-3.7-PLUS01001110 01000110 01011001QWEN-3.8-27BFRONTIERDEEPSEEK-V4-FLASHGLM-5.3-FLASHNO-FOREX-MARKUP
SPEC SHEET·FRONTIER · DEPLOY · GPUS · JOBS·BILLED IN ₹·2026

Frontier APIs, your own models,
and the GPUs to run them.

// vs a USD-billing gateway, all-in
19.6%
Same weights. Same throughput.

Buying the same model in dollars, you pay a 5–7% forex markup and 18% GST on the conversion before the model has run. We hold the provider accounts and bill you rupee-native, so neither exists on your bill.

₹95.5 interbank × 1.06 forex × 1.18 GST vs ₹96.0 flat
▸ See the full comparison

One rupee-priced account for the open frontier, any open model you want deployed, and raw GPU by the hour. No forex markup. Every response signed with the model, lane and region it actually ran on.

§ 01 — WHAT YOU CAN BUY
// Four products, one rupee-priced account

Four things to buy. One key, one wallet.

Frontier APIsLive
Six open-weight frontier models — the same weights the labs publish — on one key, at rupee-native rates. ~20% under OpenRouter's all-in cost. Full disclosure of which model answered.
See the models & rates
Deploy any open modelLive
Paste a Hugging Face repo, get an OpenAI-compatible endpoint in minutes. 500+ supported architectures, billed by the GPU-hour, idle instances stopped automatically. India-resident option.
Deploy a model
GPU rentalsLive
Rent GPUs by the hour across 30+ clouds at the cheapest available rate. Your wallet kills the instance — you cannot wake up to a bill you did not agree to. India-resident option.
Rent a GPU
Job endpointsComing soon
150 finished jobs at a flat rupee price per document — invoices parsed, statements reconciled, IDs validated. Deterministically checked; wrong output refunded, not billed.
Join the waitlist
7
Frontier models
rupee-priced inference
500+
Deployable architectures
paste a Hugging Face repo
30+
GPU clouds
cheapest card that fits
0%
Forex markup
banks take 5–7% + GST
§ 02 — PRICE ADVANTAGE
// Same model · same provider weights · lower bill

Every frontier model, cheaper than OpenRouter.

OpenRouter and every foreign gateway bill you in US dollars. By the time it hits your card you have paid a 5–7% forex markup and 18% GST on the conversion on top of the sticker price. apinfy holds the provider accounts and bills you rupee-native. Same weights, same throughput, one flat rupee rate.

PER 1M TOKENS · INDIAN DEVELOPER · ALL-IN ₹ · EXACT IN / OUTvs OpenRouter
MODELVIA OPENROUTER* (in / out)VIA APINFY (in / out)SAVE
Qwen 3.8 Max$2.00 / $6.00₹238.90 / ₹716.71₹192 / ₹57619.6%
Qwen 3.7 Plus$0.47 / $1.41₹56.14 / ₹168.43₹45 / ₹13519.6%
DeepSeek V4 Flash$0.44 / $1.32₹52.56 / ₹157.68₹42 / ₹12719.6%
Kimi K2.7 Code$0.95 / $4.00₹113.48 / ₹477.81₹91 / ₹38419.6%
Qwen 3.7 Max$1.25 / $3.75₹149.31 / ₹447.94₹120 / ₹36019.6%
GLM-5.2$1.40 / $4.40₹167.23 / ₹525.59₹134 / ₹42219.6%
GLM-5.3 Flash$0.15 / $0.50₹17.92 / ₹59.73₹14 / ₹4819.6%
*OpenRouter/direct = USD × interbank (₹95.5) × 1.06 forex markup × 1.18 GST — the real all-in cost to an Indian developer. apinfy = USD × ₹96.0 flat (interbank + ₹0.50 spread), billed in rupees with a reclaimable GST invoice; the forex markup and conversion-GST do not exist. On ₹10,00,000/mo of frontier spend that is ₹1,96,326 a month, ₹23,55,912 a year. Dedicated-lane models are rupee-native with zero FX exposure.
How to read the sources. Qwen rates are quoted against the International (Singapore) endpoint, which is the one we resell; Alibaba's Chinese Mainland (Beijing) endpoint is 60–70% cheaper and is a different rate card. DeepSeek V4 Flash is quoted at the peak rate; DeepSeek bills peak/off-peak since 16 Aug 2026, and the peak windows fall inside the Indian working day. // The whole table rests on the interbank rate (₹95.5), which moves every day; last checked 25 August 2026.
§ 03 — BENCHMARK READOUT
// Prove it, don't claim it

Where each lane lands against the frontier.

Published scores for the frontier models we serve, next to the closed frontier. Dedicated-lane scores are measured on our own harness. Gaps shown honestly.

ModelLaneSWE-bench VerifiedSWE-bench Pro₹/M out
DeepSeek V4 Flash▸FRONTIERFrontier78.055.4₹22
Kimi K2.7 Code▸FRONTIERFrontier80.258.6₹330
GLM-5.2▸FRONTIERFrontier80.658.4₹355
Claude Opus 4.8Closed88.6₹2,075
GPT-5.6 SolClosed96.2₹2,490
SRC: published cards + independent SWE-bench Pro (Scale) / Vals AI · JUL 2026. SWE-bench is scaffold-sensitive (~20pt swing) — Pro is the fairer cross-model read. ~8–16pt behind absolute frontier on Verified, at 1/6th–1/100th the price.
§ 04 — OPERATION
// Under the hood

Pick your lane. Proof on every call.

[01] ROUTE
You pick the lane
Dedicated or frontier — by the model name. One OpenAI-compatible endpoint.
[02] RUN
We serve it
Our own GPUs for dedicated; accounts we hold for frontier. Prefix-cached, failover built in.
[03] CHECK
Verified where it counts
Deterministic checks on jobs. Wrong output is refunded, not billed.
[04] SIGN
You get a receipt
Ed25519 record of model, lane and region. Verify it offline, forever.
§ 05 — INTEGRATION
// Drop-in

Change one line. Keep your code.

▸ quickstart.pyPYTHON 3 // OPENAI-COMPATIBLE
# pip install openai — then change the base_url client = OpenAI(base_url="https://apinfy.com/v1", api_key="kk_live_...") # India-resident — our GPU, in an India region r = client.chat.completions.create(model="qwen3.8:27b", messages=[...]) # or the open frontier, rupee-priced r = client.chat.completions.create(model="deepseek-v4:flash", messages=[...]) print(r.receipt.model, r.receipt.lane, r.receipt.region) # → qwen3.8:27b R in-vast ✓ signed

Dedicated or frontier. Always in rupees.

₹100 of free credit. No card. Top up by UPI when you are ready.

▸ Get your API keyRead the docs