Paste a model.
Get an endpoint.
Buying the same model in dollars, you pay a 5–7% forex markup and 18% GST on the conversion before the model has run. We hold the provider accounts and bill you rupee-native, so neither exists on your bill.
Any supported open model on Hugging Face becomes a private, OpenAI-compatible endpoint in minutes — running on a GPU we rent for you at the cheapest available rate. You pay by the hour, not by the token, and the instance stops itself when nobody is using it.
// Cheaper than a dedicated endpoint. Simpler than renting a GPU yourself.
From a repo URL to a working endpoint.
Nothing is provisioned and nothing is billed until you have seen the rate and agreed to it.
Size and precision pick the card.
Parameter count times bytes-per-weight, plus room for the KV cache, is what decides which GPU can serve your model. We do that arithmetic for you and then shop it across 30+ clouds.
| Model size | Serving precision | GPU tier | Hourly rate |
|---|---|---|---|
| Up to ~4B | bf16 / fp16 | Entry (16–24 GB) | Cheapest tier in the catalogue — quoted live |
| ~7–9B | bf16 / fp16 | Entry–mid (24 GB) | Cheapest tier in the catalogue — quoted live |
| ~13–15B | bf16, or 8-bit to fit smaller | Mid (40–48 GB) | Mid tier — quoted live |
| ~30–35B | 8-bit / 4-bit quantised | Mid–high (48–80 GB) | Mid–high tier — quoted live |
| ~70B | 4-bit quantised | High (80 GB) | Highest single-card tier — quoted live |
India, or the cheapest card anywhere.
Most capacity is international, and that is where the cheapest card usually is. If your data has to stay in India, pick an India region at launch — it is priced separately, and the region appears on the quote before the instance starts, not after.
What we will not promise.
“Any model” is marketing. Serving engines do not load every architecture. We support 500+ and tell you up front when yours is not one of them — before you are charged.
Gated models need your own token. Llama, Gemma and similar require your Hugging Face token, because the licence is between you and the publisher, not between you and us.
First start takes 10–45 minutes. The weights have to download and load. Restarts are faster. We quote the estimate before you confirm rather than after you have started waiting.
Single-GPU deployments today. Models that need two or more cards are coming. If yours needs them, the check in step one says so instead of failing halfway through a boot.
Your endpoint. Your existing code.
Paste a model. Get an endpoint.
Billed by the GPU-hour in rupees. Idle instances stop themselves.