AI / ML · LLM fine-tuning · Inference

AI Workstations in Mumbai

Workstations built for training, fine-tuning, and inference of language and diffusion models. PCIe lane budgeting, ECC tradeoffs, 1200-1500W 80+ Titanium PSUs, and VRAM tier matching for the workload. No RGB, all thermals.

AI workstation with multi-GPU setup
01

PCIe lane budgeting, on paper first

Most workstation CPUs (Threadripper, Xeon W, EPYC) advertise 48-128 PCIe lanes, but the practical allocation against a multi-GPU + NVMe + 10GbE build is tighter than the headline number. We map lanes to devices on paper before we order parts, not after.

02

ECC tradeoffs, honestly

For training and fine-tuning of LLMs and large diffusion models, ECC does not give you a measurable accuracy win. What it does give you is reliability: bit-flips in a multi-day training run are usually fatal. We tell you honestly which tier your workload needs.

03

1200-1500W PSUs, sized for transients

Modern GPUs and CPUs pull 1.5-2x TDP for milliseconds during state transitions. We add sustained power draw, multiply by 1.4 for transients, and add 100W for the rest of the system. A typical 2x RTX 4090 build needs a 1200-1500W 80+ Titanium PSU. We do not under-size PSUs for AI builds.

Workload sizing

VRAM tier guide

A practical cheat sheet for matching VRAM to model size and use case. The honest answer is usually the second-cheapest card, not the most expensive.

8 GB

Inference of small models (≤3B parameters) and local IDE assistants.

Example: RTX 4060 Ti 8GB

12-16 GB

Fine-tuning small models (≤7B) and inference of mid-size models (7-13B).

Example: RTX 4070 Ti SUPER, RTX 4080 SUPER

24 GB

Fine-tuning 7-13B and inference of 30-70B models in 4-bit. The sweet spot for most AI builders.

Example: RTX 4090, RTX 5090

48 GB

Fine-tuning 30-70B and inference of 100B+ in 4-bit. For serious ML work.

Example: RTX 6000 Ada, A6000

What we will not build

Honest constraints

  • We will not spec a single 16A Indian socket for a >1800W build. The math: a single 16A socket is fused at 16A × 230V = 3680W maximum, but sustained draw above ~1800W on a single socket is asking for a tripped breaker or a warm socket over time. Multi-GPU AI builds need a dedicated circuit.
  • We will not under-size a PSU. If your build needs 1200W sustained, you get a 1500W PSU with 80+ Titanium rating. We lose the sale sometimes to builders who say "850W will be fine" — it is not fine, and we will not sign off on it.
  • We will not skimp on cooling because the workload is "headless". A 24-hour training run is exactly the workload that cooks VRMs. We spec 360mm AIOs and positive-pressure cases as if the build were a gaming build.
  • We will not promise NVLink on consumer cards. RTX 4090 / 5090 / 6000 Ada do not have NVLink. If you need NVLink (H100, A100), that is a different class of build with very different budgeting.

AI workstation questions

What is PCIe lane budgeting and why does it matter for AI workstations?
Most workstation CPUs (Threadripper, Xeon W, EPYC) advertise 48-128 PCIe lanes, but the practical allocation against a multi-GPU + NVMe + 10GbE build is tighter than the headline number. A single high-end GPU uses 16 lanes, a second GPU wants 16 more (NVLink needs even more), a boot NVMe needs 4, a data NVMe RAID needs 4-8, and a 10GbE NIC needs 4-8. That is 48-60 lanes used. We map this out on paper before we order parts, not after.
Do I need NVLink for AI workloads?
For most fine-tuning and inference workloads on consumer / prosumer GPUs (RTX 4090, RTX 6000 Ada, RTX 5090), NVLink is not available — those cards do not have NVLink bridges. Multi-GPU works through PCIe and (where supported) NVSwitch. If your workload specifically requires NVLink (H100, A100), that is a different class of build with very different budgeting, and we will tell you that up front.
Should I get ECC memory for an AI workstation?
For training and fine-tuning of LLMs and large diffusion models, ECC does not give you a measurable accuracy win. What it does give you is reliability: bit-flips in a multi-day training run are usually fatal. The pragmatic answer: if you are running 24+ hour training jobs, ECC pays for itself the first time it saves a run. If you are running inference or fine-tuning jobs of 2-6 hours, the value of ECC is lower and you can usually redirect that budget to more VRAM.
What is the right PSU size for an AI workstation?
Add up the worst-case sustained power draw of the CPU and GPUs, multiply by 1.4 for transient spikes (modern GPUs and CPUs pull 1.5-2x TDP for milliseconds during state transitions), and add 100W for the rest of the system. A typical 2x RTX 4090 build with a Threadripper-class CPU draws ~900-1000W sustained, so a 1200-1500W 80+ Titanium PSU is the right answer. We do not under-size PSUs for AI builds.
How do I pick the right VRAM tier?
Rule of thumb: 8GB is for inference of small models (≤3B parameters) and local IDE assistants. 12-16GB is for fine-tuning small models (≤7B) and inference of mid-size models (7-13B). 24GB is the sweet spot for fine-tuning 7-13B models and inference of 30-70B models in 4-bit. 48GB (RTX 6000 Ada, A6000) opens up fine-tuning 30-70B and inference of 100B+ in 4-bit. We will tell you honestly which tier your workload actually needs.

Contact & hours

We service all 35 areas across Mumbai, Monday to Saturday.