Bare metal GPU cluster, NVLink, InfiniBand / RoCEv2
A GPU farm that works for you, not for an hourly meter
At sustained load, renting GPUs by the hour costs more than owning the hardware. We design bare metal clusters sized on your real jobs, validated with benchmarks before production.
Use cases
- Training and fine-tuning internal models on your proprietary data
- High-throughput inference at a controlled marginal cost
- 3D rendering, simulation and accelerated compute
- Continuous workloads where hourly cloud becomes a budget sinkhole
Reference architecture
- Compute
- Multi-GPU bare metal nodes with intra-node NVLink, no hypervisor: 100% of the power goes to your jobs.
- Fabric
- 400G NDR InfiniBand or tuned RoCEv2 (PFC, ECN, MTU 9000, DCQCN) depending on budget and scale, validated with ib_write_bw and NCCL tests.
- Provisioning
- Drivers, CUDA, firmware and images automated (Ansible, Terraform), reproducible at every fleet extension.
- Observability
- DCGM, Prometheus and Grafana: GPU utilization, fabric health and alerting from day one.
What you receive
- 01
Priced architecture file: topology, hardware BOM, 3-year costs
- 02
Cluster deployed, hardened and benchmarked (NCCL all-reduce, fabric bandwidth)
- 03
Operational monitoring and alerting
- 04
Complete runbook and training for your teams
FAQ
- Is InfiniBand mandatory for a GPU farm?
- No. Below 128 GPUs, a properly tuned RoCEv2 fabric delivers equivalent performance for most training workloads, at a significantly lower cost. We price both options in every audit.
- On-prem, colocation or dedicated cloud?
- All three are possible. The choice depends on your power constraints, data residency and time to service. The audit compares scenarios with numbers, not preferences.