FAQ

    Questions buyers ask before they sign.

    Straight answers on service models, GPU platforms, colocation pricing, facilities, networking, security and delivery. Where a figure depends on your configuration, site or term, we say so.

    About Helios

    Who is Helios and what do you do?

    Helios Cloud, Inc. is a GPU infrastructure company headquartered in Salt Lake City, Utah. We develop and operate data center capacity purpose-built for AI workloads and deliver it two ways: bare metal GPU compute that you rent, and colocation space for your own hardware.

    We are an operator, not a broker. We hold the power, build the facility, integrate the racks and run them.

    Why is Helios cheaper than a hyperscaler or a typical neocloud?

    Power is the dominant cost of running GPUs at scale, and our advantage is on the power side: long-term energy agreements and siting next to generation, including on-site generation at some locations. That lowers our cost floor rather than our margin, which is why pricing holds over a multi-year term instead of being an introductory rate.

    We also don't carry the cost of a general-purpose cloud platform you wouldn't use.

    Do you operate your own facilities or resell someone else's?

    Helios operates the facility environment: building, power, cooling, physical security and the modular data halls. For bare metal GPU compute we also own the hardware, integrate the racks and manage the lifecycle.

    Where an engagement involves a partner site or partner-managed hardware, we say so explicitly rather than presenting it as Helios-operated.

    Who are your hardware and technology partners?

    Builds follow NVIDIA reference architectures, with Supermicro as the primary OEM across HGX B300, GB300 NVL72 and RTX PRO 6000. Networking is NVIDIA Quantum-X800 InfiniBand, Spectrum-X Ethernet, BlueField-3 DPUs and ConnectX SuperNICs. Shared storage runs on WEKA or DDN class parallel filesystems.

    What kind of customers do you work with?

    AI-native companies running training or high-throughput inference, infrastructure platforms and compute marketplaces reselling capacity, and enterprises moving sustained AI workloads off hyperscaler pricing. Every engagement is scoped to the workload rather than sold from a fixed catalog.

    Ways to buy

    What are the different ways I can buy from Helios?

    Three. Bare metal GPU compute: we own the hardware, you rent capacity and we run it. Colocation: you own the hardware and we provide the facility, power, cooling and network around it. Build-to-spec: colocation where the data hall is engineered to your requirements before you move in.

    The right choice depends mostly on whether you want to own the asset and how long your demand runs.

    What is the practical difference between renting compute and colocating?

    With bare metal you get a working cluster and an SLA, and Helios carries hardware ownership, firmware, RMA and lifecycle risk. You pay one all-inclusive rate for GPU access.

    With colocation you carry the hardware, warranty and spares, and buy power, cooling, space and connectivity from us at a monthly per-kW rate plus metered electricity. Colocation wins on cost per effective GPU-hour if utilization is sustained and you can staff the hardware side. Bare metal wins if you'd rather not.

    Should I rent compute or colocate my own hardware?

    It comes down to sustained utilization and whether you want to carry hardware risk. Above roughly 60 to 70% sustained utilization over a multi-year horizon, owning and colocating usually wins on total cost, provided you can handle procurement, warranty, spares and lifecycle.

    Below that, or if you'd rather not staff the hardware function, bare metal is usually cheaper once you count stranded capacity and the cost of running hardware yourself. We'll model both against your actual utilization curve.

    What access modes do you support on bare metal?

    Bare metal, containers, managed Kubernetes and Slurm. Bare metal gives root-equivalent SSH with passwordless sudo and no virtualization overhead, which is what most large training customers choose. Docker and the NVIDIA Container Toolkit come installed.

    The access mode is set in the order form and determines where the responsibility boundary sits for the OS and orchestration layers.

    Do you offer on-demand or self-serve hourly capacity?

    Our model is committed, reserved capacity rather than spot or self-serve on-demand, and pricing is built around that commitment. Short bridge arrangements are possible case by case, but they are quoted rather than listed.

    Can I test the hardware before committing?

    Yes. We run paid proofs of concept on production nodes at the production per-GPU-hour rate. When a production order follows within the agreed conversion window, the POC fee is credited in full against your first production invoice. A POC doesn't oblige either side to proceed.

    Can I start small and grow?

    Yes. On compute, fabrics are engineered up front for the target scale, so growth joins the existing cluster instead of forming an isolated island. On colocation, capacity grows in modular-unit increments with expansion rights reserved in the contract.

    Can I own the hardware but have Helios buy, deploy and operate it?

    Yes, through what we call the ownership path. Helios designs the build, manages procurement and vendor relationships, integrates and commissions the racks and operates the environment, while you hold title to the hardware. Asset transfer happens at acceptance sign-off.

    GPU platforms and pricing

    Which GPUs do you offer?

    Three current NVIDIA platforms: GB300 NVL72 (rack-scale Grace Blackwell Ultra), HGX B300 (eight-GPU x86 liquid-cooled nodes) and RTX PRO 6000 Blackwell Server Edition (eight-GPU air-cooled nodes). Earlier generations can be scoped when a specific requirement calls for them.

    Compare GPU pricing
    How do I choose between GB300 NVL72, HGX B300 and RTX PRO 6000?

    By workload shape rather than headline performance. For a large unified NVLink domain for frontier-scale training, GB300 NVL72. For flexible x86 capacity that handles both training and long-context inference, HGX B300. For dense inference or generative media at the lowest cost per served token or frame, RTX PRO 6000.

    We confirm the selection with you in design review against your workload profile.

    Helios GPU platforms
    PlatformBest fitCoolingRack densityFabric
    GB300 NVL72Frontier-scale training, largest unified memory domainLiquid only132–142 kWNVLink 5 within rack, InfiniBand across
    HGX B300Mixed training and long-context inferenceLiquid (direct-to-chip)110–120 kWInfiniBand (training) or Spectrum-X (inference)
    RTX PRO 6000Dense inference, RAG, generative media, VDIAirAbout 12 kW per node400 GbE RoCE, no NVLink
    How much does reserved GPU capacity cost?

    Bare metal GPU compute is priced per GPU-hour against a committed quantity and term, invoiced monthly, and the rate falls as term and volume increase. Our published reserved schedule:

    GB300: $5.61 per GPU-hour on a 1-year term, down to $4.34 on a 5-year term.

    B300: $4.80 per GPU-hour on a 1-year term, down to $3.93 on a 5-year term.

    RTX PRO 6000: $1.35 per GPU-hour on a 1-year term, down to $1.04 on a 5-year term.

    Fees accrue only for hours the infrastructure is available and accepted. Final rates depend on platform, quantity, term, site and delivery profile.

    See every term on the pricing page
    What is included in the GPU-hour rate?

    Everything. Power, cooling, space, the compute fabric, external connectivity and data transfer, local NVMe and shared storage, monitoring, firmware and driver lifecycle, hardware replacement and 24/7 support are all inside the rate.

    There's no separate power bill, no bandwidth or egress charge, no storage line item and no setup fee. Only transaction tax is excluded.

    Do you sell fixed SKUs?

    No. We build to your bill of materials against NVIDIA and Supermicro reference designs, with separate validated builds for inference-optimized and training-optimized configurations. Fabric, host memory and NVMe endurance class all differ by workload even when the GPU is the same.

    Can I mix GPU types or OEMs in one deployment?

    Mixed GPU types across separate pods, yes. Mixed OEMs within one tightly coupled training fabric we advise against: BMC divergence, NUMA and PCIe topology differences and thermal stragglers all show up in distributed training. For loosely coupled or single-node workloads it's workable.

    I need more GPUs than a standard reference architecture covers. Is that a problem?

    Not a blocker, but it's a design review item. NVIDIA's validated HGX B300 reference architecture tops out at 1,024 GPUs. Anything larger is delivered as a deliberate multi-pod architecture with a super-spine tier, with NVIDIA and Supermicro design sign-off before we commit to it contractually.

    What should I include in a GPU capacity request?

    Start with the GPU platform, quantity, target start date, term and location. Add your interconnect and storage requirements, dataset size and throughput target, and the workload you plan to run.

    Request GPU capacity

    Colocation

    What colocation tiers do you offer?

    Three, chosen by two questions: do you need a fraction of a modular unit or whole units, and does the unit need modifying?

    Helios colocation tiers
    TierCapacity unitCommercial entryBest fit
    RetailFraction of one modular unit, sharedPer rack or per kW, shorter termsUnder about 1.5 MW, fastest time to rack
    WholesaleWhole modular units, single tenantMonthly charge per modular unit, multi-yearFull-unit capacity on standard infrastructure
    Build-to-specWhole units with modificationsOne-time plus monthly charge under a statement of workHigh-density GPU, compliance-driven or custom
    What is a Helios modular unit?

    The modular unit is our basic quantity of capacity: a factory-integrated data hall that arrives with power distribution, cooling, rack infrastructure and controls pre-installed. Each unit carries a configured rating of 1.5 to 2.5 MW of IT load, fixed per site and configuration.

    Standardizing on it is what lets us commission in months rather than years and add capacity in repeatable increments.

    Look inside the module
    What are the electrical characteristics of a unit?

    A and B feeds, 480 VAC three-phase four-wire service, 0.95 lagging reference power factor, metering and monitoring at the customer allocation, and a footprint of 480 by 198 inches. Electrical design capacity per unit is confirmed by site during design review.

    How are modular units deployed and connected?

    Units arrive weatherized and factory-integrated, so site work is foundation, utility interconnection and network entrance rather than building construction. Each unit is commissioned in place, validating A and B electrical paths, metering and alarms, cooling performance and leak detection, carrier connectivity and physical access controls before handover.

    Each unit is a self-contained pod with its own power, cooling and leaf switching. Units join at a core or super-spine tier, so adding one means adding a pod's worth of infrastructure rather than redesigning the cluster. On compute, one unit maps to a 1,024-GPU B300 pod at 16 racks.

    Can you support 100 kW-plus racks and GB300 NVL72?

    Yes, and it's the main reason customers come to us rather than to retail colocation. Direct-to-chip liquid cooling is designed for full NVL72 racks at full rated load, compared with the 20 to 30 kW per rack conventional colocation was built for.

    Floor loading for NVL72-class racks, roughly 3,000 lbs, is handled as a build-to-spec modification where reinforcement is needed.

    My hardware is air-cooled. Can you still host it?

    Yes. We support high-density contained air, direct-to-chip liquid and hybrid designs with separate thermal zones. The cooling model drives the rate, so it's a pricing input as well as an engineering one.

    How is colocation priced?

    Three components: a one-time non-recurring charge (NRC) for setup and implementation, a monthly recurring charge (MRC) billed per kW, and metered electricity passed through separately.

    Published starting MRC rates are $160–$180 per kW per month for air-cooled and $180–$220 per kW per month for liquid-cooled capacity.

    Estimate your colocation cost
    Is electricity included in the per-kW rate?

    Not on colocation. The per-kW charge covers the infrastructure to run your equipment: cooling, space, monitoring, support and capacity reservation. Metered electricity is passed through on top. This is the most common source of confusion when comparing colocation quotes, so check how any competing quote treats it.

    On bare metal the question doesn't come up, because power is inside the GPU-hour rate.

    What does the one-time setup charge cover?

    Engineering, site fit-out, rack or hall preparation, electrical work, cooling interfaces, cabling, testing and commissioning. It's invoiced once at contract initiation. Build-to-spec engagements carry a larger one-time charge because the modification scope is engineered under a statement of work.

    Which costs are included under each deployment type?

    Bare metal is a single all-inclusive rate. Colocation separates infrastructure from consumption, which is why it has more line items. Specific numbers are quoted per deployment, but the structure doesn't change.

    Cost structure by deployment type
    Cost componentBare metal GPU computeRetail / wholesale colocationBuild-to-spec colocation
    GPU and server hardwareIncludedCustomer ownsCustomer owns
    Space and power infrastructureIncludedIn $/kW monthly chargeIn $/kW monthly charge
    Electricity consumedIncludedMetered pass-throughMetered pass-through
    CoolingIncludedReflected in $/kW bandReflected in $/kW band
    Setup and fit-outIncludedOne-time chargeOne-time charge (larger, per SOW)
    Compute fabricIncludedCustomer, unless contractedCan be pre-installed in scope
    External connectivityIncludedCarrier-neutral, customer contractsCarrier-neutral, customer contracts
    Bandwidth and egressIncluded, no egress chargeOverage per order formOverage per order form
    Local NVMe and shared storageIncludedCustomer providesCustomer provides
    Hardware lifecycle and RMAHeliosCustomerCustomer
    Remote handsIncludedContracted scopeContracted scope
    TaxesExcludedExcludedExcluded
    What should I budget for beyond the headline rate?

    On bare metal, nothing beyond transaction taxes. On colocation: metered electricity, remote hands and emergency work, bandwidth overage, cross-connects, shipping, receiving and staging, spare parts storage and project labor.

    What do you need from me for a colocation design and quote?

    Two numbers to start: your megawatt target and your required date. We respond to an initial capacity and schedule submission within 48 hours.

    To go to design we need the hardware and rack list, system power and heat load, air or liquid interface, weight, network topology and carrier requirements, access and compliance model, and your growth forecast.

    Plan your deployment
    Who is responsible for what once my hardware is installed?

    Helios owns the facility, critical infrastructure, power delivery to the agreed demarcation, cooling to the contracted interface, the meet-me room and cross-connect pathway, physical access and contracted remote hands.

    You own your hardware, including warranty, firmware, spares and lifecycle, rack distribution after demarcation, your network edge and cluster fabric, and everything above the OS. Helios has no access to your software or data except for authorized support activity.

    Facilities, power and cooling

    What standard are your facilities built to?

    Helios locations are designed and operated to Tier III standards: independent power paths, N+1 redundancy with dual A and B feeds, and concurrent maintainability so planned maintenance runs without interrupting live compute.

    UPS and generator topology, fire detection and suppression, and structural limits vary by site and are shared with your facilities team during design review.

    What cooling do you support?

    High-density contained air for PCIe GPU, enterprise, storage and inference systems. Direct-to-chip liquid for HGX B300, GB300 NVL72 and HPC, with CDU and manifold integration, leak detection and a full-load design basis. And hybrid designs with dedicated thermal zones. On our own GPU builds, liquid-cooled racks use 250 kW in-rack redundant CDUs.

    In the liquid loop, coolant carries heat from chip cold plates to a coolant distribution unit, whose heat exchanger passes it to a separate facility circuit and on to dry coolers that reject it to outside air.

    See the cooling design
    How much water do your facilities use?

    Helios modular data halls run with zero operational cooling-water consumption, using closed-loop or dry heat rejection. It's a portfolio standard, not a site exception. Initial fill, maintenance, domestic and construction water use are separate.

    We don't claim a specific PUE or WUE figure. PUE depends on load and operating conditions, so review the design assumptions and reporting period for measured results.

    Is your power renewable?

    Facilities are positioned near renewable or clean-generation resources and backed by long-term power agreements, which gives both price stability and a sustainability reporting position. We deploy at the point of generation, so energy producers can monetize excess capacity without adding strain to the grid.

    We don't claim 100% renewable operation, carbon neutrality or net-zero. The renewable share varies by site and over time.

    How is power billed and metered?

    Colocation is billed on metered kW as a monthly recurring charge, with electricity passed through separately at cost. Bare metal customers aren't billed for power at all; it's inside the GPU-hour rate.

    Networking and storage

    What network fabric options do you offer?

    For training, NVIDIA Quantum-X800 InfiniBand at 1:1 non-blocking, rail-optimized for collectives. For inference, NVIDIA Spectrum-X Ethernet with RoCE. The pod plus super-spine topology is the same either way.

    For colocation we support 100, 200, 400 and 800 GbE, RoCE, InfiniBand, storage and out-of-band designs, subject to engineering review.

    How much bandwidth does each GPU get?

    On HGX B300, 800 Gb/s per GPU through a dedicated ConnectX-8 SuperNIC. On RTX PRO 6000, 400 Gb/s per node over four ConnectX-6 Dx adapters.

    Each pod stays inside a non-blocking boundary and pods join through a super-spine core tier, so a 2,048-GPU fleet is two 1,024-GPU pods presented as one fabric, and the structure extends to 4,096 GPUs without re-architecting. Per-GPU bandwidth doesn't degrade as you add pods.

    Are your facilities carrier neutral?

    Yes. You choose your own carriers, cross-connects, internet, private circuits, wavelengths or dark fiber through neutral meet-me rooms, with diverse entrances and route diversity confirmed by location. Cloud on-ramps and private circuits are supported; specific on-ramp availability is confirmed by location.

    Do you charge for bandwidth or egress?

    Not on bare metal. Network connectivity and data transfer are inside the GPU-hour rate, with no per-gigabyte egress charge, which is usually the single largest saving for teams moving off a hyperscaler. On colocation you contract your own carriers, and any Helios-supplied bandwidth and overage terms are set in the order form.

    How is management traffic separated from my data?

    A physically separate out-of-band network carries BMC and host management, isolated from both the compute fabric and your data path.

    What storage comes with a GPU node?

    Local NVMe in every node, sized and endurance-classed to the workload: roughly 61 TB per node on HGX B300 (eight 7.68 TB E1.S Gen5 drives) and about 30 TB per node on RTX PRO 6000 (four 7.6 TB E3.S drives). Inference builds use read-optimized drives and training builds use write-intensive drives. Boot is a mirrored NVMe pair.

    Do you provide shared storage?

    Yes: an all-flash parallel filesystem of WEKA or DDN class, with GPUDirect Storage over BlueField-3 for direct GPU-to-storage transfers. On bare metal it's part of the delivered cluster rather than a separate charge, sized to your dataset and checkpoint requirements in design review.

    North-south connectivity is sized separately from the GPU fabric, so bulk data movement doesn't contend with GPU-to-GPU traffic.

    Security and access

    Is my deployment single tenant?

    On reserved GPU deployments and on wholesale or build-to-spec colocation, yes: a dedicated environment not shared with other customers, with network segmentation and DPU-level offload separating compute, storage and management traffic. Retail colocation shares a modular unit with other tenants, with physical and logical separation between them.

    Do I get BMC or out-of-band access?

    It's scoped rather than granted wholesale. On NVL72 and MGX-class platforms, Redfish is the primary out-of-band interface, and our standard position is scoped operator-level Redfish access on an isolated management VLAN rather than raw IPMI credentials.

    What physical security is in place?

    Secured perimeter fencing around the full site, professional-grade surveillance and dedicated on-site security personnel. Access is available 24/7/365 for delivery, deployment and maintenance, and it's logged and audited. Visitor registration, escort requirements and card and key management apply throughout.

    Can Helios see my data or models?

    No, except as strictly necessary to supply the service, to comply with law, or with your prior authorization. You keep all rights to your workload data, model artifacts and training data. Data is encrypted at rest and in transit, and we don't sell, rent or commercially exploit it.

    Delivery and operations

    How fast can you deliver?

    Approximately 120 days from reservation to energization: reserve on day 0, scope in week 1, build-out begins in week 2 after contract execution, and energize in month 4. It's achievable because land, power pathways, modular infrastructure and critical equipment are secured before contract execution. The target date goes into the contract.

    Full GPU clusters follow the same 120-day path: site and power commissioning, first pod build and fabric bring-up with acceptance testing, then further pods, super-spine, full-fleet validation and production handover.

    What is acceptance testing, and when does billing start?

    Each tranche is tested against a mutually agreed acceptance and benchmark runbook covering GPU health, NVLink and fabric validation, thermal and power verification and burn-in. Fees start only on acceptance, not on shipment or power-on. If a tranche doesn't conform, we remedy it and resubmit it for acceptance.

    What do you need from me to hit the date?

    Timely information, access and approvals, and no mid-flight changes to the order. The schedule is a joint commitment. On colocation specifically we need the hardware list, rack elevations, power and heat load, cooling interface, weights, network topology and access model early enough to design against.

    How is availability measured?

    A GPU server counts as available when it has external connectivity, runs the agreed benchmark workloads, and every GPU has NVLink connectivity within the server and fabric connectivity to the cluster, passing health checks on all of it. It's measured per server: if any single GPU has failed, that server is unavailable.

    On bare metal, fees accrue only for periods the infrastructure is available, and service credits apply when monthly service levels fall below the committed target in your SLA.

    What monitoring and support do I get?

    A 24/7 NOC monitors GPU and fabric health, with full-stack telemetry through DCGM, Prometheus and Grafana. Support is tiered by severity with response and resolution targets for each tier, and critical issues carry a sub-hour response target.

    You get the operational metrics and logs needed to verify SLA compliance through an agreed dashboard, secure API or machine-readable export.

    Who handles hardware failure?

    On bare metal, Helios handles hardware lifecycle, RMA, replacement, firmware and BIOS, and fees stop immediately while a server is down for physical replacement. On colocation the hardware is yours, and we provide remote hands for contracted tasks.

    How is planned maintenance handled?

    You get advance written notice for any change reasonably expected to affect performance or availability. Planned maintenance is capped as a share of monthly hours, and anything beyond the cap counts against the service level.

    Growth and comparisons

    How do I add capacity later?

    On compute, additional tranches under the same order or a new order against the documented expansion path; fabrics are engineered for the target scale from the start. On colocation, additional racks or kW within a shared unit for retail, or additional whole modular units for wholesale and build-to-spec. Expansion rights are defined in the contract.

    Can I reserve capacity for a future date?

    Yes, through a capacity reservation in the agreement with the energization date written into the contract. That's how we handle phased ramps where you need certainty on future capacity without paying for it before you can use it.

    What happens when the next GPU generation ships?

    Our builds are spec-driven rather than SKU-locked, and because the site and process are pre-built, the deployment timeline holds across GPU generations. Refresh terms, residual value and migration are handled per agreement.

    How does Helios compare to a hyperscaler?

    Three differences. Cost, driven by long-term power rather than discounting, and an all-inclusive rate with no egress bill. Density, because we were built for 100 kW-plus liquid-cooled racks rather than retrofitted from 10 kW enterprise colocation. And scope: a scoped build and a named team instead of a self-serve console and a support queue.

    What you give up is the surrounding cloud service catalog, which most AI infrastructure buyers aren't using anyway.

    How does Helios compare to other GPU neoclouds?

    Most neoclouds rent capacity in facilities they don't control, so their cost floor is someone else's lease and their availability is someone else's roadmap. We hold the power and operate the facility, which is what lets us commit to both pricing and delivery dates in a contract.

    Ask any provider you're comparing whether they own the power agreement, and what happens to your rate when their lease renews.

    How do I get started?

    Tell us whether you need GPU capacity, colocation or both, the capacity you need, your location and your target date. We'll start with configuration, capacity and timing.

    Request capacity
    Is Helios hiring?

    Yes. See our open roles and apply directly on the careers page.

    View open roles

    Still have a question?

    Tell us what you're planning. We'll start with configuration, capacity and timing.

    Talk to our team