Announcing our partnership with Broadcom for the VMware VCF AI Factory.Read more

MetalSoft Fabric Manager Lab Now Available in NVIDIA DSX Air

MetalSoft Team·
MetalSoft and NVIDIA DSX Air logos on a dark rounded card over a near-black field of small green dots, squares and triangles, some joined by thin lines.

NVIDIA DSX Air is a cloud platform for simulating AI factory infrastructure before physical hardware arrives. With MetalSoft Fabric Manager inside it, enterprise AI teams and GPU cloud providers can deploy an NVIDIA Spectrum-X fabric, isolate tenants and move hosts between their networks entirely through automation, and evaluate that operating model ahead of their own build. Watch Alex Bordei, our CPO, put this into practice: deploying the fabric, provisioning isolated tenant networks and moving a host between tenants.


The MetalSoft Fabric Manager lab is now available in NVIDIA DSX Air, NVIDIA's cloud-based platform for simulating data centre infrastructure. Anyone with a DSX Air account can deploy a simulated NVIDIA Spectrum-X fabric, operate it with MetalSoft and see what it takes to run an on-premises AI factory network before the physical hardware arrives.

Three lab topologies are available: one scalability unit, two scalability units and a three-tier design, representing the network of 256-GPU and 512-GPU Spectrum-X reference architectures. Each lab builds the fabric from a clean state with the MetalSoft CLI and Terraform; no switch is configured by hand. For enterprise AI teams and GPU cloud providers, it is a practical way to evaluate fabric automation and multi-tenancy.

Automated fabric deployment

MetalSoft imports the switches, discovers their interfaces and links, assigns hostnames, ASNs and addressing, and renders the configuration from templates. Deployment then runs in stages, applying the base, underlay, overlay and QoS profiles in order and bringing up BGP and EVPN routing across every switch.

Every deployment is visible as a graph of dependent tasks showing progress, attempts and failures, and operators can inspect the resulting switch configuration to see exactly what changed.

Isolated tenant networks

MetalSoft separates administration of the shared fabric from provisioning of tenant infrastructure. Administrators prepare route domains and network profiles in Fabric Manager, while tenants define their endpoints and logical networks through Terraform or Infrastructure Designer. MetalSoft translates those definitions into configuration across the fabric.

Multiple tenants can be provisioned on the same fabric, each in its own route domain. In the walkthrough, connectivity checks confirm that hosts reach one another within their tenant while probes between the two tenants fail, so GPU clusters belonging to different teams or customers can share the physical infrastructure without sharing a network.

Moving hosts between tenants

As projects finish or capacity requirements change, hosts need to move between tenants. The lab demonstrates the network side of that move: an endpoint is removed from the first tenant's Terraform manifest and added to the second, MetalSoft updates the tenant network configuration on the fabric, and the connectivity checks that follow reflect the new membership.

Initial provisioning and later reassignment use the same infrastructure-as-code workflow, so teams can adapt tenant networks to changing allocations without touching individual switches.

Explore the lab

Watch Alex Bordei, our Chief Product Officer, deploy the fabric, provision two isolated tenants and move a host between their networks in his full walkthrough of the lab, then follow the documentation to try it yourself with one scalability unit, two scalability units or a three-tier topology.

Planning an AI factory or GPU cloud?
See how you can automate fabric deployment, provision isolated tenant networks and reallocate hosts as your teams’ needs change.

Share article

Back to all articles