For AI agents: the complete documentation index is available at https://docs.ovhcloud.com/en/llms.txt, the full documentation bundle is available at https://docs.ovhcloud.com/en/llms-full.txt, and this page is available as Markdown at https://docs.ovhcloud.com/en/guides/public-cloud/compute/deploy-8-h200-nvl-gpu-instance.md.

Deploying AI workloads on an 8x NVIDIA H200 NVL GPU instance

View as Markdown

Find out how to deploy AI workloads on an 8x NVIDIA H200 NVL GPU instance while respecting its NVLink topology

Objective

We recommend that you follow this guide if you have an h200-1920-eph or h200-1920 instance. 8x NVIDIA H200 NVL instances (h200-1920 and h200-1920-eph flavors) are organised into two groups of 4 GPUs, with NVLink connectivity only within each group and PCIe between the two groups. Taking this topology into account helps you get the best performance and stability for your AI workloads.

Info

At the moment, 8x NVIDIA H200 NVL instances are only available in the Paris region (EU-WEST-PAR).

This guide explains how to deploy and operate AI workloads on an 8x NVIDIA H200 NVL GPU instance

Requirements

  • A Public Cloud project with access to the region where the h200-1920 and h200-1920-eph instance models are available (EU-WEST-PAR)
  • An SSH key created to deploy a Linux GPU instance

OVHcloud Control Panel Access

  • Direct link:
  • Navigation path: Public Cloud > Select your project > Instances

Instructions

You will find below the information needed to deploy an 8x NVIDIA H200 NVL instance, understand its topology, then run and monitor your AI workloads on it.

Deploying the instance

In the Quick access section, click Create an instance. Then select the EU-WEST-PAR region and choose the h200-1920 or h200-1920-eph instance model in the Cloud GPU category.

Next, follow the remaining steps as detailed in the How to create a Public Cloud instance and connect to it guide. This process may take a few minutes.

Once the instance is delivered, install the NVIDIA driver as detailed in the Deploying a GPU instance guide.

Hardware overview

Characteristich200-1920h200-1920-eph
GPUs8x NVIDIA H200 NVL8x NVIDIA H200 NVL
GPU memory141 GB HBM3e per GPU (1,128 GB in total)141 GB HBM3e per GPU (1,128 GB in total)
GPU topology2 groups of 4 GPUs2 groups of 4 GPUs
Intra-group interconnectNVLink, within each group of 4 GPUsNVLink, within each group of 4 GPUs
Inter-group interconnectPCIePCIe
vCores224224
Memory (RAM)1,920 GB1,920 GB
Storage400 GB (root disk) + 4x 7.68 TB passthrough400 GB (root disk) + 1x 20 TB ephemeral
Public network25 Gbit/s25 Gbit/s
Private networkUp to 25 Gbit/sUp to 25 Gbit/s

Ephemeral disk (h200-1920-eph)

The ephemeral disk of the h200-1920-eph instance model is already formatted with an Ext4 filesystem. It may already be mounted automatically under /mnt by cloud-init.

To check whether the ephemeral disk is mounted, run the following command:

lsblk -f

If the 20 TB disk shows /mnt as its mount point, it is ready to use. Otherwise, mount it manually, replacing /dev/<device> with the device name shown by lsblk -f:

sudo mount /dev/<device> /mnt
Warning

Data stored on the ephemeral disk is not included in instance backups and is lost when the instance is deleted or shelved. We recommend backing up your important data to an external storage solution, such as Object Storage.

The 8 GPUs are organised as follows:

Diagram of two NVLink groups of 4 GPUs, linked to each other by PCIe
NVLink groupGPUsInterconnect within the groupInterconnect between groups
Group AGPU 0, 1, 2, 3NVLink (all 4 GPUs fully interconnected)PCIe
Group BGPU 4, 5, 6, 7NVLink (all 4 GPUs fully interconnected)PCIe

When allocating workloads on this instance, we recommend the following:

  • Whenever possible, run each job within a single NVLink group (4 GPUs).
  • When a single model needs all 8 GPUs, treat the two groups as two separate clusters: keep communication-intensive parallelism (such as tensor parallelism) within each group, and only use lighter communication (such as pipeline parallelism) between the two groups. An example is given in the recommended configurations below.
  • Keep in mind that NVLink performance benefits apply within each group of 4 GPUs only.
Info

Keeping communication-intensive operations within each NVLink group avoids heavy traffic on the PCIe bus, which ensures both better performance and better stability.

Before dispatching a workload, we recommend checking which GPUs belong to which NVLink group:

nvidia-smi topo -m

In the matrix, GPUs connected by NVLink show an NV# value between them (# being the number of NVLink links). GPUs connected through PCIe only show a PCIe path type instead (for example PHB, NODE or SYS). The 4 GPUs showing NV# values between them form one NVLink group.

You can also display the NVLink status of a given GPU (here GPU 0):

nvidia-smi nvlink --status -i 0

The examples below use vLLM to illustrate how to deploy a model on 4 or 8 GPUs while respecting the NVLink topology.

4 GPUs (single NVLink group)
8 GPUs (both NVLink groups)
Two replicas (one per NVLink group)

To deploy on 4 GPUs, explicitly specify the device IDs to make sure they all belong to the same NVLink group:

# Number GPUs in the same order as nvidia-smi (PCI bus order)
export CUDA_DEVICE_ORDER=PCI_BUS_ID

# Restrict the deployment to the first NVLink group (use 4,5,6,7 for the second group)
export CUDA_VISIBLE_DEVICES=0,1,2,3

# Set the tensor parallel size to 4
vllm serve <MODEL_SLUG> --tensor-parallel-size 4 ...

Operational best practices

We recommend that you:

  • confirm the GPU-to-NVLink group mapping before dispatching a workload (nvidia-smi topo -m).
  • keep all GPUs used by a single job within the same NVLink group, unless you use a cross-group configuration such as the pipeline-parallel one described above.
  • monitor NVLink status and kernel logs (dmesg) while your workloads are running.
  • regularly back up your important data.

We recommend monitoring the following metrics while your workloads are running:

MetricMethodRecommendation
NVLink status and errorsnvidia-smi nvlink --status / nvidia-smi nvlink -eRecommended
GPU detection and healthnvidia-smi -q / NVIDIA DCGM (dcgmi health)Recommended
GPU errors in kernel logssudo dmesg -T | grep -iE "xid|nvrm"Recommended
Info

GPU and NVLink errors are reported by the NVIDIA driver in the kernel logs as Xid messages. For continuous monitoring, NVIDIA DCGM can export these metrics to your monitoring stack. The smartctl and nvme commands are provided by the smartmontools and nvme-cli packages (for example sudo apt-get install smartmontools nvme-cli on Debian or Ubuntu).

Go further

Deploying a GPU instance

Join our community of users.

Was this page helpful?