
InfiniBand: Supported, GPUDirect RDMA, 8 x 200 Gigabit HDR Ultra Disks: Supported ( Learn more about availability, usage, and performance)

Azure HPC Ubuntu 18.04, 20.04 and Azure HPC CentOS 7.9 images are supported.

The Azure HPC images are strongly recommended.
Ndm stockhouse driver#
To get started with NDm A100 v4 VMs, refer to HPC Workload Configuration and Optimization for steps including driver and network configuration.ĭue to increased GPU memory I/O footprint, the NDm A100 v4 requires the use of Generation 2 VMs and marketplace images. Additionally, the scale-out InfiniBand interconnect is supported by a large set of existing AI and HPC tools that are built on NVIDIA's NCCL2 communication libraries for seamless clustering of GPUs. These instances provide excellent performance for many AI, ML, and analytics tools that support GPU acceleration 'out-of-the-box,' such as TensorFlow, Pytorch, Caffe, RAPIDS, and other frameworks. These connections are automatically configured between VMs occupying the same VM scale set, and support GPUDirect RDMA.Įach GPU features NVLINK 3.0 connectivity for communication within the VM, and the instance is backed by 96 physical 2nd-generation AMD Epyc™ 7V12 (Rome) CPU cores. Each GPU within the VM is provided with its own dedicated, topology-agnostic 200 GB/s NVIDIA Mellanox HDR InfiniBand connection. NDm A100 v4-based deployments can scale up to thousands of GPUs with an 1.6 TB/s of interconnect bandwidth per VM. The NDm A100 v4 series starts with a single VM and eight NVIDIA Ampere A100 80GB Tensor Core GPUs.

It's designed for high-end Deep Learning training and tightly coupled scale-up and scale-out HPC workloads. The NDm A100 v4 series virtual machine(VM) is a new flagship addition to the Azure GPU family. Applies to: ✔️ Linux VMs ✔️ Windows VMs ✔️ Flexible scale sets ✔️ Uniform scale sets
