PCIe DMA for high-frequency trading and packet processing

Nadi DMA

33 ns, doorbell to completion. On UltraScale+ silicon.

Every nanosecond spent in the DMA is budget taken from your algorithm. Nadi gives you the physical floor on UltraScale+ silicon.

33 ns doorbell → completion · 13 cycles @ 390.625 MHz · Alveo U280, 64-byte frames

Download the product brief

Evaluate

Try Nadi DMA on your silicon

NDA available.


Comparison with AMD/Xilinx XDMA and QDMA

Nadi DMA
33 ns · 1× 33 ns · 13 cycles at 390.625 MHz
64-byte frames on an Alveo U280
Xilinx XDMA
264 ns · 8× slower 264 ns · 8× slower than Nadi
Prefetch, poll mode, payload on card, one million samples
~1500 ns default · 45× slower ~1500 ns · 45× slower than Nadi
Same XDMA core, same card, no tuning
Xilinx QDMA
488 ns · 15× slower 488 ns · 15× slower than Nadi
Prefetch, poll mode, payload on card, one million samples
~1100 ns default · 33× slower ~1100 ns · 33× slower than Nadi
Same QDMA core, same card, no tuning

doorbell to completion (ns)

All three measured on the same Alveo U280. XDMA/QDMA tuned (descriptor pre-fetch, poll mode, one million samples). Dashed: same products, same card, no tuning.

Ultra Low-Latency DMA Engine

High-frequency trading

Nadi DMA completes a transaction in 33.2ns, the floor on AMD/Xilinx Ultrascale+ Silicon. By using Nadi DMA, you know you're getting the best performance out there.

Low-latency packet processing

On 64-byte frames, DMA latency is a large share of the per-packet budget. Lawful Interception and Low-Latency Packet processing require increasingly faster response times. With Nadi DMA, you know you're getting it.


What is included?

The IP Core is delivered as an encrypted IP.

A Linux VFIO userspace driver comes with it. For Networking applications, we include a 10/25G and 100G reference NIC with it's own Linux network device driver and Data-Plane Development Kit (DPDK) Driver.

Request an evaluation license to start integrating Nadi DMA into your design.


Frequently asked questions

What is Nadi DMA?

A PCIe DMA core for high-frequency trading and low-latency packet processing on AMD UltraScale+ (PCIE4 and PCIE4C). Doorbell to completion is 33 ns. Siliscale licenses it as encrypted IP from the Vivado catalog.

Who is it for?

Teams building high-frequency trading systems, and teams building low-latency packet pipelines, on an UltraScale+ card in a host.

What does 33 ns buy?

Tuned XDMA is 8× slower. The XDMA default is about 45× slower. Tuned QDMA is 15× slower. Same Alveo U280.

What software ships with it?

A Linux VFIO userspace driver. A 10/25G and 100G NIC reference, with a Linux network driver and a DPDK driver.

How is it licensed?

Evaluation is free and capped at about a million transactions. Production is a paid license for one project. Use the form on this page for either.


Get in touch

Related: Xilinx QDMA IP core tutorial · All FPGA IP cores