When an FPGA Belongs in a Multi-Sensor Data Path

2026.10.06
When an FPGA Belongs in a Multi-Sensor Data Path

Six 1080p cameras, a LiDAR, an IMU, and GNSS can share one compute platform and still disagree about when things happened. The cameras alone produce about 4.5 Gbit/s (RAW12, 30 fps) before any AI model runs, and each source arrives through its own interface, at its own rate, with its own timing model. Deserializers with frame sync, PTP, DMA, and zero-copy software paths already handle this in many systems. The sections below cover where they stop being enough, and what an FPGA costs at that point.

TL;DR
  • Bandwidth rarely justifies an FPGA in a sensor data path. Keeping each measurement's acquisition timing attached to its payload through aggregation is the stronger reason.
  • An FPGA timestamp applied at the interface records arrival time. When the gap between acquisition and arrival varies, a fixed calibration offset cannot fully correct it.
  • A 1 ms timing error is 2 cm of travel at 20 m/s, or roughly 3 pixels at 90°/s with a 60° field of view across 1920 pixels. The timing budget follows from motion and from how the data is used.
  • Deserializers, PTP-capable NICs, and DMA cover many systems. An FPGA adds cost and verification work unless one data path needs functions that no single fixed-function device provides.
  • RDMA, RoCEv2, and GPUDirect RDMA act on different segments of the path, and Jetson's shared DRAM means the discrete-GPU copy diagram does not apply.
  • One row in the selection table rarely justifies an FPGA. The case strengthens when several rows move to the right at once.

01 | Heterogeneous Sensors Stress Different Parts of the Data Path

Cameras account for most of the bandwidth and mainly stress throughput and buffering. The IMU and GNSS add almost no bandwidth, yet localization and fusion are anchored to their timing. The lowest-rate sensors can be among the most timing-critical in the system.

The resulting architecture has to deal with several relationships at once:

Sensor interface
→
Acquisition timing
→
Buffering
→
Aggregation
→
Transport
→
Compute

For a stable system with known interfaces and manageable traffic, standard components such as camera deserializers, PTP-capable NICs, and DMA may cover that entire path.

02 | Arrival Time vs. Acquisition Time

Several timestamps can describe the same camera frame: exposure start, exposure end, row readout, serializer transmission, arrival at the FPGA, and availability in an application buffer.

Sensor fusion needs timing information that represents when the physical measurement was taken. An FPGA timestamp applied at the interface records arrival time.

The interval between acquisition and interface arrival can depend on sensor readout, link behavior, serializer or bridge buffering, packetization, and, for timestamps taken later in the path, downstream scheduling. When that interval is variable rather than deterministic, a fixed calibration offset cannot fully correct it.

Rolling-shutter cameras add another timing dimension. If a sensor needs 10 ms to read out a full frame, the top and bottom rows represent moments 10 ms apart. At 20 m/s, the vehicle moves 20 cm during that readout, and a single frame-level timestamp cannot fully describe the image unless the fusion stage also understands the sensor's readout timing.

LiDAR packets carry timestamps tied to the sensor's time source. IMUs often timestamp samples internally. Triggered cameras can derive acquisition timing from a shared hardware trigger. The aggregation path therefore has to preserve the timing source that belongs to each measurement rather than replacing every input with a uniform arrival timestamp.

Programmable logic can participate at several points: generating or distributing a common trigger, capturing frame-start or exposure-related events close to the interface, attaching sensor-native timestamps to outgoing packets, and keeping timing metadata bound to its payload through aggregation.

Common Clock / Trigger
↓
Sensor Acquisition
↓
Sensor Timestamp or Frame Event
↓
FPGA Aggregation
↓
Timing Metadata Preserved with Payload
↓
Compute

PTP or gPTP provides a shared time domain across participating devices. Hardware triggers or frame-sync signals control acquisition timing where the sensor architecture supports them.

Estimating the Acceptable Timing Error

No single timing requirement applies across autonomous systems. The budget follows from motion and from how the sensor data is used.

Translation
At 20 m/s (72 km/h), a 1 ms timing error corresponds to 2 cm of travel. At 33 m/s (120 km/h), it corresponds to 3.3 cm.
Rotation
Consider a camera with a 60° horizontal field of view across 1920 pixels, or about 32 pixels per degree. At an angular rate of 90°/s, a 1 ms timing error corresponds to 0.09°, or roughly 3 pixels.

Whether that error is acceptable depends on object distance, image resolution, feature scale, and the fusion algorithm.

Applying the same calculation to the platform's expected translational and angular motion provides a starting point for the timing budget. That budget can then guide the choice between software timestamps, NIC or interface-level hardware timestamps, sensor-native timing, and trigger-based acquisition.

03 | Why Not Just Use SerDes, PTP, a NIC, and DMA?

In many systems, that is exactly what should be done.

A GMSL deserializer can aggregate cameras. A PTP-capable Ethernet interface can provide synchronized hardware timestamps. DMA can transfer buffers without making the CPU copy each byte, and modern frameworks can keep data in zero-copy or shared-memory paths. Adding an FPGA to a system that already meets its bandwidth, timing, and processing requirements only increases cost and verification work.

The gap appears when one data path has to combine functions that no single fixed-function device covers. A camera hub can aggregate GMSL cameras but not Ethernet sensors. A NIC can handle Ethernet and hardware timestamping but not camera acquisition. Software can bridge the two, but filtering, timing association, and routing then move back into host scheduling, buffering, and memory management.

Programmable logic provides one place where interface handling, timing events, filtering, packetization, aggregation, and routing can be composed into the same hardware data path.

If the interfaces and processing rules are already fixed, an ASIC, fixed-function bridge, SmartNIC, MCU, or SoC accelerator may be a better fit.

04 | Moving the Result Toward Compute

The aggregated stream reaches compute through PCIe when the FPGA sits on the same board or in the same chassis as the processor. Ethernet becomes the practical choice when sensors are distributed, cable runs are long, or multiple compute nodes consume the same data.

Ethernet and RDMA

An RDMA endpoint implemented in programmable logic can handle packet framing, queue management, and memory-transfer control in hardware. On the receiving side, the compute node's RNIC writes incoming payloads into registered memory without requiring the host CPU to copy each packet.

Remote Sensors
↓
FPGA / RDMA Endpoint
↓
RoCEv2 over Ethernet
↓
Compute-node RNIC
↓
Host or GPU Memory

RoCEv2 is sensitive to packet loss and congestion, so its behavior depends on the network around it. A direct point-to-point link avoids most of that risk. A switched vehicle or robot network needs its loss and congestion behavior validated first; where that effort is not justified, UDP streaming with sequence numbers and receiver-side loss detection may be simpler.

Discrete GPU Systems

Without a direct peer path, incoming data may follow:

I/O Device → Host Memory → GPU Memory

Each frame crosses PCIe twice and consumes host memory bandwidth between transfers. GPUDirect RDMA allows a supported PCIe peer device to write directly into GPU memory:

I/O Device → GPU Memory

This requires a suitable PCIe topology, a supported GPU and driver, registered memory, and IOMMU settings compatible with peer-to-peer access.

Jetson-Class Integrated Systems

On Jetson platforms, the CPU and GPU share physical DRAM. There is no separate discrete GPU memory that every frame must be copied into, so the conventional host-memory-to-GPU-memory diagram describes the wrong bottleneck.

The relevant questions are how many times a frame moves between buffers within shared memory, whether cache maintenance is required for the selected memory type, and how many CPU-managed stages such as driver handling, sockets, or user-space handoff sit between data arrival and GPU consumption.

NVIDIA documents GPUDirect RDMA support on Jetson Orin, but the Tegra implementation uses a different memory-allocation model and platform interface from discrete-GPU systems. A driver or RDMA path developed for an x86 discrete-GPU platform generally requires platform-specific adaptation before it can be used on Jetson.

RDMA, RoCEv2, and GPUDirect RDMA act on different segments of the path. Whether a particular combination works depends on the FPGA's RDMA implementation, network adapter, PCIe topology, compute platform, driver stack, IOMMU configuration, and application memory model.

05 | Selection Criteria

Requirement Usually handled without FPGA FPGA becomes relevant when
Interfaces A GMSL deserializer or NIC covers every sensor Cameras, Ethernet sensors, and proprietary links must share one hardware path
Throughput Aggregate rate fits PCIe or Ethernet with DMA Streams require filtering, packetization, or routing at line rate
Timing PTP, hardware timestamps, and sensor-native timing meet the timing budget Trigger events, acquisition timing, and timestamps from different sources must remain associated through aggregation
Rate of change Interfaces and processing rules are stable Sensor sets differ across product variants, or preprocessing and routing rules are still evolving

A single condition in the right-hand column rarely justifies adding an FPGA. If one requirement can be handled by a deserializer, NIC, MCU, SoC, or software path, that option will usually carry less engineering overhead. The case becomes stronger when several rows move to the right at the same time, especially when high-rate concurrency, timing-sensitive acquisition, and changing data-path logic have to coexist.

Programmable logic adds device cost, power and thermal load, IP licensing and integration, timing closure, bitstream version management, driver development, and verification effort. Reprogramming an FPGA takes minutes; verifying the changed data path does not.

Recommended Products

These NeuronEDGE products each cover a different stage of the sensor-to-compute path:

TALO-B1000
FPGA-based platform for heterogeneous sensor ingress and programmable data-path integration.
TALO-F1200GU
GNSS/IMU platform with PPS, ToD, PTP, and gPTP support for positioning and timing reference.
TALO-N1000
Time-aware Ethernet switch with 1000BASE-T1, 10GbE, IEEE 802.1AS, and IEEE 1588 support.
TALO-A1000
NVIDIA Jetson Thor-based AI controller for downstream perception, sensor fusion, and autonomous compute workloads.