Versal ACAP AI Engine WiFi CSI nexmon PetaLinux Edge AI

Seeing Through Walls: WiFi Human-Activity Recognition on the Versal AI Engine

A weekend hardware/software co-design that streams WiFi Channel-State-Information (CSI) from a Raspberry Pi into an AMD Versal VCK190, parses the UDP/CSI packets in the Programmable Logic, and runs the STFT + activity / breathing / fall feature extraction on the AI Engine array — turning ordinary WiFi into a device-free room sensor. Validated bit-accurate on real silicon.

5.96e-08
max abs err vs golden
3-branch
AIE feature graph
L = 256
sample window
4 ILAs
on the live datapath
10+
related publications
Act 1 — Origins
It started as a thesis

My ECE Master's at Mississippi State University was in software-defined-radio (SDR) based sensing. One project asked a deceptively simple question: can an ordinary WiFi signal tell what a person is doing in a room?

The answer is yes — and it comes from Channel-State-Information (CSI). Every OFDM WiFi receiver estimates how the channel distorted each subcarrier. When a body moves, the multipath reflections change, and those changes carry a signature of the motion. I built a WiFi-CSI human-activity recognition (HAR) pipeline, benchmarked it against a radar sensing baseline for accuracy, and published the results across several papers (see Publications).

The visual heart of that work is the spectrogram: take the Short-Time Fourier Transform of the CSI time series and each activity paints a distinct time–frequency fingerprint. These are real CSI spectrograms from the dataset — one per activity class:

Act 2 — A gentle challenge
"You could try to do this in an FPGA"

When I presented that thesis to a former manager, they gently suggested that the whole pipeline could, one day, run in an FPGA — the packet parsing in fabric, the heavy DSP and inference on an accelerator. It was said kindly, almost in passing. But it lodged in my head as a personal challenge I promised myself I'd take on someday.

"Someday" meant: get the CSI off the CPU entirely — parse the UDP packet in the Programmable Logic, and let a Versal AI Engine do the STFT and the activity inference.

The napkin sketch of that idea looked like this:

flowchart LR PI["WiFi CSI source"] -->|UDP over Ethernet| FPGA subgraph FPGA["FPGA — do it all here"] PL["PL: parse packets"] --> ACC["AI Engine: STFT + inference"] end ACC --> OUT["activity / breathing / fall"]
Act 3 — Picking it back up
From the shelf to silicon

The idea sat on the shelf while life moved on. What finally unblocked it was hardware and time: a VCK190 on the bench and a run of weekends and evenings. Modern tooling — and, honestly, LLMs for grinding through the deep Versal tool-flow issues — made a solo build of this scope tractable.

The origin

Master's thesis

  • WiFi-CSI human-activity recognition
  • benchmarked against radar
  • published papers
The nudge

The challenge

  • a gentle "do it in an FPGA"
  • filed away as a someday goal
The pause

Shelved

  • other priorities
  • idea kept warm
The rebirth

VCK190

  • PL UDP/CSI parser
  • AIE feature graph
  • bit-accurate on silicon
Act 4 — The rebirth
The thesis, now implemented in an FPGA

The pipeline now runs on real hardware. A Raspberry Pi running the nexmon_csi firmware patch turns its WiFi chip into a CSI sensor and packetizes CSI as UDP over Ethernet. That stream enters the VCK190 over an SFP link; the PL parses the UDP/CSI packet and streams the I/Q samples straight into the AI Engine, which computes the motion, breathing and phase-variance features. Only the compact results cross into Linux on the Arm cores, where an on-board web dashboard renders what's happening in the room.

AMD Versal VCK190 evaluation board
The AMD Versal VCK190 — Arm Cortex-A72 PS, Programmable Logic, and the AI Engine array on one ACAP.

Why this device is the right fit:

  • PL streams and parses the Ethernet/UDP packet at line rate.
  • The AI Engine is a vector-VLIW array built for FFTs and MAC-heavy DSP — exactly the CSI workload.
  • The NoC + DDR and Arm PS handle results and visualization.
  • It's a single-chip edge sensor: no PC in the loop.
Act 5 — How it actually works
End-to-end dataflow

The full path from WiFi chip to browser. Raw CSI (hundreds of KB/s) never touches the CPU in the target design — only the tiny feature/label vector is DMA'd to Linux. Click any diagram to open it fullscreen and zoom / pan.

flowchart LR subgraph PI["Raspberry Pi — nexmon_csi"] W["WiFi NIC
bcm43455c0"] --> CAP["CSI extractor"] --> UDP["UDP packetizer"] end UDP -->|"UDP/IP over 1000BASE-X"| SFP["SFP0"] subgraph VCK190["AMD Versal VCK190"] SFP --> GT["GTY + PCS/PMA"] --> MAC["Ethernet MAC"] MAC --> P["PL: csi_udp_parser
strip UDP → I/Q stream"] P --> MUX["csi_mux"] --> AIE subgraph AIE["AI Engine — feature_graph"] FIR["FIR band-pass"] --> ST["stats
mean / var / power"] DFT["windowed-DFT
STFT magnitude"] PH["phase-var
stats"] end AIE --> S2MM["s2mm DMA"] --> DDR[("DDR4 via NoC")] DDR --> A72["Arm A72 · PetaLinux"] A72 --> DASH["live_dashboard.py
SSE web UI"] end DASH --> BROWSER["any browser on the LAN"]
Why WiFi can see you

No camera, no wearable, often through a wall. The trick is that your body is part of the radio channel. Here is the intuition, in four steps.

📡

1. CSI is the channel's fingerprint

An OFDM receiver estimates, per subcarrier k, a complex response — amplitude and phase of how the signal arrived. nexmon_csi exposes this vector every packet.

$$H[k] = |H[k]|\,e^{\,j\angle H[k]}$$
🧍

2. You perturb the multipath

The receiver sums many rays: line-of-sight plus reflections off walls, furniture — and you. Move, and the reflected path length changes, rotating its phase and reshaping the summed H[k] over time.

🌊

3. Motion → Doppler

A body part moving at velocity v imparts a Doppler shift. Different activities have different velocity signatures, so the STFT separates walking from falling from sitting.

$$f_d = \frac{2v}{\lambda}, \quad \lambda \approx 12.5\,\text{cm at 2.4 GHz}$$
🫁

4. Breathing → slow phase tone

The chest wall moves millimeters at ~0.2–0.4 Hz. That appears as a narrow low-frequency tone in the CSI — the basis of the breathing branch.

Why STFT. Activities are non-stationary, so we take the Short-Time Fourier Transform of the CSI time series — window, DFT, magnitude — producing the time–frequency spectrograms you saw in Act 1. (This directly builds on my paper Effect of the Short-Time Fourier Transform on the Classification of Complex-Valued Mobile Signals.)

$$X[m,\,\omega] = \sum_{n} x[n]\,w[n-mR]\,e^{-j\omega n}$$
The windowed DFT: w is the analysis window, R the hop. The AI Engine computes the magnitude of this per window.

WiFi vs. radar — the thesis comparison

Radar is dedicated, high-bandwidth and precise, but it's extra hardware. WiFi CSI reuses ubiquitous infrastructure: it's device-free, works in the dark and through walls, and is privacy-preserving (no imagery). My thesis and our RadarConf 2024 paper comparing Wi-Fi-CSI and radar-based HAR quantified that accuracy trade-off — the motivation for pushing the WiFi pipeline all the way onto an edge accelerator here.

Parsing UDP/CSI in the fabric

In the target architecture the PL owns the packet. RX frames enter as an AXI4-Stream; a custom csi_udp_parser (Vitis HLS) filters the CSI UDP port, strips the Ethernet/IP/UDP headers, extracts the I/Q payload, and streams samples through a csi_mux straight into the AI Engine. This is the real Vivado block design that was built and run on the board:

Vivado block design: Ethernet → csi_udp_parser → csi_mux → AI Engine → DMA → NoC/DDR
Vivado IP-Integrator block design — Ethernet (GT + PCS/PMA + MAC) → csi_udp_parser → csi_mux → ai_engine → s2mm DMA → NoC → DDR4. Exported with write_bd_layout from the built project.

A simplified view of the same PL datapath, for reference:

flowchart LR MAC["Ethernet MAC"] --> RXW["rx_dwidth
(width conv.)"] RXW --> PAR["csi_udp_parser
(HLS)"] PAR --> MUX["csi_mux"] MUX --> SHIM["AIE ↔ PL shim
(v++ co-generated)"] SHIM --> AIE["ai_engine_0"] AIE --> S2MM["s2mm"] --> NOC["NoC"] --> DDR[("DDR4")]

Under the hood the parser is a one-byte-per-cycle state machine (Vitis HLS, II=1): it walks each Ethernet frame byte by byte, checks a handful of fixed offsets to accept or drop the frame, and reassembles the little-endian int16 I/Q payload into 32-bit words. The whole wire contract is just these offsets:

Byte(s)FieldParser action
12–13EtherTypekeep only 0x0800 (IPv4)
23IP protocolkeep only 17 (UDP)
36–37UDP dest portkeep only the CSI port (5500)
44nexmon RSSIint8 dBm → metadata
52–53sequence no.→ metadata (packet-loss tracking)
54core / spatial→ metadata
56–57chanspecchannel + bandwidth → metadata
60 →CSI payloadint16 LE real, then imag, per subcarrier (up to 256)

The complex pair. From byte 60 the parser takes four wire bytes at a time — two for the real part, two for the imaginary — sign-extends each int16, and packs them into one 32-bit AXI4-Stream beat. That packed word is the complex CSI sample H[k] = I + jQ for a single subcarrier:

hls/csi_udp_parser.cpp — reassembling one complex subcarrier from four wire bytes
// CSI payload from byte 60: int16 LE real, int16 LE imag.
ap_int<16> re; re(7,0) = b0; re(15,8) = b1;
ap_int<16> im; im(7,0) = b2; im(15,8) = b;
axis_csi o;
o.data(15,0)  = (ap_uint<16>)re;   // real -> low half
o.data(31,16) = (ap_uint<16>)im;   // imag -> high half
o.last = beat.last;                 // TLAST marks end-of-CSI for this frame
csi_out.write(o);                   // one 32-bit beat per subcarrier
n_sub++;

So the parser emits two synchronized AXI4-Streams per WiFi frame: csi_out, the run of complex subcarriers, and meta_out, a single 64-bit record carrying everything the downstream needs to normalize and label the window:

StreamWidthPayload
csi_out32-bit{ imag[31:16], real[15:0] } × up to 256 subcarriers
meta_out64-bit{ seq, rssi, n_sub, chanspec, core_spatial } — one packed record per frame
flowchart LR RX["RX byte stream
1 B / cycle · TLAST = frame"] --> FILT{"IPv4? · UDP?
port 5500?"} FILT -->|"no"| DROP["discard frame"] FILT -->|"yes"| HDR["latch header bytes
44 · 52 · 54 · 56"] HDR --> ASM["reassemble int16 I/Q
from byte 60"] ASM --> CS["csi_out
32-bit complex
{imag, real}"] HDR --> MT["meta_out
64-bit record
seq·rssi·n_sub·chanspec"]

The whole parser costs about 1.3% of the PL LUTs, zero DSPs and a few BRAMs — line-rate protocol offload is exactly what the fabric is cheapest at, which is why it lives here and not on the Arm.

From I/Q pairs to feature windows

The parser produces per-packet complex vectors, but the AI Engine branches want real-valued time windows — a stretch of many packets for one subcarrier. Bridging the two is a short conditioning chain, and deciding where each step runs — PL, AI Engine, or Arm — is the heart of the hardware/software co-design.

flowchart LR IQ["complex I/Q
per subcarrier"] --> AP["magnitude / phase
via CORDIC"] AP --> SAN["phase sanitize
linear de-trend"] SAN --> SEL["subcarrier select
variance top-K"] SEL --> WIN["window / ring-buffer
L=256 · N=64"] WIN --> AIE["→ AI Engine branches"]

Magnitude feeds the motion and breathing branches; phase feeds the phase-var branch. Phase first gets a per-packet linear de-trend across subcarriers, which cancels the carrier-frequency offset (the intercept) and the sampling/packet-delay offset (the slope) that otherwise swamp the real motion:

$$|H[k]| = \sqrt{I^2 + Q^2}, \qquad \angle H[k] = \operatorname{atan2}(Q,\,I)$$
Magnitude and phase per subcarrier — a CORDIC does both in the PL at line rate.

Where each stage of the pipeline lands is a deliberate partitioning decision:

StageRuns onWhy there
UDP/CSI parse · I/Q depacketizePLcheap streaming bit-manipulation at line rate
|H| / ∠H (CORDIC)PLCORDIC is a fabric primitive
phase sanitize · FFT / STFT · FIRAI Enginevectorized float MACs and butterflies
subcarrier select · windowingPL + DDRlight control logic + data movement
thresholds · labels · dashboardArm A72control plane and UX

Rule of thumb across the whole design: streaming & bit-twiddling → PL; FFT / GEMM / filters → AI Engine; control & UX → the Arm.

The feature graph, kernel by kernel

Those three conditioned windows are exactly what the graph waits on — each lands on its own PLIO stream input, and from here the whole pipeline stays on the AI Engine array. It runs a three-branch feature graph, one branch per window: the two amplitude windows drive motion and breathing, while the sanitized-phase window drives phase-var.

Each branch maps directly to a piece of the physics above, and every kernel is a small streaming function connected window-to-window over PLIO:

flowchart TB IN["PLIO_in
amplitude · 256"] --> FIR["kfir: FIR band-pass"] --> STA["kstats: mean/var/power"] --> O1["PLIO_out
motion {mean,var,power}"] BIN["PLIO_brt_in
amplitude · 64"] --> DFT["kdft: windowed-DFT magnitude"] --> O2["PLIO_brt_out
NB breathing bins"] PIN["PLIO_phase_in
phase · 256"] --> PHA["kphase: stats"] --> O3["PLIO_phase_out
phase-var {mean,var,power}"]
BranchKernel chainOutputPhysical meaning
MotionFIR band-pass → stats{mean, var, power}band-limit to the human-motion Doppler band, then measure energy → presence / activity intensity
Breathingwindowed-DFT magnitudeNB binsSTFT bin near 0.2–0.4 Hz → respiration rate
Phase-varstats{mean, var, power}phase variance → micro-motion / stability, feeds fall detection

Why three? Human activity shows up in the CSI at three different mechanisms and time-scales at once, so each branch is tuned to one of them and they run in parallel on separate AI Engine tiles:

🏃

Motion — fast amplitude

A 31-tap band-pass FIR keeps only the ~1–3 Hz human-Doppler band of a 256-sample amplitude window; stats then reports its variance / power. High energy means someone is moving.

🫁

Breathing — slow periodic

A 64-point Hann-windowed DFT turns a slow amplitude window into 33 magnitude bins; the peak in the 0.1–0.5 Hz band is the respiration tone (6–30 breaths/min).

📶

Phase-var — micro-motion

After sanitization, the variance of the phase over a 256-sample window is a scale-immune presence cue — it survives when amplitude is ambiguous, and it feeds fall detection.

The graph wiring is pure ADF — kernels connected stream-to-stream on the array:

aie/src/graph.cpp — the three-branch feature graph
// motion branch: FIR band-pass -> stats (kernel-to-kernel)
connect<>(in.out[0],    kfir.in[0]);
connect<>(kfir.out[0],  kstats.in[0]);
connect<>(kstats.out[0], out.in[0]);
// breathing branch: windowed-DFT magnitude
connect<>(brt_in.out[0], kdft.in[0]);
connect<>(kdft.out[0],   brt_out.in[0]);
// phase-variance branch: stats -> variance is the phase-var feature
connect<>(phase_in.out[0], kphase.in[0]);
connect<>(kphase.out[0],   phase_out.in[0]);
aie/src/stats_core.h — two-pass mean / variance / power over one window
static inline void stats_core(const float *x, float *out) {
    float s = 0.0f;
    for (int i = 0; i < L; i++) s += x[i];
    float mean = s / (float)L;
    float sd = 0.0f, s2 = 0.0f;
    for (int i = 0; i < L; i++) {
        float d = x[i] - mean;
        sd += d * d;
        s2 += x[i] * x[i];
    }
    out[0] = mean;          // DC / bias
    out[1] = sd / (float)L; // variance  == motion-band / phase-var energy
    out[2] = s2 / (float)L; // power
}

Per window the three branches emit 39 floats — mot[3], brt[33], phs[3]. Together with the RSSI from the parser metadata, those collapse into the compact 8-feature vector the classifier actually consumes:

#FeatureWhere it comes from
0presencemotion + phase-var vs. threshold
1motionmotion-branch variance / power
2breathing (bpm)breathing-branch peak DFT bin
3heart ratereserved (unreliable at 20 Hz CSI)
4phase-varphase-var-branch variance
5personsclassifier
6fallmotion transient + debounce
7RSSIparser metadata, (rssi+100)/100
flowchart LR M["motion
{mean,var,power}"] --> V["8-feature
vector"] B["breathing
33 DFT bins"] --> V P["phase-var
{mean,var,power}"] --> V R["RSSI (metadata)"] --> V V --> MLP["pretrained MLP
8→64→128 · L2"] --> LBL["activity label
+ presence / vitals"]

The classifier is deliberately tiny — a pretrained 8→64→128 MLP (~9k parameters, GELU + batch-norm, L2-normalized embedding) plus a linear presence head. All the heavy lifting is the DSP feature stage above: the vector-VLIW AI Engine chews through the FIR taps and DFT butterflies far faster than the scalar A72 could — which is the whole point of the offload.

The subtle bug that cost the most: the AIE ↔ PL shim

A plain Vivado block design that wires a PL stream into the AI Engine's S00_AXIS stalls on silicon — TREADY never asserts. The reason: a hand-built BD does not emit the paired AIE↔PL shim solution that the graph's PLIO placement requires. The binding must be co-generated by v++ --link. The fix inserts the PL stream on v++'s own post-SysLink block design via an overlay hook, leaving ai_engine_0 untouched so the correct shim solution is preserved.

Results to the Arm — NoC, DDR, dashboard

The last hop is deliberately thin. Two s2mm data movers — one for the AI Engine features, one for the parser metadata — are PL AXI-MM masters that push their results across the Versal NoC into a small reserved region of DDR4. The Arm A72 running PetaLinux reads that region and serves the dashboard. The multi-hundred-KB/s raw CSI stream never crosses into Linux — only a few hundred bytes of features and metadata per window do.

flowchart LR AIEO["AIE features
39 floats / window"] --> S["s2mm"] METAI["parser metadata"] --> SM["s2mm_meta"] S --> NOC["Versal NoC"] SM --> NOC NOC --> DDR[("DDR4
reserved no-map
@ 0x7000_0000")] DDR --> A72["Arm A72
PetaLinux"] A72 --> DASH["live_dashboard.py
SSE web UI"]

Because s2mm and s2mm_meta are driverless PL masters with no IOMMU, they write to a 1 MB no-map carve-out that Linux does not own: the packed metadata record sits at the base and the feature results 64 KB in, so the two never share a page. Everything is memory-mapped at fixed addresses that the on-target control tool (sw/csi_ctl.c — mac-init, start, mux, meta) and the reader agree on:

BlockAXI base
mm2s — DDR → AIE (golden path)0xA400_0000
s2mm — AIE features → DDR0xA401_0000
csi_udp_parser0xA402_0000
s2mm_meta — metadata → DDR0xA403_0000
csi_mux — AXIS switch0xA406_0000
ethernet MAC0xA408_0000
ai_engine0x20_0000_0000

The csi_mux is what makes both worlds testable: flip it one way and the AI Engine is fed by the live Ethernet parser; flip it the other and it's fed by mm2s replaying a known CSI window from DDR — the bit-exact golden test that anchors every silicon bring-up below.

Integrated Logic Analyzer captures

Talk is cheap — here is the datapath caught in the act on a real VCK190 (board chanterelle10). Four ILA taps were inserted on the live pipeline via the Versal debug automation: ETH_RX (nexmon UDP bytes), CSI_PARSED (parser output), AIE_IN (mux → AI Engine) and AIE_OUT (AIE features → DMA).

ILA tap map across the four datapath probes
Where each ILA tap carries valid data during a capture. AIE_IN and AIE_OUT are active; ETH_RX / CSI_PARSED light up on the live-Pi path.
256-sample WiFi CSI window entering the AI Engine, captured on ILA
The AIE_IN tap: 256 real WiFi-CSI samples streaming into the AI Engine, one AXI4-Stream beat each (float32, e.g. 0x3f0ec77a = 0.558).
AI Engine feature output matching the golden model
The AIE_OUT tap: the three feature values the AI Engine computed for that window — bit-identical to the software golden model.

The end-to-end numerical check, on hardware:

AIE result: mean=-0.002065 var=0.302189 power=0.302193 golden: mean=-0.002065 var=0.302189 power=0.302193 max_abs_err=5.96e-08 -> PASS
Act 6 — See it live
A room, sensed in real time

The VCK190 serves its own dashboard (pure-Python, standard-library only, no CDN — it runs on a bare PetaLinux rootfs). It consumes the AI Engine feature groups and derives presence, activity, breathing rate and fall alerts — and renders a live CSI spectrogram by stacking the streamed brt[33] windowed-DFT magnitude vectors over time (the same STFT view from Act 1, now updating in real time). Below is that dashboard's logic, replaying a representative feature stream right in your browser — empty room → someone walks in → sits and breathes → a fall → stillness:

loading live replay…

The widget above is a representative feature-stream replay running the same metrics and thresholds as the on-board dashboard — it is not a live capture. On the board, these tiles update from live AI-Engine features.

Live demo video
COMING SOON

A full live demo — Raspberry Pi streaming CSI into the VCK190, the room being sensed in real time — is being recorded. Until then, the interactive dashboard above is the teaser.

War-stories: the bugs worth remembering

The gap between "works in simulation" and "works on silicon" was where the real engineering happened. A few that cost days:

The AIE ↔ PL shim must be co-generated by v++
A hand-built Vivado BD wiring a PL stream into S00_AXIS stalls — TREADY never asserts — because it lacks the v++-generated shim solution (aieshim_solution.aiesol / aie_pl_intf.json) the PLIO placement requires. Fix: insert the PL stream on v++'s own post-SysLink BD through an overlay hook, leaving ai_engine_0 untouched. The resulting XSA carries the full shim solution and the DDR → mm2s → AIE → s2mm → DDR stream is bit-accurate on silicon.
The debug build crashed the placer — an unplaced GT quad
Adding the ILA debug cores crashed place_design in "Phase 5.1 Post Place Optimization". Root cause wasn't the ILAs alone: the co-generated Ethernet GT quad was left unplaced (the external SFP ports aren't top-level accessible under v++, and the OOC gt_quad can't be LOC'd from a top XDC). Fix: an implementation OPT_DESIGN.TCL.POST hook (gtloc_hook.tcl) LOCs the leaf cells after link — quad_inst → GTY_QUAD_X0Y5, IBUFDS_GTE5 → GTY_REFCLK_X0Y10. With both placed, place + route + write_device_image all pass.
Three integration fixes to make the AIE actually run under Linux
The SDT/PetaLinux flow doesn't wire everything the AI Engine needs, so three patches were required: (1) add interrupts-extended (the CU IRQs) to the zocl / zyxclmm_drm device-tree node; (2) bake the aie_image graph CDO into BOOT.BIN; (3) drop graph.wait() in the host, because the packaged CDO free-runs the graph — you drain s2mm instead.
The Ethernet image boots — it was a timing race, not a dead board
The larger Ethernet PDI appeared to hang PS Linux, but u-boot was actually fine: the failure was a race where u-boot auto-booted before the ~251 MB initramfs finished downloading over JTAG. Deterministic fix: interrupt u-boot autoboot over the console, download Image / boot.scr / dtb / rootfs, then booti. With that, the eth image boots Linux 6.12.40 and its ILAs capture live data end-to-end.
Technology & Tools
LayerTechnologyRole
CSI sourceRaspberry Pi + nexmon_csiWiFi NIC → per-packet CSI → UDP
Link1000BASE-X SFPPi Ethernet → VCK190 GTY
DeviceAMD Versal VCK190PS (A72) + PL + AI Engine ACAP
Packet parseVitis HLS csi_udp_parserUDP/CSI depacketize in PL
DSP + inferenceAI Engine (ADF graph)FIR / DFT / stats feature extraction
Data movementmm2s / s2mm + NoC + DDR4PLIO ↔ DDR, results to PS
ToolsVivado / Vitis 2025.2synth, AIE compiler, v++ link
Embedded OSPetaLinux 2025.2 + XRTA72 Linux, board host app
VisualizationPython stdlib SSE dashboardon-board live web UI
Publications

The research this build stands on. Full profile on Google Scholar.

Software Defined Radio (SDR) Based Sensing — Master's Thesis
Ajaya Dahal · Advisor: Ali C. Gurbuz
M.S. Electrical & Computer Engineering, Mississippi State University · 2024
SDRWiFi CSIRadarHAR
Comparison Between Wi-Fi-CSI and Radar-Based HAR
A. Dahal, S. Biswas, S. Z. Gurbuz, A. C. Gurbuz
IEEE International Radar Conference (RadarConf) · 2024 · doi:10.1109/RadarConf2458775.2024.10548515
WiFi vs RadarHARCSI
Robustness Analysis of Wi-Fi-Based Human Activity Recognition
A. Dahal, S. Biswas, S. Z. Gurbuz, A. C. Gurbuz
Proc. SPIE 13036, Defense + Commercial Sensing · 2024 · doi:10.1117/12.3014010
HARWiFi CSIRobustness
Wi-Fi Fingerprinting Based Room-Level Classification: Combining STFT and Imbalanced Learning
F. Islam, J. Farmer, A. Dahal, B. Tang, J. E. Ball, M. Young
Proc. SPIE 12122, Defense + Commercial Sensing · 2022 · doi:10.1117/12.2618717
STFTLocalizationML
Effect of the Short-Time Fourier Transform on the Classification of Complex-Valued Mobile Signals
L. Smith, N. Smith, S. Kodipaka, A. Dahal, B. Tang, J. E. Ball, M. Young
Proc. SPIE 11756, Defense + Commercial Sensing · 2021 · doi:10.1117/12.2587664
STFTComplex-valued
Real-Time Location Fingerprinting for Mobile Devices in an Indoor Setting
N. Smith, L. Smith, S. Kodipaka, A. Dahal, B. Tang, J. E. Ball, M. Young
Proc. SPIE 11756, Defense + Commercial Sensing · 2021 · doi:10.1117/12.2587679
LocalizationFingerprinting
Automatic Waveform Recognition from Complex RF Data with Filter-Based Deep Learning — B. Hicks, S. Biswas, A. Dahal, V. González, A. C. Gurbuz · NAECON 2024
Software Radio with MATLAB Toolbox for 5G NR Waveform Generation — W. AlQwider, A. Dahal, V. Marojevic · DCOSS 2022
Open-Source Software Radio Platform for Research on Cellular Networked UAVs: It Works! — V. Marojevic et al. · IEEE Communications Magazine 2022
RGB Pixel-Block Point-Cloud Fusion for Object Detection — T. Foster, A. Dahal, J. E. Ball · Proc. SPIE 11748 2021
Adversarial Indoor Signal Detection — S. Kodipaka, A. Dahal, L. Smith, N. Smith, B. Tang, J. E. Ball, M. Young · Proc. SPIE 11756 2021
Build it yourself

The whole project is open source. I've deliberately made it turnkey — a single project_top.tcl rebuilds the Vivado design end-to-end, and the tracked petalinux/ project directory has everything you need to build the Linux image for the board.

Fully reproducible builds — a standard I hold on every reference design

Reproducibility isn't an afterthought here; it's how I ship. Across all of my reference designs the entire Vivado project rebuilds from a single tracked TCL script and the PetaLinux image from a tracked project-spec — no GUI click-throughs, no undocumented state, so anyone can clone the repo and rebuild it bit-for-bit. You'll find the exact same discipline in my Versal Ethernet and ZCU102 Ethernet reference designs. It's the same production-grade, source-controlled methodology I apply to the AMD reference designs I author professionally.

🧩

One-command Vivado build

hw/scripts/project_top.tcl conveniently builds the inline Vivado project (Ethernet → csi_udp_parser → AI Engine → DMA) into an XSA end-to-end — no manual block-design wiring.

🐧

Ready-to-build PetaLinux

The petalinux/ directory ships the full project-spec/ (recipes + configs), the SDT flow, and the boot / rootfs bring-up scripts — everything needed to build and boot this project on the VCK190.

⚙️

AIE graph + HLS parser

The AI Engine feature graph (aie/), the Vitis HLS UDP/CSI parser (hls/), the v++ co-generation flow and the on-board dashboard (live/) are all in the repo, with a phase-by-phase build log in docs/.

Raspberry Pi + nexmon_csi → UDP over Ethernet

The Versal pipeline is only as real as the CSI feeding it. The transmitter is a Raspberry Pi 4B whose Broadcom bcm43455c0 Wi-Fi chip is reflashed with the nexmon_csi firmware patch: it turns every received OFDM frame into a per-subcarrier CSI vector and emits it as a UDP packet on port 5500. A small boot-time service then relays that stream out the Pi's Ethernet port to the VCK190's SFP.

flowchart LR subgraph PI["Raspberry Pi 4B — nexmon_csi"] W["Wi-Fi bcm43455c0
monitor · ch132/80"] --> FW["CSI firmware
7.45.189 nexmon.org/csi"] FW --> U["UDP :5500 on wlan0
10.10.10.10 → 255.255.255.255"] U --> SVC["csi-forward.service
wlan0 → eth0"] end SVC -->|"UDP/IP · 1000BASE-X SFP"| VCK["VCK190 SFP0 → csi_udp_parser → AI Engine"]

This Pi runs Raspberry Pi OS Trixie on kernel 6.18 — far newer than nexmon's classic brcmfmac-patch targets — so it uses the maintainer's recent-kernel path (Makefile.rpi + nl80211 vendor commands, no modified driver).

1 — dependencies (the 32-bit armhf cross-libs are mandatory on a 64-bit image)
sudo apt install -y git libgmp3-dev gawk qpdf bison flex make autoconf \
     libtool texinfo xxd libnl-3-dev libnl-genl-3-dev bc libssl-dev tcpdump
sudo dpkg --add-architecture armhf && sudo apt update
sudo apt install -y libc6:armhf libisl23:armhf libmpfr6:armhf libmpc3:armhf libstdc++6:armhf
sudo ln -sf /usr/lib/arm-linux-gnueabihf/libisl.so.23 /usr/lib/arm-linux-gnueabihf/libisl.so.10
sudo ln -sf /usr/lib/arm-linux-gnueabihf/libmpfr.so.6 /usr/lib/arm-linux-gnueabihf/libmpfr.so.4
# the b43 disassembler also needs python2.7 (from the Debian archive)
2 — build the firmware toolchain, nexutil, and the CSI patch
git clone --depth 1 https://github.com/seemoo-lab/nexmon.git && cd nexmon
source setup_env.sh
sed -i '1 s/$/2.7/' $NEXMON_ROOT/buildtools/b43-v3/debug/b43-beautifier
make                                                          # extract ucode / flashpatches
cd $NEXMON_ROOT/utilities/nexutil
sudo -E make install USE_VENDOR_CMD=1                         # REQUIRED on recent kernels
sudo setcap cap_net_admin+ep /usr/bin/nexutil
cd $NEXMON_ROOT/patches/bcm43455c0/7_45_189
git clone --depth 1 https://github.com/seemoo-lab/nexmon_csi.git && cd nexmon_csi
make -f Makefile.rpi install-firmware && sudo reboot         # reboot loads the CSI firmware
3 — activate CSI and confirm the UDP:5500 stream
nexutil -Iwlan0 -m1                                           # monitor mode
PARAMS=$(makecsiparams -c 132/80 -C 1 -N 1)                  # ch132, 80 MHz, core0, ss0
nexutil -Iwlan0 -s500 -b -l34 -v"$PARAMS"                    # configure the extractor
sudo tcpdump -i wlan0 -nn dst port 5500                      # → length-1042 UDP frames

Each frame is exactly what the PL csi_udp_parser reads — an 18-byte nexmon header, then 256 int16 I/Q pairs (80 MHz):

ByteFieldCaptured
42magic 0x1111—
44RSSI (int8)−47 dBm
52sequenceper-frame
54core / spatial stream0 / 0
56chanspec0xe08a = 132/80
60 …CSI — int16 LE real, int16 LE imag ×2561024 bytes

📎 Download a real 40-frame CSI capture (csi_sample.pcap, ch132/80) — the byte offsets above line up 1:1 with hls/csi_udp_parser.cpp.

The bring-up war-stories — the ones that actually cost time:

The CSI ioctl returned -EBADE — stock firmware was still loaded
nexmon stages the patched firmware via update-alternatives, but the driver had already loaded the stock brcmfmac43455-sdio.bin at boot, so nexutil -s500 was rejected with -EBADE (which presents as a hang). A reboot is what makes the driver actually download the CSI firmware (version 7.45.189 nexmon.org/csi); a plain module reload does not re-flash the chip.
nexutil needs -Iwlan0, and must be built with USE_VENDOR_CMD=1
On the recent-kernel path the driver only accepts the CSI IOCTLs as nl80211 vendor commands; without USE_VENDOR_CMD=1 they are rejected, and without an explicit -Iwlan0 the SET call blocks. An strace confirmed the sequence: vendor command sent, driver acks, done.
Skipping the armhf cross-libs silently kills monitor mode
nexmon's bundled compiler is a 32-bit ARM binary; on a 64-bit image it links against the :armhf libisl/libmpfr under their legacy sonames. Skip the symlinks and the firmware builds but monitor mode never engages.

Boot-time UDP sender (Pi → VCK190)

nexmon injects the CSI onto wlan0 (the monitor interface); it never leaves the Pi on its own. A stdlib-Python systemd service captures those frames at layer 2 and re-emits the nexmon payload as a fresh UDP datagram on eth0 — reproducing the exact ethertype@12 / proto@23 / dport@36 / nexmon@42 / CSI@60 offsets, so the Versal parser gets a byte-correct frame:

# csi-forward.service runs csi_forward.py:
#   raw AF_PACKET capture on wlan0 (udp/5500)  →  UDP sendto on eth0
sudo systemctl enable --now csi-activate.service csi-forward.service
# on the receiver:  sudo tcpdump -i eth0 -nn udp dst port 5500  → length-1042 frames

Validated live: wlan0 → eth0 at ~470 frames/s with the byte layout intact.

Status — the last mile

The Pi source and the Ethernet forwarder are done and format-verified against the parser. The Versal live front-end (GT + MAC + csi_udp_parser + csi_mux) is built in the v++ co-generation flow; on-silicon bring-up with the Pi over the SFP link — PCS link plus live parser → AI Engine — is the remaining step.

Real WiFi-HAR hardware setup with Raspberry Pi 4B and VCK190
Real hardware setup. The black enclosure is the Raspberry Pi 4B running nexmon_csi, streaming CSI over Ethernet to the VCK190.