← Back to blog

Structured Masked Diffusion for Joint Multiuser Decoding

CIDER

*Equal contribution 1The University of Texas at Austin 2Seoul National University

NeurIPS 2026

CIDER receiver overview: simultaneous users transmit through a wireless channel, a fixed symbol-level soft detector produces evidence, and CIDER reconstructs the message set.

Where We Started

This project started with Jiyoung’s earlier work on unsourced random access (Yun & Choi, 2024). We wanted to revisit the decoding problem with a different set of tools.

What is URA?

In unsourced random access (URA), many devices send short messages at once using the same codebook. The receiver sees one noisy mixture and must recover the message set, without knowing which device sent which message.

Many senders, one receiver4 users · 2 active · 12 slots per codeword
Only two users transmit. Each sends one codeword of 12 GF(64) symbols, one per slot. The receiver recovers the unordered pair.

Each message becomes a codeword, whose parity checks help detect errors. But slot-wise detector scores combine users, so picking the strongest symbol in each slot can assemble a message nobody sent.

The earlier decoder used error correction across slots and successive interference cancellation to separate the messages. Joint search and repeated belief propagation get expensive as codewords grow or more users collide, and a wrong early cancellation can affect later decisions.

We wondered whether a learned decoder could recover the message set in parallel, while keeping the messages distinct and each codeword consistent with the code.

Structured Masked Diffusion

Masked diffusion starts with a blank grid and reveals symbols over several rounds. A generic denoiser still copied messages and broke parity checks, so CIDER puts two structural modules inside each refinement step.

First, make the rows compete.

Module A makes the message rows compete for each slot-symbol candidate. Its responsibilities are soft, since two real messages can share a symbol, but one strong candidate should not be copied into every row.

Then, let the parity checks talk back.

Module B passes information along the known code’s Tanner graph, so a symbol can use context from its parity checks. The transformations are learned; the graph guides each refinement step toward code consistency.

CIDER denoiser: a partially masked grid passes through Module A demixing and Module B parity-aware propagation, then predicts symbols for the next grid. Evidence S and code matrix H condition the modules.

Watch the messages come together.

The blue reference rows below show the two codewords sent by the active users above. CIDER receives soft detector scores and reconstructs its output rows over successive steps.

Two messages · 12 slots each · GF(64)

Ground truth

Compare CIDER with

CIDER

Module A + B

Generic masked diffusion

DiT denoiser
CorrectIncorrectMaskedGround truth

Recorded appendix steps 0, 1, 4, 7, and 12; intermediate steps are not shown. Hexadecimal symbols label the GF(64) alphabet. Colors compare predictions with ground truth, which the decoder does not receive.

Accuracy, Latency, and Transfer

Removing either module drove codeword error rate close to 100% in the two-user test. Codeword error rate (CER) counts a message as wrong if any symbol is wrong, after matching the unordered output rows.

The pieces need to work together
DecoderCER ↓
Generic masked diffusion99.96%
CIDER without demixing99.93%
CIDER without parity propagation99.93%
Both modules, one-shot prediction39.79%
Full CIDER0.73%

Test settingLDPC code · 2 users (K = 2) · 12 slots (L = 12) · 64-symbol alphabet (Q = 64).

Generic diffusion had 63.5% slot overlap and 99.9% parity violations; CIDER reduced these to 1.5% and 0.7%.

And it was fast enough to be interesting.

With the same detector evidence and GPU, CIDER had lower measured latency at all four code lengths below.

Same GPU, same evidence

DecoderCER ↓ms / sample ↓
Benchmark conditions
What was timed
Decoder runtime per sample. The rest of the receiver is excluded.
Hardware and task
RTX 3090 · 2 users (K = 2) · 64 symbols (Q = 64)
Precision
BP: FP32/TF32 · CIDER: mixed precision

Could one trained decoder handle longer codewords?

We trained CIDER on 12-slot codewords, then tried it on longer codes without retraining. We supplied each new code's parity checks, kept the network weights fixed, and changed only the number of refinement steps. Its codeword error rate stayed at or below 1.10% through 24 slots, but rose to 7.62% at 96 slots. Models trained for each target setting generally did better. The 72- and 96-slot tests used only 256 examples each, so those two estimates are less certain.

More Users and Quality-Guided Remasking

With more users, the first pass gets less reliable. A small quality head scores each decoded row; we mask and re-decode low-confidence rows while keeping the others fixed.

First pass · K = 825.76%CER · 30.77 ms
With quality-guided remasking4.03%CER · 63.00 ms

Runtime scope · K = 8Reported times include quality scoring and any re-decoding performed.

For a larger system, keep each collision manageable.

We had each user choose a preamble that assigns its message to a bin, then used a load-specific CIDER decoder within each bin. With about four users per bin on average, the experiment reached 100 total users at a 6.19% per-user error rate. This test supplied the decoder with the true bin load and counted bins above eight users as erasures.

What Still Needs Work

CIDER decodes soft evidence from a fixed detector; it does not learn the full wireless receiver. Reducing detector iterations barely hurt in the matched test, but a 4 dB SNR drop raised CER to 30.87%.

Without estimating channel gains or retraining for fading, CER rose from 0.73% on AWGN to 40.70% on Rayleigh fading; SIC-BP reached 8.13% in that fading test. Heavy codeword overlap was another weakness. Better channel-aware evidence and fading-aware training are natural next steps.

Paper & earlier work

  1. Taekyun Lee, Jiyoung Yun, Jeffrey G. Andrews, and Hyeji Kim. CIDER paper on arXiv, 2026. Accepted to NeurIPS 2026.
  2. Jiyoung Yun and Wan Choi. Erasure Correcting Blind Detection in Unsourced Random Access for Grant-Free Massive Connections. IEEE Transactions on Wireless Communications, 23(3):2428–2439, 2024.
  3. Jaeyeon Kim, Seunggeun Kim, Taekyun Lee, David Z. Pan, Hyeji Kim, Sham Kakade, and Sitan Chen. Fine-Tuning Masked Diffusion for Provable Self-Correction, 2025.

BibTeX

@misc{lee2026structuredmaskeddiffusion,
  title   = {Structured Masked Diffusion for Joint Multiuser Decoding},
  author  = {Lee, Taekyun and Yun, Jiyoung and Andrews, Jeffrey and Kim, Hyeji},
  year    = {2026},
  eprint  = {2605.26580},
  archivePrefix = {arXiv},
  url     = {https://arxiv.org/abs/2605.26580}
}