NeurIPS
University of Kentucky Washington University in St. Louis

TRACE: Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning

1 University of Kentucky 2 Washington University in St. Louis
NeurIPS 2026Main Track
Prior single-frame attacks invert each gradient independently and lose scene structure. TRACE passes gradient tokens through a causal transformer with visual history, so each reconstruction carries context from the previous one.
Schematic: as sequence length grows, pixel error falls and reconstruction fidelity rises relative to a single-frame attack.
Single-frame attacks treat each gradient on its own. TRACE reads the gradient stream as a sequence, so every step inherits context from the last, and reconstruction gets better the longer the trajectory runs.
18.8dB
PSNR on held-out scenes
100%
Action accuracy, PPO and A2C
3–4.5ms
Per reconstructed frame
>6,900×
Faster than optimization attacks
Abstract

Distributed learning in embodied reinforcement-learning agents offers a degree of privacy by retaining raw sensor data on-device and transmitting only policy gradients to the server. Yet temporal structure can amplify this leakage beyond single-frame attacks. We introduce Temporal Reconstruction Attack on Consecutive Encodings (TRACE), an amortized temporal gradient-inversion attack that autoregressively reconstructs the sequence of private observation-action trajectories from per-step policy-learning gradients.

The attack exploits two structural signals ignored by prior single-frame methods: (i) cross-time correlation between successive embodied gradients, which we formalize via a conditional mutual-information bound, and (ii) closed-form action recovery from policy-head gradient structure, which we prove exact when standard entropy regularization is sufficiently small. On held-out embodied scenes, TRACE reaches 18.8 dB PSNR with near-perfect action recovery at 3–4.5 ms per reconstructed frame, dominating the learning-based baseline across all reconstruction metrics and exceeding optimization attacks while running orders of magnitude faster.

Further evaluation demonstrates TRACE's broader applicability across recurrent, residual, and compact transformer victim architectures, multi-modal inputs, and larger discrete action spaces. Defense experiments suggest that protecting temporal gradient streams may require sequence-aware privacy mechanisms.

Method

One gradient in, one frame and one action out

A trained inverter reads the gradient stream step by step. No per-sample optimization, so attacking a new trajectory is a single forward pass.

TRACE pipeline: a gradient encoder compresses each gradient to a latent, an image encoder embeds prior observations, a causal temporal transformer contextualizes the interleaved sequence, and a residual decoder reconstructs the observation and predicts the action.
The gradient encoder compresses each per-step gradient to a latent and the image encoder embeds earlier reconstructions. A causal temporal transformer contextualizes the interleaved sequence, and a residual decoder emits the observation and action for each step.
Signal 1

Cross-time correlation

Successive embodied gradients share information about the scene. We bound what the past adds about the present with a conditional mutual-information argument.

Signal 2

Closed-form action recovery

The policy-head gradient gives the action away directly. Recovery is provably exact when entropy regularization is small enough.

Training

Autoregressive, exposure-aware

The inverter conditions on its own previous reconstructions, trained so that errors do not compound over long rollouts.

Results

What the agent saw, recovered from its gradients

Point-goal navigation in AI2-THOR. The victim sees 84×84 RGB frames and picks one of five actions; the attacker only sees per-step policy gradients.

t = 0MoveAheadt = 1LookUpt = 2RotateLeftt = 3LookUpt = 4LookUpt = 5RotateLeftt = 6LookUpt = 7MoveAhead
Ground truth
TRACE
Ground truth (top) and TRACE reconstructions (bottom) over T = 8 PPO steps. TRACE recovers every action, shown under each step.
Method MSE ↓ PSNR ↑ SSIM ↑ LPIPS ↓ Act. Acc. ↑ Time / frame
DLG 0.289 5.47 0.043 1.224 16.6% 31.0 s
Inverting Gradients 0.161 8.73 0.166 0.766 31.5% 39.3 s
Learning to Invert 0.023 16.79 0.529 0.671 100% 1.2 ms
TRACE 0.014 18.77 0.627 0.362 100% 4.5 ms

Baselines are adapted from supervised classification with an actor-critic surrogate loss. PSNR in dB. Means over held-out sequences; standard deviations are in the paper.

Generality

Not tied to one victim network

Each PPO victim gets its own separately trained attacker, with no data augmentation. K is the number of actions.

Victim K MSE ↓ PSNR ↑ SSIM ↑ LPIPS ↓ Act. Acc. ↑
CNN 10 0.013 18.92 0.610 0.372 99.5%
Tiny ViT 5 0.048 13.99 0.456 0.543 99.0%
IMPALA-style CNN 5 0.007 21.85 0.725 0.244 97.4%
Wider IMPALA CNN 5 0.007 22.36 0.773 0.208 98.5%
CNN + GRU 5 0.015 18.62 0.565 0.390 97.5%
Multi-modal CNN 5 0.011 19.78 0.697 0.335 97.4%

More capacity does not protect the victim: the wider IMPALA network gives the best reconstructions. Action recovery stays above 97% across every architecture.

Defenses

Per-step defenses barely slow it down

Pruning and light noise leave the reconstruction almost untouched. Only aggressive quantization, strong noise and DP-SGD break it.

t = 1t = 2t = 3t = 4t = 5t = 6t = 7t = 8
Ground truth
No defense
Pruning 50%
Noise σ = 0.01
DP-SGD
The same trajectory under no defense, 50% pruning, noise σ = 0.01, and DP-SGD (C = 1.0, σ = 1.0).
Defense Setting MSE ↓ PSNR ↑ SSIM ↑ LPIPS ↓ Act. Acc. ↑
No defense — 0.018 18.9 0.628 0.374 99.9%
Quantization
8-bit 0.020 18.5 0.619 0.384 99.9%
4-bit 0.075 12.4 0.422 0.645 37.9%
2-bit 0.091 11.3 0.380 0.689 19.4%
Pruning (keep %)
90% 0.018 18.9 0.628 0.374 99.9%
50% 0.018 18.9 0.628 0.374 99.9%
10% 0.018 18.8 0.624 0.371 99.9%
Gaussian noise (σ)
0.001 0.018 18.8 0.626 0.376 99.8%
0.01 0.023 18.1 0.599 0.401 94.9%
0.1 0.049 14.2 0.438 0.572 59.9%
DP-SGD (ε, δ = 10⁻⁵)
ε = 10 0.069 12.2 0.345 0.658 21.4%
ε = 5 0.069 12.1 0.342 0.663 21.4%
ε = 1 0.071 12.0 0.346 0.658 17.6%

Highlighted cells mark settings where the attack clearly degrades. Protecting a gradient stream likely needs sequence-aware privacy mechanisms rather than per-step perturbation.

Cite

BibTeX

@inproceedings{bhujel2026trace,
  title         = {Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning},
  author        = {Bhujel, Sudip and Shi, Shanghao and Huang, Ruiquan and Zhang, Ning and Xiao, Yang},
  booktitle     = {Advances in Neural Information Processing Systems (NeurIPS)},
  year          = {2026},
  eprint        = {2609.30258},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  url           = {https://trace-rl.github.io/}
}