Abstract
Recent video transformers commonly adopt factorised spatial-temporal attention to reduce computational complexity. However, such factorisation limits temporal modelling by constraining attention to fixed spatial locations across frames, ignoring the motion-dependent spatial context essential for accurate prediction. To address this limitation, we propose Trajectory-Aware Deformable Attention for the video prediction task. This novel mechanism integrates motion-aware sampling into temporal attention by dynamically aligning keys and values along motion paths to restore spatial context. Patch trajectories are estimated in continuous pixel coordinates using a lightweight convolutional flow estimator with forward-backwards consistency checking for occlusion detection. For each query, a set of learnable deformable sample points are placed around the trajectory-predicted position at every source time step, and features are extracted via differentiable bilinear interpolation. The proposed method is trained and evaluated on the KITTI Caltech dataset. It achieves consistent improvements across all metrics, including a 7.6% reduction in MSE and a significant 20.1% reduction in LPIPS, and improvements in PSNR, MAE and SSIM, demonstrating the benefit of motion-aware sampling. These gains are achieved with only a 14.3% increase in FLOPs, demonstrating that it can be integrated into any video transformer at a modest additional cost.
| Original language | English |
|---|---|
| Title of host publication | 2026 IEEE Conference on Artificial Intelligence (CAI) |
| Publisher | IEEE |
| Pages | 844-849 |
| Number of pages | 6 |
| ISBN (Electronic) | 979-8-3315-6039-3 |
| ISBN (Print) | 979-8-3315-6040-9 |
| DOIs | |
| Publication status | Published - 8 May 2026 |
| Event | 2026 IEEE Conference on Artificial Intelligence (CAI) - Granada, Spain Duration: 8 May 2026 → 10 May 2026 |
Publication series
| Name | 2026 IEEE Conference on Artificial Intelligence (CAI) |
|---|---|
| Publisher | IEEE |
Conference
| Conference | 2026 IEEE Conference on Artificial Intelligence (CAI) |
|---|---|
| Country/Territory | Spain |
| City | Granada |
| Period | 8/05/26 → 10/05/26 |
Bibliographical note
Publisher Copyright:© 2026 IEEE.
Funding
This research was funded by Coventry University grant number BDN68KDUX.
| Funders | Funder number |
|---|---|
| Coventry University | BDN68KDUX |
Fingerprint
Dive into the research topics of 'Trajectory-Aware Deformable Transformer for Contextual Video Prediction'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS