NeurIPS 2026 · Project page
ReSCUE: Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation
Abstract
Simultaneous Sign Language Translation (SLT) is critical for real-time communication, yet existing methods remain largely confined to sentence-level, offline settings that assume pre-segmented inputs. We present ReSCUE, a unified framework for simultaneous SLT on unsegmented long-form sign language videos. ReSCUE combines inference-aware training for partial inputs, non-signing pauses, and multi-sentence contexts; stabilized re-translation for low-latency, revisable predictions with less flicker; and sentence commitment for online segmentation and memory management. On long-form unsegmented datasets, ReSCUE approaches oracle offline translation quality while operating at substantially lower latency.
Demo
The real-time system demonstration will appear here.
Method
Overview of the proposed simultaneous SLT pipeline. The unsegmented video stream is processed in fixed-length chunks and accumulated in a dynamic input buffer. At each step, the SLT module performs re-translation over the buffered video segment to produce a revisable hypothesis for the current sentence while encouraging prefix consistency with the previous hypothesis. Finalized translations of previous sentences are stored in the committed history and concatenated with the current hypothesis to form the user-facing output. A sentence commitment mechanism detects sentence completion, extracts the finalized translation, appends it to the committed history, and advances the input buffer accordingly.
Sentence-level Translation
Faster translation without sacrificing quality
ReSCUE achieves low translation latency while maintaining strong translation quality. We evaluate simultaneous SLT on sentence-level datasets, showing that ReSCUE has higher translation quality at lower latency compared to existing methods.
BLEU measures similarity to reference translations (higher is better), while Average Lagging (AL) measures how far the output trails the incoming video (lower is better). The upper-left region is preferred.
Loading CSL-Daily results…
Long-form Translation
Multiple sentences, translated continuously
ReSCUE translates unsegmented long-form sign language videos containing multiple sentences. We evaluate it against the offline state-of-the-art Uni-Sign and an oracle offline system that is given ground-truth sentence boundaries (marked as O-R).
ReSCUE reaches translation quality close to the oracle system while reducing latency from tens of seconds to about 1.6 seconds. The oracle setting is unavailable in practical inference.
Loading long-form results…
Citation
Coming soon.
Citation coming soon.