NeurIPS 2026 · Project page

ReSCUE: Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation

Authors will be announced
Affiliations coming soon
Teaser video · Coming soon

The streaming teaser video will appear here.

To protect privacy, the video has been processed using AI and may contain slight distortions in some movements.

Abstract

Simultaneous Sign Language Translation (SLT) is critical for real-time communication, yet existing methods remain largely confined to sentence-level, offline settings that assume pre-segmented inputs. We present ReSCUE, a unified framework for simultaneous SLT on unsegmented long-form sign language videos. ReSCUE combines inference-aware training for partial inputs, non-signing pauses, and multi-sentence contexts; stabilized re-translation for low-latency, revisable predictions with less flicker; and sentence commitment for online segmentation and memory management. On long-form unsegmented datasets, ReSCUE approaches oracle offline translation quality while operating at substantially lower latency.

Demo

Demo video · Coming soon

The real-time system demonstration will appear here.

Method

Overview of ReSCUE: chunks enter a dynamic input buffer, an SLT model performs stabilized re-translation, and stable terminal punctuation triggers sentence commitment.

Overview of the proposed simultaneous SLT pipeline. The unsegmented video stream is processed in fixed-length chunks and accumulated in a dynamic input buffer. At each step, the SLT module performs re-translation over the buffered video segment to produce a revisable hypothesis for the current sentence while encouraging prefix consistency with the previous hypothesis. Finalized translations of previous sentences are stored in the committed history and concatenated with the current hypothesis to form the user-facing output. A sentence commitment mechanism detects sentence completion, extracts the finalized translation, appends it to the committed history, and advances the input buffer accordingly.

Sentence-level Translation

Faster translation without sacrificing quality

ReSCUE achieves low translation latency while maintaining strong translation quality. We evaluate simultaneous SLT on sentence-level datasets, showing that ReSCUE has higher translation quality at lower latency compared to existing methods.

BLEU measures similarity to reference translations (higher is better), while Average Lagging (AL) measures how far the output trails the incoming video (lower is better). The upper-left region is preferred.

Loading CSL-Daily results…

CSL-Daily · BLEU-4 versus Average Lagging

Long-form Translation

Multiple sentences, translated continuously

ReSCUE translates unsegmented long-form sign language videos containing multiple sentences. We evaluate it against the offline state-of-the-art Uni-Sign and an oracle offline system that is given ground-truth sentence boundaries (marked as O-R).

ReSCUE reaches translation quality close to the oracle system while reducing latency from tens of seconds to about 1.6 seconds. The oracle setting is unavailable in practical inference.

Loading long-form results…

Long-form translation results · O-R uses ground-truth sentence boundaries

Citation

Coming soon.

BibTeX
Citation coming soon.
Copied to clipboard