Glossless sign language translation method and system based on latent visual-text alignment

By employing a latent visual-text alignment method and utilizing the InfoNCE loss function for modal in-modal and out-of-modal contrastive learning, combined with masked text reconstruction and autoregressive translation tasks, the problem of visual-text alignment in glossless sign language translation is solved, achieving high-precision video-to-text translation and improving translation quality and robustness.

CN120954105BActive Publication Date: 2026-02-03ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511484788.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-02-03
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing gloss-free sign language translation methods struggle to achieve high-precision visual-text alignment without human annotation, and also suffer from significant computational overhead and deployment difficulties.

Method used

By employing a latent visual-text alignment method, the InfoNCE loss function is used for intra-modal and cross-modal contrastive learning to align sign language video frames with text sentence data in the latent embedding space. This is then combined with masked text reconstruction and autoregressive translation tasks for joint optimization training to construct a sign language translation model.

Benefits of technology

It achieves high-precision, end-to-end sign language video-to-text translation, improving translation quality and robustness while reducing annotation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954105B_ABST
    Figure CN120954105B_ABST
Patent Text Reader

Abstract

The application discloses a gloss-free sign language translation method and system based on potential visual-text alignment, relates to the field of computer vision and natural language processing, and comprises the following steps: acquiring sign language video frame sequence data and corresponding text sentence data; performing feature extraction on potential visual segments to generate potential visual representations; performing feature extraction on text sub-word units to generate potential text representations; mapping the potential visual representations and the corresponding potential text representations to the same potential embedding space; aligning the potential visual segments and the text sub-word units in the potential embedding space to obtain aligned data; inputting the aligned data into an initial sign language translation model, taking a masked text reconstruction task and a sign language video to text translation task as a joint optimization target, training the sign language translation model, and obtaining a sign language translation model; and acquiring target sign language video data, inputting the target sign language video data into the sign language translation model, and obtaining a translation result. The application improves the translation quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and natural language processing technology, and particularly relates to a gloss-free sign language translation method and system based on latent visual-text alignment. Background Technology

[0002] In recent years, thanks to the rapid development of deep learning and large-scale pre-trained models, computer vision and natural language processing have made significant progress. In most cross-modal tasks, researchers have improved the performance of downstream tasks by aligning visual and text representations through self-supervised pre-training and contrastive learning. Sign Language Translation (SLT), as a typical vision-language conversion problem, requires capturing the temporal semantic information of both visual space—hand and facial movements—and language sequences, making cross-modal alignment extremely challenging.

[0003] Traditional sign language translation relies heavily on manual gloss (manual annotation of sign language videos or gestures). After mapping the video into intermediate discrete labels, text is generated through CTC (Connectionist Temporal Classification) or sequence-to-sequence models. This can narrow the gap between visual and linguistic granularity and significantly improve translation accuracy. However, gloss annotation is costly and difficult to transfer across domains.

[0004] To reduce reliance on gloss, several gloss-free solutions have emerged in recent years, mainly including: large-scale visual self-supervised pre-training, which learns spatiotemporal features through tasks such as reconstruction, mask prediction, and contrastive learning; and cross-modal contrastive learning, which borrows from CLIP (Contrastive Language–Image Pre-training) to minimize the InfoNCE (Info Noise-Contrastive Estimation) loss on video-text pairs.

[0005] Large language model fusion maps video features to discrete token-based cueing LLMs (Large Language Models), or employs parameter-efficient fine-tuning techniques to construct multimodal adapters. While the aforementioned gloss-free methods show potential in zero-shot or few-shot scenarios, they still struggle to achieve the detail alignment and translation quality of the gloss baseline due to the lack of fine-grained intramodal structural constraints. Furthermore, they often rely on multi-stage training or complex large models, increasing computational overhead and deployment difficulty. Summary of the Invention

[0006] To address the aforementioned technical issues, this invention proposes a gloss-free sign language translation method and system based on latent visual-text alignment, which significantly improves the translation quality and robustness of the model and substantially reduces annotation costs.

[0007] To achieve the above objectives, this invention provides a gloss-free sign language translation method based on latent visual-text alignment, comprising:

[0008] Acquire sign language video frame sequence data and corresponding text sentence data, and preprocess the sign language video frame sequence data and the text sentence data respectively to obtain potential visual segments and text sub-word units;

[0009] Feature extraction is performed on the latent visual segments to generate latent visual representations; feature extraction is performed on the text sub-word units to generate latent text representations; the latent visual representations and the corresponding latent text representations are mapped to the same latent embedding space;

[0010] In the latent embedding space, based on InfoNCE intramodal contrastive learning and crossmodal contrastive learning, the latent visual fragments are aligned with the text word units to obtain aligned data;

[0011] An initial sign language translation model is constructed by inputting the aligned data into the initial sign language translation model and training the sign language translation model with the masked text reconstruction task and the sign language video to text translation task as joint optimization objectives to obtain the sign language translation model.

[0012] Obtain the target sign language video data, input it into the sign language translation model, and obtain the translation result.

[0013] On the other hand, to achieve the above objectives, the present invention also provides a gloss-free sign language translation system based on latent visual-text alignment, comprising:

[0014] The module includes a data acquisition module, a feature extraction module, a data alignment module, a model training module, and a sign language translation module.

[0015] The data acquisition module is used to acquire sign language video frame sequence data and corresponding text sentence data, and to preprocess the sign language video frame sequence data and the text sentence data respectively to obtain potential visual segments and text sub-word units.

[0016] The feature extraction module is used to extract features from the latent visual segments to generate latent visual representations; extract features from the text sub-word units to generate latent text representations; and map the latent visual representations and the corresponding latent text representations to the same latent embedding space.

[0017] The data alignment module is used to align the latent visual fragments with the text word units in the latent embedding space based on intra-modal contrastive learning and cross-modal contrastive learning of InfoNCE, so as to obtain aligned data.

[0018] The model training module is used to construct an initial sign language translation model. The aligned data is input into the initial sign language translation model, and the sign language translation model is trained using the masked text reconstruction task and the sign language video to text translation task as joint optimization objectives to obtain the sign language translation model.

[0019] The sign language translation module is used to acquire target sign language video data, input it into the sign language translation model, and obtain the translation result.

[0020] Technical Advantages of this Invention: This invention discloses a gloss-free sign language translation method and system based on latent visual-text alignment, achieving high-precision, end-to-end sign language video-to-text conversion without manual gloss adjustment. By constructing frame-level and sub-word-level contrastive learning within the modality and introducing symmetric cross-entropy alignment at the cross-modal level, while integrating masked text reconstruction and autoregressive translation tasks, this invention significantly improves the model's translation quality and robustness, and substantially reduces annotation costs. This invention enables video-to-text translation tasks without gloss. The method of this invention improves the accuracy and efficiency of sign language translation through precise alignment of visual and textual information. Attached Figure Description

[0021] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0022] Figure 1 This is a flowchart illustrating the gloss-free sign language translation method based on latent visual-text alignment in an embodiment of the present invention.

[0023] Figure 2 This is a schematic diagram of the structure of a gloss-free sign language translation system based on latent visual-text alignment in an embodiment of the present invention. Detailed Implementation

[0024] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0025] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0026] like Figure 1 As shown, this embodiment provides a gloss-free sign language translation method based on latent visual-text alignment, including:

[0027] Acquire sign language video frame sequence data and corresponding text sentence data, and preprocess the sign language video frame sequence data and the text sentence data respectively to obtain potential visual segments and text sub-word units;

[0028] Feature extraction is performed on the latent visual segments to generate latent visual representations; feature extraction is performed on the text sub-word units to generate latent text representations; the latent visual representations and the corresponding latent text representations are mapped to the same latent embedding space;

[0029] In the latent embedding space, based on InfoNCE intramodal contrastive learning and crossmodal contrastive learning, the latent visual fragments are aligned with the text word units to obtain aligned data;

[0030] An initial sign language translation model is constructed by inputting the aligned data into the initial sign language translation model and training the sign language translation model with the masked text reconstruction task and the sign language video to text translation task as joint optimization objectives to obtain the sign language translation model.

[0031] Obtain the target sign language video data, input it into the sign language translation model, and obtain the translation result.

[0032] Furthermore, acquiring sign language video frame sequence data and corresponding text sentence data, and preprocessing the sign language video frame sequence data and the text sentence data respectively to obtain potential visual segments and text sub-word units includes:

[0033] First, obtain a training dataset where each data item includes a sequence of sign language video frames. and the corresponding text sentence Specifically, sign language video frame sequences Depend on The frame composition represents the various frames of the i-th sign language video. Based on the inter-frame similarity, the sign language video is divided into several potential visual segments; while the text sentence... This represents the text sentence corresponding to the video, and each text sentence is broken down into sub-word units. These sub-word units It represents the basic components of a sentence. By decomposing sentences at the sub-word level, the semantic units of each sentence can be processed separately for subsequent processing.

[0034] Further, feature extraction is performed on the latent visual segments to generate latent visual representations; feature extraction is performed on the text sub-word units to generate latent text representations; mapping the latent visual representations and corresponding latent text representations to the same latent embedding space includes:

[0035] via visual encoder Sign language video frame sequence Perform feature extraction to generate latent visual representations. :

[0036] ;

[0037] in, It is a visual encoder that extracts features from video frames using convolutional neural networks (CNNs) or other deep learning models to generate latent visual representations. The dimension is ,in d is the number of video frames, and d is the dimension of the latent representation.

[0038] via text encoder Text sentence data Extracting latent text representations :

[0039] ;

[0040] in, It is a text encoder, typically employing a Transformer-based architecture (such as BERT). This encoder processes each text subword... Convert to the corresponding embedded representation Ultimately, this forms the latent representation of the text. Its dimensions are ,in It is the number of subwords in the text sentence.

[0041] By using a visual encoder and a text encoder, this embodiment can map sign language video frames and corresponding text sentences into the same latent space, thereby laying the foundation for subsequent alignment learning.

[0042] Furthermore, within the latent embedding space, based on InfoNCE-based intra-modal contrastive learning and cross-modal contrastive learning, the latent visual segments are aligned with the text word units, yielding aligned data including:

[0043] Learning is achieved by constructing a set of positive and negative samples between each video frame and other video frames for comparative analysis. Specifically, for each video frame... We construct a positive sample set and a negative sample set. The positive sample set includes frames similar to the current frame, and the negative sample set includes frames dissimilar to the current frame. By minimizing the similarity difference between the positive and negative samples, we optimize the model so that similar frames cluster in the latent space. That is, for latent video segments... Construct a set of positive and negative gloss For text sub-word units Construct a set of positive and negative gloss ,in For cosine similarity, All are preset thresholds.

[0044] For each video frame and text sub-word units The similarity between the video frames and text word units in their latent embedding space is calculated, and cross-modal contrastive learning is performed. This process optimizes the parameters of the visual encoder and text encoder by minimizing the contrastive loss between video frames and text word units, enabling the latent representations of video frames and text word units to align in the same space. Intra-modal contrastive learning is performed using InfoNCE loss.

[0045] ;

[0046] ;

[0047] in, and These are frame-level anchor representation and word-level anchor representation, respectively. These are the corresponding positive samples and the i-th negative sample, respectively. for The corresponding positive sample and the i-th negative sample; For temperature coefficient, M,K The number of negative samples. For visual modality InfoNCE loss; For text modality InfoNCE loss; The transpose of the frame-level anchor point representation is used to calculate the dot product similarity. Positive sample sub-words The vector representation obtained after text encoder; negative sample sub-words The vector representation obtained after text encoding.

[0048] By optimizing the alignment between video frames and text word units, a cross-modal alignment loss is calculated to enhance the alignment accuracy between video frames and text word units. The cross-modal alignment loss takes the following form:

[0049] ;

[0050] ;

[0051] Where N is the batch size. Represents video frames and text sentences similarity, For temperature coefficient, To extend the modal alignment loss; Let be the potential vector representation of the i-th sign language video segment after passing through the visual encoder;

[0052] Furthermore, an initial sign language translation model is constructed. The aligned data is input into the initial sign language translation model, and the sign language translation model is trained using the masked text reconstruction task and the sign language video-to-text translation task as joint optimization objectives. The resulting sign language translation model includes:

[0053] During the training phase, this invention employs a multi-task joint training strategy, using the masked text reconstruction task and the sign language video-to-text translation task as joint optimization objectives, so that the model can complete both tasks simultaneously during training, thereby improving overall performance.

[0054] Specifically, masked text reconstruction loss The goal is to force a model to predict masked words using contextual information by randomly masking a portion of the text. This task helps the model better understand and generate accurate text.

[0055] ;

[0056] Where N is the batch size. This represents the text sentence after the mask, where w is the sub-word unit after the mask. is the embedding of the masked text generated by the text encoder, and h is the text decoder.

[0057] Meanwhile, sign language translation task This is used to map sign language video V to corresponding text S, and the difference between the translation predicted by the model and the actual text is calculated using cross-entropy loss.

[0058] ;

[0059] in, It is the uth word after translation. This represents the first u-1 words of the sentence, and V is the input sign language video sequence. It is the probability of the word at that position predicted by the model. The sign language translation loss for end-to-end sign language video-to-text translation tasks.

[0060] The joint masked text reconstruction loss and the video-text translation loss constitute the overall goal of end-to-end training.

[0061] ;

[0062] in, All are weighting coefficients. The total loss function for joint optimization, The alignment loss between the potential visual and textual representations in the embedding space. For the loss of the masked text reconstruction task, For the translation loss of sign language video to text, For the InfoNCE contrast loss within the visual modality, For the InfoNCE contrast loss within the text modality, For the latent visual representation of the input, The potential text representation of the input.

[0063] Furthermore, the hyperparameters used for training are set as follows: temperature coefficient. This is used to control the smoothness of the similarity distribution in the contrastive loss. The number of negative samples is set to M = K = 5, meaning 5 negative samples are selected for each contrastive learning iteration. The number of negative samples directly affects the effectiveness of contrastive learning; an appropriate number of negative samples helps to balance learning efficiency and model robustness.

[0064] The batch size is set to N = 4. This relatively small value is suitable for few-shot learning tasks, effectively reducing memory burden and improving model training stability. The masking rate is set to 15%, meaning that 15% of the text subwords will be randomly masked in each training iteration. A higher masking rate helps the model better learn textual context information, enhancing the model's predictive ability.

[0065] The training rounds were set to 80 epochs. Through multiple iterations, the model was ensured to gradually converge, achieving a balance between accuracy and computational efficiency. The Adam optimizer was chosen, combining the advantages of momentum and adaptive learning rate, making it suitable for large-scale deep learning tasks and effectively accelerating training convergence. By setting these hyperparameters, this invention effectively improves performance in sign language translation tasks, exhibiting high computational efficiency and training stability, and adapting to different types of tasks and datasets.

[0066] like Figure 2 As shown, this embodiment provides a gloss-free sign language translation system based on latent visual-text alignment, including: a data acquisition module, a feature extraction module, a data alignment module, a model training module, and a sign language translation module;

[0067] The data acquisition module is used to acquire sign language video frame sequence data and corresponding text sentence data, and to preprocess the sign language video frame sequence data and the text sentence data respectively to obtain potential visual segments and text sub-word units.

[0068] The feature extraction module is used to extract features from the latent visual segments to generate latent visual representations; extract features from the text sub-word units to generate latent text representations; and map the latent visual representations and the corresponding latent text representations to the same latent embedding space.

[0069] The data alignment module is used to align the latent visual fragments with the text word units in the latent embedding space based on intra-modal contrastive learning and cross-modal contrastive learning of InfoNCE, so as to obtain aligned data.

[0070] The model training module is used to construct an initial sign language translation model. The aligned data is input into the initial sign language translation model, and the sign language translation model is trained using the masked text reconstruction task and the sign language video to text translation task as joint optimization objectives to obtain the sign language translation model.

[0071] The sign language translation module is used to acquire target sign language video data, input it into the sign language translation model, and obtain the translation result.

[0072] This invention discloses a gloss-free sign language translation method and system based on latent visual-text alignment, achieving high-precision, end-to-end sign language video-to-text conversion without manual gloss analysis. By constructing frame-level and word-level contrastive learning within the modality and introducing symmetric cross-entropy alignment at the cross-modal level, while integrating masked text reconstruction and autoregressive translation tasks, this invention significantly improves the model's translation quality and robustness, and substantially reduces annotation costs. This invention enables video-to-text translation tasks without gloss. The method of this invention improves the accuracy and efficiency of sign language translation through precise alignment of visual and textual information.

[0073] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A gloss-free sign language translation method based on latent visual-text alignment, characterized in that, include: Acquire sign language video frame sequence data and corresponding text sentence data, and preprocess the sign language video frame sequence data and the text sentence data respectively to obtain potential visual segments and text sub-word units; Feature extraction is performed on the potential visual segments to generate potential visual representations; Feature extraction is performed on the text sub-word units to generate latent text representations; the latent visual representations and the corresponding latent text representations are mapped to the same latent embedding space; In the latent embedding space, based on InfoNCE intramodal contrastive learning and crossmodal contrastive learning, the latent visual segments are aligned with the text word units to obtain aligned data; An initial sign language translation model is constructed by inputting the aligned data into the initial sign language translation model and training the sign language translation model with the masked text reconstruction task and the sign language video to text translation task as joint optimization objectives to obtain the sign language translation model. Acquire the target sign language video data, input it into the sign language translation model, and obtain the translation result; The aligned data includes: Construct a set of positive and negative samples between each video frame and other video frames for comparative learning; For each video frame and text word unit, the similarity in the latent embedding space is calculated, and cross-modal contrastive learning is performed. The parameters of the visual encoder and text encoder are optimized by minimizing the contrastive loss between the video frame and the text word unit to obtain the aligned data. The joint optimization objectives include the masked text reconstruction task and the sign language video-to-text translation task. ; in, All are weighting coefficients. The total loss function for joint optimization, The alignment loss between the potential visual and textual representations in the embedding space. For the loss of the masked text reconstruction task, For the translation loss of sign language video to text, For the InfoNCE contrast loss within the visual modality, For the InfoNCE contrast loss within the text modality, For the latent visual representation of the input, The potential text representation of the input.

2. The gloss-free sign language translation method based on latent visual-text alignment as described in claim 1, characterized in that, Preprocessing the sign language video frame sequence data includes: The sign language video frame sequence data is divided into several potential visual segments based on inter-frame similarity.

3. The gloss-free sign language translation method based on latent visual-text alignment as described in claim 1, characterized in that, Preprocessing the text sentence data includes: The text sentence data is decomposed into sub-word units to obtain text sub-word units.

4. The gloss-free sign language translation method based on latent visual-text alignment as described in claim 1, characterized in that, Generating latent visual representations includes: ; in, For potential visual representation, For visual encoders, This is a sequence of sign language video frames. d represents the number of video frames and d represents the dimension of the latent representation.

5. The gloss-free sign language translation method based on latent visual-text alignment as described in claim 1, characterized in that, Generating potential text representations includes: ; in, For potential text representation, For text encoders, For text sentence data, d represents the number of subwords in the text sentence, and d represents the dimension of the latent representation.

6. The system of the gloss-free sign language translation method based on latent visual-text alignment according to any one of claims 1-5, characterized in that, include: The module includes a data acquisition module, a feature extraction module, a data alignment module, a model training module, and a sign language translation module. The data acquisition module is used to acquire sign language video frame sequence data and corresponding text sentence data, and to preprocess the sign language video frame sequence data and the text sentence data respectively to obtain potential visual segments and text sub-word units. The feature extraction module is used to extract features from the latent visual segments and generate latent visual representations; Feature extraction is performed on the text sub-word units to generate latent text representations; the latent visual representations and the corresponding latent text representations are mapped to the same latent embedding space; The data alignment module is used to align the latent visual fragments with the text word units in the latent embedding space based on intra-modal contrastive learning and cross-modal contrastive learning of InfoNCE, so as to obtain aligned data. The model training module is used to construct an initial sign language translation model. The aligned data is input into the initial sign language translation model, and the sign language translation model is trained using the masked text reconstruction task and the sign language video to text translation task as joint optimization objectives to obtain the sign language translation model. The sign language translation module is used to acquire target sign language video data, input it into the sign language translation model, and obtain the translation result.

7. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the gloss-free sign language translation method based on latent visual-text alignment as described in any one of claims 1-5.

8. A computer storage medium, characterized in that, The computer storage medium stores computer program instructions, which, when executed by a processor, implement the gloss-free sign language translation method based on latent visual-text alignment as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Sign language translation method and device based on vision and word feature pre-training alignment

    CN119785439A

  • Continuous sign language recognition method based on multi-granularity cross-modal comparative learning

    CN119863842A