Sign Language Translation Method and Device Based on Pre-Training Alignment of Visual and Lexical Features

Through the sign language translation method with pre-training and aligned alignment of visual and word features, the problem of insufficient alignment of visual and word features in sign language translation is solved, and the accuracy and robustness of sign language recognition and translation are improved.

CN119785439BActive Publication Date: 2025-07-11ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510283064.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-11
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

In the existing sign language translation technology, the alignment of visual features and word features is insufficient, resulting in low translation accuracy.

Method used

Position and motion features are extracted and fused through visual encoder, and compared learning and mask prediction tasks are combined with text encoder, visual and word features are pre-trained, sign language recognition models are constructed, and large language models are translated.

Benefits of technology

It improves the accuracy and robustness of sign language recognition and translation, and achieves efficient and accurate sign language translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785439B_ABST
    Figure CN119785439B_ABST
Patent Text Reader

Abstract

The present invention discloses a sign language translation method and device based on pre-training alignment of visual and lemma features, belonging to the technical field of sign language translation, including: extracting visual features of a sign language video by using a visual encoder, extracting lemma text features by using a text encoder, and performing contrastive learning based on the visual and lemma text features to obtain a pre-trained visual encoder; performing pre-training on a text decoder for lemma text mask prediction; constructing the pre-trained visual encoder and text decoder into a sign language recognition model to recognize a lemma text sequence from the sign language video; connecting a large language model pre-trained within the domain to the sign language recognition model to construct a sign language translation model and jointly fine-tuning it to translate the lemma text sequence into a natural language text. The present invention can achieve more efficient, accurate and reliable sign language recognition and translation, and is applied to fields such as intelligent sign language translation, barrier-free communication, sign language education, etc., providing a more accurate and natural language interaction experience for the hearing-impaired group.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of sign language translation, and particularly relates to a sign language translation method and device based on visual and gloss feature pre-training alignment. Background Art

[0002] Sign language, as the main communication method among the hearing-impaired, plays a crucial role in information exchange and social interaction. With the rapid development of information technology and artificial intelligence, sign language recognition and translation technology has increasingly become an important technical means for building a barrier-free society. Sign language recognition can be divided into the following two categories according to different tasks. The first category is isolated sign language recognition (ISLR), which generally refers to recognizing individual sign language glosses or single characters from a sign language video. The second category is continuous sign language recognition (CSLR), and this task aims to recognize a sequence of sign language glosses whose word order is consistent with the sign language video from a sign language video. Sign language translation is a more complex task, which requires combining computer vision perception technology to perceive the deep features corresponding to the sign language video image and natural language processing to understand the text information corresponding to the sign language video. According to different research paradigms, sign language translation frameworks can be divided into two types: sign language video to text (sign2-text, S2T) and sign language video to gloss to text (sign2gloss2text, S2G2T). The S2T framework directly translates a continuous sign language video into a natural language sentence, while the S2G2T framework first extracts the gloss sequence in the sign language video through a continuous sign language recognition model, and then uses a pre-trained Gloss2Text network to translate the gloss sequence into a natural language text.

[0003] In terms of cutting-edge technologies, Gloss-free algorithms and Gloss-based algorithms each exhibit unique advantages. Gloss-free algorithms generally adopt a combination of a visual encoder and a text decoder, and optimize the model performance through pre-training and joint training. Such algorithms continue to innovate in the model structure, such as adopting combinations like ViT+GPT, ViT+T5 or CLIP+Transformer to better capture the deep features of sign language videos and generate accurate natural language texts. In addition, unsupervised algorithms and training enhancement strategies are also widely used to improve the generalization ability and translation accuracy of the model. Although Gloss-free algorithms reduce the dependence on sign language glosses, it is more difficult to align visual features and text features in the absence of explicit gloss supervision.

[0004] Gloss-based algorithms use the supervised information of words to strengthen the connection between sign language videos and natural language. This type of algorithm usually adopts the framework structure of visual encoder and text decoder, and combines the connection temporal classification (CTC) loss and text generation loss to achieve joint training of sign language recognition and translation. During the training process, a variety of enhancement strategies are used to improve model performance, such as using external corpus for pre-training, step-by-step optimization of visual and translation models, and cross-task learning. In addition, technologies such as fusing skeleton key points, introducing multimodal features, optimizing model structure, and introducing new loss functions are also used to further improve the performance of Gloss-based algorithms. Although the Gloss-based algorithm uses word supervision information, the accurate extraction of word sequences in continuous sign language recognition still faces challenges. In addition, due to the high cost of collecting and annotating word data, data scarcity problems arise, which further limits the generalization ability of the translation model.

[0005] In summary, due to the complexity and diversity of sign language videos, existing methods still have the problem of insufficient alignment of visual features and word features in the process of sign language video recognition and translation, which in turn affects the alignment from word to text, making the accuracy of sign language recognition and translation need to be further improved. Summary of the invention

[0006] In view of the above, the purpose of the present invention is to provide a sign language translation method and device based on pre-training alignment of visual and word features, and feature learning is performed through two core pre-training tasks: one is to compare and learn visual features with word text features, aiming to align the semantic representation of sign language video and word text; the other is a mask prediction task based on word text, aiming to enhance the generation ability of word text. In visual feature learning, human posture features and motion characteristics are fully integrated to fully capture the fine-grained motion information and global semantic relationships in sign language videos, aiming to improve the accuracy of sign language video recognition. After pre-training is completed, the language translation capability of a large language model is used based on words to achieve end-to-end training from sign language video to natural language, which effectively solves the bottleneck problems of existing methods in semantic alignment and temporal modeling, and significantly improves the translation accuracy and robustness of the model, thereby achieving more efficient, accurate and reliable sign language recognition and translation services. The present invention has the characteristics of strong versatility, high training efficiency and excellent performance. It provides an innovative solution for the development of sign language translation technology and can be widely used in intelligent sign language translation, barrier-free communication, sign language education and other fields, providing the hearing-impaired group with a more accurate and natural language interaction experience.

[0007] In order to achieve the above-mentioned invention object, the technical solution provided by the present invention is as follows:

[0008] An embodiment of the present invention provides a sign language translation method based on pre-training alignment of visual and word features, comprising the following steps:

[0009] Use a visual encoder to separately extract pose features and motion features from a sign language video and fuse them into visual features, use a text encoder to extract entry text features from the original entry text, and perform contrastive learning based on the visual features and the entry text features to obtain a pre-trained visual encoder;

[0010] Use the text encoder and the text decoder to perform pre-training on a masked prediction task based on the original entry text to obtain a pre-trained text decoder;

[0011] Construct a pre-trained visual encoder and a pre-trained text decoder into a sign language recognition model to recognize an entry text sequence from a sign language video, and train the sign language recognition model;

[0012] Use a large language model pre-trained within the domain based on a public dataset and a constructed sign language translation dataset as a language translation model to convert the entry text sequence into natural language text, connect the language translation model to the trained sign language recognition model to construct a sign language translation model and perform joint fine-tuning to achieve the translation of sign language videos into natural language text;

[0013] Input a new sign language video into the jointly fine-tuned sign language translation model for sign language translation and output natural language text.

[0014] Preferably, the use of the visual encoder to separately extract pose features and motion features from a sign language video and fuse them into visual features includes:

[0015] Construct a visual encoder including the Sapiens large model, the VideoMAE large model, and a feature fusion network. Use the Sapiens large model to capture the dynamic changes of the signer in the spatial dimension to extract the human pose features from the sign language video. The pose features include the signer's skeleton structure, limb movements, and facial expression changes. Use the VideoMAE large model to capture the dynamic changes of the signer in the time series dimension to extract the human motion features from the sign language video. The motion features include the motion trajectories between video frames, gesture changes, and rate features. Use the feature fusion network to fuse the pose features and the motion features to obtain the visual features of the sign language video.

[0016] Preferably, the feature fusion in the feature fusion network includes:

[0017] The pose features and motion features are respectively mapped to the same feature dimension space through the pose feature linear layer and the motion feature linear layer to achieve the unification of feature dimensions. The pose features and motion features after dimension unification are concatenated in the feature dimension to obtain a new feature vector. The new feature vector is input into a two-layer convolutional neural network for high-dimensional feature extraction. Each layer of the convolutional neural network includes a convolutional layer, a normalization layer, an activation layer, and a pooling layer. The high-dimensional features extracted by the two-layer convolutional neural network are input into a multi-layer perceptron for further feature fusion and transformation to obtain visual features with rich semantic information.

[0018] Preferably, the contrastive learning based on visual features and lemma text features to obtain a pre-trained visual encoder includes:

[0019] Using the CLIP algorithm to construct a CLIP loss function by calculating the similarity between visual features and lemma text features to achieve feature alignment at the video and sentence levels, and using the SoftDTW algorithm to construct a SoftDTW loss function by calculating the cost matrix between visual features and lemma text features to achieve feature alignment at the continuous video frame and lemma levels;

[0020] Taking the CLIP loss function and the SoftDTW loss function as the total alignment loss function in contrastive learning to pre-train the visual encoder to obtain a pre-trained visual encoder.

[0021] Preferably, the pre-training of the text decoder for the masked prediction task based on the original lemma text using a text encoder and a text decoder includes:

[0022] Randomly mask the original lemma text to obtain a masked lemma text, input the masked lemma text into the text encoder for encoding to obtain an encoded lemma text, then input the encoded lemma text into the text decoder for decoding to obtain a decoded lemma text, freeze the parameters of the text encoder, use the original lemma text as a label, and pre-train the text decoder through a masked prediction loss function. Train the generation ability of the text decoder by reconstructing the masked content to obtain a pre-trained text decoder.

[0023] Preferably, both the text encoder and the text decoder are constructed using the mBART model pre-trained on a large-scale corpus containing multiple languages.

[0024] Preferably, using the large language model pre-trained in the domain based on the public dataset and the constructed sign language translation dataset as a language translation model includes:

[0025] Construct a sign language translation dataset based on lemmas, strictly following the general sign language dictionary standard, covering sign language lemmas in multiple fields, providing the corresponding relationship between lemmas and natural language, and in the sign language translation dataset, each lemma text sequence corresponds to a natural language sentence;

[0026] Based on the obtained public dataset and the constructed sign language translation dataset, perform in-domain pre-training on the large language model through the Lora fine-tuning method to obtain a language translation model.

[0027] Preferably, the large language model uses the Qwen model.

[0028] Preferably, use the data-augmented sign language video dataset to train the sign language recognition model and jointly fine-tune the sign language translation model. Data augmentation includes: cropping, rotating, translating the sign language video, and adjusting color, brightness, contrast, and sharpening effects, and then normalizing the pixel values.

[0029] To achieve the above invention purpose, the embodiment of the present invention also provides a sign language translation device based on visual and lemma feature pre-training alignment, which is implemented by using the above-mentioned sign language translation method based on visual and lemma feature pre-training alignment, including: an alignment pre-training module, a masked pre-training module, a sign language recognition training module, a language translation fine-tuning module, and a sign language translation module;

[0030] The alignment pre-training module is used to respectively extract pose features and motion features from the sign language video by using a visual encoder and fuse them into visual features, extract lemma text features from the original lemma text by using a text encoder, and perform contrastive learning based on the visual features and lemma text features to obtain a pre-trained visual encoder;

[0031] The masked pre-training module is used to perform pre-training on the masked prediction task based on the original lemma text by using a text encoder and a text decoder to obtain a pre-trained text decoder;

[0032] The sign language recognition training module is used to construct the pre-trained visual encoder and the pre-trained text decoder into a sign language recognition model to recognize the lemma text sequence from the sign language video, and train the sign language recognition model;

[0033] The language translation fine-tuning module is used to use the large language model after in-domain pre-training based on the public dataset and the constructed sign language translation dataset as a language translation model to convert the lemma text sequence into natural language text, connect the language translation model to the trained sign language recognition model to construct a sign language translation model and perform joint fine-tuning to achieve translating the sign language video into natural language text;

[0034] The sign language translation module is used to input a new sign language video into the jointly fine-tuned sign language translation model for sign language translation and output natural language text.

[0035] Compared with the prior art, the beneficial effects of the present invention at least include:

[0036] (1) The present invention constructs a feature fusion network by combining human body pose features and motion features, enhancing the model's ability to capture signer's spatial position information and temporal relationship information, comprehensively capturing fine-grained motion information and global semantic relationships in sign language videos, which helps to improve the accuracy of sign language recognition, and proposes a sign language recognition and translation architecture based on pre-training alignment of visual and gloss features and then from gloss to natural language, combining the visual and gloss feature pre-training alignment method and the method of sign language translation based on gloss, improving the overall performance and robustness, and enabling more efficient, accurate and reliable sign language recognition and translation.

[0037] (2) The present invention performs two-stage pre-training. In the first stage, the visual encoder is pre-trained by means of contrastive learning of visual features and gloss text features to align sign language video features with gloss features, aligning the video and gloss text through a common vector space, enabling the visual encoder to efficiently perform cross-modal understanding to adapt to the sign language recognition task. In the second stage, the text decoder is pre-trained by means of mask prediction based on the original gloss text to further enhance the generation ability of the gloss text, so that the constructed sign language recognition model has the accurate recognition ability from video to gloss text.

[0038] (3) The present invention optimizes feature alignment by using CLIP loss and SoftDTW loss in the contrastive learning stage based on visual and gloss text features, precisely matching the dynamic actions of sign language videos with the semantics of gloss text, making the video information highly consistent with the language expression, thereby improving the quality of sign language recognition.

[0039] (4) The present invention constructs a sign language translation dataset based on glosses, provides a high-quality correspondence from glosses to natural language, reduces the dependence on large-scale labeled data through in-domain pre-training, and enhances the language translation ability and generalization ability of large language models. After connecting to the sign language recognition model to construct a sign language translation model and performing joint fine-tuning, it realizes the accurate translation from sign language video to gloss text sequence and finally to natural language text. Description of the Drawings

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 is a schematic flowchart of the sign language translation method based on pre-training alignment of visual and lexeme features provided by an embodiment of the present invention;

[0042] Figure 2 is a schematic diagram of the feature fusion network provided by an embodiment of the present invention;

[0043] Figure 3 is a schematic diagram of the framework based on pre-training alignment of visual and lexeme features provided by an embodiment of the present invention;

[0044] Figure 4 is a schematic diagram of the training data path from lexemes to the natural language domain implemented by an embodiment of the present invention based on a large language model;

[0045] Figure 5 is a schematic diagram of the overall framework of the sign language translation stage provided by an embodiment of the present invention;

[0046] Figure 6 is a schematic structural diagram of the sign language translation device based on pre-training alignment of visual and lexeme features provided by an embodiment of the present invention. Detailed implementation manners

[0047] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0048] The inventive concept of the present invention is as follows: Aiming at the problem of insufficient alignment of visual features and lexeme features in the process of sign language video recognition and translation in the prior art, which leads to insufficient translation accuracy, embodiments of the present invention provide a sign language translation method and device based on pre-training alignment of visual and lexeme features. For a sign language video dataset, a combination of human body pose features and motion characteristics is adopted. After pre-training alignment through visual features and lexeme text features, sign language recognition is achieved. Then, through a large language model, language translation is realized from the lexeme text sequence to the natural language text. The sign language recognition and language translation processes are integrally constructed into a sign language translation framework, and combined with various optimization technologies such as data augmentation, feature fusion, and masked text prediction, effectively improving the semantic representation ability and translation accuracy of sign language videos, overcoming the deficiencies in feature modeling and semantic alignment in the prior art, and achieving accurate and fluent translation from sign language videos to natural language, providing an efficient and robust solution for sign language translation.

[0049] The goal of sign language translation (SLT) is to accurately predict the corresponding natural language sentence according to a given sequence of sign language video frames. Input: A sign language video containing frames = , where is the th video frame, and are the length and width of the video frame. Output: a natural language sentence containing word segments = , where is the th word segment.

[0050] Figure 1 is a schematic flowchart of the sign language translation method based on pre-training alignment of visual and lemma features provided by an embodiment of the present invention. As Figure 1 shown, the embodiment provides a sign language translation method based on pre-training alignment of visual and lemma features, including the following steps:

[0051] S1. Use a visual encoder to separately extract pose features and motion features from a sign language video and fuse them into visual features, use a text encoder to extract lemma text features from the original lemma text, and perform contrastive learning based on the visual features and lemma text features to obtain a pre-trained visual encoder.

[0052] S1.1. Preprocessing of sign language video data augmentation. For a sign language video containing frames, save it in the format of an image and normalize the RGB values to the range of [0, 1]. In the training phase, horizontally flip the input sign language video with a probability of 50% in the spatial dimension, including operations such as cropping, rotation, translation, and adjusting color, brightness, contrast, and sharpening effects, and then normalize the pixel values of the input image data. Define the mean and standard deviation of each channel (RGB) of the image, and normalize each pixel value according to the following formula:

[0053] ,

[0054] where is the normalized pixel value, is the original pixel value, and are the mean and standard deviation of the pixel values respectively. By subtracting the mean operation, the pixel values are centered so that their mean is 0, eliminating the influence of brightness differences. By dividing by the standard deviation operation, the pixel value range is normalized, avoiding the problem that feature values at different scales cause instability in model optimization. Through standardization, the convergence of the model can be accelerated, and the problems of gradient disappearance or explosion can be avoided, improving the training stability.

[0055] By generating diverse training samples through these data augmentation methods, the robustness and generalization ability of the dataset can be improved, providing rich data support for subsequent model training and avoiding overfitting problems.

[0056] S1.2, Human Pose Feature Encoding. Use the pre-trained Sapiens large model as the human pose feature encoder Extract the human pose features from the sign language video The pose features include the skeleton structure, limb movements, and facial expression changes of the signer. Freeze the parameters of the Sapiens large model and extract the pose features through the ability of the Sapiens large model to capture spatial features , is the number of video frames, is the feature dimension of, and its formula is defined as:

[0057] ,

[0058] where is the human pose feature encoder.

[0059] In the embodiment, the Sapiens large model is based on the Masked Autoencoder (MAE) architecture, draws on the image encoding method of Vision Transformer (ViT), and through pre-training on the large-scale dataset Humans-300M, can extract fine-grained gestures, arms, upper body and other parts of the representation from images, including the human pose, contour, structure, and the relationship information of each part. These deep representations provide the visual information of the signer's pose, and by capturing the dynamic changes of the signer in space, the understanding ability of the model for sign language expression is improved, and the visual encoding process is optimized. The encoder of the Sapiens large model processes the image by dividing it into non-overlapping image patches of a fixed size. Different from the traditional Convolutional Neural Network (CNN), the Sapiens large model adopts the self-attention mechanism, which can capture the long-range dependencies and global information in the image, randomly select a part of the image patches for masking, and the remaining image patches remain visible. Through this masking strategy, the model can learn to recover the complete image from the partially visible information, so as to effectively extract the latent representation of the image. In the embodiment, a frozen Sapiens large model is used to extract the features of the signer. First, the video of the signer is used as the input, and each frame of the image is interpolated to a size of 1024×1024 pixels . Then, the image is divided into non-overlapping image patches of a fixed size of 16×16 , where the number of image patches seq is 4096, , and no masking is performed on any image patches during this process to comprehensively extract the latent human representation of the image. The sequence is input into the ViT architecture, and the feature dimension after passing through the linear projection layer is And add it to the features of the positional encoding. The feature dimension of the picture block is mapped to in the Embedding layer. After passing through multiple Transformer encoders in Vit, the obtained feature size is Among them, according to different Spaiens model parameters, in models with sizes of 0.3b, 0.6b, 1b, and 2b they are 1024, 1280, 1536, and 1920 respectively, and the number of Transformer layers is 24, 32, 40, and 48 respectively. Perform average pooling to obtain features Finally, T Concatenate the features of the frames to obtain the human body pose features of this sign language video

[0060] S1.3, Human motion feature encoding. Use the pre-trained VideoMAE large model as the human motion feature encoder Extract motion features from the sign language video . The motion features include the motion trajectories, gesture changes, and rate features between video frames to supplement the deficiencies of the pose features. Freeze the parameters of the VideoMAE large model, process the sign language segments segmented from the video, and use the sliding window method to capture the implicit vocabulary-level representation. First, segment the sign language video into fixed-length overlapping segments, and then input each segment into to extract the implicit vocabulary-level motion features Among them, N= T / t t is the number of frames between the start frames of adjacent segments, is the feature dimension of is to represent N a sequence of segment-level features, and its formula is defined as:

[0061]

[0062] Among them, is the human motion feature encoder.

[0063] In the embodiment, the VideoMAE large model is pre-trained on the Kinetics-400 dataset, and uses the Masked Autoencoder technology to perform time series modeling on the sign language video, extract the motion trajectories, gesture changes, and rate features between video frames, enhance the ability to model the time series of sign language actions, so that the translation system can more accurately capture and understand the dynamic information in the sign language video. In the embodiment, by freezing its parameters, it is used to extract the motion features of the sign language video, and the input video is represented as , where each frame is cropped to a fixed size , first, the input video is divided into N video segments using a sliding window of size 16 frames, and the step size of the sliding window movement is t . N = T / t , a single video is divided into N segments , where , VideoMAE uses a spatio-temporal block embedding of size 2×16×16×3 to sample the video frames, and the resulting feature representation is , where seq is the sequence length, with a size of 1568 ( ), and at the same time, the dimension of each spatio-temporal block is mapped to , and the finally obtained encoded feature is , and then this feature is pooled into , and finally N segment features are concatenated to obtain .

[0064] S1.4, Feature fusion. The feature fusion network is used to fuse the pose feature and the motion feature . As Figure 2 shown, the pose feature and the motion feature are respectively mapped to the same feature dimension space through the pose feature linear layer and the motion feature linear layer to achieve the unification of the feature dimensions, and the pose feature and the motion feature are obtained, where is the unified feature dimension, which is set to 1024 in the embodiment. The pose feature and the motion feature with unified dimensions are concatenated in the feature dimension to obtain a new feature vector , , which is expressed as:

[0065] ,

[0066] where, represents feature concatenation.

[0067] The concatenated feature vector It is input into a two-layer convolutional neural network for high-dimensional feature extraction. Each layer of the convolutional neural network (CNN) includes a convolutional layer (K5), a normalization layer (BN), an activation layer (ReLU), and a pooling layer (P2). The structure of the two-layer CNN is represented as {K5+BN+ReLU→P2→K5+BN+ReLU→P2}. Among them, the convolutional kernel size of the convolutional layer is 5 and the stride is 1. The pooling layer is a max-pooling layer, the pooling kernel size is 1 and the stride is 2. By extracting short-term spatio-temporal features, the length of the feature sequence is reduced by the max-pooling layer. The high-dimensional features extracted by the two-layer convolutional neural network are input into a multi-layer perceptron (MLP) for further feature fusion and transformation to obtain visual features with rich semantic information. , where is the sequence length after feature fusion. The above feature fusion strategy can effectively retain the spatial structure information and temporal dynamic features of gesture actions, thereby optimizing the overall representation ability of sign language videos.

[0068] S1.5, Text feature encoding. To effectively process natural language text data, a text encoder in the pre-trained mBART model is adopted. . The mBART model is a multi-lingual neural machine translation (NMT) model based on an autoregressive model, which has been pre-trained on a large-scale corpus CC25 containing 25 languages. Its encoder is composed of 12 layers of Transformers stacked together, and can effectively capture the semantics and context information of the input sentence. The text encoder can extract semantic information from the original lemma text and effectively encode the features of the language to obtain lemma text features , is the number of lemmas, is 's feature dimension, and its formula is defined as:

[0069] ,

[0070] where, is the text encoder.

[0071] S1.6, Alignment of visual features and lemma text features. As Figure 3As shown in the figure, the Sapiens large model, the VideoMAE large model, and the feature fusion network are constructed as a visual encoder. By simultaneously using the CLIP algorithm and the SoftDTW algorithm to align visual features and gloss text features, the model's ability to model the semantic and temporal dependency relationships of sign language videos is enhanced, the matching accuracy of sign language video actions and gloss text is improved, and thus the accuracy of sign language translation is optimized. Among them, the CLIP algorithm constructs a CLIP loss function by calculating the similarity between visual features and gloss text features to achieve feature alignment at the video and sentence levels. The SoftDTW algorithm constructs a SoftDTW loss function by calculating the cost matrix between visual features and gloss text features to achieve feature alignment at the continuous video frame and gloss levels. The CLIP loss function and the SoftDTW loss function are used as the total alignment loss function in contrastive learning to pre-train the visual encoder, thereby obtaining the pre-trained visual encoder.

[0072] The core idea of the CLIP algorithm is to pre-train with a large amount of image-text pair data to learn a vision-language representation model that can map images and text to the same semantic space, enabling efficient cross-modal retrieval and understanding between images and text. CLIP emphasizes the advantages of learning from natural language over other task-agnostic pre-training methods, making it particularly suitable for sign language tasks. The characteristic of CLIP is the use of the contrastive learning method, which makes the visual feature and gloss text feature representations in the same semantic space, and optimizes the model parameters by calculating the cosine similarity between visual features and gloss text features. The training objective of CLIP is to maximize the cosine similarity of positive sample pairs (i.e., samples where visual features and gloss text features match), while minimizing the cosine similarity of negative sample pairs (i.e., samples where visual features and gloss text features do not match). This method can effectively enhance the similarity between sign language videos and corresponding texts in the semantic space, ensure that the visual features of sign language videos can accurately express the meaning of natural language, and achieve global alignment at the video level and text level. Input the visual features and gloss text features into their respective head networks and map them to a joint multi-modal semantic space through linear projection for similarity calculation. Both head networks in CLIP consist of a linear layer, and this process is expressed as:

[0073] ,

[0074] where, is the linear layer, is the last layer of the visual encoder at <cls>The activation value at is the last layer of the text encoder at <eos>The activation value at. Subsequently, and are layer-normalized and pair-wise scaled to generate a pair of features for the video and the sentence { }, and then the CLIP loss is calculated through a symmetric cross-entropy loss function , which is expressed by the formula:

[0075] ,

[0076] where and represent the -th visual feature and the -th lemma text feature respectively, and represent the -th visual feature and the -th lemma text feature respectively, represents the cosine similarity, represents the size of the batch, represents a learnable scaling factor.

[0077] The SoftDTW algorithm is an improved version of the dynamic time warping (DTW) algorithm, which processes potential temporal differences in video and lemma text sequences through a flexible alignment strategy. This algorithm is applicable to calculating the distance between two sequences of different lengths of visual features and lemma text features. According to the characteristics that the alignment path is continuous and unidirectional, it can effectively reflect the similarity between two time series. This method performs fine-grained alignment between the key frame sequence of the sign language video and the corresponding text lemma, enabling the temporal dynamic changes of the sign language actions to more precisely match the text sequence, solving the differences in gesture execution speed and action styles among different sign language users, ensuring that the model can capture the time dependence of sign language expressions, thereby optimizing the correspondence between video frames and text lemmas and achieving precise alignment at the picture level and the lemma level. By summing the distances at each point on the alignment path, the obtained dynamic time warping distance, as the DTW distance measuring the similarity of time series, can be backpropagated during the neural network training process to optimize the network parameters. In the embodiment, for the two feature sequences of the input visual feature and the lemma text feature , SoftDTW introduces a smooth distance function to measure the distance between each pair of elements ( The distance between , and the calculation method of the cost matrix is as follows:

[0078] ,

[0079] ,

[0080] where is the th visual feature, the th lemma text feature, is the Euclidean distance, is the Softmin function, is the matrix element at the th row and the th column in the cost matrix, , and are used in the formula to respectively refer to , and . The parameter is used to control the smoothness of alignment. When , it approaches the Euclidean distance, while when , it approaches the traditional DTW. In the embodiment, ablation experiments are carried out on the value of .

[0081] Finally, the total alignment loss function in the contrastive learning is expressed as:

[0082] ,

[0083] where is the matrix element at the th row and the th column in the cost matrix. Based on the total alignment loss function , the visual encoder is optimized through contrastive learning to obtain the pre-trained visual encoder.

[0084] S2. Use the text encoder and text decoder to perform pre-training on the masked prediction task based on the original lemma text to obtain the pre-trained text decoder.

[0085] In the embodiment, the text encoder and text decoder in the pre-trained mBART model on a large-scale corpus containing multiple languages are adopted, and the tokenizer of Mbart is pruned using the dataset containing the lemma text sequence . The text encoder and text decoder are trained through the masked self-supervised learning strategy, and the goal is to recover the masked words in the input sentence. According to the name of the training sign language video, find the corresponding lemma text and perform masking operation on the text. In the embodiment, a text perturbation strategy based on noise injection is used, aiming to enhance the robustness of the text decoder by simulating a noisy environment. Specifically, first randomly discard some words in the original lemma text according to a preset noise rate, and use a special token <mask>Instead, the masked lemma text is obtained, and these missing positions are dynamically selected by the noise rate to introduce uncertainty and randomness. For the input sequence of lemma texts , retain the first words out of a total of words. The calculation formula of is:

[0086] ,

[0087] where is a random value uniformly distributed within . Replace the non-retained words in the lemma text sequence with <mask>, the masked gloss text is obtained. As Figure 3 shown, the masked gloss text is tokenized by a tokenizer and then input into a text encoder (using a 12-layer Transformer structure) for encoding to obtain the encoded gloss text. Then, the encoded gloss text is input into a text decoder for decoding to obtain the decoded gloss text. The parameters of the text encoder are frozen, and the original gloss text is used as a label, and the text decoder is pre-trained through a masked prediction loss function (i.e., the cross-entropy loss function) to train the generation ability of the text decoder through the above reconstruction of the masked content, thereby obtaining the pre-trained text decoder.

[0088] S3. The pre-trained visual encoder and the pre-trained text decoder are constructed into a sign language recognition model to recognize the gloss text sequence from the sign language video, and the sign language recognition model is trained.

[0089] In the sign language recognition stage (SLR), the pre-trained visual encoder and the pre-trained text decoder are constructed into a sign language recognition model for converting the input sign language video frame sequence into the corresponding gloss text sequence.

[0090] The pre-trained visual encoder is tasked with extracting features from the video frames to obtain visual features . The visual encoder has optimized its representation ability for visual information through co-training with the text decoder in the pre-training stage. After the visual encoder, a language classifier and CTC loss are used for training. The CTC loss considers all possible alignment paths between the visual features and the gloss text sequence and minimizes the error simultaneously. By traversing all possible alignment paths between the input video and the corresponding gloss text sequence to calculate the conditional probability :

[0091] ,

[0092] where is one of the paths, is all possible alignment paths, and the probability is calculated by the language classifier (independent of the visual encoder, including a linear layer and a Softmax layer, used to map the visual features to frame-level gloss probabilities), and the CTC loss is expressed as:

[0093] .

[0094] The pre-trained text decoder is tasked with based on the visual features Generate the corresponding sequence of lemma texts The text decoder gradually predicts the next token until the complete sequence of lemma texts is generated. This process is based on modeling the conditional probability of the previous tokens, expressed by the formula:

[0095] ,

[0096] ,

[0097] ,

[0098] where the first word of the sentence is artificially set as a special flag word <bos>, until the generation of the flag word <eos>, the text decoder will end the generation. The decoded vector in the linear layer and the Softmax layer then generates the predicted sequence of lemma texts , is the weight of the linear layer, is the bias of the linear layer. By minimizing the cross-entropy loss of the generated lemma text sequence to optimize the network.

[0099] In the sign language recognition stage, the total optimization loss for training the sign language recognition model is expressed as:

[0100] ,

[0101] where, and are weight parameters, and in the embodiment, is taken.

[0102] S4. Use the large language model after in-domain pre-training based on the public dataset and the constructed sign language translation dataset as the language translation model to convert the sequence of lemma texts into natural language text. Connect the language translation model to the trained sign language recognition model to construct a sign language translation model and perform joint fine-tuning to achieve the translation of sign language videos into natural language text.

[0103] In the sign language translation stage (STR), it realizes sign language translation based on the pre-training alignment of video and lemma text features, and further explores the translation scheme from the sequence of lemma texts to natural language sentences based on the fine-tuning of the large language model, which not only reduces the dependence on large-scale labeled data, but also improves the effect and scalability of the sign language translation task.

[0104] In the embodiment, the large language model uses Qwen2.5-14B-Instruction. After in-domain pre-training through Lora fine-tuning on the constructed sign language translation dataset and the public dataset, it is used as the language translation model, enabling the model to effectively capture the semantic association between the lemmas in the sign language video and natural language. As Figure 4 shown, it realizes the translation from the sequence of lemma texts to natural language sentences. As Figure 5 As shown, after connecting the large language model (language translation model) pre-trained within the domain to the trained sign language recognition model, the overall structure is built into a sign language translation model and jointly fine-tuned to achieve sign language translation based on lexical items. In the embodiment, the sign language translation dataset built based on lexical items strictly follows the general sign language dictionary standard, covers sign language lexical items in multiple fields, provides the corresponding relationship between lexical items and natural language, has a wide coverage and accurate annotation. In the sign language translation dataset, each lexical item text sequence corresponds to a natural language sentence, providing high-quality data support for model training. Through in-domain pre-training on the self-built sign language translation dataset and public datasets and combined with fine-tuning of the large language model, the language translation from sign language lexical items to natural language is achieved, which can reduce the dependence on large-scale labeled data and improve the generalization ability of the sign language translation task.

[0105] The process of the language translation stage is expressed as:

[0106] ,

[0107] ,

[0108] Among them, is the language translation model, is the conditional probability of generating the natural language sentence from the lexical item text sequence . By minimizing the cross-entropy loss of generating the natural language sentence to optimize the network, it is expressed as follows:

[0109] .

[0110] In the sign language translation stage, the total optimization loss of jointly fine-tuning the sign language translation model including the language translation model and the trained sign language recognition model is expressed as:

[0111] ,

[0112] Among them, and are weight parameters. In the embodiment, is taken.

[0113] S5. Input the new sign language video into the jointly fine-tuned sign language translation model for sign language translation and output the natural language text.

[0114] The goal of sign language translation is to convert the input sequence of sign language video frames into the corresponding natural language sentence . Specifically, input the sign language video frames into the jointly fine-tuned sign language recognition model to obtain the predicted gloss text sequence, and then map the gloss text sequence through the jointly fine-tuned language translation model to obtain the natural language sentence formed by the word segmentation in the gloss text sequence.

[0115] Based on the same inventive concept, as Figure 6 shown, an embodiment of the present invention further provides a sign language translation device 600 based on pre-training alignment of visual and gloss features, including: an alignment pre-training module 610, a mask pre-training module 620, a sign language recognition training module 630, a language translation fine-tuning module 640, and a sign language translation module 650.

[0116] The alignment pre-training module 610 is used to respectively extract pose features and motion features from the sign language video by using a visual encoder and fuse them into visual features, extract gloss text features from the original gloss text by using a text encoder, and perform contrastive learning based on the visual features and gloss text features to obtain the pre-trained visual encoder.

[0117] The mask pre-training module 620 is used to perform pre-training on the mask prediction task based on the original gloss text by using a text encoder and a text decoder to obtain the pre-trained text decoder.

[0118] The sign language recognition training module 630 is used to construct the pre-trained visual encoder and the pre-trained text decoder into a sign language recognition model to recognize the gloss text sequence from the sign language video and train the sign language recognition model.

[0119] The language translation fine-tuning module 640 is used to use the large language model pre-trained within the domain based on the public dataset and the constructed sign language translation dataset as the language translation model to convert the gloss text sequence into natural language text, connect the language translation model to the trained sign language recognition model to construct a sign language translation model and perform joint fine-tuning to achieve the translation of the sign language video into natural language text.

[0120] The sign language translation module 650 is used to input the new sign language video into the jointly fine-tuned sign language translation model for sign language translation and output the natural language text.

[0121] It should be noted that the above-mentioned sign language translation device based on pre-training alignment of visual and gloss features belongs to the same inventive concept as the sign language translation method based on pre-training alignment of visual and gloss features. For the specific implementation process, please refer to the embodiment of the sign language translation method based on pre-training alignment of visual and gloss features, which will not be elaborated here.

[0122] The specific embodiments described above have elaborated in detail the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention shall be included within the protection scope of the present invention.< / eos> < / bos> < / mask> < / mask> < / eos> < / cls>

Claims

1. A sign language translation method based on pre-training alignment of visual and lemma features, characterized in that It includes the following steps: Use a visual encoder to separately extract pose features and motion features from a sign language video and fuse them into visual features, use a text encoder to extract lemma text features from the original lemma text, and perform contrastive learning based on the visual features and lemma text features to obtain a pre-trained visual encoder. The contrastive learning process includes: using the CLIP algorithm to construct a CLIP loss function by calculating the similarity between the visual features and the lemma text features to achieve feature alignment at the video and sentence levels, using the SoftDTW algorithm to construct a SoftDTW loss function by calculating the cost matrix between the visual features and the lemma text features to achieve feature alignment at the continuous video frame and lemma levels, and using the CLIP loss function and the SoftDTW loss function as the total alignment loss function in the contrastive learning to pre-train the visual encoder to obtain a pre-trained visual encoder; Use the text encoder and text decoder to perform pre-training on the masked prediction task based on the original lemma text to obtain a pre-trained text decoder, including: randomly masking the original lemma text to obtain a masked lemma text, inputting the masked lemma text into the text encoder for encoding to obtain an encoded lemma text, then inputting the encoded lemma text into the text decoder for decoding to obtain a decoded lemma text, freezing the parameters of the text encoder, using the original lemma text as a label and pre-training the text decoder through a masked prediction loss function, and training the generation ability of the text decoder by reconstructing the masked content to obtain a pre-trained text decoder; Construct a sign language recognition model with the pre-trained visual encoder and the pre-trained text decoder to identify a lemma text sequence from a sign language video, and train the sign language recognition model; Use a large language model pre-trained within the domain based on a public dataset and a constructed sign language translation dataset as a language translation model to convert the lemma text sequence into natural language text, connect the language translation model to the trained sign language recognition model to construct a sign language translation model and perform joint fine-tuning to achieve the translation of sign language videos into natural language text; Input a new sign language video into the jointly fine-tuned sign language translation model for sign language translation and output natural language text.

2. The sign language translation method based on pre-training alignment of visual and lemma features according to claim 1, characterized in that, The use of the visual encoder to separately extract pose features and motion features from a sign language video and fuse them into visual features includes: Construct a visual encoder including the Sapiens large model, the VideoMAE large model, and a feature fusion network. Use the Sapiens large model to capture the dynamic changes of the signer in the spatial dimension to extract the pose features of the human body from the sign language video. The pose features include the skeleton structure, limb movements, and facial expression changes of the signer. Use the VideoMAE large model to capture the dynamic changes of the signer in the time series dimension to extract the motion features of the human body from the sign language video. The motion features include the motion trajectory between video frames, gesture changes, and rate features. Use the feature fusion network to fuse the pose features and motion features to obtain the visual features of the sign language video.

3. The sign language translation method based on pre-training alignment of visual and lemma features according to claim 2, wherein Performing feature fusion in the feature fusion network includes: The pose features and motion features are respectively mapped to the same feature dimension space through the pose feature linear layer and the motion feature linear layer to achieve the unification of feature dimensions. The pose features and motion features after dimension unification are concatenated in the feature dimension to obtain a new feature vector. The new feature vector is input into a two-layer convolutional neural network for high-dimensional feature extraction. Each layer of the convolutional neural network includes a convolutional layer, a normalization layer, an activation layer, and a pooling layer. The high-dimensional features extracted by the two-layer convolutional neural network are input into a multi-layer perceptron for further feature fusion and transformation to obtain visual features with rich semantic information.

4. The sign language translation method based on pre-training alignment of visual and lemma features according to claim 1, characterized in that, Both the text encoder and the text decoder are constructed using the mBART model pre-trained on a large-scale corpus containing multiple languages.

5. The sign language translation method based on pre-training alignment of visual and lexeme features according to claim 1, wherein The large language model pre-trained within the domain based on the public dataset and the constructed sign language translation dataset is used as the language translation model, including: Construct a sign language translation dataset based on lemmas, strictly following the general sign language dictionary standard, covering sign language lemmas in multiple fields, providing the correspondence between lemmas and natural language, and each lemma text sequence in the sign language translation dataset corresponds to a natural language sentence. Based on the obtained public dataset and the constructed sign language translation dataset, the large language model is pre-trained within the domain through the Lora fine-tuning method to obtain the language translation model.

6. The sign language translation method based on pre-training alignment of visual and lemma features according to claim 1 or 5, characterized in that The large language model uses the Qwen model.

7. The sign language translation method based on pre-training alignment of visual and lexicographic features according to claim 1, wherein The sign language recognition model is trained and the sign language translation model is jointly fine-tuned using the data-augmented sign language video dataset. Data augmentation includes: cropping, rotating, translating the sign language video, and adjusting the color, brightness, contrast, and sharpening effects, and then normalizing the pixel values.

8. A sign language translation device based on pre-training alignment of visual and lemma features, which is implemented by using the sign language translation method based on pre-training alignment of visual and lemma features according to any one of claims 1-7, characterized in that Including: An alignment pre-training module, a masked pre-training module, a sign language recognition training module, a language translation fine-tuning module, and a sign language translation module; The alignment pre-training module is used to respectively extract pose features and motion features from the sign language video using the visual encoder and fuse them into visual features, extract lemma text features from the original lemma text using the text encoder, and perform contrastive learning based on the visual features and lemma text features to obtain the pre-trained visual encoder; The masked pre-training module is used to perform masked prediction task pre-training based on the original lemma text using the text encoder and the text decoder to obtain the pre-trained text decoder; The sign language recognition training module is used to construct a sign language recognition model from the pre-trained visual encoder and the pre-trained text decoder to recognize the lemma text sequence from the sign language video, and train the sign language recognition model; The language translation fine-tuning module is used to use the large language model pre-trained within the domain based on the public dataset and the constructed sign language translation dataset as the language translation model to convert the lemma text sequence into natural language text, connect the language translation model to the trained sign language recognition model to construct a sign language translation model and perform joint fine-tuning to achieve the translation of sign language video into natural language text; The sign language translation module is used to input a new sign language video into the jointly fine-tuned sign language translation model for sign language translation and output natural language text.

Citation Information

Patent Citations

  • Continuous sign language recognition method fusing cross-modal alignment auxiliary task

    CN116311522A

  • Sign language translation method and device, electronic equipment and storage medium

    CN118609202A