Cross-frame interactive continuous sign language recognition method based on extended context space
By constructing a cross-frame interaction recognition method that extends the context space, using context-related and differential perception modules to extract features, the problem of insufficient recognition accuracy in the prior art is solved, and efficient and accurate sign language recognition is achieved, which is suitable for real-time communication and other fields such as human-computer interaction and virtual reality.
Patent Information
- Application Number
- CN202510384529.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
Existing continuous sign language recognition methods ignore extended context information, resulting in insufficient recognition robustness and accuracy, making it difficult to effectively capture long-term symbolic information in sign language videos.
The trained recognition network model is adopted to extract the context-dependent characteristics and motion characteristics of sign language videos through the context-dependent perception module and the context-differential perception module, and combine residual connection and self-distillation loss function to construct a cross-frame interactive recognition method that extends the context space.
It improves the accuracy and robustness of sign language recognition, can maintain a high recognition rate under different environments and conditions, is suitable for real-time communication scenarios, and has personalized recognition capabilities.
Smart Images

Figure CN120340125A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of continuous sign language recognition, and particularly relates to a cross-frame interaction continuous sign language recognition method based on an extended context space. Background Art
[0002] Continuous Sign Language Recognition (CSLR) focuses on recognizing sign language video sequences without clear boundaries and converting them into annotation sequences. Traditional CSLR methods usually adopt frame-level feature extraction to avoid the need for explicit symbol segmentation, and at the same time use CTC (Connectionist Temporal Classification) for label alignment. However, these methods ignore the influence of cross-frame information (such as body trajectories), resulting in low accuracy.
[0003] Currently, there are already methods to improve traditional methods by establishing various cross-frame interactions. For example, SEN captures motion features by calculating the differences between consecutive frames, while CorNet evaluates the correlation between adjacent frames to align shifted regions. CoSign calculates the coordinate differences of key points to obtain the dynamic information between two consecutive frames. In addition, SignGraph uses a graph-based method to establish edges between regions in consecutive frames to achieve cross-frame feature learning. Generally speaking, existing methods mainly focus on calculating cross-frame information within a context space composed of 2-3 frames.
[0004] However, the above methods ignore the importance of extended context for CSLR, which has an adverse impact on the robustness and accuracy of recognition. Generally speaking, a single symbol in a video spans 8-13 frames, but existing methods are difficult to extend to such a long context space, limiting their ability to capture complete symbols. Although 3D CNN promotes multi-frame interaction by adjusting the kernel size, they often have deficiencies in precisely defining sign language boundaries. Summary of the Invention
[0005] The purpose of this application is to provide a cross-frame interaction continuous sign language recognition method based on an extended context space to overcome the deficiencies of existing technologies in recognition robustness and accuracy.
[0006] To achieve the above purpose, the technical solution of this application is as follows:
[0007] A cross-frame interaction continuous sign language recognition method based on an extended context space uses a trained recognition network model for sign language recognition. The recognition network model includes an encoder, a decoder, and a classifier. The cross-frame interaction continuous sign language recognition method based on an extended context space includes:
[0008] Input the sign language video to be recognized into the encoder of the recognition network model. For the features of each stage extracted by the encoder, extract the context correlation features and motion features through the context-dependent perception module and the context difference perception module respectively, then concatenate the context correlation features and motion features and perform a residual connection to obtain the features finally extracted at the current stage, and input them into the next stage;
[0009] Input the features output by the encoder into the decoder, and then obtain the final recognition result through the classifier.
[0010] Further, the context-dependent perception module extracts context correlation features, including:
[0011] The input features of the context-dependent perception module are processed by convolution for channel downsampling;
[0012] Use a sliding window along the time dimension for the features obtained by channel downsampling to construct a context space;
[0013] Calculate the affinity matrix of each frame in the context space with the features of the remaining frames in the context space;
[0014] Perform normalization processing on the affinity matrix;
[0015] Use the normalized affinity as the weight to aggregate spatio-temporal information, and then use convolution to restore the channel dimension to obtain the context correlation features.
[0016] Further, the context difference perception module extracts motion features, including:
[0017] The input features of the context difference perception module are processed by convolution for channel downsampling;
[0018] Use a sliding window along the time dimension for the features obtained by channel downsampling to construct a context space;
[0019] Subtract each frame in the context space from the frames with different time spans in the context to obtain motion features of different speeds;
[0020] Process the features of different motion speeds sequentially using convolution, adaptively enhance the regions that need to be concerned, and obtain the enhanced features;
[0021] Concatenate the enhanced motion features of each frame in the context space along the channel dimension, then use convolution to adaptively aggregate the motion features of different speeds, and restore the channel dimension to obtain the motion features.
[0022] Further, when training the recognition network model, the total loss function includes the CTC loss and the self-distillation loss.
[0023] Further, the decoder includes a one-dimensional convolutional neural network and a bidirectional long short-term memory neural network.
[0024] Further, for the self-distillation loss, the output of the entire recognition network model is regarded as the teacher, and the output of the one-dimensional convolutional neural network is regarded as the student, and it is calculated using the KL divergence.
[0025] This application proposes a cross-frame interaction continuous sign language recognition method based on an extended context space. By introducing the extended context space, it can better capture the context information of sign language, thereby maintaining a high recognition accuracy under different environments and conditions. This method can effectively cope with the noise and interference in sign language expression and improve the stability of the system in practical applications. Using the mechanism of cross-frame interaction, it can more comprehensively analyze the dynamic features of sign language and reduce recognition errors caused by insufficient single-frame information. Through the analysis of continuous gestures, the system can more accurately understand the semantics of sign language and improve the overall recognition effect. In the design, the requirements of real-time processing are considered. It can achieve fast response while ensuring recognition accuracy, is suitable for real-time communication scenarios, and improves the user experience. It can adapt to the sign language habits and expression methods of different users, has good personalized recognition ability, and can be dynamically adjusted according to the characteristics of users to further improve the recognition effect. The technical solution of this application is not only applicable to sign language recognition, but also can be extended to other fields, such as human-computer interaction, virtual reality, etc., providing new ideas and methods for the development of related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a flowchart of the cross-frame interaction continuous sign language recognition method based on the extended context space of this application.
[0027] Figure 2 It is a schematic diagram of the structure of the recognition network model of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0029] The goal of continuous sign language recognition is to convert a video sequence with T frames into an annotation sequence where N represents the length of the annotation sequence. The annotation sequence refers to the lexical-level representation of each sign language action in the sign language video. Traditional methods use a two-dimensional convolutional neural network (2D-CNN) to extract frame-by-frame features These features are then processed by a one-dimensional convolutional neural network (1D-CNN) and a bidirectional long short-term memory neural network (BiLSTM) for local and global temporal modeling. Finally, the output is classified and optimized through the CTC loss.
[0030] This application improves the traditional method. After each stage of the encoder (ResNet34), the context correlation features are captured by a context correlation awareness module, and the motion features of different speed movements are captured by a context difference awareness module. Finally, the output features of the two modules are integrated by learnable parameters and input into the next stage of the encoder.
[0031] An embodiment of this application, as Figure 1 shown, proposes a cross-frame interaction continuous sign language recognition method based on an extended context space, and uses a trained recognition network model for sign language recognition. The recognition network model includes an encoder, a decoder, and a classifier. The cross-frame interaction continuous sign language recognition method based on the extended context space includes:
[0032] Step S1: Input the sign language video to be recognized into the encoder of the recognition network model. For the features of each stage extracted by the encoder, the context correlation features and motion features are respectively extracted by a context correlation awareness module and a context difference awareness module, and then the context correlation features and motion features are concatenated and then residual-connected to obtain the finally extracted features of the current stage, which are input into the next stage.
[0033] To recognize continuous sign language, this application constructs and trains a recognition network model, as Figure 2 shown, including an encoder, a decoder, and a classifier.
[0034] When training the recognition network model, the sign language videos are sourced from the publicly available large German sign language dataset PHOENIX14, PHOENIX14-T, and the Chinese dataset CSL-Daily. The video frames in the dataset are obtained as training samples through data augmentation and size adjustment, and the recognition network model is trained.
[0035] After training the recognition network model, for the sign language video to be recognized, each frame of the sign language video is uniformly cropped to the same size and input into the recognition network model to obtain the recognition result.
[0036] This application improves the encoder. The encoder uses a two-dimensional convolutional neural network, such as lightweight convolutional neural networks like ResNet34, ResNet18, ResNet50, SqueezeNet, etc.
[0037] In an encoder, feature extraction is usually carried out in several stages. For example, the residual block layer of ResNet34 includes four stages. In this application, the same context-related calculation is performed on the current stage features extracted in each stage and then input to the next stage. After the context-related calculation on the current stage features extracted in the last stage, they are used as the output of the entire encoder.
[0038] For the current stage features extracted in each stage of the encoder, the same context-related calculation is performed, that is, context correlation features and motion features are respectively extracted through a context correlation perception module and a context difference perception module, and then the context correlation features and motion features are concatenated and residual connection is performed to obtain the finally extracted features of the current stage.
[0039] Among them, extracting context correlation features through the context correlation perception module includes:
[0040] Step 1.1.1: The input features of the context correlation perception module are processed by convolution for channel downsampling.
[0041] The input features of the context correlation perception module are processed by convolution for channel downsampling, and the channels are downsampled to C hid to obtain thus reducing the computational burden in subsequent operations.
[0042] Step 1.1.2: The features obtained by channel downsampling are used to construct a context space along the time dimension by using a sliding window.
[0043] In this step, the features X obtained by channel downsampling hid are used to construct a context space along the time dimension by using a sliding window:
[0044]
[0045] In the formula represents a sliding window operation with a window size of 2n + 1. where m ∈ [-n, n] represents the m-th frame, and |n| > 1.
[0046] It should be noted that the context space composed of 2 - 3 frames in the prior art refers to m ∈ [-1, 1], which is limited to [X (-1) , X (0) , [X (0) , X (1) or [X (-1) , X (0) , X (1)The existing methods do not explore the broader context space for the space constituted by
[0047] Step 1.1.3: Calculate the affinity matrix of the features of each frame in the context space with the remaining frames in the context space.
[0048] In this embodiment, all frames X (m) are split into multiple heads in the channel dimension, and the splitting formula is as follows:
[0049]
[0050] where the Split(.) operation represents splitting X in the channel dimension (m) to focus on different subspaces, head_num represents the total number of heads split into, represents the k-th head, and C hid = head_num × D.
[0051] Then, the subspace representation of the current frame of the k-th head at the (i, j) position is represented as and its affinity is calculated as:
[0052]
[0053] In the formula, A represents the affinity matrix, whose shape is T×Context×Head_num×H×W×H×W, where Context is the context dimension (excluding the current frame here), Head_num represents the number of heads, (i′, j′) represents the query position. k represents the k-th head, d represents the d-th channel, and D represents the dimension of each head, i.e., the number of channels. represents all the values of the current frame of the d-th channel of the k-th head at the (i, j) position, represents the distance of all the values of the m-th frame of the current frame of the d-th channel of the k-th head at the (i ′ , j ′ ) position, represents the value at the (i, j, i ′ , j ′ ) position on the k-th head of the affinity matrix for all the m-th frames of the current frame.
[0054] Step 1.1.4: Perform normalization processing on the affinity matrix.
[0055] In this embodiment, the sigmoid function is used to limit the numerical range of the affinity matrix to (0, 1), and 0.5 is subtracted to suppress redundant connections. To avoid excessive differences in affinity, the sigmoid function is applied to map them to the range (0, 1) and subtract 0.5 to suppress redundant associations:
[0056]
[0057] Step 1.1.5: Aggregate spatio-temporal information using the affinity after normalization as weights, and then use convolution to restore the channel dimension to obtain context correlation features.
[0058] Using the affinity matrix as weights, then aggregate spatio-temporal information, the calculation is as follows:
[0059]
[0060] where E k (i, j) represents the enhanced feature at the (i, j) position of the k-th head, and the shape of E k is D×T×H×W. represents the value at the (i, j, i ′ , j ′ ) position on the k-th head at m frames from all current frames in the affinity matrix, as weights. represents the values of the k-th head at the (i ′ , j ′ ) position at all current frames m frames away.
[0061] Finally, all heads are concatenated along the channel dimension, and at the same time, convolution is used to restore the channels to C in to obtain the context correlation feature X CCA :
[0062] X CCA = Conv 1×1×1 (Concat[E 1 , E 2 ,..., E n )
[0063] where, E k represents the enhanced feature obtained in the previous step, Concat(.) represents the concatenation operation which mainly concatenates all heads, and Conv 1×1×1 is a 3D convolution with a convolution kernel of 1×1×1 which is mainly used to restore the channel dimension.
[0064] Among them, motion features are extracted through the context difference perception module, including:
[0065] Step 1.2.1: The input features of the context difference perception module are processed by convolution for channel downsampling.
[0066] Input features of the context difference perception module Processed by convolution to downsample the channels to C hid Obtain Thereby reducing the computational burden in subsequent operations.
[0067] Step 1.2.2: Use a sliding window along the time dimension for the features obtained by channel downsampling to construct a context space.
[0068] For feature X hid Use a sliding window along the time dimension to construct a context space:[[]]
[0069]
[0070] Where Represents a sliding window operation with a window size of 2n + 1. Where m ∈ [-n, n] represents a bias of m frames.
[0071] Step 1.2.3: Subtract each frame in the context space from frames with different time spans within the context to obtain motion features at different speeds.
[0072] Define the subspace at the current frame position (i, j) as X (0) (i, j), calculate the inter-frame difference to obtain motion information:[[]]
[0073] X′ (m) (i, j) = X (0) (i, j) - X (m) (i, j)
[0074] If there is no motion change at position (i, j), the result X′ (m) (i, j) tends to 0, effectively suppressing the interference caused by static redundancy (such as the background environment, relatively static torso parts) to model recognition. In addition, larger |m| values are better at capturing slower motion events, while smaller |m| values are more effective in capturing fast motion events.
[0075] Step 1.2.4: Use convolution to process the features at different motion speeds in sequence, adaptively enhance the regions that need to be concerned, and obtain enhanced features.
[0076] Use convolution to calculate the extracted motion features and adaptively enhance the regions that need to be concerned:[[]]
[0077] X″ (m) = Conv_Block(X′ (m) )
[0078] Among them, Conv_Block(·) mainly consists of three convolutions with a kernel size of 1×3×3.
[0079] Step 1.2.5: Concatenate the enhanced motion features of each frame in the context space along the channel dimension, then adaptively aggregate the motion features at different speeds using convolution, and restore the channel dimension to obtain the motion features.
[0080] Concatenate along the channel dimension {X″ (-n) ,..., X″ (n)}, and use CNN to restore the channels to C in to obtain the motion feature X CVA :
[0081] X CVA = Conv 1×1×1 (Concat[X″ (-n) ,..., X″ (n) )
[0082] Among them, the context correlation feature and the motion feature are concatenated and then residual connection is performed to obtain the finally extracted feature at the current stage.
[0083] Specifically, multiply the context correlation feature X CCA and the motion feature X CVA by learnable hyperparameters respectively and add them to input to the next stage of the feature extraction module.
[0084] X out = α * X CCA + β * X CVA + X in
[0085] Among them, α and β respectively represent two learnable parameters, which are used to control the contributions of the CCA and CVA output features. The initial values of α and β are 0 to maintain the original features of the feature extraction module.
[0086] In this embodiment, the context correlation perception module enhances the model's utilization of context correlation, while the context difference perception module improves the model's perception ability of motions at different speeds.
[0087] Step S2: Input the features output by the encoder into the decoder, and then obtain the final recognition result through the classifier.
[0088] The features finally output by the encoder are successively passed through a one-dimensional convolutional neural network (1D CNN) and a bidirectional long short-term memory neural network (BiLSTM), and then the final recognition result is obtained through the classifier.
[0089] Another embodiment of the present application, when training the recognition network model, the total loss function used is as follows:
[0090]
[0091] where γ1 and γ2 are two hyperparameters used to balance the loss, for example, they can be set to 1 and 25 respectively. represents the CTC loss, represents the self-distillation loss. The CTC loss is the most widely used loss in the CSLR task. It uses dynamic programming to solve the alignment problem between the model output and the label, thereby realizing end-to-end training.
[0092] The self-distillation loss regards the output of the entire recognition network model as the teacher and the output of the one-dimensional convolutional neural network (1D CNN) as the student, and is calculated using the KL divergence. Among them, the calculation of the loss using the KL divergence is a relatively mature technology in this field and will not be elaborated here.
[0093] To verify the technical effects of the present application, performance comparisons are made with the current state-of-the-art methods on the datasets PHOENIX14 and PHOENIX14-T. Table 1 shows the experimental results.
[0094] Table 1
[0095]
[0096]
[0097] Table 1 shows the performance comparison of the method of this application with other methods on the PHOENIX14 and PHOENIX14-T datasets. Among them, * indicates the case of using an additional network or pre-extracted heatmaps to extract other clues (such as facial or hand features). Del / ins represent the deletion and insertion error rates respectively, and the accuracy of the prediction results is expressed as the word error rate (WER). WER is used to measure the difference between the predicted text and the standard text. The smaller the value, the lower the error rate and the better the performance. Dev is the validation set, and Test is the test set. The existing technologies for comparison include: SFL (Stochastic Fine-grained Labeling), VAC (Visual Alignment Constraint), SMKD (Self-Mutual distillation learning), TLP (Temporal Lift Pooling), SEN (Self-Emphasizing Network), AdaBrowse+ (AdaptiveVideo Browser), CorrNet (Correlation Network), CoSign (Co-occurrence Signals), SignGraph (ASign Sequence is Worth Graphs of Nodes), TCNet (Trajectories andCorrelated Regions), DNF (Deep Neural Frame Work), STMC (Spatial Temporal Multi-Cue), C 2 SLR (Consistency-Enhanced Continuous), TwoStream (Two-Stream Network).
[0098] The technical solution of this application only relies on a single video input, without the need for expensive supervision information, and its performance exceeds all existing methods. Compared with the best single-cue model TCNet, the word error rate (WER) of this application on the PHOENIX14 and PHOENIX14-T test sets is reduced by 0.9% and 0.8% respectively. In addition, compared with the multi-cue method Two-Stream that utilizes additional key-point information, this application still achieves a 0.8% and 0.7% reduction in WER on the test sets of the two datasets. This improvement benefits from the unique attention of this application to extended context information, setting a new state-of-the-art recognition accuracy. In summary, this application demonstrates significant performance advantages in the field of sign language recognition, proving that it can still achieve efficient and accurate recognition without complex supervision information, providing a new direction for future research and applications.
[0099] This application also conducts a performance comparison with existing technical methods on the CSL-Daily dataset, and the comparison results are shown in Table 2.
[0100] Table 2
[0101] Method Dev(%) Test(%) LS-HAN 39.0 39.4 SEN 31.1 30.7 AdaBrowse+ 31.2 30.7 CorrNet 30.6 30.1 CoSign 28.1 27.2 TCNet 29.7 29.3 SignGraph 27.3 26.4 TwoStream-SLR* 25.4 25.3 The method of this application 25.7 24.7
[0102] Table 2 shows the performance comparison between the method of this application and other methods on the CSL-Daily dataset. Dev is the validation set, Test is the test set, and the existing technical methods include: LS-HAN (Video-based sign language recognition without temporal segmentation), SEN (Self-emphasizing network), AdaBrowse+ (Adaptive Video Browser), CorrNet (correlation network), CoSign (Co-occurrence Signals), TCNet (Trajectories and Correlated Regions), SignGraph (A Sign Sequence is Worth Graphs of Nodes), TwoStream (Two-Stream Network).
[0103] The method of this application demonstrates state-of-the-art performance on the CSL-Daily dataset. Compared with the best single-stream model, SignGraph, which only focuses on adjacent frames and has difficulty capturing the complete context, the method of this application solves this problem and reduces the WER by 1.6% and 1.7% on the development set and the test set, respectively. In addition, compared with the best multi-cue method, Two-Stream, the WER of the method of this application is still reduced by 0.6% on the test set. These results illustrate the strong generalization ability of this application on the largest Chinese Sign Language dataset.
[0104] The above-described embodiments merely represent several implementation manners of this application, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all fall within the protection scope of this application. Therefore, the protection scope of this application patent shall be subject to the appended claims.
Claims
1. A cross-frame interaction continuous sign language recognition method based on an extended context space, which uses a trained recognition network model for sign language recognition. The recognition network model includes an encoder, a decoder, and a classifier, and is characterized in that The cross-frame interaction continuous sign language recognition method based on an extended context space includes: Input the sign language video to be recognized into the encoder of the recognition network model. For the features at each stage extracted by the encoder, the context correlation features and motion features are respectively extracted through the context correlation perception module and the context difference perception module, and then the context correlation features and motion features are concatenated and then residual connection is performed to obtain the finally extracted features at the current stage, which are input to the next stage; Input the features output by the encoder into the decoder, and then obtain the final recognition result through the classifier.
2. The cross-frame interaction continuous sign language recognition method based on an extended context space according to claim 1, wherein The context correlation perception module extracts context correlation features, including: The input features of the context correlation perception module are processed by convolution for channel downsampling; The features obtained by channel downsampling are used to construct a context space along the time dimension by using a sliding window; Calculate the affinity matrix of each frame in the context space with the features of the remaining frames in the context space; Perform normalization processing on the affinity matrix; Use the normalized affinity as the weight to aggregate spatio-temporal information, and then use convolution to restore the channel dimension to obtain the context correlation features.
3. The cross-frame interaction continuous sign language recognition method based on an extended context space according to claim 1, wherein The context difference perception module extracts motion features, including: The input features of the context difference perception module are processed by convolution for channel downsampling; The features obtained by channel downsampling are used to construct a context space along the time dimension by using a sliding window; Subtract each frame in the context space from the frames with different time spans in the context to obtain motion features with different speeds; Process the features with different motion speeds sequentially by convolution, adaptively enhance the regions that need to be concerned, and obtain the enhanced features; Concatenate the enhanced motion features of each frame in the context space along the channel dimension, and then use convolution to adaptively aggregate the motion features with different speeds and restore the channel dimension to obtain the motion features.
4. The cross-frame interaction continuous sign language recognition method based on an extended context space according to claim 1, wherein When training the recognition network model, the total loss function includes the CTC loss and the self-distillation loss.
5. The cross-frame interaction continuous sign language recognition method based on an extended context space according to claim 4, characterized in that, The decoder includes a one-dimensional convolutional neural network and a bidirectional long short-term memory neural network.
6. The cross-frame interaction continuous sign language recognition method based on an extended context space according to claim 5, wherein The self-distillation loss is calculated using the KL divergence, taking the output of the entire recognition network model as the teacher and the output of the one-dimensional convolutional neural network as the student.