Continuous sign language recognition method based on direction-sensitive long-time motion decoupling
By decoupling sign language motion features into horizontal and vertical components and purifying and coupling, combining one-dimensional convolutional neural networks and bidirectional long and short-term memory neural networks, the problem of difficult to capture medium- and long-term and multi-directional motion features is solved, and the recognition accuracy and robustness are improved.
Patent Information
- Application Number
- CN202510384530.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
The existing continuous sign language recognition technology fails to effectively capture long-term and multi-directional motion characteristics in sign language, resulting in insufficient recognition accuracy and robustness.
Using a method based on direction-sensitive long-term motion decoupling, the motion features are decoupled into horizontal and vertical components through the feature extraction module, motion purification and stage coupling are performed, and a one-dimensional convolutional neural network and a two-way long and short-term memory neural network are used for identification.
It improves the accuracy of continuous sign language recognition, reduces word error rate, enhances the model's cross-scene generalization ability, and can effectively deal with the speech speed differences and morphological changes of different sign language users.
Smart Images

Figure CN120340126A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of sign language recognition technology, and specifically relates to a continuous sign language recognition method based on direction-sensitive long-term motion decoupling. Background Art
[0002] Continuous sign language has a unique grammatical structure and vocabulary, which makes it quite challenging for hearing people without relevant sign language knowledge to learn. In recent years, many studies have been dedicated to continuous sign language recognition (CSLR), aiming to facilitate smooth communication between deaf and hearing people, thereby significantly narrowing the gap between them.
[0003] Recent CSLR studies have mainly focused on using the inherent motion characteristics of sign language to improve recognition accuracy. For example, the Self-Emphasizing Network (SEN) enhances the discriminative ability of key frames by leveraging the differences caused by motion; the Correlation Network (CorrNet) calculates the correlation between adjacent frames to match the motion displacement area; the Trajectory and Correlation Region Network (TCNet) captures motion trajectories through an optical flow-inspired module.
[0004] However, previous studies have overlooked the long-term and multi-directional characteristics of motion in sign language, which has a great impact on the accuracy and robustness of recognition. Summary of the Invention
[0005] The purpose of this application is to provide a continuous sign language recognition method based on direction-sensitive long-term motion decoupling, to solve the problem that it is difficult to capture the long-term and multi-directional motion features in sign language, and to improve the efficiency of continuous sign language recognition.
[0006] To achieve the above purpose, the technical solution of this application is as follows:
[0007] A continuous sign language recognition method based on direction-sensitive long-term motion decoupling uses a trained recognition network model for sign language recognition. The recognition network model includes a feature extraction module. The continuous sign language recognition method based on direction-sensitive long-term motion decoupling includes:
[0008] Sequentially select one frame from the current stage features extracted by the feature extraction module as the current frame, extract the motion features of each frame in the context space of the current frame, and then aggregate to obtain an aggregated feature;
[0009] Decouple the aggregated feature to obtain a horizontal motion feature component and a vertical motion feature component;
[0010] Perform motion purification on the decoupled horizontal motion feature component and vertical motion feature component respectively to obtain the purified horizontal motion feature component and vertical motion feature component;
[0011] Perform stage coupling on the purified horizontal motion feature component and vertical motion feature component, then add them to the current stage feature, obtain the updated current stage feature, and enter the next stage of the feature extraction module to extract the next stage feature;
[0012] Pass the feature finally output by the feature extraction module through a one-dimensional convolutional neural network to obtain the first feature;
[0013] Perform cross-stage coupling on the purified horizontal motion feature component and vertical motion feature component in each stage of the feature extraction module, and then pass them through a one-dimensional convolutional neural network to obtain the second feature;
[0014] Add the first feature and the second feature and input them into a bidirectional long short-term memory neural network, and then obtain the final recognition result through a classifier.
[0015] Further, successively select one frame from the current stage features extracted by the feature extraction module as the current frame, extract the motion features of each frame within the context space of the current frame, and then aggregate them to obtain the aggregated feature, including:
[0016] Successively select one frame from the current stage features extracted by the feature extraction module as the current frame, and establish the motion context space of the current frame;
[0017] Calculate the difference between the current frame and the remaining frames within the context space to obtain the motion region relative to the current frame;
[0018] Process the motion region of each frame sequentially through a convolutional block to obtain enhanced motion features;
[0019] Concatenate the enhanced motion features of each frame in the context space along the channel dimension, then use convolution to adaptively aggregate long-term motion information, and apply residual connection to retain the original information to obtain the aggregated feature.
[0020] Further, perform motion purification on the decoupled horizontal motion feature component and vertical motion feature component respectively to obtain the purified horizontal motion feature component and vertical motion feature component, including:
[0021] Add the decoupled horizontal motion feature component and vertical motion feature component of the current stage to the purified horizontal motion feature component and vertical motion feature component of the previous stage after downsampling respectively to obtain the horizontal motion feature component and vertical motion feature component to be purified in the current stage;
[0022] Perform convolution operations and residual connections on the horizontal motion feature component and vertical motion feature component to be purified in the current stage respectively to obtain the intermediate features of the horizontal motion feature component and vertical motion feature component;
[0023] Nonlinear transformations are respectively performed on the intermediate features of the horizontal motion feature component and the vertical motion feature component to obtain the purified horizontal motion feature component and vertical motion feature component.
[0024] Furthermore, the stage coupling of the purified horizontal motion feature component and the vertical motion feature component includes:
[0025] Upsampling is respectively performed on the purified horizontal motion feature component and the vertical motion feature component, and then weighted summation is carried out.
[0026] Furthermore, the cross-stage coupling of the purified horizontal motion feature component and the vertical motion feature component in each stage of the feature extraction module includes:
[0027] Global average pooling is performed on the purified horizontal motion feature component and the vertical motion feature component in each stage, and then the purified horizontal motion feature component and the vertical motion feature component after global average pooling are added together. After passing through a fully connected layer, the purified motion features of each stage are obtained;
[0028] The purified motion features of each stage are added together to obtain the cross-stage coupling feature.
[0029] Furthermore, the total loss function adopted by the recognition network model includes the CTC loss and the self-distillation loss.
[0030] A continuous sign language recognition method based on direction-sensitive long-term motion decoupling proposed in this application. In contrast to traditional models that usually capture all motions as a whole, which can lead to feature confusion and reduce accuracy. In contrast, the method of this application decouples motion features into their horizontal and vertical components, enhancing the model's ability to recognize complex patterns. It achieves a breakthrough improvement in the recognition accuracy of large vocabulary continuous sign language. Verified by experiments, the word error rate (WER) on the standard test set is absolutely reduced by 1.6% compared with the existing optimal solution; it demonstrates enhanced environmental adaptability and maintains stable optimal recognition performance in both laboratory-controlled environments and complex real-world scenarios; it has enhanced cross-scenario generalization ability and can effectively handle the speech rate differences and morphological changes of different sign language users. Its technical advantages are confirmed through multi-scenario cross-validation. Description of the Drawings
[0031] Figure 1 It is a flowchart of the continuous sign language recognition method based on direction-sensitive long-term motion decoupling of this application.
[0032] Figure 2 It is a schematic diagram of the structure of the recognition network model of this application.
[0033] Figure 3 It is a schematic diagram of the structure of the feature extraction module in the embodiment of this application. Detailed Embodiments
[0034] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0035] The purpose of continuous sign language recognition is to convert a video with T frames into sign language annotations Where N represents the length of the gloss sequence, and the gloss sequence refers to the lexical-level representation of each sign language action in the sign language video. Traditional methods use a two-dimensional convolutional neural network (2D-CNN) to extract frame-by-frame features Then, these features are processed by a one-dimensional convolutional neural network (1D-CNN) and a bidirectional long short-term memory neural network (BiLSTM) for local and global temporal modeling. Finally, the output is classified and optimized through the CTC loss.
[0036] This application improves the traditional method. When extracting features, it uses a method of decoupling and then coupling. After each stage of the feature extraction module (ResNet34), the features first aggregate long-term motion information, and then decouple it into horizontal and vertical feature components, enhance the directional motion perception through motion purification, enhance the features through stage coupling, and finally enrich the multi-scale information through cross-stage coupling.
[0037] An embodiment of this application, as Figure 1 shown, proposes a continuous sign language recognition method based on direction-sensitive long-term motion decoupling, and uses a trained recognition network model for sign language recognition. The recognition network model includes a feature extraction module. The continuous sign language recognition method based on direction-sensitive long-term motion decoupling includes:
[0038] Step S1: Sequentially select one frame from the current stage features extracted from the feature extraction module as the current frame, extract the motion features of each frame in the context space of the current frame, and then aggregate to obtain the aggregated features.
[0039] To recognize continuous sign language, this application also constructs and trains a recognition network model, as Figure 2 shown, which also includes a feature extraction module, a one-dimensional convolutional neural network, a bidirectional long short-term memory neural network, and a classifier.
[0040] When training the recognition network model, the sign language videos are from the publicly available large German sign language dataset PHOENIX14, PHOENIX14-T, and the Chinese dataset CSL-Daily. The video frames in the dataset are processed through data augmentation and size adjustment to obtain training samples, and the recognition network model is trained.
[0041] After training the recognition network model, for the sign language video to be recognized, each frame of the sign language video is uniformly cropped to the same size and input into the recognition network model to obtain the recognition result.
[0042] This application mainly improves the feature extraction module. The feature extraction module uses a two-dimensional convolutional neural network, such as lightweight convolutional neural networks like ResNet34, ResNet18, ResNet50, SqueezeNet, etc.
[0043] In the feature extraction module, features are usually extracted in several stages. For example, the residual block layer of ResNet34 includes four stages, and this application performs the same operation on the current stage features extracted in each stage.
[0044] This step is to perform motion aggregation on the current stage features extracted. From the current stage features extracted from the feature extraction module, a frame is sequentially selected as the current frame, and the motion features of each frame within the context space of the current frame are extracted, and then the aggregated features are obtained, including:
[0045] Step 1.1: Sequentially select a frame from the current stage features extracted from the feature extraction module as the current frame, and establish the motion context space of the current frame.
[0046] This embodiment uses to represent the current stage feature sequence of T frames, where the feature of each frame in the current stage is represented as To reduce the computational cost, the channel dimension can be compressed to C / r through convolution. After that, a motion context space is established for each frame through the following operations:
[0047]
[0048] In the formula represents the sliding window operation with a window size of 2n + 1, represents the feature sequence with a motion context of 2n + 1.
[0049] Unfold X (2n+1) along the motion context to be where represents the i-th frame offset from the current frame, is the current frame.
[0050] Step 1.2: Calculate the difference between the current frame and the remaining frames within the context space to obtain the motion area relative to the current frame.
[0051] This embodiment calculates the motion area of the i-th offset frame relative to the current frame, and the formula is as follows:
[0052]
[0053] Among them, (h, w) represents the spatial position. If the feature at this position remains unchanged (such as the background area), the result tends to 0.
[0054] This application focuses on processing key regions without being affected by redundancy. Redundant information is mainly manifested as the interference of static regions (such as the background environment, relatively static torso parts) on the model recognition process. These regions not only cannot provide effective discriminant features for the model, but may also reduce the recognition performance due to factors such as high-brightness backgrounds. Therefore, the model needs to focus on the hand regions with significant motion features and facial expression changes to achieve accurate sign language recognition. Frame difference calculation can effectively suppress the relatively static regions between frames and retain the motion regions, thereby reducing the interference of redundant information on recognition.
[0055] Step 1.3: Sequentially process the motion regions of each frame through a convolutional block to obtain enhanced motion features.
[0056] To aggregate discrete spatial motion regions and enhance the discriminant ability of features, a shared convolutional block is used to sequentially process each frame:
[0057]
[0058] Among them, ConvBlock(·) mainly consists of 3 convolutions, and the size of each convolution kernel is 1×3×3. Represents the feature of a single position (h, w), This step of the convolutional block processes multiple spatial positions.
[0059] It should be noted that this step enhances the motion features of each frame. It is also possible not to include this step but directly proceed to the next step of processing.
[0060] Step 1.4: Concatenate the enhanced motion features of each frame in the context space along the channel dimension, then use convolution to adaptively aggregate long-term motion information, and apply residual connection to retain the original information to obtain aggregated features.
[0061] This step concatenates along the channel dimension, then uses convolution to adaptively aggregate long-term motion information, and at the same time reduces the channels to C / r:
[0062]
[0063] Among them,
[0064] Then, use convolution to restore the channels, restore the channels back to C, obtain the features after aggregating long-term motion, and then apply residual connection to retain the original information to obtain the aggregated features:
[0065] X agg = Conv 1×1×1 (X′) + X.
[0066] Among them, X agg represents the finally obtained aggregated features.
[0067] Step S2: Decouple the aggregated features to obtain the horizontal motion feature component and the vertical motion feature component.
[0068] The aggregated contains multi-directional motion. To simplify these complex motions, the motion features can be projected into a low-dimensional space. Specifically, to extract the vertical motion, all horizontal motion components are excluded, and vice versa.
[0069] This step can be expressed as:
[0070]
[0071] Among them, respectively represent the decoupled horizontal component and vertical component. represents the decoupling operation, which can be obtained by, for example, using the max pooling operation (Maxpool) or the average pooling operation (Avgpool).
[0072] Step S3: Perform motion purification on the decoupled horizontal motion feature component and vertical motion feature component respectively to obtain the purified horizontal motion feature component and vertical motion feature component.
[0073] In this step, motion purification is performed on the decoupled horizontal motion feature component and vertical motion feature component respectively.
[0074] In a specific embodiment, performing motion purification on the decoupled horizontal motion feature component and vertical motion feature component respectively to obtain the purified horizontal motion feature component and vertical motion feature component includes:
[0075] Step 3.1: Add the decoupled horizontal motion feature component and vertical motion feature component at the current stage to the purified horizontal motion feature component and vertical motion feature component at the previous stage after downsampling respectively to obtain the horizontal motion feature component and vertical motion feature component to be purified at the current stage.
[0076] Specifically, as Figure 3 shown, add the purified horizontal motion feature component and vertical motion feature component X at the previous stage ′h ′ / v Perform downsampling, add it to the features decoupled at the current stage, gradually enrich the multi-scale information, and increase the network depth of OMP:
[0077]
[0078] Where X h / v represents the horizontal motion feature component and the vertical motion feature component decoupled at the current stage, represents downsampling.
[0079] Step 3.2: Perform convolution operations and residual connections on the horizontal motion feature component and the vertical motion feature to be purified at the current stage, respectively, to obtain intermediate features of the horizontal motion feature component and the vertical motion feature component.
[0080] As Figure 3 shown, the motion purification OMP in this embodiment includes horizontal perception motion purification (HMP) and vertical perception motion purification (VMP) for purifying the horizontal motion feature component and the vertical motion feature component, respectively. In HMP and VMP, the decoupled feature X h , X v is regarded as a special two-dimensional image. MPBlock (two-dimensional convolution) is used to extract the motion information in a specific direction, and then a residual connection is performed:
[0081] X′ h / v = MPBlock(X h / v ) + X h / v ,
[0082] where, in the formula represents the intermediate feature of the horizontal motion feature component or the vertical motion feature component.
[0083] The MPBlock (two-dimensional convolution) in this embodiment includes a convolution layer (Conv), a normalization layer (BatchNormalization), an activation function (Relu), and a convolution layer (Conv). These layers are all relatively mature technologies in the art and will not be elaborated here. MPBlock is mainly composed of two-dimensional convolution, which can play a role in the operation of cross-frame information, and realizes more effective information purification and expression by extracting and strengthening the cross-frame motion information in the horizontal or vertical direction.
[0084] Step 3.3: Perform non-linear transformation on the intermediate features of the horizontal motion feature component and the vertical motion feature component, respectively, to obtain the purified horizontal motion feature component and the vertical motion feature component.
[0085] Nonlinear transformations are respectively performed on the intermediate features of the horizontal motion feature component and the vertical motion feature component to obtain a more discriminative feature representation.
[0086] Specifically, a feedforward neural network FFN is used to perform a nonlinear transformation on the channel to enhance the modeling ability of OMP:
[0087] X″ h / v = max(0, Norm(X′ h / v W1))W2,
[0088] where represents two-dimensional convolutions with a kernel size of 1, Norm(·) represents layer normalization, and X ′ h ′ / v represents the purified horizontal motion feature component X″ h or the vertical motion feature component X″ v .
[0089] Step S4: Perform stage coupling on the purified horizontal motion feature component and the vertical motion feature component, and then add them to the current stage feature to obtain an updated current stage feature, and enter the next stage of the feature extraction module to extract the next stage feature.
[0090] Specifically, perform an upsampling operation on X″ h , X″ v , and then perform weighted summation, that is, perform coupling using learnable parameters, and update the current stage feature through a residual connection in the current stage of the feature extraction module to obtain an updated current stage feature X″:
[0091]
[0092] where US(·) represents the upsampling operation, and α and β are learnable parameters used to control the contributions of motion features in different directions. They are initialized to zero and are adaptively adjusted by the model during training.
[0093] Step S5: Pass the feature finally output by the feature extraction module through a one-dimensional convolutional neural network to obtain a first feature.
[0094] In the feature extraction module, features are sequentially extracted through each stage, and finally the feature finally output by the feature extraction module is obtained. Then, through one-dimensional convolution (1D-CNN), a first feature is obtained.
[0095] In this embodiment, the one-dimensional convolutional neural network (1D-CNN) is not modified, which belongs to a relatively mature technology in the art and will not be elaborated here.
[0096] Step S6: The purified horizontal motion feature components and vertical motion feature components at each stage of the feature extraction module are coupled across stages, and then passed through a one-dimensional convolutional neural network to obtain the second feature.
[0097] Cross-stage coupling utilizes multi-scale information by coupling features of different depths, and its input comes from the purified horizontal motion feature components and vertical motion feature components at different stages. First, global average pooling is performed on the purified horizontal motion feature components and vertical motion feature components at each stage, then the global average pooled horizontal motion feature components and vertical motion feature components are added together, and after passing through a fully connected layer, the purified motion features at each stage are obtained; then the purified motion features at each stage are added together to obtain the cross-stage coupling feature.
[0098] It is expressed by the formula as follows:
[0099]
[0100] where n represents the number of stages of the feature extraction module, and are the purified horizontal motion feature component and vertical motion feature component at the i-th stage, represents the global average pooling operation in the height or width dimension. W i represents the i-th fully connected layer. The fully connected layer maps the channels to the same dimension, facilitating subsequent feature coupling, aligning the differential information between different scales, and thus enhancing the model's representation ability and the ability to model complex patterns. represents the feature obtained through cross-stage coupling. In this step, features of different depths are coupled to obtain a cross-stage coupling feature containing multi-scale information.
[0101] Then, the cross-stage coupling feature passes through one-dimensional convolution (1D-CNN) to obtain the second feature.
[0102] Step S7: The first feature and the second feature are added together and input into a bidirectional long short-term memory neural network, and then the final recognition result is obtained through a classifier.
[0103] In this implementation, after adding the first feature and the second feature, they are input into a bidirectional long short-term memory neural network, and then the final recognition result is obtained through a classifier, which will not be elaborated here.
[0104] In another embodiment of the present application, when training the recognition network model, the total loss function used is as follows:
[0105]
[0106] where γ1 and γ2 are two hyperparameters used to balance the loss, for example, they can be set to 1 and 25 respectively. Represents the CTC loss, represents the self-distillation loss.
[0107] The CTC loss is the most widely used loss in the CSLR task. It uses dynamic programming to solve the alignment problem between the output and the label, thus enabling end-to-end training. The CTC loss in this embodiment is expressed as:
[0108]
[0109] where, represents the supervision of the outputs of two 1D-CNNs, represents the supervision of the output of the BiLSTM. represents calculating the CTC loss (optimizing the sum of probabilities of all feasible paths) using the classification result of the first feature supervised by the label, represents calculating the CTC loss using the classification result of the second feature supervised by the label, represents calculating the CTC loss using the classification result of the entire network supervised by the label.
[0110] For the self-distillation loss, the output of the entire network is regarded as the teacher, and the outputs of the two 1D-CNNs are regarded as the students. The loss is calculated using the KL divergence:
[0111]
[0112] represents calculating the KL divergence (Kullback-Leibler Divergence) using the classification result of the first feature supervised by the output features of the entire network regarded as the teacher, represents calculating the KL divergence using the classification result of the second feature supervised by the output features of the entire network regarded as the teacher.
[0113] The experimental data of this application show that the method of this application has better recognition accuracy compared with other methods of the prior art. The experimental data are shown in Table 1 and Table 2 below:
[0114] Table 1
[0115]
[0116] Table 1 shows the performance comparison of the method of this application with other methods on the PHOENIX14 and PHOENIX14-T datasets. Among them, * indicates the situation of using an additional network or pre-extracted heatmaps to extract other clues (such as facial or hand features). Del / ins represent the deletion and insertion error rates respectively, and the accuracy of the prediction results is expressed by the word error rate (WER). WER is used to measure the difference between the predicted text and the standard text. The smaller the value, the lower the error rate and the better the performance. Dev is the validation set, Test is the test set, and the full English names of other methods in the prior art are as follows: SFL is Stochastic Fine-grained Labeling, VAC is Visual Alignment Constraint, SMKD is Self-Mutual distillation learning, TLP is Temporal Lift Pooling, SEN is Self-Emphasizing Network, AdaBrowse+ is Adaptive Video Browser, CorrNet is Correlation Network, CoSign is Co-occurrence Signals, SignGraph is A Sign Sequence is Worth Graphs of Nodes, TCNet is Trajectories and Correlated Regions, DNF is Deep Neural Frame Work, STMC is Spatial Temporal Multi-Cue, C 2 SLR is Consistency-Enhanced Continuous, TwoStream is Two-Stream Network
[0117] The results in Table 1 show the effectiveness of the method of this application in capturing and utilizing complex actions in sign language.
[0118] Table 2
[0119] Method Dev(%) Test(%) LS-HAN 39.0 39.4 SEN 31.1 30.7 AdaBrowse+ 31.2 30.7 CorrNet 30.6 30.1 CoSign 28.1 27.2 TCNet 29.7 29.3 SignGraph 27.3 26.4 TwoStream-SLR* 25.4 25.3 The method of this application 25.8 24.7
[0120] Table 2 shows the performance comparison with prior art methods on CSL-Daily. Dev is the validation set and Test is the test set. The full English names of other prior art methods are as follows: LS-HAN is Video-based sign language recognition without temporal segmentation, SEN is Self-emphasizing network, AdaBrowse+ is Adaptive Video Browser, CorrNet is correlation network, CoSign is Co-occurrence Signals, TCNet is Trajectories and Correlated Regions, SignGraph is A Sign Sequence is Worth Graphs of Nodes, and TwoStream is Two-Stream Network.
[0121] Table 2 shows the superior performance of the method of the present application. Compared with the best single-cue method SignGraph before, the WER of the method of the present application is reduced by 1.5% (Dev) and 1.7% (Test). In addition, we even show that the WER drops by 0.6% in the test compared with the best multi-cue method TwoStream. The above results demonstrate the advancement of the method of the present application and its adaptability to large-scale Chinese datasets.
[0122] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A continuous sign language recognition method based on direction-sensitive long-term motion decoupling, which uses a trained recognition network model for sign language recognition. The recognition network model includes a feature extraction module, and is characterized in that The continuous sign language recognition method based on direction-sensitive long-term motion decoupling includes: Sequentially select one frame from the current-stage features extracted by the feature extraction module as the current frame, extract the motion features of each frame within the context space of the current frame, and then aggregate them to obtain an aggregated feature; Decouple the aggregated feature to obtain a horizontal motion feature component and a vertical motion feature component; Respectively perform motion purification on the decoupled horizontal motion feature component and vertical motion feature component to obtain the purified horizontal motion feature component and vertical motion feature component; Perform stage coupling on the purified horizontal motion feature component and vertical motion feature component, and then add them to the current-stage features to obtain the updated current-stage features, and enter the next stage of the feature extraction module to extract the next-stage features; Pass the features finally output by the feature extraction module through a one-dimensional convolutional neural network to obtain the first feature; Perform cross-stage coupling on the purified horizontal motion feature component and vertical motion feature component of each stage of the feature extraction module, and then pass them through a one-dimensional convolutional neural network to obtain the second feature; Add the first feature and the second feature and input them into a bidirectional long short-term memory neural network, and then obtain the final recognition result through a classifier.
2. The continuous sign language recognition method based on direction-sensitive long-term motion decoupling according to claim 1, characterized in that, The step of sequentially selecting one frame from the current-stage features extracted by the feature extraction module as the current frame, extracting the motion features of each frame within the context space of the current frame, and then aggregating them to obtain an aggregated feature includes: Sequentially select one frame from the current-stage features extracted by the feature extraction module as the current frame, and establish a motion context space for the current frame; Calculate the difference between the current frame and the remaining frames within the context space to obtain the motion region relative to the current frame; Sequentially process the motion region of each frame through a convolutional block to obtain enhanced motion features; Concatenate the enhanced motion features of each frame in the context space along the channel dimension, then adaptively aggregate long-term motion information using a convolution, and apply a residual connection to retain the original information to obtain an aggregated feature.
3. The continuous sign language recognition method based on direction-sensitive long-term motion decoupling according to claim 1, characterized in that The step of respectively performing motion purification on the decoupled horizontal motion feature component and vertical motion feature component to obtain the purified horizontal motion feature component and vertical motion feature component includes: Respectively add the decoupled horizontal motion feature component and vertical motion feature component of the current stage to the purified horizontal motion feature component and vertical motion feature component of the previous stage after downsampling to obtain the horizontal motion feature component and vertical motion feature component to be purified in the current stage; Respectively perform convolution operations and residual connections on the horizontal motion feature component and vertical motion feature component to be purified in the current stage to obtain intermediate features of the horizontal motion feature component and vertical motion feature component; Respectively perform non-linear transformations on the intermediate features of the horizontal motion feature component and vertical motion feature component to obtain the purified horizontal motion feature component and vertical motion feature component.
4. The continuous sign language recognition method based on direction-sensitive long-term motion decoupling according to claim 1, wherein The step of performing stage coupling on the purified horizontal motion feature component and vertical motion feature component includes: Respectively perform upsampling on the purified horizontal motion feature component and vertical motion feature component, and then perform weighted summation.
5. The continuous sign language recognition method based on direction-sensitive long-term motion decoupling according to claim 1, characterized in that, The purified horizontal motion feature components and vertical motion feature components at each stage of the feature extraction module are cross-stage coupled, including: Performing global average pooling on the purified horizontal motion feature components and vertical motion feature components at each stage, then adding the globally averaged horizontal motion feature components and vertical motion feature components, and passing through a fully connected layer to obtain the purified motion features at each stage; Adding the purified motion features at each stage to obtain cross-stage coupled features.
6. The continuous sign language recognition method based on direction-sensitive long-term motion decoupling according to claim 1, wherein The total loss function adopted by the recognition network model includes CTC loss and self-distillation loss.