Skeleton sign language recognition method of double-flow space-time dynamic graph convolutional network fused with residual learning
The dual-stream space-time dynamic graph convolution network separates hand posture and wrist motion trajectory, combined with the optimal transmission theory of residual learning and geometric driving, solves the problems of insufficient feature expression and low computational efficiency in skeleton sign language recognition, and achieves efficient and accurate sign language recognition.
Patent Information
- Application Number
- CN202510486459.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-08
AI Technical Summary
The existing skeleton sign language recognition methods lack comprehensive modeling of hand posture and motion trajectory, low computing efficiency, insufficient feature expression ability, difficult to meet the real-time recognition needs, and lack effective multimodal information fusion.
A dual-stream space-time dynamic graph convolution network with fusion residual learning is used to separate the hand posture and wrist motion trajectory data flow, and features are extracted through space-time dynamic graph convolution network and differential geometric modeling, and feature fusion is performed by combining residual learning and geometric-driven optimal transmission theory.
It significantly improves the accuracy and real-time performance of sign language recognition, improves computing efficiency, and enhances robustness and generalization capabilities in complex environments.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and gesture recognition technology, and in particular relates to a sign language recognition method and system based on deep learning, specifically a skeleton sign language recognition method of a dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning. Background Art
[0002] Sign language is the primary means of communication for the hearing-impaired. Automatic sign language recognition technology can effectively facilitate communication between the hearing-impaired and hearing individuals, and has important applications in areas such as accessible communication, educational assistance, and remote conferencing. With the advancement of skeletal point tracking technology, skeleton-based sign language recognition methods offer advantages over video-based methods, such as reduced data usage, improved privacy protection, and enhanced anti-interference capabilities.
[0003] Currently, mainstream approaches to skeleton sign language recognition are based on models such as convolutional neural networks (CNNs), long short-term memory networks (LSTMs), and graph convolutional networks (GCNs). While these methods have achieved some success in sign language recognition, they still suffer from the following shortcomings: First, they typically focus solely on hand posture features, ignoring the rich motion trajectory information in sign language; second, existing methods lack effective modeling of spatial structure and temporal evolution, making it difficult to capture complex spatiotemporal dependencies; third, traditional graph convolutional networks are computationally inefficient when processing time-series skeleton data, making them difficult to apply in real-time scenarios; fourth, their limited feature representation capabilities result in low recognition accuracy for similar gestures and rapid movements; and fourth, they rarely directly address the relationship between hand and facial posture.
[0004] In recent years, graph neural networks have demonstrated their advantages in processing non-Euclidean data. Dynamic graph convolutional neural networks (DGCNNs) are capable of effectively modeling point cloud data. However, when directly applied to skeletal sign language recognition, they struggle to capture temporal changes and details. Furthermore, traditional feature fusion methods, often based on simple concatenation or weighted summation, fail to effectively integrate complementary information from different modalities. Furthermore, insufficiently optimized network structures lead to low training efficiency and prone to overfitting.
[0005] Therefore, how to design a sign language recognition method that can simultaneously capture hand posture and motion trajectory information, effectively model spatiotemporal dependencies, be computationally efficient, and have strong feature expression capabilities is an urgent problem to be solved in current research. Summary of the Invention
[0006] The purpose of this invention is to propose a skeleton sign language recognition method based on a dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning, so as to solve the problems of insufficient feature expression ability, insufficient spatiotemporal modeling, and lack of effective fusion of multimodal information in existing sign language recognition methods, and improve recognition accuracy and real-time performance.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] A skeleton sign language recognition method based on a dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning comprises the following steps: S1, inputting gesture skeleton sequence data and dividing the data into two data streams: detailed hand posture and wrist motion trajectory; S2, extracting mainstream features from the hand posture data through the spatiotemporal dynamic graph convolutional network; S3, extracting auxiliary stream features from the wrist trajectory data through differential geometry and spatiotemporal causal convolution; S4, realizing dual-stream feature interaction so that the mainstream and auxiliary stream features enhance each other; S5, performing dual-stream feature adaptive fusion and integrating the two streams of information in combination with the residual learning mechanism; S6, outputting the sign language recognition result through the classifier.
[0009] Furthermore, in S1, the dimension of the input gesture skeleton sequence data relative to the facial posture is [batch_size, frames, hands, joints, coordinates], where joints represents the number of joint points and coordinates represents the three-dimensional coordinate value; the detailed hand posture data is obtained by subtracting the wrist point coordinates to obtain the relative position, and the wrist motion trajectory data directly selects the wrist point coordinates.
[0010] Furthermore, in S2, the spatiotemporal dynamic graph convolutional network includes: a. constructing a k-nearest neighbor dynamic graph to extract spatial features; b. applying a residual convolution module to process graph features; c. performing convolution in the time dimension to capture temporal changes; d. multi-scale feature aggregation; e. applying a self-attention mechanism to enhance key features.
[0011] Furthermore, the residual convolution module adopts a convolution structure with skip connections, including: a. Using ResidualBlockCov to implement the main path convolution; b. Using 1×1 convolution to implement identity mapping; c. Adding the main path output and the identity mapping output to form a residual connection; d. Applying batch normalization and ReLU activation function.
[0012] Furthermore, the time dimension convolution adopts grouped convolution and multi-scale structure to reduce the amount of computation while maintaining expressiveness.
[0013] Furthermore, in S3, auxiliary stream feature extraction includes: a. using FinslerMetric to implement trajectory metric learning based on Finsler geometry; b. implementing spatiotemporal causal convolution through SpatioTemporalCausalConv; c. using a bidirectional LSTM network to model long-term dependencies; d. weighting features according to Finsler energy.
[0014] Furthermore, the implementation process of FinslerMetric includes: a. encoding position and velocity direction information separately; b. constructing Finsler metric basis functions; c. implementing homogeneity constraints; d. calculating Finsler energy and normalizing it.
[0015] Furthermore, the implementation process of spatiotemporal causal convolution includes: a. processing the spatial convolution of the left and right wrists separately; b. applying multi-scale dilated convolution to capture dependencies in different time ranges; c. fusing multi-scale features to form the final output.
[0016] Furthermore, in S4, the dual-stream feature interaction is implemented through the CrossFeatureInteraction module, including: a. Projecting the mainstream and auxiliary stream features into the latent space of the same dimension; b. Using the cross-attention mechanism, the mainstream features guide the auxiliary stream features, and the auxiliary stream features guide the mainstream features; c. Feature fusion and residual connection; d. Layer normalization processing to enhance stability.
[0017] Furthermore, in S5, the adaptive fusion of dual-stream features is implemented through the AdaptiveFeatureFusion module, including: a. Geometry-driven optimal transport (Geo-OT) fusion; b. Attention fusion; c. Dynamic weight adjustment and residual connection; d. Feature calibration and dimension alignment.
[0018] Furthermore, the implementation process of geometry-driven optimal transmission fusion includes: a. constructing a geometry-aware cost matrix; b. solving the regularized optimal transmission problem through the Sinkhorn algorithm; c. using adaptive entropy regularization parameters; d. enhancing feature alignment through geometric consistency constraints.
[0019] Furthermore, the dynamic weight adjustment implementation process includes: a. using learnable parameters to adaptively adjust the weights of OT fusion and attention fusion; b. dynamically adjusting the fusion strategy according to the different distribution characteristics of the input features; c. retaining the original feature information through residual connections.
[0020] Furthermore, in S6, the classifier is implemented using a multi-layer feed-forward network, which includes batch normalization, dropout regularization, and ReLU activation function to map the fused features to the sign language category space.
[0021] The present invention also discloses a skeleton sign language recognition system based on a dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning, which is characterized by including: a. a data preprocessing module: used to divert the input skeleton sequence data; b. a mainstream feature extraction module: used to extract hand posture features through ST-DGCNN; c. a secondary feature extraction module: used to extract wrist trajectory features through geometric modeling; d. a feature interaction module: used to achieve mutual enhancement of dual-stream features; e. a feature fusion module: used to integrate dual-stream information; f. a classification and recognition module: used to output the final sign language recognition result.
[0022] Furthermore, the mainstream feature extraction modules include: a. ST-DGCNN: implements gesture feature extraction; b. SpatioTemporalBlock: implements spatiotemporal feature processing; c. MultiHeadAttention: implements multi-head attention mechanism; d. Residual connection structure: enhances feature learning capabilities.
[0023] Furthermore, the auxiliary stream feature extraction module includes: a. FinslerMetric: implements trajectory geometry modeling; b. SpatioTemporalCausalConv: implements spatiotemporal causal modeling; c. LSTM network: implements long-term dependency modeling.
[0024] Furthermore, the feature fusion module is implemented using AdaptiveFeatureFusion, which dynamically fuses two features through geometry-driven optimal transmission theory and attention mechanism.
[0025] The beneficial effects of the present invention are as follows:
[0026] (1) The dual-stream network architecture proposed in this paper captures hand gesture details through the main stream and analyzes wrist motion trajectory through the auxiliary stream, thus achieving comprehensive modeling of sign language movements. Compared with the single feature modeling method, the recognition accuracy is improved by more than 10%, especially in complex sign language vocabulary and fast action scenes.
[0027] (2) The spatiotemporal dynamic graph convolutional network (ST-DGCNN) designed in this paper incorporates a residual learning mechanism, effectively solving the vanishing gradient problem of deep networks, enhancing feature expression capabilities, and improving training stability. Experimental verification shows that compared with traditional DGCNN, the model convergence speed is increased by 40%, and the final test accuracy reaches over 95%.
[0028] (3) The trajectory metric learning method based on Finsler geometry proposed in this paper can effectively capture the dynamic trajectory features in sign language, especially high-order dynamic characteristics such as speed and acceleration, and has significant advantages in distinguishing similar gestures.
[0029] (4) This invention innovatively applies the geometry-driven optimal transmission theory to feature fusion. Compared with traditional fusion methods, it can more effectively integrate the complementary information of different modalities, especially in noisy environments and occlusion conditions. The recognition robustness is significantly enhanced and the ability to resist interference is greatly improved by %.
[0030] (5) The present invention applies optimized computing strategies in each module, including group convolution, JIT compilation, feature pooling and other technologies, which significantly improves the computing efficiency. Compared with traditional methods, the inference speed is 17ms, which can meet the needs of real-time applications.
[0031] (6) The proposed method realizes multi-level and multi-scale fusion of features through feature interaction and residual connection, obtains richer semantic expression, and has significantly better generalization ability on large-scale sign language datasets than existing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flow chart of the overall architecture of the method of the present invention;
[0033] Figure 2 It is a schematic diagram of the structure of the mainstream feature extraction module of the present invention;
[0034] Figure 3 Schematic diagram of the auxiliary stream feature extraction module of the present invention;
[0035] Figure 4 This is a schematic diagram of the feature interaction module structure and the geometry-driven optimal transmission fusion module structure of the present invention;
[0036] Figure 5 This is the accuracy and loss data of the present invention during the training process and the inference time diagram under deployment;
[0037] Figure 6 This is the output of the present invention before and after t-SNE dimensionality reduction results. DETAILED DESCRIPTION
[0038] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.
[0039] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0040] like Figure 1 As shown, an embodiment of the present invention discloses a skeleton sign language recognition method of a dual-stream spatiotemporal dynamic graph convolutional network integrating residual learning, comprising the following steps.
[0041] S1: Inputs gesture skeleton sequence data relative to the face and divides the data into two data streams: detailed hand posture and wrist motion trajectory relative to the face;
[0042] S2: Mainstream feature extraction of hand posture data through spatiotemporal dynamic graph convolutional network;
[0043] S3: Extract auxiliary flow features from wrist trajectory data through differential geometry and spatiotemporal causal convolution;
[0044] S4: Implement dual-stream feature interaction to enhance the features of the main and auxiliary streams.
[0045] S5: Perform dual-stream feature adaptive fusion and integrate the two-way information with the residual learning mechanism;
[0046] S6: The classifier outputs the sign language recognition result.
[0047] As a preferred embodiment of the present invention, in step S1, the process of processing the gesture skeleton sequence data is as follows:
[0048] (1-1) The input gesture skeleton sequence data dimension is [batch size, number of frames, number of hands, number of joints, coordinate dimension], where the number of hands is usually 2 (corresponding to the left and right hands), the number of joints is usually 21 points (including key positions such as wrist, base of finger, middle finger, fingertip), and the coordinate dimension is 3 (corresponding to the three-dimensional space coordinates x, y, z), where the position is relative to the face;
[0049] (1-2) For mainstream hand posture analysis, the influence of the overall wrist movement needs to be eliminated. Therefore, the coordinates of the corresponding wrist point are subtracted from the coordinates of all joint points of each hand to obtain the local coordinates relative to the wrist, forming a hand posture data stream;
[0050] (1-3) For auxiliary stream wrist trajectory analysis, the 3D coordinate sequence of the left and right wrist points is directly extracted and reorganized into the form of [batch size, number of frames, 6], where 6 represents the 3D coordinates of the left and right wrists, forming a wrist trajectory data stream relative to the face.
[0051] As a preferred embodiment of the present invention, refer to Figure 2 ,In step S2, the mainstream feature extraction adopts the ,spatial temporal dynamic graph convolutional network (ST-DGCNN), ,which efficiently models the hand posture sequence in ,time, through components such as k-nearest neighbor dynamic graph construction, ,spatial graph convolution, temporal dimension convolution, and multi-scale feature ,aggregation.
[0052] As a preferred embodiment of the present invention, the k-nearest-neighbor dynamic graph construction process in ST-DGCNN includes: first, treating the hand joint points as point clouds, and calculating the Euclidean distance between points for the point cloud data of each time step; then selecting k nearest neighbor points for each point (k is set to 10 in this embodiment) to construct a dynamic graph structure; finally, generating edge features for each edge, including the coordinate difference between adjacent points and the splicing of the original coordinates, to form an edge feature tensor of [batch size × number of frames, number of channels × 2, number of points, k] dimensions.
[0053] As a preferred embodiment of the present invention, the graph convolution processing process in ST-DGCNN includes: inputting edge features into the residual graph convolution block, which contains main path convolution and skip connection; the main path first processes the edge features through 1×1 convolution, and then applies batch normalization and ReLU activation function; the skip connection performs 1×1 convolution on the original input to achieve channel mapping; finally, the main path output is added to the skip connection output to form a residual learning structure, which effectively alleviates the gradient disappearance problem in deep networks.
[0054] As a preferred embodiment of the present invention, the time dimension convolution processing process in ST-DGCNN includes: first, performing maximum pooling on the features after spatial convolution in the point dimension to obtain global features of each time step; then arranging these global features in chronological order; then using one-dimensional grouped convolution to process the time series, where the number of groups is adaptively selected according to the number of channels, usually 4 or less, to reduce the amount of computation while maintaining expressiveness; finally, applying batch normalization and ReLU activation function to obtain time series features.
[0055] As a preferred embodiment of the present invention, in step S3, auxiliary stream feature extraction includes the following steps:
[0056] Step (3-1) Finsler geometry-based trajectory metric learning: First, the wrist trajectory data is input into the FinslerMetric module, which encodes the position and velocity direction information separately; then the Finsler metric basis function is constructed through the multi-layer perceptron; then the homogeneity constraint is introduced to ensure that the metric satisfies the geometric properties; finally, the Finsler energy of the trajectory segment is calculated and normalized to obtain the geometric characteristic representation of the trajectory
[0057] Step (3-2) Spatiotemporal Causal Convolution Processing: The left and right wrist trajectories are fed into the spatial convolution layer for feature extraction. Multi-scale dilated convolution is then applied to the extracted spatial features to capture dependencies across different timeframes, with dilation rates of 1, 2, and 4, respectively, to capture temporal patterns from short-term to long-term. Finally, features at different scales are fused to form a complete spatiotemporal representation.
[0058] Step (3-3) Bidirectional LSTM long-term dependency modeling: The output of the spatiotemporal causal convolution is fed into a two-layer bidirectional LSTM network with a hidden layer size of 512. The LSTM network processes the sequence in both the forward and backward directions to effectively capture long-term contextual relationships. Finally, the Finsler energy obtained in step (3-1) is used to weight the LSTM output features to highlight the key parts of the trajectory.
[0059] As a preferred embodiment of the present invention, in step S4, the dual-stream feature interaction process includes the following steps:
[0060] Step (4-1) Feature Projection: Use linear transformation to map the main stream features (dimension is batch size × hidden layer size × 2) and auxiliary stream features (dimension is batch size × number of frames × 1024) to a common latent space (dimension is 512), so that the features of the two streams can be compared in the same semantic space;
[0061] Step (4-2) Cross-Attention Calculation: Based on a multi-head attention mechanism (the number of heads is set to 8), the main stream features are used as queries and the auxiliary stream features as key-value pairs to calculate the attention weights. At the same time, the auxiliary stream features are used as queries and the main stream features as key-value pairs to implement reverse attention calculation. This two-way interaction enables the two streams to enhance each other's expression.
[0062] Step (4-3) Feature fusion and residual connection: The original features are concatenated with the attention-enhanced features, and mapped back to the original feature dimension through a fully connected layer. The mapping result is then added to the original features through a residual connection, and layer normalization is applied to enhance stability. Finally, a dropout layer (with a dropout rate of 0.1) is used to improve the generalization ability of the model.
[0063] As a preferred embodiment of the present invention, in step S5, the dual-stream feature adaptive fusion process includes the following steps:
[0064] Step (5-1) Geometry-driven optimal transmission fusion: First, a geometry-aware cost matrix combined with temporal relative position encoding is constructed. This matrix represents the matching cost between the main and auxiliary stream features. Then, the Sinkhorn algorithm is used to solve the regularized optimal transmission problem. The number of iterations is set to 10, and the regularization parameters are adaptively adjusted using learnable parameters. Finally, the auxiliary stream features are aligned to the main stream feature space based on the optimal transmission plan.
[0065] Step (5-2) Attention Fusion: Expand the main stream features to the same time dimension as the auxiliary stream features, and use the multi-head attention mechanism to calculate the correlation between them; the output of the attention layer is average pooled to obtain the attention-driven fusion feature;
[0066] Step (5-3) Dynamic weight adjustment: Two learnable parameters are introduced to control the weights of optimal transmission fusion and attention fusion respectively, so that the model can adaptively adjust the importance of different fusion strategies according to the data characteristics; then the two weighted fusion results are spliced with the original mainstream features, and the dimension and distribution are adjusted through the feature calibration network to enhance the expressiveness and consistency of the features.
[0067] The geometry-driven optimal transport fusion in step 5 uses the Sinkhorn algorithm to calculate the regularized optimal transport plan: First, a geometric perception cost matrix combined with temporal relative position coding is constructed to represent the matching cost between the main and auxiliary stream features. Then, the Sinkhorn iterative algorithm is used to solve the regularized optimal transmission problem. The iterative steps include: - Calculate the kernel matrix K = exp(-cost / epsilon), where epsilon is generated by parameterizing the network and is limited to the range of 0.1-5.0 - Initialize row and column scaling vectors u and v - Alternate row normalization and column normalization - Stop when convergence or maximum number of iterations is reached - Construct the final transmission plan Gamma Finally, based on the optimal transmission plan, the auxiliary stream features are aligned to the mainstream feature space
[0068] As a preferred embodiment of the present invention, in step S6, the classification output process includes: inputting the fused features into a multi-layer feedforward neural network, the network contains two hidden layers, and each layer is followed by a batch normalization layer and a ReLU activation function; the middle layer uses a Dropout layer with a dropout rate of 0.5 to prevent overfitting; the last layer maps the features to an output space of the dimension of the number of categories; the model is trained using a cross entropy loss function, and the output is converted into a probability distribution using a Softmax function, and the category with the highest probability is selected as the recognition result.
[0069] To make the purpose, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention are fully and clearly described below in conjunction with the accompanying drawings of the present invention:
[0070] The dataset used in this paper is based on sign language skeleton data obtained by MediaPipe hand tracking technology, and contains action sequences of commonly used sign language vocabulary. The LSA64 dataset contains a total of 2000 samples, covering 64 sign language categories, with approximately 31 samples in each category. WLASL_100 and WLASL_300 contain 100 and 300 sign language categories respectively, and each sample contains a complete sign language action sequence. Each frame contains the three-dimensional coordinates of 21 key points of the left and right hands, which are extracted in real time by the MediaPipe hand tracking model. The dataset is divided into training set, validation set and test set in a ratio of 8:1:1.
[0071] The network architecture of the present invention mainly includes three key modules: mainstream feature extraction module, auxiliary feature extraction module and feature interaction fusion module. Figure 1 shown.
[0072] The mainstream feature extraction module uses the spatiotemporal dynamic graph convolutional network (ST-DGCNN) to process hand posture data, such as Figure 2 As shown in the figure, this module first constructs a dynamic graph structure using the k-nearest neighbor algorithm, with the k value set to 12. It then extracts spatial features through multiple layers of residual graph convolution, with the number of convolutional layers being 32, 64, 128, and 256 channels, respectively. Each layer is followed by batch normalization and a LeakyReLU activation function (with a negative slope of 0.2). The extracted spatial features are max-pooled, and then captured through one-dimensional grouped convolution with a kernel size of 3 to capture temporal variations. Finally, multi-scale feature aggregation and a self-attention mechanism enhance expressiveness. ST-DGCNN's innovation lies in the introduction of residual connections and grouped convolutions, which significantly improve the model's expressiveness and computational efficiency.
[0073] The auxiliary flow feature extraction module processes wrist trajectory data based on differential geometry and spatiotemporal causal convolution, such as Figure 3 As shown in the figure, this module first constructs a Finsler geometric model using FinslerMetric, which uses a 64-dimensional hidden layer to encode position and velocity information respectively. It then generates the Finsler metric function using a multi-layer perceptron. It then implements multi-scale spatiotemporal causal convolution using the SpatioTemporalCausalConv module, using convolutional layers with three different expansion rates (1, 2, and 4) to capture motion features across different timeframes. Finally, it further processes temporal features using a bidirectional LSTM network (with a hidden layer size of 512 and a number of layers of 2), and uses Finsler energy to weight and enhance the output features.
[0074] The feature interaction fusion module includes two key sub-modules: CrossFeatureInteraction and AdaptiveFeatureFusion. Figure 4As shown in the figure, CrossFeatureInteraction achieves mutual guidance between two-stream features through multi-head cross-attention, using an 8-head attention mechanism and a 512-dimensional hidden layer. AdaptiveFeatureFusion innovatively introduces geometry-driven optimal transmission theory to achieve feature fusion, using the Sinkhorn algorithm (10 iterations, learnable regularization strength) to solve the optimal transmission problem, and adaptively adjusts the weights of different fusion strategies through dynamic weight parameters.
[0075] The training process used the following configuration: batch size was set to 8, initial learning rate was 0.0001, weight decay was 0.0001, and the maximum number of training epochs was 50. The AdamW optimizer with a cosine annealing learning rate schedule was used, with a minimum learning rate of 0.00001. The loss function consisted of a cross-entropy loss and a geometric consistency loss, with the total loss function being L = L_ce + λ·L_geo, where λ was 0.1. Mixed-precision training and gradient clipping (value 1.0) were used during training to improve training stability and efficiency. On an NVIDIA RTX 4090 graphics card, a single epoch took approximately 40 seconds, and the total training time was approximately 0.56 hours.
[0076] During the test, the performance of the method of the present invention on the test set is as follows: the overall accuracy reaches 97.65%, the recall rate is 96.8%, and the F1 score is 95.8%. Comparative experiments show that the dual-stream architecture of the present invention has a significantly higher accuracy of 10.5% than the single-stream architecture, especially when processing similar gestures and fast movements. Compared with traditional sign language recognition methods, the accuracy of the method of the present invention on the LSA-64 sign language dataset is improved by 7.5 percentage points. In the robustness test under complex environments, the accuracy of the method is significantly lower than that of the benchmark method, showing stronger generalization ability. In terms of inference speed, the average processing time per frame of the optimized model is only 8.2 milliseconds, which meets the needs of real-time applications. At the same time, the accuracy on WLASL_100 and WLASL_300 reached 94% and 91% respectively, reaching advanced levels on all benchmarks. As Figure 5 As shown, the average inference speed is around 17ms, which meets the real-time standard. Figure 6 It can be seen that the classification results after feature dimensionality reduction before and after processing are significant.
[0077] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, can make equivalent replacements or changes based on the technical solution and inventive concept of the present invention, which should be within the scope of protection of the present invention.
Claims
1. A skeleton sign language recognition method based on a dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning, characterized by: The steps include: S1, inputs the skeleton sequence data of the hand gesture relative to the face, and divides the data into two data streams: the detailed hand posture and the wrist motion trajectory relative to the face; S2, extracting features containing spatial topological relationships from the hand gesture stream through a spatiotemporal dynamic graph convolutional network (ST-DGCNN), where the ST-DGCNN integrates dynamic residual convolution blocks; S3, the Finsler trajectory dynamics encoder (FTDE) is used to perform differential geometry modeling on the wrist motion manifold trajectory flow, and a multi-scale causal convolutional network is combined to capture direction-sensitive features; S4, geometry-driven interaction of two-stream features is achieved through bidirectional cross-feature enhancer BCFE, which includes the manifold alignment process; S5, uses the geometry optimal transport fuser Geometry-OT to perform adaptive fusion of heterogeneous feature spaces; S6, output the recognition result through the multi-layer regularization classifier, and the classifier is configured with dropout probability p∈(0.3,0.5).
2. The skeleton sign language recognition method of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 1 is characterized in that: In S1, the input gesture skeleton sequence data dimension is [batch_size, frames, hands, joints, coordinates], where joints represents the number of joint points and coordinates represents the three-dimensional coordinate value; The detailed hand posture data is obtained by subtracting the wrist point coordinates to obtain the relative position of the wrist itself. The wrist motion trajectory data directly selects the wrist point coordinates to retain the position relative to the face.
3. The skeleton sign language recognition method of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 1 is characterized in that: In S2, the spatiotemporal dynamic graph convolutional network includes: a. Construct a k-nearest neighbor dynamic graph and extract spatial features; b. Apply the residual convolution module to process graph features; c. Perform convolution on the time dimension to capture timing changes; d. Multi-scale feature aggregation; e. Apply self-attention mechanism to enhance key features.
4. The skeleton sign language recognition method of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 3 is characterized in that: The residual convolution module adopts a convolution structure with skip connections, including: a. Use ResidualBlockCov to implement main path convolution; b. Use 1×1 convolution to implement identity mapping; c. Add the main path output and the identity map output to form a residual connection; d. Apply batch normalization and ReLU activation function.
5. The skeleton sign language recognition method of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 3 is characterized in that: The time dimension convolution uses grouped convolution and multi-scale structure to reduce the amount of computation while maintaining expressiveness.
6. The skeleton sign language recognition method of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 1 is characterized in that: In S3, auxiliary stream feature extraction includes: a. Implement trajectory metric learning inspired by Finsler geometry using FinslerMetric; b. Implement spatiotemporal causal convolution through SpatioTemporalCausalConv; c. Use bidirectional LSTM networks to model long-term dependencies; d. Weight the features according to Finsler energy.
7. The skeleton sign language recognition method of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 6 is characterized in that: The implementation process of FinslerMetric includes: a. Encode position and speed direction information separately; b. Construct Finsler metric basis function; c. Implement homogeneity constraints; d. Calculate the Finsler energy and normalize it.
8. The skeleton sign language recognition method of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 6 is characterized in that: The implementation process of spatiotemporal causal convolution includes: a. Spatial convolution for each wrist; b. Apply multi-scale dilated convolution to capture dependencies at different time scales; c. Fuse multi-scale features to form the final output.
9. The skeleton sign language recognition method of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 1 is characterized in that: In S4, dual-stream feature interaction is implemented through the CrossFeatureInteraction module, including: a. Project the main and auxiliary stream features into the latent space of the same dimension; b. Using the cross-attention mechanism, mainstream features guide auxiliary features, and auxiliary features guide mainstream features; c. Feature fusion and residual connection; d. Layer normalization enhances stability.
10. The skeleton sign language recognition method of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 1 is characterized in that: In S5, dual-stream feature adaptive fusion is implemented through the AdaptiveFeatureFusion module, which includes: a. Geometry-driven Optimal Transport (Geo-OT) fusion; b. Attention fusion; c. Dynamic weight adjustment and residual connection; d. Feature calibration and dimensional alignment.
11. The skeleton sign language recognition method of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 10 is characterized in that: The geometry-driven optimal transmission fusion implementation process includes a. Constructing a geometry-aware cost matrix; b. Solve the regularized optimal transmission problem using the Sinkhorn algorithm; c. Use adaptive entropy regularization parameter; d. Enhance feature alignment through geometric consistency constraints.
12. The skeleton sign language recognition method of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 10 is characterized in that: The dynamic weight adjustment implementation process includes: a. Use learnable parameters to adaptively adjust the weights of OT fusion and attention fusion; b. Dynamically adjust the fusion strategy based on the different distribution characteristics of input features; c. Retain the original feature information through residual connection.
13. The skeleton sign language recognition method of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 1 is characterized in that: In S6, the classifier is implemented using a multi-layer feed-forward network with batch normalization, dropout regularization, and ReLU activation function to map the fused features to the sign language category space.
14. A skeleton sign language recognition system based on a dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning, characterized by include: a. Data preprocessing module: used to divert the input skeleton sequence data; b. Mainstream feature extraction module TSSN: used to extract hand posture features through ST-DGCNN; c. FTDE (Assisted Flow Feature Extraction) module: used to extract wrist trajectory features through geometric modeling; d. Feature interaction module: used to achieve mutual enhancement of dual-stream features; e. Feature fusion module: used to integrate dual-stream information; f. Classification and recognition module: used to output the final sign language recognition results.
15. The skeleton sign language recognition system of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 14 is characterized in that: Mainstream feature extraction modules include: a. ST-DGCNN: implements gesture feature extraction; b. SpatioTemporalBlock: implements spatiotemporal feature processing; c. MultiHeadAttention: implements the multi-head attention mechanism; d. Residual connection structure: enhances feature learning capabilities.
16. The skeleton sign language recognition system of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 14 is characterized in that: The auxiliary stream feature extraction module includes: a. FinslerMetric: implements trajectory geometry modeling; b. SpatioTemporalCausalConv: implements spatiotemporal causal modeling; c. LSTM network: realizes long-term dependency modeling.
17. The skeleton sign language recognition system of the dual-stream spatiotemporal dynamic graph convolutional network integrated with residual learning according to claim 14 is characterized in that: The feature fusion module is implemented using AdaptiveFeatureFusion, which dynamically fuses two features through geometry-driven optimal transmission theory and attention mechanism.
Citation Information
Cited By
Gesture recognition method based on sparse millimeter wave radar point cloud
CN120853268A
A gesture recognition method based on sparse millimeter wave radar point cloud
CN120853268B
Skeleton gesture recognition method based on difference attention graph convolutional network
CN121281136A
Virtual digital human generation method and system based on modular parameters
CN121437698A