Sign language translation method based on SRGT double-flow structure
By adopting the SRGT dual-stream structure method in sign language translation, combining the timing improvement pooling and BERT-CRF model, the problem of feature extraction difficulties and incomplete information in sign language translation is solved, and more efficient and accurate sign language translation effect is achieved.
Patent Information
- Application Number
- CN202510323210.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art has problems in sign language translation with difficulty in feature extraction, incomplete information extraction and slow translation inference speed.
The sign language translation method based on the SRGT dual-stream structure is adopted, and the pooling is improved by designing the SRGT feature extraction method and timing improvement, combining the dual-stream structure to fusion of the skeleton flow and RGB flow characteristics, and sign language translation is performed using the BERT-CRF model.
It improves the feature extraction ability and information integrity of sign language translation, improves the accuracy and efficiency of translation, and is suitable for actual sign language translation tasks.
Smart Images

Figure CN120198963A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and specifically relates to a sign language translation method based on the SRGT dual-stream structure. Background Art
[0002] Sign language translation is a complex task, the purpose of which is to enable deaf people to communicate with people with normal hearing without barriers by recognizing gestures and converting them into text form. This process involves the intersection of multiple disciplines and technical fields, including but not limited to machine vision, text mining, and feature extraction. Extracting effective sign language action information from sign language videos is a crucial step, involving factors such as processing dynamic information, inter-frame relationships, and temporal modeling. At the same time, to accurately map sign language actions to text representations, it is necessary to overcome the expression differences and semantic gaps between sign language and natural language. The challenges of sign language translation also include dealing with complex gesture combinations, considering the grammatical characteristics of sign language, and the influence of long-process contexts. Traditional sign language recognition and translation methods mainly focus on manually extracting and expressing the characteristics of sign language data, and these methods often rely on complex algorithms and engineered feature designs. Deep learning-based methods, on the other hand, can automatically learn features from raw data, not only simplifying the feature engineering process but also greatly improving the generalization ability and recognition accuracy of the model.
[0003] In the field of continuous sign language recognition and translation, skeleton data, as an effective representation, has many advantages. Compared with the original video images, skeleton data is more adaptable to the diversity of video backgrounds, changes in lighting, changes in body scale, and changes in camera perspectives. This is because skeleton data provides more abstract and stable pose information, which can better retain the essential characteristics of sign language actions. In addition, skeleton data has a lower dimension, which helps to reduce the consumption of computing resources and storage space and improve the efficiency and performance of the algorithm. Graph convolutional networks are widely used in processing skeleton data due to their adaptability to graph-structured data, relationship modeling ability, and task flexibility. However, traditional GCNs cannot capture long-range dependencies due to their local aggregation nature, and the information transmission efficiency is low. SRGT (Spatiotemporal Residual GraphTransformer) is proposed, which has better long-range dependency capture ability compared with traditional GCNs. Introducing Transformer and constructing a residual structure can alleviate the over-smoothing and over-squeezing phenomena caused by information transmission.
[0004] As the number of nodes increases, the input information becomes more complex, and it is necessary to use pooling to extract key information to enhance the robustness and generalization ability of the model. However, traditional pooling methods have some drawbacks, such as information loss, spatial invariance, resolution reduction, fixed-size pooling windows, and uncontrollable computational complexity. Therefore, temporal boosting pooling is proposed, which can balance factors such as information loss and spatial invariance, ensure the balance of model performance and efficiency, and effectively process temporal data while retaining discriminative features.
[0005] SRGT focuses on skeleton data processing but ignores the local information in sign language videos. To make full use of this information, a sign language translation model based on a two-stream structure is proposed. Introducing the ResNet-50 network can effectively extract the local information of images, enrich the feature representation, and improve the translation performance. This combination fills the gap in SRGT's processing of image local information and provides an effective solution for the improvement of sign language translation models. At the same time, a sign language translation model based on BERT-CRF is proposed to address the limitations of the Transformer model in processing long sequences and the need for labeled data. By changing the output layer of the BERT model to a CRF layer, combining semantic understanding and sequence labeling capabilities, it can more accurately capture the entity boundaries and context dependencies of sign language sequences. Compared with the Transformer model, the BERT-CRF-based model has advantages in improving translation accuracy and efficiency while processing long sequences. Parameter sharing and end-to-end advantages make the model more lightweight and efficient, suitable for actual sign language translation tasks. Summary of the Invention
[0006] To solve some problems existing in the prior art, the present invention proposes a sign language translation method based on the SRGT two-stream structure to better solve the problems of difficult feature extraction, incomplete information extraction, and slow sign language translation inference speed. The technical solution adopted by the present invention is a sign language translation method based on the SRGT two-stream structure, and the method includes the following steps:
[0007] S10, design a feature extraction method named SRGT to effectively capture the spatio-temporal features of sign language actions;
[0008] S20, design temporal boosting pooling to effectively process temporal data while retaining discriminative features;
[0009] S30, use a two-stream structure to fuse the skeleton stream features processed by S10 and the RGB stream features processed by ResNet-50, and finally use the BERT-CRF model for sign language translation.
[0010] In step S10, the design of a feature extraction method named SRGT to effectively capture the spatio-temporal features of sign language actions is as follows:
[0011] S101, Extract the spatial skeleton graph. Use MMPose to extract 50 key points of the upper body and hand nodes, and by default, the extraction of face, body, and hand nodes is dense.
[0012] S102, By calculating the differences of joint points between adjacent frames, obtain the joint motion information, and concatenate the joint information and its motion information along the channel dimension as the input of the network. Among them, at the t-th frame, the representation of the motion information of joint point i is as shown in Equation , where x, y, and score respectively represent the horizontal and vertical coordinates and confidence score of joint point i in the image frame.
[0013] S103, The human joint coordinates extracted from each video frame form a skeleton sequence. To represent this hierarchical structure, construct a spatio-temporal skeleton graph.
[0014] S104, Use SRGT to extract features from the constructed spatio-temporal skeleton graph, as shown in Figure 1 , and the specific steps are as follows:
[0015] S1041, In the Graph Transformer, use the multi-head attention mechanism to learn the edge features in the spatio-temporal skeleton graph. The multi-head attention mechanism divides the node features into multiple subspaces, calculates the attention weights on each subspace respectively, and integrates the attention weights of different subspaces through concatenation to obtain the final node representation. In this way, the model can learn the spatio-temporal correlations between nodes at different focus points, so as to better extract the features of the spatio-temporal skeleton graph. Among them, the input node feature is defined as Equation , and through the calculated multi-head attention, the information aggregation calculation from j to i is as Equation , , where, and are trainable parameters used to transform the feature of node j into the value vector , is the representation of node j at the l-th layer of the c-th head, || represents the concatenation operation of C-head attention, and C is the number of attention heads.
[0016] S1042, To prevent the model from over-smoothing, introduce Residual Granph Transformer, that is, use gated residual connections between Graph Transformer layers, combined with the spatio-temporal skeleton graph. The specific calculation process is as Equation , , .
[0017] In step S20, the designed temporal boosting pooling effectively processes temporal data while retaining the discriminative features. The process of temporal boosting pooling is as follows Figure 2 as shown, and the specific steps are as follows:
[0018] S201, Temporal boosting. Decompose the one-dimensional time signal x into a downscaled approximation signal s and a difference signal d. The specific calculation is as shown in Equation , where consists of three functions, is the function composition operator. Thus, the boosting process can be divided into three sub-processes: splitting, prediction, and update.
[0019] S2011, In the splitting process, first divide the input signal x into two disjoint sets and using even and odd indices to reduce the generation of signals and make them closely related in time. The calculation formula is as shown in Equation .
[0020] S2012, In the prediction process, the specific calculation process is as shown in Equation . Given a selected set, such as , uses the predictor to predict the other set . The difference signal d is obtained through a high-pass coefficient. The specific calculation is as shown in Equation .
[0021] S2013, In the update process, the specific calculation is as shown in Equation . To avoid information loss, use the update function to take the difference signal d as the compensation input and generate a smoothed downscaled value s. The specific calculation is as shown in Equation .
[0022] S2014, For the prediction function and the update function , the specific calculation is as shown in Equation , , where represents the size of the convolution kernel, represents the number of groups of convolutions. First, deploy a one-dimensional convolution with a kernel size of to aggregate local time patterns, and then perform ReLU activation. Use a normal convolution to achieve channel aggregation, and finally use for feature prediction.
[0023] S202, Component weighting. Design a component weighting module to dynamically emphasize or suppress or Some components in generate a specific coefficient for each channel at the timestamp to generate a total weight matrix. A mini fully convolutional network is used to instantiate and then function is used to dynamically determine the weights during the end-to-end optimization process. Each value in W is in the range of (0,1), representing the importance of a certain component generated during the lifting process. Component weighting is performed in a residual manner, and the specific calculation formula is as shown in Equation where is a matrix of all 1s.
[0024] S203, pooling fusion. A tiny convolutional network composed of a convolution sequence with a kernel of 1, BatchNorm, and ReLU is used to combine and and the calculation is as shown in Equation .
[0025] In step S30, the two-stream structure is used to fuse the skeleton stream features processed by S10 and the RGB stream features processed by ResNet-50, and finally the BERT-CRF model is used for sign language translation. The schematic diagram of this structure is as shown in Figure 3 and the specific steps are as follows:
[0026] S301, preprocess the RGB stream data, including segmentation, cropping, and elimination, etc.
[0027] S302, use ResNet-50 as the deep learning model to extract features of the sign language in the RGB stream.
[0028] S303, linearly fuse the extracted RGB stream sign language features with the skeleton stream features to further improve the accuracy and robustness of sign language recognition.
[0029] S304, input the fusion features obtained in S303 into BERT-CRF. When the BERT-CRF model is applied to the fusion data of sign language translation, by using the bidirectional encoding ability, more comprehensive information can be obtained, and translation can be performed without a large amount of sign language data, thus reducing the dependence on sign language corpora; combining semantic understanding and sequence annotation capabilities, accurately capturing the entity boundaries and context dependencies of sign language sequences.
[0030] Compared with the prior art, the advantages of the present invention are:
[0031] (1) The present invention proposes SRGT, which can better capture long - distance dependencies while alleviating the over - smoothing and over - squeezing phenomena caused by information transmission, solving the problems that traditional GCNs have local aggregation properties, cannot capture long - distance dependencies, and have low information - transmission efficiency. Through temporal enhancement pooling, while retaining discriminative features, it can effectively process temporal data, solving the disadvantages of traditional pooling methods such as information loss, spatial invariance, resolution reduction, and fixed - size pooling windows.
[0032] (2) The present invention adopts a two - stream structure and introduces a ResNet - 50 network to effectively extract local information of images, which can more completely extract sign - language expression information in sign - language videos, solving the problem that SRGT only focuses on skeleton data processing but ignores local information in images. By changing the output layer of the BERT model to a CRF layer, compared with the Transformer model, the model of the present invention has advantages in improving translation accuracy and efficiency while processing long sequences. Parameter sharing and end - to - end advantages make the model lighter and more efficient, suitable for actual sign - language translation tasks. Brief Description of the Drawings
[0033] Figure 1 It is the structure diagram of the SRGT network;
[0034] Figure 2 It is the process diagram of temporal enhancement pooling;
[0035] Figure 3 It is the overall structure diagram. Detailed Embodiment
[0036] The following will clearly and completely describe the concept, specific structure, and technical effects generated by the present invention in combination with embodiments and drawings to fully understand the purpose, features, and effects of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0037] Refer to Figure 3 As shown, the present invention discloses a sign - language translation method based on the two - stream structure of SRGT, and the method includes the following steps:
[0038] S10, design a feature - extraction method called SRGT to effectively capture the spatio - temporal features of sign - language actions. The specific steps are as follows:
[0039] S101, extract the spatial skeleton graph, use MMPose to extract 50 key points of the upper body and hand nodes, and by default, the extraction of face, body, and hand nodes is dense.
[0040] S102. Obtain joint motion information by calculating the differences between joint points in adjacent frames, and concatenate the joint information and its motion information along the channel dimension as the input of the network. Among them, at the t-th frame, the representation of the motion information of joint point i is as shown in Equation , where x, y, and score respectively represent the horizontal and vertical coordinates and confidence score of joint point i in the image frame.
[0041] S103. The human joint coordinates extracted from each video frame form a skeleton sequence. To represent this hierarchical structure, a spatio-temporal skeleton graph is constructed.
[0042] S104. Use SRGT to extract features from the constructed spatio-temporal skeleton graph, as shown in Figure 1 . The specific steps are as follows:
[0043] S1041. In the Graph Transformer, the multi-head attention mechanism is used to learn the edge features in the spatio-temporal skeleton graph. The multi-head attention mechanism divides the node features into multiple subspaces, calculates the attention weights on each subspace respectively, and integrates the attention weights of different subspaces through concatenation to obtain the final node representation. In this way, the model can learn the spatio-temporal correlations between nodes at different focus points, so as to better extract the features of the spatio-temporal skeleton graph. Among them, the input node feature is defined as Equation . Through the calculated multi-head attention, the information aggregation calculation from j to i is as shown in Equation , , where and are trainable parameters used to convert the feature of node j into the value vector , is the representation of node j at the l-th layer of the c-th head, || represents the concatenation operation of C-head attention, and C is the number of attention heads.
[0044] S1042. To prevent the model from over-smoothing, a Residual Granph Transformer is introduced, that is, gated residual connections are used between Graph Transformer layers, combined with the spatio-temporal skeleton graph. The specific calculation process is as shown in Equation , , .
[0045] S20. Design a temporal boosting pooling to decompose the input signal into sub-bands of different frequencies, effectively process the temporal data while retaining the discriminative features. The temporal boosting pooling process is as shown in Figure 2 . The specific steps are as follows:
[0046] S201, Temporal Enhancement. Decompose the one-dimensional time signal x into a downscaled approximation signal s and a difference signal d. The specific calculation is as shown in Equation , where consists of three functions, is the function composition operator. Thus, the enhancement process can be divided into three sub-processes: splitting, prediction, and update.
[0047] S2011, During the splitting process, first divide the input signal x into two disjoint sets using even and odd indices and to reduce the generation of the signal and be closely related in time. The calculation formula is as shown in Equation .
[0048] S2012, During the prediction process, the specific calculation process is as shown in Equation . Given a selected set, for example , uses the predictor to predict the other set . The difference signal d is obtained through the high-pass coefficient. The specific calculation is as shown in Equation .
[0049] S2013, During the update process, the specific calculation is as shown in Equation . To avoid information loss, use the update function to take the difference signal d as the compensation input and generate a smoothed downscaled value s. The specific calculation is as shown in Equation .
[0050] S2014, For the prediction function and the update function , the specific calculation is as shown in Equation , , where represents the size of the convolution kernel, represents the number of groups of convolution. First, deploy a one-dimensional convolution with a kernel size of to aggregate local time patterns, and then perform ReLU activation. Use normal convolution to achieve channel aggregation, and finally use for feature prediction.
[0051] S202, Component Weighting. Design a component weighting module to dynamically emphasize or suppress or certain components in. At the timestamp , functions to generate a specific coefficient for each channel, thus generating a total weight matrix Instantiate using a tiny fully convolutional network , and then use function to dynamically determine the weights during the end-to-end optimization process. Each value in W is in the range (0, 1), representing the importance of a certain component generated during the lifting process. Component weighting is performed in a residual manner, and the specific calculation formula is as shown in Equation , where is a matrix of all 1s.
[0052] S203, Pooling fusion. Use a tiny convolutional network composed of a convolutional sequence with a kernel of 1, BatchNorm, and ReLU to combine and , and the calculation is as shown in Equation .
[0053] S30, Use a two-stream structure to fuse the skeleton stream features processed by S10 and the RGB stream features processed by ResNet-50, and finally use the BERT-CRF model for sign language translation. The schematic diagram of this structure is as shown in Figure 3 , and the specific steps are as follows:
[0054] S301, Preprocess the RGB stream data, including segmentation, cropping, and elimination, etc.
[0055] S302, Use ResNet-50 as a deep learning model to extract features of the sign language in the RGB stream.
[0056] S303, Linearly fuse the extracted RGB stream sign language features with the skeleton stream features to further improve the accuracy and robustness of sign language recognition.
[0057] S304, Input the fusion features obtained in S303 into BERT-CRF. When applying the BERT-CRF model to the fusion data of sign language translation, utilize the bidirectional encoding ability to obtain more comprehensive information, and translation can be performed without a large amount of sign language data, thus reducing the dependence on sign language corpora; combine semantic understanding and sequence annotation capabilities to accurately capture the entity boundaries and context dependencies of sign language series.
Claims
1. A sign language translation method based on SRGT dual-stream structure, characterized in that: The method comprises the following steps: S10, design a feature extraction method called SRGT to effectively capture the spatiotemporal characteristics of sign language movements; S20, design time series improvement pooling to effectively process time series data while retaining identification features; S30, uses a dual-stream structure to fuse the skeleton stream features processed by S10 and the RGB stream features processed by ResNet-50, and finally uses the BERT-CRF model for sign language translation.
2. The sign language translation solution based on the SRGT dual-stream structure according to claim 1 is characterized in that: In step S10, a feature extraction method named SRGT is designed to effectively capture the spatiotemporal features of sign language movements. The specific steps are as follows: S101, extract the spatial skeleton graph, use MMPose to extract 50 key points of the upper body and hand nodes, and default the face, body and hand node extraction is dense; S102, by calculating the difference of joint points between adjacent frames, the joint motion information is obtained, and the joint information and its motion information are connected according to the channel dimension as the input of the network; wherein, at the tth frame, the motion information of the joint point i is represented as follows: , where x, y and score represent the horizontal and vertical coordinates and confidence score of joint point i in the image frame respectively; S103, the human body joint coordinates extracted from each video frame form a skeleton sequence, and a spatiotemporal skeleton graph is constructed to express this hierarchical structure; S104, using SRGT to extract features from the constructed spatiotemporal skeleton graph, as shown in FIG1 , the specific steps are as follows: S1041, in Graph Transformer, a multi-head attention mechanism is used to learn edge features in the spatiotemporal skeleton graph. The multi-head attention mechanism divides the node features into multiple subspaces, and calculates the attention weights on each subspace separately, and integrates the attention weights of different subspaces by splicing to obtain the final node representation; in this way, the model can learn the spatiotemporal associations between nodes at different attention points, so as to better mine the features of the spatiotemporal skeleton graph; the input node feature is defined as follows , through the calculation of multi-head attention, the information aggregation calculation from j to i is as follows , ,in, and is a trainable parameter used to transform the features of node j Convert to a vector of values , is the representation of node j of the c-th head at layer l, || represents the connection operation of C-head attention, and C is the number of attention heads; S1042, in order to prevent the model from being over-smoothed, the Residual Granph Transformer is introduced, that is, the gated residual connection is used between the Graph Transformer layers, combined with the spatiotemporal skeleton graph. The specific calculation process is as follows: , , .
3. The sign language translation method based on the SRGT dual-stream structure according to claim 1 is characterized in that: In step S20, the timing improvement pooling is designed to effectively process the timing data while retaining the identification features. The timing improvement pooling process is shown in FIG2 , and the specific steps are as follows: S201, time series lifting; decompose the one-dimensional time signal x into a reduced-scale approximate signal s and a difference signal d. The specific calculation is as follows: ,in, It consists of three functions: ; is the function combination operator; hence, the lifting process can be divided into three sub-processes: segmentation, prediction, and update; S2011, during the segmentation process, the input signal x is first divided into two disjoint sets using even and odd indices , , in order to reduce the generation of signals and to be closely related in time, the formula is as follows ; S2012, during the prediction process, the specific calculation process is as follows ; Given a selected set, such as , Using the Predictor Predict another set ; The differential signal d is obtained by taking the high-pass coefficient, and the specific calculation is as follows: ; S2013, during the update process, the specific calculation is as follows , to avoid information loss, use the update function The differential signal d is used as the compensation input to generate a smoothed reduced value s. The specific calculation is as follows: ; S2014, for the prediction function and update function , the specific calculation is as follows , ,in represents the size of the convolution kernel, Represents the number of convolution groups; first deploy a kernel size of One-dimensional convolution of to aggregate local temporal patterns, followed by ReLU activation; using normal Convolution is used to achieve channel aggregation, and finally use To make feature predictions; S202, component weighting; design a component weighting module to dynamically emphasize or suppress or Some components in the time stamp Down, The role of is to generate a specific coefficient for each channel, thus producing a total weight matrix ; Instantiate using a tiny fully convolutional network , and then use The function dynamically determines the weights during the end-to-end optimization process; Each value in is in the range of (0,1), indicating the importance of a component generated by the lifting process; the component weighting is performed using the residual method, and the specific calculation formula is as follows ,in, is a matrix of all 1s; S203, pooling fusion; using a convolution sequence with a kernel of 1, a small convolution network consisting of BatchNorm and ReLU to combine and , calculated as .
4. The sign language translation method based on the SRGT dual-stream structure according to claim 1 is characterized in that: In step S30, the dual-stream structure is used to fuse the skeleton stream features processed by S10 and the RGB stream features processed by ResNet-50, and finally the BERT-CRF model is used to perform sign language translation. The schematic diagram of the structure is shown in FIG3 , and the specific steps are as follows: S301, preprocessing RGB stream data, including segmentation, cropping and culling; S302, using ResNet-50 as a deep learning model to extract features of sign language in the RGB stream; S303, linearly fusing the extracted RGB stream sign language features with the skeleton stream features to further improve the accuracy and robustness of sign language recognition; S304, input the fusion features obtained in S303 into BERT-CRF; when the BERT-CRF model is applied to the fusion data of sign language translation, the bidirectional encoding capability is utilized to obtain more comprehensive information, and translation can be performed without a large amount of sign language data, thereby reducing the dependence on sign language prediction; combined with semantic understanding and sequence labeling capabilities, the entity boundaries and context dependencies of the sign language series are accurately captured.