A continuous sign language recognition method and device

Through the Spatial-Temporal Transformer network, the problem of low sign language recognition rate in complex contexts is solved, and higher recognition accuracy and robustness are achieved through blocking and feature extraction technology.

CN115393949BActive Publication Date: 2025-07-08HEBEI UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210827343.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-14
Publication Date
2025-07-08
Estimated Expiration
2042-07-14

AI Technical Summary

Technical Problem

The existing sign language recognition method has low recognition rate in complex backgrounds, making it difficult to effectively extract sign language video features within long time intervals, resulting in inaccurate semantic recognition.

Method used

Using the Spatial-Temporal Transformer-based method, the feature extraction is enhanced by randomly deleting redundant frames, blocking operations and vectorization processing, and the encoder combining time and spatial attention mechanisms, and feature decoding and prediction are used to enhance the feature extraction and recognition of sign language videos.

Benefits of technology

It improves the accuracy and robustness of sign language recognition, can effectively extract dynamic spatial and long-term temporal features in complex contexts, and improves the recognition rate of continuous sign language recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393949B_ABST
    Figure CN115393949B_ABST
Patent Text Reader

Abstract

The present invention relates to a continuous sign language recognition method and device. The recognition method deletes redundant frames from the original sign language video through a random deletion method to obtain a continuous sign language video sequence; performs a chunking operation and a vectorization process on the obtained sign language video sequence to obtain a sign language sequence vector; extracts features from the obtained sign language sequence vector through a spatio-temporal encoder to obtain a sign language sequence feature vector; decodes the obtained sign language sequence feature vector; predicts a sequence for the decoded features; calculates the WER between the predicted sequence of the obtained sign language video and the sign language text sequence; performs network-level training and outputs the final sign language recognition result. The present invention has strong robustness and can obtain a higher recognition rate in the cases of multiple sign language users, multiple sentences, and multiple language inputs, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a human-computer interaction method, and more particularly to a continuous sign language recognition method and apparatus. Background Art

[0002] Sign language is an important expression of human body language, which contains a large amount of information and is the main way of communication between deaf-mute people and hearing people. Due to the rich semantics of sign language, the movement amplitude has locality and detail compared with other human behaviors, and it is also affected by light, background, movement speed, etc. The accuracy and robustness achieved by using traditional pattern recognition or machine learning methods have reached a bottleneck period, and are often limited to static gesture recognition or isolated dynamic gesture recognition, while continuous sign language recognition can better meet the communication needs of deaf-mute people.

[0003] Continuous sign language recognition is different from isolated word recognition in that the video sequence is longer and more complex, and feature and semantic learning need to be carried out in the continuous frame sequence of the sign language video. In actual scenarios, sign language videos contain complex life scenes, so there are long-term semantic dependencies in the videos. Each video frame is related not only to adjacent video frames, but also to distant video frames. However, existing methods are difficult to capture the detailed temporal dynamics within a long time interval using simple video representations. The reason is still that feature extraction is not sufficient.

[0004] Patent No. CN202010083258.8 discloses a method for continuous sign language recognition based on an encoding-decoding network that fuses multi-modal image sequence features and self-attention mechanisms. This method first obtains an optical flow image sequence, and through the extraction of spatio-temporal features of the original sign language image sequence and the optical flow image sequence, the fusion of spatio-temporal features of the multi-modal image sequence, and the extraction of the text feature sequence of the sign language sentence label, the fused spatio-temporal features of the multi-modal image sequence and the extracted text feature sequence of the sign language sentence label are input into the encoding-decoding network based on the self-attention mechanism for sign language label prediction output. However, currently, for sign language recognition methods based on deep learning, under long-sequence continuous sign language sequences, the existing networks have a slow convergence speed and the sign language recognition rate is not high.

[0005] Due to the variability of sign language, the results of hand detection are prone to generate rich backgrounds, which interfere with sign language recognition and reduce interactivity. Sign language in complex backgrounds also has problems such as long sequences and large computational amounts. In addition, sign language videos contain rich context semantic information, and insufficient feature extraction leads to inaccurate semantic recognition and reduces the recognition effect. Summary of the Invention

[0006] The objective of the present invention is to provide a continuous sign language recognition method and device based on Spatial-Temporal Transformer, so as to solve the problem that the existing methods have a low recognition rate for continuous sign language in complex backgrounds.

[0007] The present invention is implemented as follows: A continuous sign language recognition method based on Spatial-Temporal Transformer includes the following steps:

[0008] S1. The original sign language video deletes redundant frames through a random deletion method to obtain a continuous sign language video sequence;

[0009] S2. Perform chunking operations and vectorization processing on the obtained sign language video sequence to obtain a sign language sequence vector;

[0010] S3. Use the encoder in the Spatial-Temporal Transformer network as a spatio-temporal encoder to extract features from the obtained sign language sequence vector to obtain a sign language sequence feature vector; the encoder is an encoder with dual channels of time and space;

[0011] S4. Decode the features of the obtained sign language sequence feature vector;

[0012] S5. Perform sequence prediction on the decoded features to obtain a prediction sequence of the sign language video;

[0013] S6. Calculate the WER between the obtained prediction sequence of the sign language video and the sign language text sequence;

[0014] S7. Perform network-level training on steps S3, S4, S5, and S6, and output the final sign language recognition result.

[0015] The Spatial-Temporal Transformer network includes a vectorization module, an encoder, and a decoder. The vectorization module includes patch operations, Patch-embedding, and Positional Encoding operations; the encoder includes a time attention calculation mechanism, a spatial attention mechanism, and a feed-forward neural network; the decoder includes a self-attention calculation mechanism, a cross-attention mechanism, and a feed-forward neural network.

[0016] Furthermore, in step S2 of the present invention, in order to facilitate processing of the input (T-frame sign language video) with a dimension of f ∈ R B×T×N×DThe sign language video sequence vector reshapes each frame in the sign language video frames of T frames (B is the batch-size, H and W are the resolutions of the original images, and C is the number of channels) into a 2D patch of dimension (h×w)×(p1×p2×C), where H = h×p1 and W = w’×p2. h×w is the number of patches each frame is divided into, which directly affects the length of the input sequence (the choice of p can be used for experimental comparison), and a constant hidden vector d is used on all layers model The patches are flattened and projected onto d model = the size of D. This projected output is the patch embedding. At this time, the size of the feature map is B×T×N×D, N = h×w, and the vector after patch embedding is denoted as: X (p,t) .

[0017] After obtaining the feature map f0 ∈ R B×T×N×D dimension, Positional Encoding is also required. Since the self-attention of the original transformer does not contain position information, but there is very strong sequence information in the sign language video frames. To prevent the loss of the concept of front and back frames and to facilitate the subsequent extraction of features in the time dimension, position information is added to the feature map. The position encoding requires that each position has a unique positional encoding and the relationship between two positions can be modeled by the affine transformation between their positional encodings. And through experimental verification:

[0018]

[0019]

[0020] Formulas (1, 2) can exactly meet these two requirements, that is, Positional Encoding (PE), where pos corresponds to the position of the token in the sequence, the starting token position is denoted as 0, 2i and 2i + 1 represent the dimensions of the Positional Encoding, and d model is the dimension after position encoding, and the value range of i is [0, d model / 2]. The position encoding information is marked as

[0021] Further, the encoder used in step S3 of the present invention is an encoder structure that takes into account both time and space, including a spatial attention module and a temporal attention module; the sign language video vectors input into the encoder enter the temporal attention module and the spatial attention module through two channels respectively, and then the features extracted by the temporal attention module and the spatial attention module are concatenated, and the dynamic spatial correlation and long-term temporal correlation are utilized to improve the extraction and encoding of the sign language video frame features by the network.

[0022] Further, the encoding process of the encoder is as follows: the output result of the vectorization module is rearranged in dimensions. First, the time dimension t is placed on the first dimension batch, and the dynamic spatial correlation attention calculation is performed on the n vector blocks in the spatial dimension; then the spatial dimension n is placed on the first dimension batch, and the long-term temporal correlation attention calculation is performed on the t frame sequence in the time dimension; then the temporal attention calculation result and the spatial attention calculation result are fused; finally, after passing through the linear normalization layer and the feed-forward neural network, the output is performed.

[0023] Further, the temporal and spatial attention calculation processes of the encoder are as follows:

[0024] (1) The Spatial Self-Attention Block only performs MSA calculation on different tokens of the same frame. The attention value calculation of the query Q vector in the spatial dimension is as shown in formula (3) below.

[0025]

[0026]

[0027] Among them, space refers to performing attention calculation in the spatial dimension, time refers to performing attention calculation in the temporal dimension, softmax refers to the activation function, l refers to the l-th layer, a refers to the a-th attention head, p refers to the p-th block in each frame, and t refers to the t-th frame. D h = D / A refers to the dimension value of the corresponding attention head, D is the dimension value of the vector, A is the total number of heads, q refers to the query vector, and k is the weight matrix corresponding to q.

[0028] (2) The Temporal Self-Attention Block only performs MSA calculation on tokens at the same position in different frames. The attention calculation of the query Q vector in the temporal dimension is as shown in formula (4). After calculating the attention in the temporal and spatial dimensions respectively, cat concatenation is performed:

[0029] Further, the specific operation method of step S3 of the present invention includes the following steps:

[0030] S3-1. Use the training data in step S2 as the input video, and first perform Embedding and Positional Encoding operations. Among them, the calculation of Positional Encoding is based on the following two formulas:

[0031]

[0032]

[0033] where: pos corresponds to the position of the token in the sequence, the starting token position is recorded as 0, 2i and 2i + 1 represent the dimensions of Positional Encoding, and the value range of i is [0, d model / 2], mark the position encoding information as

[0034] S3-2. Use the vectors after Embedding and Positional Encoding as the input P of the STT encoding module. After the vector enters the STTN encoding module, first use the dimension transformation operation to put the number of frames T of the vector P to the first dimension batch-size. The spatial encoding module performs spatial attention calculation on the vector P0 after dimension transformation, and the calculated vector is recorded as Z0; then use the dimension transformation operation to put the number of blocks N of the vector P to the first dimension batch-size, and perform temporal attention calculation on the vector P1 after dimension transformation through the temporal encoding module, and the calculated vector is recorded as Z1; the calculation methods of the temporal attention and spatial attention are both Self Attention calculations; the three matrices W Q , W K , W V perform three linear transformations on all P0 / P1 vectors respectively, and all vectors derive three new vectors q t , k t , v t ; all vectors q t are assembled into the query matrix Q, all vectors k t are assembled into the key matrix K, and all vectors v t are assembled into the value matrix V;

[0035] The calculation formulas are as follows:

[0036] Q = Linear(X) = XW Q (8)

[0037] K = Linear(X) = XW K (9)

[0038] V = Linear(X) = XWV (10)

[0039] X attention = Self Attention(Q, K, V) (11)

[0040] where X is the input sequence, and W Q , W K , W V are three matrices;

[0041] Perform a feature concatenation operation on vector Z0 and vector Z1, and after Layer Normalization and FeedForward operations, perform the output of the encoder. The formula is as follows:

[0042] X attention = Layer Norm(X attention ) (12)

[0043] where FeedForward is two-layer linear mapping and activated by an activation function. The activation function uses ReLU;

[0044] X hidden = Layer(ReLU(Linear(X attention ))) (13)

[0045] where X hidden ∈R batch_size*seq_len*embed_dim ;

[0046] S3-3. Perform three parts of operations on the decoder in sequence: ① Masked Multi-Head Self-Attention, ② Multi-Head Encoder-Decoder Attention, ③ FeedForward Network. Each part of the operation is followed by a Layer Normalization; the attention calculation of the decoder includes the self-attention calculation between sign language texts and the cross-attention calculation between the attention output and the encoder output.

[0047] Furthermore, the overall decoding operation of the decoder includes the following steps:

[0048] S3-3-1. First, perform a Word-Embedding operation on the sign language text, map it to a D-dimensional vector through a Matrix, denoted as matX, and then perform a Positional-Encoding operation to encode different position information for different words in the sign language text, with a dimension of D, denoted as matP. Add them to obtain the input of the decoder: matDec = matP + matX;

[0049] S3-3-2. When calculating the self-attention between sign language texts, the three inputs of Multi-head attention are Q, K, and V, and linear transformations are performed on V, K, and Q respectively; Q is then divided into num_heads segments on the last dimension, and the divided matrices are concat-linked on axis=0; the same operation is performed on V and K; the matrices after the operation are recorded as Q_, K_, V_; Q_ matrix is ​​multiplied by the transpose of K_, and the result is recorded as outputs.

[0050] S3-3-3, scale outputs once and update them to outputs; perform sofimax operation on outputs and update outputs;

[0051] S3-3-4. When calculating the cross attention between the attention output and the encoder output, Q is the output of the encoder, K=V=matDec, and the calculation process is the same as the calculation of the self-attention between the sign language text; the operation in the Add&Norm layer is the same as ResNet, and the initial input is superimposed with its corresponding output once, that is, outputs=outputs+Q, so that the network can be effectively superimposed to avoid gradient disappearance. After the Add&Norm layer and Feed Forward, the outputs are normalized and linearly transformed; after completing Nx times, the vector generated by the decoder stack is projected to a larger vector through the linear layer to become a logits vector, and then converted into a probability by the softmax layer, and the highest probability unit is selected to generate the word associated with it as the output of the current time step. At this point, the decoding of the model is completed.

[0052] Furthermore, step S4 of the present invention is to use a decoding network to perform feature decoding, and the decoding network includes a multi-head attention mechanism, a multi-head cross-attention mechanism and a feedforward neural network. First, self-attention calculation is performed on the sign language text, and then cross-attention calculation is performed with the feature vector generated by the encoding.

[0053] Furthermore, the decoder used in step S4 of the present invention includes three sub-layers, the first sub-layer includes a multi-head self-attention layer containing a mask, a normalization layer, and a residual connection layer, the second sub-layer includes a multi-head cross-attention layer, a normalization layer, and a residual connection layer, and the third sub-layer includes a feedforward neural network, a normalization layer, and a residual connection layer.

[0054] The calculation formula for the three layers is:

[0055]

[0056] Among them, Q i-1 Refers to the output of the previous layer, F refers to the output of the encoder, F with the positional encoding operation added to it.

[0057] Furthermore, in step S5 of the present invention, positional encoding is also required before entering the decoding process to add positional information to the sign language text. After that, it will first pass through a masked self-attention layer. The purpose of using the mask is to solve the problem of information leakage during the training process encountered in the decoding process, that is, to avoid the problems of model cheating and inconsistent model architectures during the prediction process. This is because using sequence masked in the prediction stage can keep the prediction results of repeated sentences the same, which not only conforms to the rules but also can be updated incrementally. At the same time, it can also be consistent with the model architecture during training and the way of forward propagation. In addition, there is only one output at the encoder end, and each decoding layer passed to the decoder part acts as the K and V in the multi-head attention mechanism of the second sub-layer. At the end of the decoder is a linear layer, which is a simple fully connected neural network that projects the vector generated by the decoder stack into a larger vector to become a logits vector.

[0058] The logits vector is converted into probabilities by the softmax layer, and the highest probability unit is selected to generate the word associated with it as the output at the current time step.

[0059] In order to well input the sign language video into the network for training, the model regards the sign language video as a spatio-temporal sequence containing many image patches. Each video frame is divided into multiple patches, and the semantics of each patch are captured by attention-weighting each patch with other patches in the video, and it can well capture the short-term dependencies between adjacent patches and the context dependencies of distant patches. The specific implementation is as follows: The spatial encoding part extracts all tokens of the entire sign language sequence to calculate the spatial attention, and then calculates the temporal attention for the tokens of the same spatial sequence (i.e., the patches divided from the same frame).

[0060] The sign language recognition method of the present invention first uses the patch operation to preprocess the sign language video frames to make their size and dimensions convenient for direct processing by the model, while reducing the computational complexity. Secondly, the context time dimension and dynamic spatial dimension features of the sign language video frames are respectively extracted and encoded by combining the spatio-temporal dual-channel encoder, and the dynamic features of the sign language video are fully extracted through the two dimensions. In addition, the encoded features are fused, and the decoder is used to predict and output the fused feature vector. Finally, the predicted sequence is aligned and recognized with the sign language text sequence. Through the present invention, the feature extraction of the video frames by the model can be improved, thereby improving the recognition rate of continuous sign language recognition.

[0061] In a complex background, the proposed dense segmentation network is used to filter out redundant backgrounds and segment gesture images. The located gesture regions are input into the gesture recognition network, and an improved algorithm is used for recognition. The present invention improves the segmentation performance of gesture images, thereby improving the recognition rate of gesture images.

[0062] The sign language recognition method of the present invention is a sign language recognition network improved based on Transformer. Through dual-channel encoding of time and space, it truly achieves the fusion of dynamic local features and long-term global features, enriching the feature expression. The present invention has strong robustness and can obtain a higher recognition rate in cases of multiple sign language users, multiple sentences, and multiple language inputs.

[0063] The present invention can also be implemented as follows: A continuous sign language recognition device based on Spatial-Temporal Transformer, including the following modules:

[0064] A sign language video acquisition module, connected to the preprocessing module, for extracting sign language video frames from a sign language video to obtain RGB sign language video frames;

[0065] A preprocessing module, respectively connected to the sign language video extraction module and the sign language recognition network training module, for performing patch operations on the color sign language video frames to provide serialized sign language video blocks for the Spatial-Temporal Transformer network;

[0066] A sign language recognition network training module, including an STT encoding unit, a decoding unit, a cross-entropy loss function, and a backpropagation unit, for performing feature extraction encoding and decoding prediction on sign language video frames; and

[0067] An output module, connected to the sign language recognition network training module, for outputting the final sign language recognition result.

[0068] Furthermore, the preprocessing module includes a patch operation unit and an embedding network; the embedding network includes a patch-embedding network and a positional-encoding network.

[0069] Furthermore, the STT encoding unit in the sign language recognition network training module includes multiple layers of encoders, and each layer of encoder includes a dual-channel encoder of time and space, a linear normalization layer, and a feed-forward neural network; the STT encoding unit uses the input frames output by the preprocessing module to perform feature extraction and encoding of dynamic spatial correlation and long-term time correlation through dual channels of time and space.

[0070] Further, the decoding part in the sign language recognition network training module includes multiple layers of decoders, and each layer of decoder includes three sub-layers; among them, the first sub-layer includes: a masked multi-head self-attention layer, a normalization layer, and a residual connection layer; the second sub-layer includes: a multi-head cross-attention layer, a normalization layer, and a residual connection layer; the third sub-layer includes: a feed-forward neural network, a normalization layer, and a residual connection layer.

[0071] The sign language recognition device of the present invention designs a patch operation in the preprocessing module to serialize the sign language video frames, reducing the computational complexity of the network and making it more convenient for the network to process. The dual-channel encoder in the sign language recognition training network can comprehensively obtain the dynamic spatial features and long-term temporal features of the sign language video, fuse rich action semantics and context information together, and obtain a more complete feature expression. Therefore, the Spatial-Temporal Transformer Network (STTN) combines global and high-level semantic features with local and detailed semantic features to filter out redundant information in the background, which helps to improve the recognition effect.

[0072] The present invention utilizes the acquisition of dual-channel sign language dynamic and context feature information in time and space to obtain more accurate sign language recognition results. The overall performance of the present invention is better than that of general mainstream algorithms and is more suitable for human-machine products. The beneficial effect of the improved sign language recognition network is that by comprehensively extracting visual features, it improves the network processing ability and is better than the convolutional-based sign language recognition method. Brief Description of the Drawings

[0073] Figure 1 is the structural block diagram of the continuous sign language recognition device of the present invention.

[0074] Figure 2 is the training flow block diagram of the sign language recognition network.

[0075] Figure 3 is the comparison diagram of the position encoding effect.

[0076] Figure 4 is the framework diagram of the STTN network.

[0077] Figure 5 is the framework diagram of the ST encoder.

[0078] Figure 6 is the framework diagram of the decoder.

[0079] Figure 7 is the generation diagram of the attention Q, K, V.

[0080] Figure 8 is the prediction output diagram at the decoder end.

[0081] Figure 9 It is an example diagram of the Chinese Sign Language dataset CSL100.

[0082] Figure 10 It is the training effect diagram of the STTN network on the RWTHPHOENIX-Weather-2014 (PHOENIX14) dataset.

[0083] Figure 11 It is an example diagram of the RWTHPHOENIX-Weather-2014 (PHOENIX14) dataset. Detailed implementation manners

[0084] Example 1: A continuous sign language recognition method based on Spatial-Temporal Transformer.

[0085] As Figure 1 shown, the sign language recognition method of the present invention includes the following steps:

[0086] Step S1: The original sign language video deletes redundant frames by a random deletion method to obtain a continuous sign language video sequence.

[0087] Input the RGB sign language video. The RGB sign language video input in the embodiment of the present invention is selected from the public datasets CSL100 and RWTHPHOENIX-Weather-2014 (PHOENIX14). The input RGB sign language video is the basis for subsequent training and validating the network model. The Spatial-Temporal Transformer network includes a vectorization module, an encoder, and a decoder. The vectorization module includes patch operations, Patch-embedding, and Positional Encoding operations; the encoder includes a temporal attention calculation mechanism, a spatial attention mechanism, and a feed-forward neural network; the decoder includes a self-attention calculation mechanism, a cross-attention mechanism, and a feed-forward neural network.

[0088] Step S2: Perform a chunking operation and vectorization processing on the obtained sign language video sequence to obtain a sign language sequence vector.

[0089] Preprocess the input image to make it reach a fixed dimension and perform a chunking operation. In this step, the number of preprocessed videos in the CSL100 dataset is 25,000, among which 20,000 videos are used as the training set and 5,000 videos are used as the validation set. The number of preprocessed videos in the RWTHPHOENIX-Weather-2014 (PHOENIX14) dataset is 6,841, among which 5,672 videos are used as the training set, 540 videos are used as the validation set, and 629 videos are used as the test set. Uniformly extract frames from the videos and randomly delete some. The remaining sign language videos are default adjusted (cropped, resized) to a size of 224×224 pixels, and further put into the patch module for chunking operation, which is default divided into chunks of size 16×16 pixels. To enrich the experiment, the sign language video frame sizes are designed as: 112×112 pixels, 224×224 pixels, 256×256 pixels; the sizes of the chunks are: 8×8 pixels, 16×16 pixels, 32×32 pixels.

[0090] To facilitate the processing of the sign language video sequence vector with the input (T-frame sign language video) dimension of f∈R B×T×N×D in the sign language video, each frame in the T-frame (B is the batch-size, H, W are the resolutions of the original image, and C is the number of channels) sign language video frames is reshaped into a 2D tile of dimension (h×w)×(p1×p2×C), where H = h×p l , W = w×p2. h×w is the number of tiles each frame is divided into, which directly affects the length of the input sequence (the choice of p can be used as an experimental comparison), and a constant hidden vector d model is used on all layers, and the tiles are flattened and projected onto d model = D in size, and this projection output is the patch embedding. At this time, the size of the feature map is B×T×N×D, N = h×w, and the vector after patch embedding is denoted as: X (p,t) .

[0091] After obtaining the feature map of dimension f0∈R B×T×N×D , Positional Encoding is also required. Because the self-attention of the original transformer does not contain position information, but there is very strong sequential information in the sign language video frames. To prevent the loss of the concept of the front and back frames and to facilitate the subsequent extraction of features in the time dimension, position information is added to the feature map. The positional encoding requires that each position has a unique positional encoding and the relationship between two positions can be modeled by the affine transformation between their positional encodings. And through experimental verification:

[0092]

[0093]

[0094] Formulas (1, 2) can exactly meet these two requirements, that is, Positional Encoding (PE), where pos corresponds to the position of the token in the sequence, the starting token position is denoted as 0, 2i and 2i + 1 represent the dimensions of the Positional Encoding, and d model is the dimension after position encoding, and the value range of i is [0, d model / 2), and the position encoding information is marked as Figure 3 are the effect diagrams before and after position encoding. The left figure is the effect without position encoding, where the dimension values corresponding to each position are the same, and the information of different positions cannot be distinguished. The right figure is the effect with position encoding, and it can be seen that the dimension values at each position are unique, so the information of each position can be marked.

[0095] Step S3: Use the encoder in the Spatial-Temporal Transformer network as the spatio-temporal encoder to extract features from the obtained sign language sequence vector to obtain a sign language sequence feature vector; the encoder is a dual-channel encoder for time and space.

[0096] Specifically, first construct the STTN network. The set STTN network is specifically designed for sign language video recognition close to daily life. As Figure 4 shown, the specific structure of the neural network consists of a video frame processing part, a text embedding part, an STT encoder, and a decoder. As Figure 5 shown, the STT encoding module structure is a dual-channel encoding structure of time attention and space attention. As Figure 6 shown, the structure of the decoder consists of two layers of multi-head attention layers, a feed-forward neural network, a linear connection layer, a softmax function, and multiple summation and normalization layers.

[0097] The encoder used in step S3 of the present invention is an encoder structure that takes both time and space into account, including a space attention module and a time attention module; the sign language video vector passed into the encoder enters the time attention module and the space attention module through two channels respectively, and then the features extracted by the time attention module and the space attention module are concatenated, using dynamic spatial correlation and long-term time correlation to improve the extraction and encoding of the sign language video frame features by the network.

[0098] Further, the encoding process of the encoder is as follows: Rearrange the dimensions of the output result of the vectorization module. First, place the time dimension t on the first dimension batch, and perform dynamic spatial correlation attention calculation on the n vector blocks in the spatial dimension; then place the spatial dimension n on the first dimension batch, and perform long-term temporal correlation attention calculation on the t-frame sequence in the time dimension; then fuse the temporal attention calculation result and the spatial attention calculation result; finally, output after passing through a linear normalization layer and a feed-forward neural network.

[0099] Further, the temporal and spatial attention calculation processes of the encoder are as follows:

[0100] (1) The Spatial Self-Attention Block only performs MSA calculation on different tokens in the same frame. The attention value calculation of the query Q vector in the spatial dimension is as shown in formula (3) below.

[0101]

[0102]

[0103] Among them, space refers to performing attention calculation in the spatial dimension, time refers to performing attention calculation in the time dimension, softmax refers to the activation function, l refers to the l-th layer, a refers to the a-th attention head, p refers to the p-th block in each frame, and t refers to the t-th frame. Dh = D / A refers to the dimension value of the corresponding attention head, D is the dimension value of the vector, A is the total number of heads, q refers to the query vector, and k is the weight matrix corresponding to q.

[0104] (3) The Temporal Self-Attention Block only performs MSA calculation on tokens at the same position in different frames. The attention calculation of the query Q vector in the time dimension is as shown in formula (4). After calculating the attention in the time and spatial dimensions respectively, perform cat concatenation:

[0105] The specific operation method of step S3 of the present invention is:

[0106] S3-1. Use the training data in step S2 (here only use the preprocessed training set) as the input video of step 3. First perform Embedding and Positional Encoding operations on the input video. These two operations have the same dimension. The Positional Encoding calculation formula is as follows:

[0107]

[0108]

[0109] where: pos corresponds to the position of the token in the sequence, the starting token position is denoted as 0, 2i and 2i + 1 represent the dimensions of the Positional Encoding, and the value range of i is [0, d model / 2], and mark the position encoding information as

[0110] S3-2. Use the vectors after Embedding and Positional Encoding as the input P of the STT encoding module. After the vector enters the STTN encoding module, first, use the dimension transformation operation to put the number of frames T of the vector P to the first dimension batch-size, and perform spatial attention calculation on the vector P0 after dimension transformation by the spatial encoding module. The calculated vector is denoted as Z0; then, use the dimension transformation operation to put the number of blocks N of the vector P to the first dimension batch-size, and perform temporal attention calculation on the vector P1 after dimension transformation by the temporal encoding module. The calculated vector is denoted as Z1; where the temporal and spatial attention calculation methods are the same, both are Self Attention calculations, and the three matrices W Q , W K , W V , use these three matrices to perform three linear transformations on all P0 / P1 vectors respectively. Thus, all vectors derive three new vectors q t , k t , v t . Concatenate all the vectors q t into a large matrix, denoted as the query matrix Q, concatenate all the vectors k t into a large matrix, denoted as the key matrix K, and concatenate all the vectors v t into a large matrix, denoted as the value matrix V (see the "query", "key", and "value" matrices in Figure 7 respectively).

[0111] The calculation formulas are as follows:

[0112] Q = Linear(X) = XW Q (8)

[0113] K = Linear(X) = XW K (9)

[0114] V = Linear(X) = XW V (10)

[0115] X attention = Self Attention(Q, K, V) (11)

[0116] Among them, X refers to the input sequence, and W Q , W K , W V are three matrices.

[0117] After that, a feature concatenation operation is performed on Z0 and Z1. After the Layer Normalization and FeedForward operations, the output formula of the encoder is as follows:

[0118] X attention = Layer Norm(X attention ) (12)

[0119] Among them, FeedForward is a two-layer linear mapping activated by an activation function. The activation function ReLU is selected.

[0120] X hidden = Layer(ReLU(Linear(X attention ))) (13)

[0121] Among them, X hidden ∈R batch-size*seq_len*embed_dim .

[0122] S3-3. For the decoder, similar to the encoder, three parts of operations are required in sequence: ① Masked Multi-Head Self-Attention, ② Multi-Head Encoder-Decoder Attention, ③ FeedForward Network. Similarly, each part of the operation is followed by a Layer Normalization. In order to recover more detailed features during the decoding process, the decoder includes two types of attention calculations, namely the self-attention calculation between sign language texts and the cross-attention calculation between the attention output and the encoder output. The calculation method is similar to the multi-head attention in the encoder part, but with an additional masked operation. Specifically, in the traditional Seq2Seq, the RNN model is used in the decoder. Therefore, during the training process, when inputting the word at time t, the model can never see the words at future times. Because the recurrent neural network is time-driven, only when the operation at time t is completed can the word at time t+1 be seen. However, the Transformer Decoder abandons the RNN and uses Self-Attention instead. As a result, a problem arises. During the training process, the entire ground truth is exposed in the decoder, which is obviously incorrect. Some processing needs to be performed on the input of the decoder, and this processing is called Mask.

[0123] The overall decoding of the decoder includes the following steps:

[0124] S3-3-1. First, perform Word-Embedding operation on the sign language text, map it to a D-dimensional vector through Matrix, denoted as matX, and then perform Positional-Encoding operation to encode different position information for different words in the sign language text, with the same dimension D, denoted as matP. At this time, the shapes of matX and matP are the same, and add them to get the decoder input matDec = matP + matX.

[0125] S3-3-2. The three inputs of Multi-head attention during self-attention calculation are Q, K, and V respectively. At this time, Q = K = V = matDec; first, perform linear transformation on V, K, and Q respectively, that is, input them into three single-layer neural network layers respectively, and select relu as the activation function to output new V, K, and Q (the shapes of the three are the same as the original shape, that is, the output dimension is the same as the input dimension during linear transformation); then slice Q into num_heads (assumed to be 8) segments on the last dimension, and then concatenate the sliced matrix on the axis = 0 dimension; perform the same operations on V and K as on Q; the matrices after the operation are denoted as Q, K, and V; multiply the Q matrix by the transpose of K (for the last two dimensions), and the generated result is denoted as outputs.

[0126] S3-3-3. Update outputs by scaling once; this matrix multiplication is to calculate the correlation between words, and slicing into multiple num_heads for calculation is to achieve the calculation of the deep correlation between words; perform softmax operation on outputs to update outputs, that is, outputs = softmax(outputs).

[0127] S3-3-4. Perform cross-attention calculation, Q is the output of the encoder, K = V = matDec, and the calculation process is the same as that in the sub-attention stage. After that is the Add&Norm layer, and the operation of this layer is similar to ResNet, which superimposes the initial input and its corresponding output once, that is, outputs = outputs + Q, making the network effectively superimposed to avoid gradient disappearance. After passing through the Add&Norm layer and Feed Forward, normalize and linearly transform outputs; after completing Nx times, project the vector generated by the decoder stack to a larger vector through a linear layer (which is a simple fully connected neural network) to become a logits vector, and then convert it to a probability by the softmax layer, and select the highest probability unit to generate the associated word as the output at the current time step (Figure 8 ) The decoding part of the model is completed.

[0128] Step S4: Perform feature decoding on the obtained sign language sequence feature vectors.

[0129] In step S4 of the present invention, feature decoding is performed using a decoding network. The decoding network includes a multi-head attention mechanism, a multi-head cross-attention mechanism, and a feed-forward neural network. First, self-attention calculation is performed on the sign language text, and then cross-attention calculation is performed with the feature vectors generated by the encoding.

[0130] The present invention proposes a strategy of patch operation + STTN network for sign language recognition. The patch operation can greatly reduce the computational complexity problem brought by long sequences and reduce the network burden. The STTN network can obtain dynamic spatial features while processing long-term context semantic features, thereby improving the accuracy of sign language recognition. The STTN network in step S4 mainly consists of two parts: an STT encoder and a decoder.

[0131] The decoder used in step S4 of the present invention includes three sub-layers. The first sub-layer includes a multi-head self-attention layer with a mask, a normalization layer, and a residual connection layer. The second sub-layer includes a multi-head cross-attention layer, a normalization layer, and a residual connection layer. The third sub-layer includes a feed-forward neural network, a normalization layer, and a residual connection layer.

[0132] The calculation formula for the three layers is:

[0133]

[0134] Among them, Q i-1 refers to calculating the output of the previous layer, F refers to the output of the encoder, refers to F with the position encoding operation added.

[0135] Such as Figure 5 shown, the input of the encoder network in step S4 is a sequence of T×N×D obtained by dividing the 224×224×3 RGB images of T frames into blocks, where N is the number of divided blocks and D is the vector dimension, default set to 512; the encoding part consists of a temporal encoder and a spatial encoder module. Among them, the temporal encoder encodes and extracts features of the long-term correlation in the temporal dimension for all T frames, and the spatial encoder encodes and extracts features of the dynamic correlation in the spatial dimension for all N blocks of each frame. After temporal encoding and spatial encoding, they are fused, and then after passing through the Add&Norm, Feed Forward, and Add&Norm operations in sequence, the encoding is completed.

[0136] Such as Figure 6As shown, the input to the decoder network in step S4 is the output of the decoder and the sign language text. The decoder includes two attention calculations, namely the self-attention calculation between sign language texts and the cross-attention calculation between the attention output and the encoder output. The calculation method is similar to the multi-head attention in the encoder part, but there is an additional masked operation. Because in the decoding part, decoding is performed sequentially from left to right. When the first character is decoded, the first character can only calculate the correlation with the first character. When the second character is decoded, only the correlation between the second character and the first and second characters can be calculated. Therefore, a masked operation is required. Specifically, the decoding process is as follows:

[0137] S4-1. First, perform a Word-Embedding operation on the sign language text, map it to a D-dimensional vector through a Matrix, denoted as matX, and then perform a Positional-Encoding operation to encode different position information for different words in the sign language text, with the same dimension D, denoted as matP. At this time, the shapes of matX and matP are the same, and they are added together to obtain the decoder input matDec = matP + matX.

[0138] S4-2. When calculating self-attention, the three inputs of the Multi-head attention are Q, K, and V. At this time, Q = K = V = matDec; perform a linear transformation on V, K, and Q respectively, that is, input them into three single-layer neural network layers respectively, and select the relu activation function to output new V, K, and Q (the shapes of the three are the same as the original shape, that is, the output dimension is the same as the input dimension during the linear transformation).

[0139] S4-3. Split Q into num_heads (assumed to be 8) segments on the last dimension, and then concatenate the split matrices on the axis = 0 dimension; perform the same operations on V and K as on Q; the matrices after the operation are denoted as Q_, K_, and V_; multiply the Q_ matrix by the transpose of K- (for the last two dimensions), and denote the generated result as outputs, and then update outputs by scale once; this matrix multiplication is to calculate the correlation between words, and splitting into multiple num_heads for calculation is to achieve the calculation of the deep correlation between words.

[0140] S4-4. Perform a softmax operation on the outputs and update the outputs, i.e., outputs = softmax(outputs); multiply the latest outputs (i.e., the correlation matrix of K and Q) by V, and update its value to outputs; finally, split the outputs into num_heads segments along the axis = 0 dimension, and then merge them along the axis = 2 dimension to restore the original dimension of Q.

[0141] S4-5. In the cross-attention stage, Q is the output of the encoder, K = V = matDec, and the calculation process is the same as that in the sub-attention stage. After that is the Add&Norm layer, and the operation of this layer is similar to ResNet, which superimposes the initial input and its corresponding output once, i.e., outputs = outputs + Q, enabling the network to effectively superimpose and avoid gradient disappearance;

[0142] S4-6. Perform a normalization correction once. Calculate the mean and variance of the last dimension of outputs, update the value of outputs by subtracting the mean and dividing by the variance + spsilon, and then variable gamma × outputs + variable beta. The next layer is FeedForward, which is two convolutional operations. Perform the first convolutional operation on outputs, and update the result to outputs (the convolutional kernel is 1×1, and each convolutional operation is performed on the vector elements corresponding to a word. The number of convolutional kernels is the length of the last-dimensional vector, that is, the dimensionality of the vector corresponding to a word); perform the second convolutional operation on the latest outputs, and the convolutional kernel is still 1×1, and the number of convolutional kernels is N.

[0143] S4-7. Then perform the Add&Norm layer, which is the same as the Add&Norm in e. After the above operations, the shape of the latest output is the same as that of matEnc at this time; let matEnc = outputs, complete one loop, and then return to S4-3 to start the second loop. A total of Nx (user-defined; the structure of each loop is the same, but the corresponding parameters are different, that is, they are independently trained) loops are performed. After completing Nx times, the decoding part of the model is completed.

[0144] Step S5: Perform sequence prediction on the decoded features to obtain the predicted sequence of the sign language video.

[0145] In step S5 of the present invention, position encoding is also required before entering the decoding process to add position information to the sign language text. Subsequently, it will first pass through a masked self-attention layer. The purpose of using the mask is to solve the problem of information leakage during the training process encountered in the decoding process, that is, to avoid model cheating and the problem of inconsistent model architectures during the prediction process. This is because using sequence masking in the prediction stage can keep the prediction results of repeated sentences the same, which not only conforms to the rules but also can be incrementally updated. At the same time, it can also be consistent with the model architecture of the training and the way of forward propagation. In addition, there is only one output at the encoder end, and each decoding layer passed to the decoder part acts as the K and V of the multi-head attention mechanism in the second sub-layer. At the end of the decoder is a linear layer, which is a simple fully connected neural network that projects the vector generated by the decoder stack into a larger vector to become the logits vector.

[0146] The logits vector is converted into probabilities by the softmax layer, and the highest probability unit is selected to generate the word associated with it as the output at the current time step.

[0147] Step S6: Calculate the WER for the predicted sequence of the sign language video and the sign language text sequence obtained.

[0148] Step S7: Perform network-level training on steps S3, S4, S5, and S6, and output the final sign language recognition result.

[0149] Embodiment 2: A continuous sign language recognition device based on Spatial-Temporal Transformer.

[0150] As Figures 1 to 3 shown, the continuous sign language recognition device based on Spatial-Temporal Transformer of the present invention includes a sign language video acquisition module, a preprocessing module, a sign language recognition network training module, and an output module, etc. Among them, the sign language video acquisition module is connected to the preprocessing module and is used to extract sign language video frames from the sign language video to obtain RGB sign language video frames. The preprocessing module is respectively connected to the sign language video extraction module and the sign language recognition network training module and is used to perform patch operations on the color sign language video frames to provide serialized sign language video blocks for the Spatial-Temporal Transformer network. The sign language recognition network training module includes an STT encoding part, a decoding part, a cross-entropy loss function, and a backpropagation part, and is used to perform feature extraction encoding and decoding prediction on the sign language video frames. The output module is connected to the sign language recognition network training module and outputs the final sign language recognition result.

[0151] The preprocessing module includes a patch operation unit and an embedding network; the embedding network includes a patch-embedding network and a positional-encoding network. The STT encoding unit in the sign language recognition network training module includes multiple layers of encoders, and each layer of encoder includes a two-channel encoder for time and space, a layer normalization layer, and a feed-forward neural network; the STT encoding unit uses the input frames output by the preprocessing module to perform feature extraction and encoding of dynamic spatial correlation and long-term temporal correlation through the two channels of time and space. The decoding unit in the sign language recognition network training module includes multiple layers of decoders, and each layer of decoder includes three sub-layers; among them, the first sub-layer includes: a masked multi-head self-attention layer, a normalization layer, and a residual connection layer; the second sub-layer includes: a multi-head cross-attention layer, a normalization layer, and a residual connection layer; the third sub-layer includes: a feed-forward neural network, a normalization layer, and a residual connection layer.

[0152] To further prove the effectiveness of the STTN model proposed by the present invention, the embodiments of the present invention conduct sign language recognition experiments on the CSL100 and RWTHPHOENIX-Weather-2014 (PHOENIX14) public datasets and compare them with other deep learning-based recognition algorithms. The experimental comparison results of CSL100 are shown in Table 1.

[0153] Table 1: Recognition rates on the CSL100 dataset

[0154]

[0155] As can be seen from Table 1, the recognition error rate of the STTN network proposed by the present invention on the CSL100 dataset is reduced to 1.2%, which is improved compared with other algorithms. Therefore, the recognition algorithm proposed by the present invention can greatly improve the accuracy of sign language recognition.

[0156] The experimental comparison results of the present invention on the RWTHPHOENIX-Weather-2014 (PHOENIX14) public dataset are shown in Table 2.

[0157] Table 2: Comparison results between the present invention and deep learning methods on the RWTHPHOENIX-Weather-2014 (PHOENIX14) dataset

[0158]

[0159] As can be seen from Table 2, compared with the previous method using convolution, the sign language recognition method of the present invention has advantages, and it is compared with the multi-network fusion method. Compared with pure convolution, the multi-network fusion method helps the network notice more information. However, the sign language recognition method of the present invention extracts information hierarchically at the time and space levels and uses a transformer that can accurately record context information. The present invention can extract more sufficient information, which shows that the sign language recognition method of the present invention is superior to the comparative algorithms in all aspects.

[0160] As Figure 9 shown, to better understand the learning process, a random data sample was selected from the RWTH-PHOENIX-Weather multi-signer2014 dataset, and the figure shows the continuous action postures of signers in a continuous sentence.

[0161] As Figure 10 shown, the visualized training effect (the WER change curve during training, testing, and validation) shows that the WER drops faster during the training process, and the curves of testing and validation basically remain unchanged after 7 epochs, and the curves tend to be more stable, and the best result is obtained at the 13th epoch.

[0162] As Figure 11 shown, in an example of the sentence recognition result from the Chinese Sign Language dataset (CSL100), five predicted sentences are shown, and the error rate decreases from prediction 1 to 5. The first row represents the input frame sequence. The boxes preceded by S, D, etc. in the text indicate incorrect predictions. The operations of deletion, substitution, and insertion are represented by "D", "S", and "I" respectively.

Claims

1. A continuous sign language recognition method, characterized in that, It includes the following steps: S1. The original sign language video deletes redundant frames through a random deletion method to obtain a continuous sign language video sequence; S2. The obtained sign language video sequence is subjected to block operation and vectorization processing to obtain a sign language sequence vector; S3. Using the encoder in the Spatial-Temporal Transformer network as a spatio-temporal encoder, feature extraction is performed on the obtained sign language sequence vector to obtain a sign language sequence feature vector; the encoder is a dual-channel encoder for time and space; S4. Feature decoding is performed on the obtained sign language sequence feature vector; S5. Sequence prediction is performed on the decoded features to obtain a predicted sequence of the sign language video; S6. Calculate the WER between the obtained predicted sequence of the sign language video and the sign language text sequence; S7. Perform network-level training on steps S3, S4, S5, and S6, and output the final sign language recognition result; The encoder used in step S3 is an encoder structure that takes both time and space into account, including a spatial attention module and a temporal attention module; the sign language video vector input into the encoder enters the temporal attention module and the spatial attention module through two channels respectively, and then the features extracted by the temporal attention module and the spatial attention module are concatenated; The specific operation method of step S3 includes the following steps: S3-1. Take the training data in step S2 as the input video, and first perform Embedding and PositionalEncoding operations; S3-2. Take the vectors after Embedding and Positional Encoding as the input P of the STT encoding module. After the vector enters the STTN encoding module, first use the dimension transformation operation to put the number of frames T of the vector P into the first dimension batch-size, and the spatial encoding module performs spatial attention calculation on the vector P0 after dimension transformation, and the calculated vector is denoted as Z0; then use the dimension transformation operation to put the number of blocks N of the vector P into the first dimension batch-size, and the temporal encoding module performs temporal attention calculation on the vector P1 after dimension transformation, and the calculated vector is denoted as Z1; perform feature concatenation operation on the vector Z0 and the vector Z1, and perform the output of the encoder after Layer Normalization and FeedForward operations; S3-3. Perform three parts of operations on the decoder in sequence: ① Masked Multi-Head Self-Attention, ② Multi-Head Encoder-Decoder Attention, ③ FeedForward Network, and each part of the operation is followed by a Layer Normalization; the attention calculation of the decoder includes self-attention calculation between sign language texts and cross-attention calculation between the attention output and the encoder output; The decoder used in step S4 includes three sub-layers. The first sub-layer includes a multi-head self-attention layer with a mask, a normalization layer, and a residual connection layer. The second sub-layer includes a multi-head cross-attention layer, a normalization layer, and a residual connection layer. The third sub-layer contains a feedforward neural network, a normalization layer, and a residual connection layer.

2. The continuous sign language recognition method according to claim 1, characterized in that In step S2, each frame in the sign language video frames of T frames is reshaped into a 2D tile of (h×w)×(p1×p2×C) dimensions, where H = h×p1, W = w×p2; h×w is the number of tiles each frame is divided into; a constant latent vector d is used on all layers model , and the flattened projection of the tile is mapped to d model = the size of D, and this projection output is the patch embedding; at this time, the size of the feature map is B×T×N×D, N = h×w, and the vector after patch embedding is denoted as: X (p,t) .

3. The continuous sign language recognition method according to claim 1, characterized in that In step S3-1, PositionalEncoding is calculated based on the following two formulas: where: pos corresponds to the position of the token in the sequence, the starting token position is denoted as 0, 2i and 2i + 1 represent the dimensions of PositionalEncoding, and the value range of i is [0, d model / 2], and mark the positional encoding information as The calculation methods of both the temporal attention and the spatial attention in step S3-2 are Self Attention calculations; the three matrices W Q , W K , W V perform three linear transformations on all the P0 / P1 vectors respectively, and all the vectors derive three new vectors q t , k t , v t ; all the vectors q t are assembled into a query matrix Q, all the vectors k t are assembled into a key matrix K, and all the vectors v t are assembled into a value matrix V; The calculation formula is as follows: Q = Linear(X) = XW Q (8) K = Linear(X) = XW K (9) V = Linear(X) = XW V (10) X attention = Self Attention(Q, K, V) (11) where X is the input sequence, W Q , W K , W V are three matrices; The feature concatenation operation is performed on vector Z0 and vector Z1, and the encoder output is obtained after Layer Normalization and FeedForward operations. The formula is as follows: X attention = Layer Norm(X attention ) (12) Among them, FeedForward is a two-layer linear mapping and activated by an activation function, and the activation function is ReLU; X hidden = Layer(ReLU(Linear(X attention ))) (13) where X hidden ∈R batch_size*seq_len*embed_dim .

4. The continuous sign language recognition method according to claim 3, characterized in that The overall decoding operation of the decoder includes the following steps: S3-3-1. First, perform Word-Embedding operation on the sign language text, map it into a D-dimensional vector through Matrix, denoted as matX, and then perform Positional-Encoding operation to encode different position information of different words in the sign language text, with a dimension of D, denoted as matP. Add them together to get the decoder input: matDec=matP+matX; S3-3-2. When calculating the self-attention between sign language texts, the three inputs of Multi-head attention are Q, K, and V, respectively. Linear transformations are performed on V, K, and Q respectively. Then Q is divided into num_heads segments on the last dimension, and the divided matrices are concat-linked on axis=0. The same operation is performed on V and K. The matrices after the operation are recorded as Q_, K_, V_. The Q_ matrix is ​​multiplied by the transpose of K_, and the result is recorded as outputs. S3-3-3, scale the outputs once and update them to outputs; perform softmax operation on the outputs and update the outputs; S3-3-4. When calculating the cross attention between the attention output and the encoder output, Q is the output of the encoder, K=V=matDec, and the calculation process is the same as the calculation of the self-attention between the sign language text; the operation in the Add&Norm layer is the same as ResNet, and the initial input is superimposed with its corresponding output once, that is, outputs=outputs+Q, so that the network can be effectively superimposed to avoid gradient disappearance. After the Add&Norm layer and Feed Forward, the outputs are normalized and linearly transformed; after completing Nx times, the vector generated by the decoder stack is projected to a larger vector through the linear layer to become a logits vector, and then converted into a probability by the softmax layer, and the highest probability unit is selected to generate the word associated with it as the output of the current time step. At this point, the decoding of the model is completed.

5. The continuous sign language recognition method according to claim 4, characterized in that, Step S4 is to perform feature decoding using a decoding network, which includes a multi-head attention mechanism, a multi-head cross-attention mechanism, and a feed-forward neural network; first, self-attention calculation is performed on the sign language text, and then cross-attention calculation is performed with the feature vector generated by encoding.

6. The continuous sign language recognition method according to claim 1, characterized in that The calculation formulas for the three sub-layers in step S4 are as follows: Among them, Q i-1 refers to the output of the previous layer, F refers to the output of the encoder, and refers to F with the positional encoding operation added.

7. A continuous sign language recognition device, characterized in that, This device is used to implement the method described in any one of claims 1 to 6, and the device includes: A sign language video acquisition module, connected to the preprocessing module, for extracting sign language video frames from the sign language video to obtain RGB sign language video frames; A preprocessing module, connected to the sign language video extraction module and the sign language recognition network training module respectively, for performing patch operations on the color sign language video frames to provide serialized sign language video blocks for the Spatial-Temporal Transformer network; the sign language recognition network training module includes an STT encoding part, a decoding part, a cross-entropy loss function, and a backpropagation part, for performing feature extraction encoding and decoding prediction on the sign language video frames; and An output module, connected to the sign language recognition network training module, for outputting the final sign language recognition result; The preprocessing module includes a patch operation part and an embedding network; the embedding network includes a patch-embedding network and a positional-encoding network; The STT encoding part in the sign language recognition network training module includes multiple layers of encoders, and each layer of encoder includes a two-channel encoder for time and space, a layer normalization layer, and a feed-forward neural network; the STT encoding part uses the input frames output by the preprocessing module to perform feature extraction and encoding of dynamic spatial correlation and long-term temporal correlation through the two channels of time and space; The decoding part in the sign language recognition network training module includes multiple layers of decoders, and each layer of decoder includes three sub-layers; among them, the first sub-layer includes: a masked multi-head self-attention layer, a normalization layer, and a residual connection layer; the second sub-layer includes: a multi-head cross-attention layer, a normalization layer, and a residual connection layer; the third sub-layer includes: A feed-forward neural network, a normalization layer, and a residual connection layer.

Citation Information

Patent Citations

  • Continuous sign language recognition method

    CN111339837A

  • Sign language video generation method based on improved Transform model

    CN115393948A