Continuous sign language recognition method based on layered space-time enhancement

By integrating the timing causal module and alignment module in the ResNet34 network, the spatiotemporal challenges of existing methods when dealing with sign language recognition are solved, and more efficient and accurate continuous sign language recognition is achieved, improving recognition performance and robustness.

CN120340124APending Publication Date: 2025-07-18ZHEJIANG UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510384526.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing continuous sign language recognition methods fail to effectively and independently process the spatial appearance information and temporal motion information of sign language, making it difficult to balance the accuracy and calculation cost, especially in long video recognition.

Method used

The hierarchical spatiotemporal enhancement method based on ResNet34 network is adopted, combining the timing causal module and the alignment module, and through multi-stage feature extraction and hierarchical alignment supervision, the spatial and temporal features of sign language videos are captured, and future and past information are aggregated to improve recognition accuracy.

Benefits of technology

It significantly improves the accuracy and robustness of continuous sign language recognition, solves the problem of underfitting in spatial and temporal contexts, and achieves efficient and accurate recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340124A_ABST
    Figure CN120340124A_ABST
Patent Text Reader

Abstract

The invention discloses a continuous sign language recognition method based on layered space-time enhancement. The method comprises the following steps: acquiring a sign language video; and inputting the sign language video into the trained sign language recognition model to obtain a first recognition result and a second recognition result, and taking the first recognition result as a final sign language recognition result. According to the continuous sign language recognition method based on hierarchical space-time enhancement, multi-stage output of a ResNet34 network is captured through an alignment module, and additional hierarchical alignment supervision is provided; according to the method, the alignment module and the timing causal module are integrated into the ResNet34 network, future and past information is aggregated through the timing causal module, so that more accurate vocabulary boundary perception is achieved, the alignment module and the timing causal module are integrated into the ResNet34 network, good balance between accuracy and calculation cost is achieved, an efficient and accurate solution is provided for a continuous sign language recognition task, and the continuous sign language recognition efficiency is improved. The recognition performance and robustness are remarkably improved, and the problem of under-fitting of the ResNet34 network in space and time contexts is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of sign language recognition, and particularly relates to a continuous sign language recognition method based on hierarchical spatio-temporal enhancement. Background Art

[0002] As the main communication method for the hearing-impaired group, sign language poses great challenges for ordinary people to understand due to its specific grammar and vocabulary system, thus forming significant communication barriers in aspects such as education, medical care, and social interaction. To solve this problem, sign language recognition promotes efficient communication between the hearing-impaired population and the general population by converting visual gestures into understandable text. Video-based sign language recognition tasks can be roughly divided into isolated sign language recognition (ISLR) and continuous sign language recognition (CSLR). Isolated sign language recognition focuses on word-level gesture recognition, where each short video segment corresponds to a specific vocabulary annotation; while continuous sign language recognition deals with sentence-level annotation information and usually adopts a weakly supervised learning strategy during training to achieve sequence recognition modeling. Early CSLR methods mainly relied on manually designed features to represent sign language information and used a recognition system based on the hidden Markov model to segment sign language videos to cope with the limited amount of data.

[0003] Benefiting from the powerful feature representation ability of deep neural networks, the success of convolutional neural networks (CNNs) and recurrent neural networks (RNNs) in feature representation and temporal modeling has attracted increasing attention and provided powerful tools for continuous sign language recognition. Methods based on convolutional neural networks usually solve the problem of continuous sign language recognition through a backbone framework that combines a feature extractor and an alignment module. The feature extractor consists of Conv3D or Conv2D+Conv1D and is used to extract discriminative visual features from the input video. Research shows that because the number of frames in sign language videos is long and the computational cost is high, compared with Conv3D, the structure of Conv2D+Conv1D performs better in terms of performance and efficiency. Therefore, existing methods tend to use Conv2D+Conv1D as the main structure. In this framework, the feature extractor is used to capture the visual features of the input video, while the alignment module focuses on modeling long-term context relationships, achieving sequence learning, and aligning the predicted distribution with the target annotation sequence. The alignment module is usually based on bidirectional long short-term memory networks (BiLSTMs) or Transformers to efficiently model long-term and short-term context features and complete sequence alignment. In contrast, early methods used hidden Markov models (HMMs) to align video frames to text annotations, while modern methods generally adopt a more efficient connectionist temporal classification (CTC) loss as the supervision mechanism for the entire CSLR framework.

[0004] Although the above methods have achieved certain improvements, they are not specifically optimized for the independent modeling of spatial appearance information and temporal motion information. Specifically, the spatial features of sign language mainly include the position and shape of the hands, facial expressions, and body postures, which are crucial for understanding and recognizing sign language. From a temporal perspective, sign language has variable-length gestures and complex temporal dependencies, and the ambiguous distinction of sign language boundaries directly affects the accuracy of recognition. Existing methods fail to independently handle these two aspects, and due to the complexity of long videos, it is difficult to balance accuracy and computational cost. Therefore, addressing spatio-temporal challenges remains a key issue in CSLR. Summary of the Invention

[0005] An object of the present invention is to propose a continuous sign language recognition method based on hierarchical spatio-temporal enhancement to solve the problems raised in the background technology.

[0006] To achieve the above object, the technical solutions adopted by the present invention are as follows:

[0007] A continuous sign language recognition method based on hierarchical spatio-temporal enhancement proposed by the present invention includes:

[0008] Obtain a sign language video;

[0009] Input the sign language video into a trained sign language recognition model to obtain a first recognition result and a second recognition result, and use the first recognition result as the final result of sign language recognition;

[0010] The sign language recognition model includes a ResNet34 network, three temporal causal modules, and an alignment module. The ResNet34 network includes a first stage, a second stage, a third stage, and a fourth stage arranged in sequence. The three temporal causal modules are respectively arranged between the first stage and the second stage, between the second stage and the third stage, and between the third stage and the fourth stage of the ResNet34 network. When the sign language video is input into the trained sign language recognition model, it is first input into the first stage and forward-inferred. The outputs of the four stages are all input into the alignment module to obtain the second recognition result, and the output of the fourth stage is then sequentially passed through the fourth convolutional layer, bidirectional long short-term memory network, and the first fully connected layer in the ResNet34 network to obtain the first recognition result.

[0011] Preferably, the input features of each of the temporal causal modules are first reshaped in shape, and then zero-padding with a preset number is performed on both sides of the reshaped input features along the time dimension to obtain two padded features;

[0012] The two padded features respectively pass through an atrous grouped convolutional layer to obtain corresponding first features, and the two first features are concatenated along the channel dimension to obtain a second feature;

[0013] The second feature passes through the first convolutional layer to obtain a third feature;

[0014] The third feature is fused with the input of the temporal causal module to obtain the output of the temporal causal module.

[0015] Preferably, the alignment module includes a spatial alignment module and a temporal alignment module connected in sequence, wherein the spatial alignment module includes four parallel spatial branches, and the first spatial branch, the second spatial branch, and the third spatial branch all include a second convolutional layer, a pooling layer, and a third fully connected layer connected in sequence, the fourth spatial branch includes a pooling layer and a third fully connected layer connected in sequence, and the temporal alignment module includes four parallel temporal branches, and each temporal branch includes a temporal convolutional network and a second fully connected layer connected in sequence;

[0016] The outputs of the first stage, the second stage, the third stage, and the fourth stage are sequentially and one-to-one corresponding as the inputs of the first spatial branch, the second spatial branch, the third spatial branch, and the fourth spatial branch respectively, and four fourth features are obtained, and the spatial dimensions and the number of channels of the four fourth features are the same;

[0017] The four fourth features are then respectively and one-to-one input into the four temporal branches, and third recognition results corresponding to the four fourth features are respectively obtained, and the four third recognition results together constitute the second recognition result.

[0018] Preferably, the temporal convolutional network includes two TCN BLOCK modules connected in sequence, and each TCN BLOCK module includes a third convolutional layer, an activation layer, a batch normalization layer, and a max pooling layer connected in sequence.

[0019] Preferably, the calculation formula of the loss function L of the sign language recognition model is as follows:

[0020]

[0021] wherein, L1 is the loss calculated by the first fully connected layer for the first recognition result it outputs and the true label, L i is the loss calculated by the second fully connected layer for the third recognition result of the i-th stage it outputs and the true label, n is the total number of stages in the ResNet34 network, which is 4, and both L1 and L i adopt the CTC loss.

[0022] Preferably, the formula for fusing the third feature and the input of the temporal causal module is as follows:

[0023]

[0024] wherein, denotes the output obtained by passing the output of the $i$-th stage through the corresponding temporal causal module, where $i\in\{1,2,3\}$, and $A$ i denotes the output of the $i$-th stage, i.e., the input to the corresponding temporal causal module, denotes the third feature corresponding to the output of the $i$-th stage passing through the first convolutional layer in the corresponding temporal causal module, $\alpha$ i denotes the weight of the third feature corresponding to the output of the $i$-th stage passing through the first convolutional layer in the corresponding temporal causal module.

[0025] Preferably, the shape reshaping merges the width dimension and the height dimension in the input features of each of the temporal causal modules.

[0026] Preferably, the dilated grouped convolutional layer first performs grouped convolution on the input padded features, divides the number of channels into multiple groups, and then performs dilated convolution on each group.

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0028] This continuous sign language recognition method based on hierarchical spatio-temporal enhancement captures the multi-stage outputs of the ResNet34 network through the alignment module and provides additional hierarchical alignment supervision; and aggregates future and past information through the temporal causal module, thereby achieving more accurate lexical boundary perception. By integrating the alignment module and the temporal causal module into the ResNet34 network, a good balance is achieved between accuracy and computational cost, providing an efficient and accurate solution for the continuous sign language recognition task, significantly improving the recognition performance and robustness, and solving the underfitting problem of the ResNet34 network in spatial and temporal contexts. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a schematic diagram of the sign language recognition model result of the continuous sign language recognition method based on hierarchical spatio-temporal enhancement of the present invention;

[0030] Figure 2 is a block diagram of the module for the filling process of the present invention;

[0031] Figure 3 is a schematic structural diagram of the temporal causal module of the present invention;

[0032] Figure 4 is a schematic structural diagram of the alignment module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0034] As shown Figures 1 - 4 in the figure, a continuous sign language recognition method based on hierarchical spatio-temporal enhancement is provided, including:

[0035] Step 1, obtain a sign language video;

[0036] In this embodiment, the sign language video is an RGB image including T frames, and the sign language video is represented as where x t is the image of the t-th frame, and R T×3×H×W is a matrix of dimension T×3×H×W (specifically R T×3×224×224 ), with 3 channels, and W and H are the width and height of the image respectively.

[0037] Step 2, input the sign language video into the trained sign language recognition model to obtain a first recognition result and a second recognition result, and use the first recognition result as the final result of sign language recognition;

[0038] As shown Figure 1 in the figure, the sign language recognition model includes a ResNet34 network, three temporal causal modules, and an alignment module. The ResNet34 network includes a first stage, a second stage, a third stage, and a fourth stage arranged in sequence. The three temporal causal modules are respectively arranged between the first stage and the second stage, between the second stage and the third stage, and between the third stage and the fourth stage of the ResNet34 network. When the sign language video is input into the trained sign language recognition model, it is first input into the first stage and forward-inferred. The outputs of the four stages are all input into the alignment module to obtain the second recognition result. The output of the fourth stage then passes through the fourth convolutional layer, bidirectional long short-term memory network, and the first fully connected layer in the ResNet34 network in sequence to obtain the first recognition result.

[0039] It should be noted that the ResNet34 network includes a first stage (stage1), a second stage (stage2), a third stage (stage3), a fourth stage (stage4, the structures of each stage are not described in this solution and are the same as those of the ResNet34 network in the prior art), a fourth convolutional layer (using 1DCNN), a bidirectional long short-term memory network (Bi-LSTM), and a first fully connected layer, arranged in sequence from the data input direction to the output direction. Three temporal causal modules are respectively arranged between the first stage and the second stage, between the second stage and the third stage, and between the third stage and the fourth stage of the ResNet34 network (for convenience of representation, the temporal causal module between the first stage and the second stage is called the first temporal causal module, the temporal causal module between the second stage and the third stage is called the second temporal causal module, and the temporal causal module between the third stage and the fourth stage is called the third temporal causal module). The outputs of the four stages are parallel as the inputs of the alignment module. The alignment module respectively outputs the third recognition results for the four stages, and the four third recognition results together constitute the second recognition result. The four third recognition results and the first recognition are both used to calculate the loss to train the sign language recognition model.

[0040] Step 2.1, as Figure 3 shown, the input features of each temporal causal module (that is, the output of the first stage is used as the input feature of the first temporal causal module, the output of the second stage is used as the input feature of the second temporal causal module, and the output of the third stage is used as the input feature of the third temporal causal module) are first reshaped in shape, and then zero-padding with a preset number is performed on both sides of the reshaped input features along the time dimension to obtain two padded features;

[0041] The two padded features respectively pass through an atrous grouped convolutional layer (the two padded features are input into two atrous grouped convolutions one by one, and the weight parameters of the two atrous grouped convolutions are shared) to obtain corresponding first features (with a dimension of R T×Ci×HiWi ), and the two first features are concatenated along the channel dimension to obtain a second feature (with a dimension of );

[0042] The second feature then passes through a first convolutional layer to obtain a third feature (so that the dimension is adjusted to the dimension before concatenation, that is, the dimension is );

[0043] The third feature is fused with the input of the temporal causal module to obtain the output of the temporal causal module.

[0044] Among them, each stage outputs a 4D tensor, and the outputs of each stage are expressed as A i represents the output of the i-th stage, is T×Ci ×H i ×W i matrix of dimension C i 、H i and W i are the number of channels, height, and width at the i-th stage respectively. For example, the output A1 of the first stage belongs to R T×64×56×56 , the output A2 of the second stage belongs to R T×128×28×28 , the output A3 of the third stage belongs to R T×256×14×14 , and the output A4 of the fourth stage belongs to R T×512×7×7 ; Shape reshaping is to merge the width dimension and height dimension in the input features of each temporal causal module (i.e., obtaining ); As shown in Figure 2 , zero-padding is performed on the reshaped input features along the left and right sides of the time dimension (e.g., T = 5), and the specific number of padded features is set according to actual needs. In this embodiment, four features are padded on both the left and right sides;

[0045] Among them, the dilated grouped convolutional layer first performs grouped convolution on the input padded features, divides the number of channels into multiple groups, and then performs dilated convolution on each group to further expand the receptive field;

[0046] Among them, the formula for the fusion of the third feature and the input of the temporal causal module is as follows:

[0047]

[0048] Among them, represents the output obtained by passing the output of the i-th stage through the corresponding temporal causal module, and i is {1, 2, 3}, A i represents the output of the i-th stage, that is, the input of the corresponding temporal causal module (e.g., the output of the first stage is the input of the first temporal causal module, the output of the second stage is the input of the second temporal causal module, and the output of the third stage is the input of the third temporal causal module), represents the third feature corresponding to the output of the i-th stage passing through the first convolutional layer in the corresponding temporal causal module, and α i represents the weight of the third feature corresponding to the output of the i-th stage passing through the first convolutional layer in the corresponding temporal causal module.

[0049] The temporal causal module can effectively capture the bidirectional semantic relationship in sign language learning. Continuous sign language recognition, as a typical sequence task, has highly dynamic communication characteristics. Each gesture not only conveys its own meaning but is also closely related to the actions before and after, so it has a very strong dependence on the context. Therefore, the temporal causal module precisely captures the temporal dependence and causal relationship of action transitions by restricting the effective time context to past or future information.

[0050] Step 2.2, as Figure 4 shown, the alignment module includes a spatial alignment module and a temporal alignment module connected in sequence. The spatial alignment module includes four parallel spatial branches. The first, second, and third spatial branches each include a second convolutional layer, a pooling layer, and a third fully connected layer connected in sequence. The fourth spatial branch includes a pooling layer and a third fully connected layer connected in sequence (where the pooling layer of each spatial branch uses a size of 7×7, and the weights of the third fully connected layers of each spatial branch are shared). The temporal alignment module includes four parallel temporal branches, and each temporal branch includes a temporal convolutional network and a second fully connected layer connected in sequence;

[0051] The outputs of the first stage, the second stage, the third stage, and the fourth stage are sequentially and respectively used as the inputs of the first, second, third, and fourth spatial branches, and four fourth features (R T×1024 ) are obtained, and the spatial dimensions and the number of channels of the four fourth features are the same. Among them, after the outputs of the first stage, the second stage, and the third stage pass through the second convolutional layers of the first, second, and third spatial branches respectively, the spatial dimensions and the number of channels of the features output by each second convolutional layer are all matched with the output of the fourth stage (i.e., R T×512×7×7 ). Specifically, when the output of the first stage (64 channels) passes through the corresponding second convolutional layer, it first passes through 3D average pooling of (1, 4, 4) and then uses 3D convolution of (1, 2, 2), completing a total of 8-fold spatial scaling and dimension increasing to 512 channels. When the output of the second stage (128 channels) passes through the corresponding second convolutional layer, it uses convolutional resolution downsampling of (1, 4, 4) once to downsample 4 times and is also mapped to 512 channels. When the output of the third stage (256 channels) passes through the corresponding second convolutional layer, it only needs to use convolution of (1, 2, 2) to downsample 2 times and map the channels to 512. The fourth stage does not require any operation. Among them, 1 in (1, 4, 4) represents the convolutional operation in the time dimension (T), and (4, 4) represents the convolutional operation of size 4×4 in the spatial height (H) and width (W). Similarly for (1, 2, 2).

[0052] The four fourth features are then respectively input into the four temporal branches, and third recognition results corresponding to the four fourth features are respectively obtained and the four third recognition results together constitute the second recognition result.

[0053] Among them, the temporal convolutional network includes two sequentially connected TCN BLOCK modules, and each TCN BLOCK module includes a third convolutional layer (using one-dimensional convolution with a kernel size of 5), an activation layer (using the ReLU activation function), a batch normalization layer, and a max pooling layer (with a kernel size of 2) connected in sequence from the data input to the output direction. By extracting features at different stages of the ResNet34 network, expressions with different granularities can be obtained, including high-level semantic features and low-level detail features. The alignment module can capture and align spatial features of all granularities, effectively integrate features from each level, and enhance their consistency and complementarity. At the same time, combining the temporal causal module further enhances the modeling of temporal relationships, thereby achieving a more comprehensive understanding of dynamic scenes.

[0054] Among them, during the training stage of the sign language recognition model, the first recognition result and four third recognition results are jointly used to calculate the loss and train the entire sign language recognition model. The calculation formula of the loss function L of the sign language recognition model is as follows:

[0055]

[0056] Among them, L1 is the loss calculated by the first fully connected layer by comparing its output first recognition result with the true label, and L i is the loss calculated by the second fully connected layer by comparing its output third recognition result at the i-th stage with the true label. n is the total number of stages in the ResNet34 network, which is 4, and both L1 and L i adopt the CTC loss.

[0057] Sign language is a complex visual language that contains various information forms, including gestures, facial expressions, and body postures. Subtle changes in gestures are manifested as differences in local scales, while overall gesture movements are expressed on a larger scale. By utilizing cross-scale features, the sign language recognition model can narrow the semantic gap between features of different scales, enrich spatial semantic features, enabling it to more effectively focus attention on the gesture itself and reduce the impact of background interference.

[0058] This continuous sign language recognition method based on hierarchical spatio-temporal enhancement captures the multi-stage outputs of the ResNet34 network through the alignment module and provides additional hierarchical alignment supervision; and aggregates future and past information through the temporal causal module to achieve more precise lexical boundary perception. By integrating the alignment module and the temporal causal module into the ResNet34 network, a good balance is achieved between accuracy and computational cost, providing an efficient and accurate solution for the continuous sign language recognition task, significantly improving the recognition performance and robustness, and solving the underfitting problem that appears in the ResNet34 network in spatial and temporal contexts.

[0059] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A continuous sign language recognition method based on hierarchical spatio-temporal enhancement, characterized in that: The continuous sign language recognition method based on hierarchical spatio-temporal enhancement includes: Obtain a sign language video; Input the sign language video into the trained sign language recognition model to obtain a first recognition result and a second recognition result, and use the first recognition result as the final result of sign language recognition; The sign language recognition model includes a ResNet34 network, three temporal causal modules, and an alignment module. The ResNet34 network includes a first stage, a second stage, a third stage, and a fourth stage arranged in sequence. The three temporal causal modules are respectively arranged between the first stage and the second stage, between the second stage and the third stage, and between the third stage and the fourth stage of the ResNet34 network. When the sign language video is input into the trained sign language recognition model, it is first input into the first stage and forward-inferred. The outputs of the four stages are all input into the alignment module to obtain the second recognition result, and the output of the fourth stage is then sequentially passed through the fourth convolutional layer, bidirectional long short-term memory network, and first fully connected layer in the ResNet34 network to obtain the first recognition result.

2. The continuous sign language recognition method based on hierarchical spatio-temporal enhancement according to claim 1, wherein: The input features of each temporal causal module are first reshaped in shape, and then zero-padding with a preset number is performed on both sides of the reshaped input features along the time dimension to obtain two padded features; The two padded features respectively pass through an atrous grouped convolutional layer to obtain corresponding first features, and the two first features are concatenated along the channel dimension to obtain a second feature; The second feature then passes through a first convolutional layer to obtain a third feature; The third feature is fused with the input of the temporal causal module to obtain the output of the temporal causal module.

3. The continuous sign language recognition method based on hierarchical spatio-temporal enhancement according to claim 1, wherein: The alignment module includes a spatial alignment module and a time alignment module connected in sequence. The spatial alignment module includes four parallel spatial branches, and the first spatial branch, the second spatial branch, and the third spatial branch all include a second convolutional layer, a pooling layer, and a third fully connected layer connected in sequence. The fourth spatial branch includes a pooling layer and a third fully connected layer connected in sequence. The time alignment module includes four parallel time branches, and each time branch includes a temporal convolutional network and a second fully connected layer connected in sequence; The outputs of the first stage, the second stage, the third stage, and the fourth stage are sequentially and respectively used as the inputs of the first spatial branch, the second spatial branch, the third spatial branch, and the fourth spatial branch to obtain four fourth features, and the spatial dimensions and channel numbers of the four fourth features are the same; The four fourth features are then respectively and respectively input into the four time branches to obtain third recognition results corresponding to the four fourth features respectively, and the four third recognition results together constitute the second recognition result.

4. The continuous sign language recognition method based on hierarchical spatio-temporal enhancement according to claim 3, characterized in that: The temporal convolutional network includes two TCN BLOCK modules connected in sequence, and each TCN BLOCK module includes a third convolutional layer, an activation layer, a batch normalization layer, and a max pooling layer connected in sequence.

5. The continuous sign language recognition method based on hierarchical spatio-temporal enhancement according to claim 3, wherein: The calculation formula of the loss function L of the sign language recognition model is as follows: Among them, L1 is the loss calculated by the first fully connected layer between the first recognition result it outputs and the true label, and L i is the loss calculated by the second fully connected layer between the third recognition result of the i-th stage it outputs and the true label. n is the total number of stages in the ResNet34 network, which is 4, and both L1 and L i adopt the CTC loss.

6. The continuous sign language recognition method based on hierarchical spatio-temporal enhancement according to claim 2, characterized in that: The formula for fusing the third feature and the input of the temporal causal module is as follows: Among them, represents the output obtained by passing the output of the i-th stage through the corresponding sequential causal module, and i is {1, 2, 3}, A i represents the output of the i-th stage, that is, the input of the corresponding sequential causal module, represents the third feature corresponding to the output of the i-th stage passing through the first convolutional layer in the corresponding sequential causal module, α i represents the weight of the third feature corresponding to the output of the i-th stage passing through the first convolutional layer in the corresponding sequential causal module.

7. The continuous sign language recognition method based on hierarchical spatio-temporal enhancement according to claim 2, wherein: The reshaping in shape is to merge the width dimension and height dimension in the input features of each temporal causal module.

8. The continuous sign language recognition method based on hierarchical spatio-temporal enhancement according to claim 2, wherein: The dilated grouped convolutional layer first performs grouped convolution on the input padded features, divides the number of channels into multiple groups, and then performs dilated convolution on each group.

Citation Information

Cited By

  • Sign language recognition method and device based on cross-stage focus distillation and multi-branch time sequence learning

    CN122049994A

  • Sign language recognition method and device based on cross-stage focus distillation and multi-branch temporal learning

    CN122049994B