A method for constructing a sign language data set and a method for classifying and recognizing sign language words
By defining meta-motions and using dual-dimensional recognition with signal preprocessing and advanced models, the method addresses data collection and hardware compatibility issues in hand gesture recognition, ensuring efficient and accurate gesture recognition with flexible sensor arrangements.
Patent Information
- Application Number
- CN202510632062.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The existing sign language recognition solution requires a large amount of duplicate data acquisition when adding support words and adjusting sensor positions, resulting in high data acquisition costs and low hardware compatibility, hindering the flexible adjustment of the sign language recognition system.
By defining hand shape and positional element actions, the mapping relationship between element actions and composite actions is constructed, and combining FIR low-pass filtering, adaptive feature extraction and deep convolutional bidirectional LSTM model to realize the construction and recognition of sign language data sets.
It reduces the data acquisition cost of the sign language recognition system, improves hardware compatibility and the expansion capabilities of the recognition system, and realizes efficient sign language word classification recognition.
Smart Images

Figure CN120148123B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sign language recognition, and specifically relates to a method for constructing a sign language data set and a method for classifying and recognizing sign language words. Background Art
[0002] For people who cannot clearly express their meanings by speaking, sign language is an important bridge for communicating with the outside world. However, due to the differences between the sign language system and the spoken language system, ordinary people who have not undergone sign language learning and training usually cannot communicate smoothly with sign language users. Researchers have noticed this problem and explored various solutions to facilitate the communication between sign language users and ordinary people.
[0003] Today's mainstream solutions can be divided into three categories. The first category is to use a camera to collect the actions of sign language users in real time, and after processing the obtained images using methods in the field of computer vision, perform semantic output; the second category is to measure the signal characteristics of the corresponding parts during movement through a sensor array placed at specific parts, and after processing through means such as machine learning and deep learning, perform semantic output; the third category is a fusion scheme of the two.
[0004] Compared with the optical recognition scheme, the sensor scheme shows significant environmental robustness, deployment portability, and cost advantages. However, due to the differences in the types of sensors, placement positions, and selected features in different studies, they usually cannot share the same database. When researchers hope to verify whether a new sensor scheme is feasible, even if they only adjust the position of one sensor based on the original scheme, they need to collect completely new data from scratch.
[0005] The "National General Sign Language Common Word List" released in 2018 contains more than 5,000 national standard sign language words, and research in the field of dialect sign language also needs to consider thousands of local sign language words. In order to improve the accuracy of the scheme, each sign language word needs to be sampled at least more than a hundred times. The large amount of data collection work has hindered researchers from exploring sign language translation sensor schemes that have both sufficient accuracy and coverage.
[0006] In addition, when endowing a sensor device with the ability to recognize new sign language words, the existing method is to wear the corresponding sensing device, repeat the sampling of the sign language word several times, and input the obtained data into a neural network for training until an acceptable accuracy is obtained.
[0007] In summary, the existing methods have the following defects:
[0008] 1. Adding support for new words is cumbersome; when researchers hope to add translation support for a sign language word, they need to wear the device and repeat the collection of the sign language word action several times, and input this data into a neural network to ensure that the network can recognize this word without affecting the recognition of other words.
[0009] 2. Low hardware compatibility; when researchers modify the placement position of the sensor and need to adjust the neural network accordingly to ensure the recognition ability of the new device, they need to resample and retrain all the sign language words in the original dataset, which hinders the flexible adjustment of the solution. Summary of the Invention
[0010] In view of the deficiencies in the prior art, the present invention proposes a method for constructing a sign language dataset and a method for classifying and recognizing sign language words. The present invention can reduce the data acquisition cost of a sensor-based sign language recognition solution, improve the expansion ability of sign language recognition work, and reduce the hardware compatibility cost of the sign language recognition system.
[0011] The technical solution adopted by the present invention is as follows:
[0012] In the first aspect, the present invention provides a method for constructing a sign language dataset, including the following steps:
[0013] S1. Define the "hand shape" primitive action and the "posture" primitive action to obtain the corresponding category list and number list; wherein, the "posture" primitive action includes the "orientation", "position", and "motion" primitive actions of both hands.
[0014] Use Shape directly as the category list of the "hand shape" primitive action.
[0015] Synthesize the "position", "motion", and "orientation" primitive actions to obtain the category list PosState of the "posture" primitive action:
[0016] ,
[0017] wherein, represents the category list of the relative position of the hand to the body part, represents the category list of the relative position between the two hands, represents the category list of the overall motion mode of the hand, represents the category list of the hand orientation.
[0018] Merge and number the elements in the two category lists to form the number lists PosState_class and Shape_class.
[0019] S2. Collect primitive action information.
[0020] Collect the hand shape signals corresponding to all the "hand shape" primitive actions, and use the numbers in the number list Shape_class corresponding to the "hand shape" primitive action as labels.
[0021] Collect the position signals corresponding to all "position state" primitive actions, and use the numbers in the list PosState_class corresponding to the "position state" primitive actions as labels.
[0022] S3. Establish the mapping relationship between primitive actions and composite actions.
[0023] Regard the sign language word as a composite action formed by arranging primitive actions in combination, and realize sign language word recognition during the process of splitting and recombining the composite action.
[0024] Take the composite action to be classified as the value of the dictionary. According to the correspondence between each primitive action in Shape and PosState and each number in Shape_class and PosState_class, represent the composite action as a sequence of primitive actions presented in the form of numbers :
[0025] ,
[0026] where , are the numbers corresponding to the primitive actions that make up the composite action in Shape_class and PosState_class, the subscript , N is the number of classifiable composite actions, and the subscript , is the maximum value of the number of primitive actions in the Shape dimension and PosState dimension included in the composite action .
[0027] Take as the key, as the value, and store them in the dictionary in pairs, thus obtaining the mapping relationship between the composite action and the primitive action:
[0028] ;
[0029] Based on the mapping relationship between the composite action and the primitive action, combined with the primitive action signals collected in step S2, construct a sign language data set containing N sign language words.
[0030] In the second aspect, based on the above sign language data set, the present invention also provides a sign language word classification and recognition method, including the following steps:
[0031] S4. Perform data preprocessing on the sign language data set.
[0032] S41. Perform preprocessing on the handshape signal:
[0033] A. Anti-aliasing filtering is performed on the hand shape signal using a FIR low-pass filter to suppress high-frequency noise and prevent signal aliasing from affecting subsequent processing.
[0034] B. Downsampling processing is performed on the filtered hand shape signal using a phased downsampling strategy to reduce the data volume and improve the calculation efficiency.
[0035] C. Band-pass filtering is performed on the downsampled hand shape signal to remove noise interference in irrelevant frequency bands.
[0036] D. Low-pass filtering is performed on the band-pass filtered hand shape signal to further smooth the signal and reduce the influence of high-frequency noise.
[0037] E. L hand shape signals within a preset time are used as one frame of data, and a sliding window with a fixed window length and step size is used to segment the hand shape signal obtained in step D to obtain the preprocessed hand shape signal, which is used as the training data and test data for the hand shape recognition model.
[0038] S42. Preprocess the pose signal.
[0039] An adaptive feature extraction module is used to extract features from the pose signal as the input data for the pose recognition model.
[0040] Specifically, the original pose signal is segmented through a sliding window to generate subsequences of a fixed length; then wavelet transform is used to extract multi-scale features to significantly enhance the data representation ability; and then an autoencoder model is used to learn the low-dimensional representation of the data and reduce the dimension of the high-dimensional multi-scale features to obtain the preprocessed pose signal, which is used as the training data and test data for the pose recognition model.
[0041] S5. Construct a hand shape recognition model.
[0042] The hand shape recognition model uses a mean template matching model based on majority voting, including a feature extraction module, a mean template matching module, and a multi-level voting decision module set in sequence.
[0043] Further, the feature extraction module is used to extract multi-dimensional feature vectors layer by layer;
[0044] The mean template matching module is used to construct a mean template based on the multi-dimensional feature vectors extracted by the feature extraction module, perform similarity matching, and generate a recognition result;
[0045] The multi-level voting decision module is used to perform multi-round voting fusion on the recognition results of continuous time series based on spatio-temporal consistency constraints through a queue mechanism, and finally output a hand shape classification result that meets the preset confidence threshold.
[0046] S6. Construct a pose recognition model.
[0047] The pose recognition model is implemented using the DeepConvBiLSTM model.
[0048] S7. Integrate the two-dimensional information and output the sign language semantics.
[0049] Collect the handshape signal and pose signal of the sign language to be recognized. After preprocessing the handshape signal, input it into the handshape recognition model to obtain the handshape recognition result. After preprocessing the pose signal, input it into the pose recognition model to obtain the pose recognition result, which is expressed as:
[0050] ,
[0051] where respectively represent the handshape recognition result and the pose recognition result; when t = 1, that is, when the signal is first backpropagated and the recognition result is output for the first time, traverse the key list of the dictionary described in step S3 , to obtain a new key list , where, , and satisfy ; for , repeat the operation, continuously traverse the new key list obtained last time until finally obtaining a list with unique values ; query the corresponding , to obtain the final output result , that is, the sign language word.
[0052] Advantages of the present invention:
[0053] 1. The present invention ensures the scalability of sign language recognition. Only need to maintain the recognition ability of "elementary actions", and the recognition method based on this can support an infinite number of "compound actions". At the same time, it also avoids the possibility that the accuracy of the original "compound actions" may decrease when adding new "compound actions" in the existing direct recognition scheme. That is, the present invention keeps the parameters of the elementary action classifier constant and only updates the lightweight mapping dictionary (computational complexity O(1)).
[0054] 2. Ensure the low complexity of the dataset reconstruction work when changing the sensor layout. Only need to re-measure the signals of all elementary actions in the finite "elementary action" library, rather than re-measuring all the recognizable "compound actions" originally, ensuring an acceptable complexity. That is, the existing scheme needs to perform data acquisition of the O(N) order of magnitude, and this scheme achieves a complexity of the O(1) order of magnitude through the elementary action combination mechanism.
[0055] 3. Recognition Method. The posture recognition algorithm and the gesture recognition algorithm have demonstrated excellent performance in practical applications, fully reflecting their strong capabilities in complex classification tasks. The posture recognition algorithm achieved an accuracy rate of 90% in the 22-classification problem, indicating that the algorithm can effectively process complex data with high dimensions and multiple categories and perform well in feature extraction and pattern recognition. The gesture recognition algorithm achieved an accuracy rate of 95% in the 8-classification problem, further demonstrating its excellent performance in specific classification tasks. Gesture recognition usually involves capturing and analyzing dynamic features, and the algorithm can complete the classification task with extremely high accuracy, indicating its advantages in feature extraction, processing of time-series data, and classifier design. Brief Description of the Drawings
[0056] Figure 1 is the overall flowchart;
[0057] Figure 2 is the process diagram for outputting sign language semantics;
[0058] Figure 3 is the comparison diagram before and after preprocessing of the handshape signal. Detailed Implementation Modes
[0059] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0060] This embodiment provides a method for constructing a sign language data set and a method for classifying and recognizing sign language words, and the overall process is as Figure 1 shown. Among them, the method for constructing a sign language data set includes the following steps:
[0061] S1. Define the "handshape" elementary actions and "posture" elementary actions to obtain the corresponding category list and number list; wherein, the "posture" elementary actions include the "orientation", "position", and "motion" elementary actions of both hands.
[0062] We consider the "sign language actions" of complete semantic sign language words as "composite actions", and the "basic handshapes" and "basic hand movements" that make up these sign language actions as elementary actions. The sign language is recognized by using the composite action equal to the superposition of elementary actions (i.e., establishing the mapping relationship between elementary actions and composite actions). The composite action will increase with the increase of sign language semantics, but the elementary action library can be considered relatively unchanged.
[0063] Therefore, regarding the "handshape" elementary actions, considering the occurrence frequencies of various handshapes in actual sign language practice, 12 types of single-handed handshapes are defined: Shape = ["thumb straight", "thumb dotting", "1", "2", "3", "4", "5", "bent 1", "bent 2", "bent 3", "bent 4", "bent 5"].
[0064] Regarding the "position" elementary action, the common hand placement positions are listed and classified into the relative position between the hand and other body parts and the relative position between the two hands. Define 9 types of relative positions between the hand and body parts: PosB = ["hand on top of the head", "hand on the side of the head", "hand close to the cheek", "hand close to the earlobe", "hand close to the forehead", "hand close to the lips", "hand on the chest", "hand on the side of the body", "hand on the abdomen"]; Define 5 types of relative positions between the hands: PosH = ["both hands facing forward", "right hand on the left hand", "right hand in front of the left hand", "right hand behind the left hand", "right hand below the left hand"].
[0065] Regarding the "motion" elementary action, based on the basic motion lines of moving, shaking, and rotating, combined with the direction and route length, the overall motion modes of the hand are defined and classified. Define 10 types of motions: PosMotion = ["rotate 90° clockwise", "rotate 90° counterclockwise", "rotate 180° clockwise", "rotate 180° counterclockwise", "rotate 90°", "move forward", "move backward", "move left", "move right", "move up", "move down"].
[0066] Regarding the "orientation" elementary action, define 6 basic hand orientations: PosToward = ["standing horizontally", "standing sideways", "extending horizontally", "extending obliquely", "extending sideways", "extending across"].
[0067] Combine the "position" elementary action, "motion" elementary action, and "orientation" elementary action to obtain the "posture" elementary action category list PosState:
[0068] ,
[0069] Use Shape directly as the "hand shape" elementary action category list.
[0070] Adopt two-digit decimal numbers to merge and number the elements in the two category lists to obtain the number list:
[0071] , ,
[0072] Among them, number 00 and number 31 represent NULL, numbers 01 - 30 have a one-to-one correspondence with PosState, and numbers 32 - 43 have a one-to-one correspondence with Shape.
[0073] S2. Collect elementary action information.
[0074] Collect the hand shape signals corresponding to all "hand shape" elementary actions through the hand sensors, and use the numbers in the number list Shape_class corresponding to the "hand shape" elementary actions as labels.
[0075] Collect the position state signals corresponding to all "position state" elementary actions through the hand sensor, and use the numbers in the list PosState_class corresponding to the "position state" elementary actions as labels.
[0076] S3. Establish the mapping relationship between elementary actions and composite actions.
[0077] Regard the sign language word as a composite action formed by arranging elementary actions in combination, and realize sign language word recognition during the process of splitting and recombining the composite action.
[0078] The composite action to be classified As the value of the dictionary, according to the corresponding relationship between each elementary action in Shape and PosState and each number in Shape_class and PosState_class, the composite action Is represented as a sequence of elementary actions in the form of numbers :
[0079] ,
[0080] Among them, , Are the numbers corresponding to the elementary actions that make up the composite action In Shape_class and PosState_class, ; Subscript , N is the number of classifiable composite actions, subscript , Is the maximum value of the number of elementary actions in the Shape dimension and PosState dimension included in the composite action .
[0081] For example, = Thank you. The elementary actions that make up the sign language word "Thank you" are: "Thumb dotting" in the list of "hand shape" elementary actions Shape, corresponding to the number 33 in the list Shape_class; "Hands placed in front of the chest" and "Hands facing each other" in the list of "position state" elementary actions PosState, corresponding to the numbers 07 and 10 in the list PosState_class in sequence.
[0082] Therefore, "Thank you" contains one "hand shape" elementary action and two "position state" elementary actions. Therefore , The value of the unknown in the overall formula is: , That is, the sequence of elementary actions corresponding to "Thank you" is .
[0083] Take As the key, As values, they are stored in pairs in a dictionary, from which the mapping relationship between composite actions and primitive actions is obtained:
[0084] ,
[0085] Based on the mapping relationship between composite actions and primitive actions, combined with the handshape signals and postures signals collected in step S2, a sign language data set containing N sign language words is constructed.
[0086] In addition, based on the sign language word formation and spelling method, for derivative words, we adopt the method of calling the defined words and then combining them to generate new words, which simplifies the definition steps. For example, if there is a combination of primitive actions for "mountain", then "Shanxi" = ["mountain", "4", "standing horizontally"].
[0087] The present invention can significantly reduce the acquisition cost of sign language data and improve hardware compatibility through primitive action decomposition and two-dimensional recognition, providing technical support for the development of efficient sign language translation devices.
[0088] Based on the above sign language data set, this embodiment also provides a sign language word classification and recognition method, including the following steps:
[0089] S4. Perform data preprocessing on the sign language data set.
[0090] S41. Perform preprocessing on the handshape signals;
[0091] A. Use a FIR low-pass filter to perform anti-aliasing filtering on the handshape signals to suppress high-frequency noise and prevent signal aliasing from affecting subsequent processing; during the filtering process, set the cut-off frequency to the Nyquist frequency of 50 Hz of the target sampling rate.
[0092] B. Use a phased downsampling strategy to perform downsampling on the filtered handshape signals to reduce the amount of data and improve the calculation efficiency; specifically, select the optimal downsampling factor through the factorization method to ensure data integrity, and the target sampling rate is 100 Hz.
[0093] C. Use a Butterworth sixth-order band-pass filter to perform band-pass filtering on the downsampled handshape signals, and the filtering range is 0.1 - 10 Hz to remove noise interference in irrelevant frequency bands.
[0094] D. Use a fourth-order, 10 Hz low-pass filter to perform low-pass filtering on the band-pass filtered handshape signals to further smooth the signals and reduce the influence of high-frequency noise. After filtering and downsampling processing, the handshape signals as shown in Figure 3 are obtained.
[0095] E. Since the hand gesture signals are continuous, the data at a single time point has no practical significance. Therefore, L hand gesture signals within a preset time are used as one frame of data, and a sliding window with a fixed window length and step size is used to segment the hand gesture signals obtained in step D, resulting in preprocessed hand gesture signals to be used as the training data and test data for the hand gesture recognition model.
[0096] S42. Preprocess the pose signals;
[0097] Use an adaptive feature extraction module to extract features from the pose signals to be used as the input data for the pose recognition model.
[0098] Specifically, segment the pose signals through a sliding window to generate subsequences of a fixed length; then use wavelet transform to extract multi-scale features to significantly enhance the data representation ability; and then use an autoencoder model to learn the low-dimensional representation of the data and reduce the dimension of the high-dimensional multi-scale features to obtain the preprocessed pose signals to be used as the training data and test data for the pose recognition model.
[0099] S5. Construct a hand gesture recognition model.
[0100] The hand gesture recognition model uses a mean template matching model based on majority voting, including a feature extraction module, a mean template matching module, and a multi-level voting decision module set up in sequence.
[0101] The feature extraction module is used to extract multi-dimensional feature vectors layer by layer. Specifically, the feature extraction module includes 3 convolutional layers, and each convolutional layer is followed by a batch normalization layer, a ReLU non-linear activation layer, and a pooling layer, and the final output is a 32-dimensional feature vector.
[0102] The mean template matching module is used to construct a mean template based on the multi-dimensional feature vectors extracted by the feature extraction module, perform similarity matching, and generate a recognition result.
[0103] The multi-level voting decision module is used to perform multi-round voting fusion on the recognition results of continuous time series based on spatio-temporal consistency constraints through a queue mechanism, and finally output a hand gesture classification result that meets the preset confidence threshold.
[0104] Use the hand gesture training data and test data to train and verify the hand gesture recognition model to obtain a trained hand gesture recognition model.
[0105] S6. Construct a pose recognition model.
[0106] The pose recognition model uses an improved model based on deep convolutional bidirectional LSTM, including an adaptive feature extraction module and a deep convolutional bidirectional LSTM module set up in sequence.
[0107] The adaptive feature extraction module is used to preprocess and extract features from the input data. Specifically, the input data is segmented into fixed-length subsequences by the sliding window segmentation method; then multi-dimensional features are extracted from the time domain and frequency domain by wavelet transform to enhance the data representation ability; subsequently, an autoencoder is constructed to reduce the dimension of the multi-dimensional features extracted by wavelet transform, and the reduced multi-dimensional features are obtained, providing efficient feature input for subsequent processing.
[0108] The deep convolutional bidirectional LSTM module is constructed based on the traditional DeepConvBiLSTM network. The traditional DeepConvBiLSTM network includes a convolutional layer and a bidirectional LSTM layer. The convolutional layer consists of one-dimensional convolutional kernels to slide and extract local features; the bidirectional LSTM layer processes data from both forward and backward directions simultaneously to capture context relationships. In the present invention, a residual connection module is added between the convolutional layer and the bidirectional LSTM layer to enhance the feature transfer ability; a self-attention mechanism module is added after the bidirectional LSTM layer to dynamically allocate feature weights and improve the attention to key information.
[0109] The posture recognition model is trained and verified using the posture training data and test data to obtain a trained posture recognition model.
[0110] S7. Integrate the two-dimensional information and output the sign language semantics.
[0111] Collect the handshape signal and posture signal of the sign language to be recognized; after preprocessing the handshape signal, input it into the handshape recognition model to obtain the handshape recognition result, and after preprocessing the posture signal, input it into the posture recognition model to obtain the posture recognition result, expressed as:
[0112] ,
[0113] where respectively represent the handshape recognition result and the posture recognition result; Figure 2 is the process diagram of outputting the sign language semantics. When t = 1, that is, when the signal is transmitted back and the recognition result is output for the first time, traverse the key list of the dictionary described in step S3 to obtain a new key list , where, , and satisfy ; for , repeat the operation, continuously traverse the new key list obtained last time until finally obtaining a list with unique values ; query the corresponding to obtain the final output result , that is, the sign language word.
Claims
1. A method for constructing a sign language dataset, characterized in that, It includes the following steps: S1. Define the "hand shape" elementary actions and "posture" elementary actions to obtain the hand shape category list Shape, the hand shape number list Shape_class, the posture category list PosState, and the posture number list PosState_class. Among them, the "posture" elementary actions include the "orientation", "position", and "motion" elementary actions of both hands; S2. Collect elementary action information; Collect the hand shape signals corresponding to all "hand shape" elementary actions, and use the numbers in the number list Shape_class corresponding to the "hand shape" elementary actions as labels; Collect the posture signals corresponding to all "posture" elementary actions, and use the numbers in the number list PosState_class corresponding to the "posture" elementary actions as labels; S3. Establish the mapping relationship between elementary actions and composite actions; Regard the sign language word as a composite action formed by arranging meta-actions in combination; regard the composite action to be classified As the value of the dictionary, according to the correspondence between each meta-action in Shape and PosState and each number in Shape_class and PosState_class, represent the composite action as a sequence of meta-actions presented in the form of numbers : , Among them, and are the corresponding numbers of the primitive actions that make up the composite action in Shape_class and PosState_class. The subscript , N is the number of classifiable composite actions, and the subscript , is the maximum value of the number of primitive actions in the Shape dimension and the PosState dimension included in the composite action ; Take as the key, as the value, and store them in pairs in a dictionary, thus obtaining the mapping relationship between composite actions and meta-actions: , Through the mapping relationship between composite actions and elementary actions, combined with the elementary action signals collected in step S2, construct a sign language data set containing N sign language words.
2. The method for constructing a sign language data set according to claim 1, wherein In step S1, the representation methods of the category list and the number list are as follows: Take Shape as the "hand shape" elementary action category list; Synthesize the "position" elementary action, the "motion" elementary action, and the "orientation" elementary action to obtain the "posture" elementary action category list PosState: , Among them, A list of categories representing the relative positions of the hand and body parts, A list of categories representing the relative positions between the two hands, A list of categories representing the overall movement patterns of the hand, A list of categories representing the hand orientations; Merge and number the elements in the two category lists to form the number lists PosState_class and Shape_class.
3. A method for classifying and recognizing sign language words, characterized in that, It includes the following steps: S4. Perform data preprocessing on the sign language data set constructed by the method for constructing a sign language data set described in claim 2; S5. Construct a hand shape recognition model; The hand shape recognition model is implemented using a mean template matching model based on majority voting, and includes a feature extraction module, a mean template matching module, and a multi-level voting decision module set in sequence; S6. Construct a posture recognition model; The posture recognition model is implemented using a DeepConvBiLSTM model; S7. Synthesize the two-dimensional information and output the sign language semantics; Collect the hand shape signal and the posture signal of the sign language to be recognized, input the hand shape signal into the hand shape recognition model after the data preprocessing described in step S4 to obtain the hand shape recognition result, and input the posture signal into the posture recognition model after preprocessing to obtain the posture recognition result, expressed as: , Among them, respectively represent the handshape recognition result and the posture recognition result; when t = 1, that is, when the signal is transmitted back and the recognition result is output for the first time, traverse the key list of the dictionary described in step S3 , to obtain a new key list , where , and satisfy ; for , repeat the operation, continuously traverse the new key list obtained last time until finally obtaining a list with unique values ; query the corresponding , to obtain the final sign language word output result .
4. The sign language word classification and recognition method according to claim 3, characterized in that, In step S4, the method for preprocessing the hand shape signal includes the following steps: A. Use a FIR low-pass filter to perform anti-aliasing filtering on the hand shape signal; B. Use a phased downsampling strategy to perform downsampling on the filtered hand shape signal; C. Perform band-pass filtering on the downsampled hand shape signal; D. Perform low-pass filtering on the band-pass filtered hand shape signal; E. Take L hand shape signals within a preset time as one frame of data, and use a sliding window with a fixed window length and step size to segment the hand shape signal obtained in step D to obtain the preprocessed hand shape signal.
5. The sign language word classification and recognition method according to claim 3, wherein In step S4, the method for preprocessing the posture signal includes the following steps: Use an adaptive feature extraction module to perform feature extraction on the posture signal as the input data of the posture recognition model; Specifically, the original state signal is segmented by a sliding window to generate subsequences of a fixed length; then wavelet transform is used to extract multi-scale features; and then an autoencoder model is used to learn the low-dimensional representation of the data, reducing the dimension of the high-dimensional multi-scale features to obtain the preprocessed state signal.
6. The sign language word classification and recognition method according to claim 4, wherein In step S5, the feature extraction module is used to extract multi-dimensional feature vectors layer by layer; the mean template matching module is used to construct a mean template based on the multi-dimensional feature vectors extracted by the feature extraction module, perform similarity matching, and generate an identification result; the multi-level voting decision module is used to perform multi-round voting fusion on the identification results of consecutive time series based on spatio-temporal consistency constraints through a queue mechanism, and finally output a hand shape classification result that meets the preset confidence threshold.
7. The sign language word classification and recognition method according to claim 5, characterized in that In step S6, the state recognition model includes a traditional DeepConvBiLSTM network, a residual connection module, and a self-attention mechanism module; Among them, the traditional DeepConvBiLSTM network includes a convolutional layer and a bidirectional LSTM layer. The convolutional layer consists of one-dimensional convolutional kernels and slides to extract local features; the bidirectional LSTM layer processes data from both the forward and reverse directions simultaneously to capture context relationships; the residual connection module is arranged between the convolutional layer and the bidirectional LSTM layer to enhance the feature transfer ability; the self-attention mechanism module is arranged after the bidirectional LSTM layer to dynamically allocate feature weights and improve the attention to key information.
Citation Information
Patent Citations
Combined action recognition method and system based on multi-level feature interactive fusion
CN114333057A
Gesture action video generation method and device based on fine-grained semantic description
CN119444943A