Sign language recognition method based on Leap Motion sensor and deep learning
Through the combination of Leap Motion sensor and deep learning, the keyframe and complement frame technology and a two-way long and short-term memory network are used to solve the real-time and cross-language adaptability problems of the existing sign language recognition system, and efficient and stable sign language recognition is achieved.
Patent Information
- Application Number
- CN202510281919.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-22
AI Technical Summary
Existing sign language recognition systems have shortcomings in real-time, computing resource consumption, and cross-language adaptability, especially in low-light or inclement weather conditions, and rely on high-cost sensors and complex deep learning algorithms.
The Leap Motion sensor is used to collect dynamic sign language data, combine keyframe and complement frame technology, and build a dynamic sign language recognition model through two layers of bidirectional long and short-term memory networks, extract single-finger and double-finger features, optimize the training process, and improve recognition accuracy and real-timeness.
It improves the accuracy and real-time nature of sign language recognition, reduces computing costs, enhances cross-language adaptability, and achieves stable recognition in different environments.
Smart Images

Figure CN120356258A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sign language recognition, and particularly relates to a sign language recognition method based on a Leap Motion sensor and deep learning. Background Art
[0002] In recent years, the emergence of high-precision sensors such as Leap Motion has provided important support for the development of sign language recognition technology. These sensors can capture hand movements and provide accurate feature data. The introduction of deep learning technology has significantly improved the accuracy and robustness of sign language recognition. Although these technological advancements have enhanced the performance of sign language recognition systems, there are still some deficiencies in the existing technologies: First, many existing sign language recognition systems rely on high-precision sensors such as LeapMotion and Kinect. These devices are expensive and may perform poorly in some environments. For example, in low-light or adverse weather conditions, the recognition ability of the sensors may be affected, thus affecting the practical application of the system; Second, although the existing deep learning algorithms have improved the recognition accuracy, the processing speed is slow and the computational resources consumption is large. Especially when dealing with complex dynamic sign languages, a large amount of calculation and training data are required, which poses high requirements on the hardware and computing capabilities; Finally, the existing sign language recognition technologies lack strong cross-language and cross-cultural adaptability. Many systems can only be optimized for sign languages in specific regions or languages, so they often perform unstably in applications in different regions or different languages.
[0003] Generally speaking, although the existing technologies have made some progress in the accuracy and effect of sign language recognition, there are still many deficiencies in universality, real-time performance, adaptability, and the ability to process complex sign language sequences. Therefore, how to improve the real-time performance of sign language recognition systems, reduce the consumption of computational resources, and enhance their cross-language adaptability remains an important topic in the current technological development. Summary of the Invention
[0004] In order to overcome the defects and deficiencies of the existing technologies, the present invention provides a sign language recognition method based on a Leap Motion sensor and deep learning. The present invention uses a Leap Motion sensor to collect dynamic sign language data, and improves the adaptability and stability of the system through key frame and supplementary frame technologies. An efficient dynamic sign language recognition model is constructed based on deep learning algorithms, and the parameters in the training process are optimized. The present invention improves the recognition accuracy and real-time performance, and provides a new solution for dynamic sign language recognition technology.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] The present invention provides a sign language recognition method based on a Leap Motion sensor and deep learning, including the following steps:
[0007] Collect a dynamic sign language data set based on the Leap Motion sensor;
[0008] Extract features from the dynamic sign language data set, extract single-finger features and double-finger features, and form a dynamic sign language feature data set;
[0009] Divide the dynamic sign language feature data set into a training data set and a test data set;
[0010] Build a dynamic sign language recognition model based on a two-layer bidirectional long short-term memory network;
[0011] Train the dynamic sign language recognition model based on the training data set, test the dynamic sign language recognition model based on the test data set, and use the dynamic sign language recognition model with the optimal accuracy as the trained dynamic sign language recognition model;
[0012] Perform dynamic sign language recognition based on the trained dynamic sign language recognition model and output the corresponding sign language category.
[0013] As a preferred technical solution, collecting a dynamic sign language data set based on the Leap Motion sensor specifically includes:
[0014] Each frame of data includes the palm normal vector L N, the palm center position L C, the finger tip position L F i ;
[0015] When collecting each dynamic sign language sample, copy a frames of data before the key frame as the first a frames of data of the n + a + b frames of data, and copy the last frame of the consecutive n frames of key frames into b copies as the last b frames of data of the n + a + b frames of data;
[0016] Add labels to the collected dynamic sign language samples to form a dynamic sign language data set.
[0017] As a preferred technical solution, extracting single-finger features and double-finger features from the dynamic sign language data set specifically includes:
[0018] Through the coordinate transformation formula, transform the vector description in the Leap Motion device center coordinate system L into the vector description in the palm coordinate system H;
[0019] The single-finger gesture features include: the distance between the finger tip and the palm center, the angle formed by the line connecting the finger tip and the palm center and the palm plane, and the distance from the finger tip to the palm plane;
[0020] The double-finger gesture features include the distance between the fingertips of two adjacent fingers and the angle between the lines connecting the fingertips of two adjacent fingers to the palm center.
[0021] As a preferred technical solution, the distance between the fingertip and the palm center is expressed as:
[0022] D i =|| L F i - L C|| / D max
[0023] Where D i represents the Euclidean distance from the fingertip L F i to the palm center position L C in the central coordinate system L of the Leap Motion device, and D max represents the maximum distance value from all fingertip positions to the palm center position. The distance value between the fingertip and the palm center is normalized to the interval [0, 1];
[0024] The angle formed by the line connecting the fingertip and the palm center and the palm plane is expressed as:
[0025] A i =π / 2 - ∠( H F i - H C, H N)
[0026] Where A i represents the angle formed by the line connecting the fingertip H F i and the palm center H C and the palm plane H N in the palm coordinate system H, and ∠( H F i - H C, H N) represents the angle between the line connecting the fingertip H F i and the palm center H C and the normal vector, i = 1,..., 5;
[0027] The distance from the fingertip to the palm plane is expressed as:
[0028] H i =D i ·D max cos(∠( H F i - H C, H N))
[0029] Among them, H i represents the distance from the fingertip to the palm plane, and D i ·D max represents the distance calculated from the fingertip to the center position of the palm.
[0030] As a preferred technical solution, the distance between the fingertips of two adjacent fingers is expressed as:
[0031] DD i = || H F i - H F i+1 || / DD max
[0032] Among them, DD i represents the Euclidean distance between the fingertips of two adjacent fingers, H F i+1 represents the fingertip of the finger in the palm coordinate system H H F i of the adjacent fingertip, and DD max represents the maximum distance between the fingertips of two adjacent fingers;
[0033] The included angle between the line connecting the fingertips of two adjacent fingers and the palm center is expressed as:
[0034] AA i = ∠( H F i - H C, H F i+1 - H C)
[0035] Among them, AA i represents the included angle between the line connecting the fingertips of two adjacent fingers and the palm center, H and C represents the palm center in the palm coordinate system H.
[0036] The present invention also provides a sign language recognition system based on a Leap Motion sensor and deep learning, including: a dynamic sign language dataset acquisition module, a feature extraction module, a dataset division module, a dynamic sign language recognition model construction module, a dynamic sign language recognition model training module, and a dynamic sign language recognition module;
[0037] The dynamic sign language dataset acquisition module is used to acquire a dynamic sign language dataset based on the Leap Motion sensor;
[0038] The feature extraction module is used to extract features from the dynamic sign language dataset, extract single-finger features and double-finger features, and form a dynamic sign language feature dataset;
[0039] The dataset partitioning module is used to partition the dynamic sign language feature dataset into a training dataset and a test dataset;
[0040] The dynamic sign language recognition model construction module is used to construct a dynamic sign language recognition model based on a two-layer bidirectional long short-term memory network;
[0041] The dynamic sign language recognition model training module is used to train the dynamic sign language recognition model based on the training dataset, test the dynamic sign language recognition model based on the test dataset, and use the dynamic sign language recognition model with the optimal accuracy as the trained dynamic sign language recognition model;
[0042] The dynamic sign language recognition module is used to perform dynamic sign language recognition based on the trained dynamic sign language recognition model and output the corresponding sign language category.
[0043] As a preferred technical solution, the dynamic sign language dataset acquisition module is used to acquire the dynamic sign language dataset based on the Leap Motion sensor, specifically including:
[0044] Each frame of data includes the palm normal vector of the Leap Motion device center coordinate system L The palm center position The finger tip position L F i ;
[0045] When collecting each dynamic sign language sample, copy a frame of data before the key frame a times as the first a frames of the n + a + b frames of data, and copy the last frame of the continuous n frames of key frames b times as the last b frames of the n + a + b frames of data;
[0046] Add labels to the collected dynamic sign language samples to form a dynamic sign language dataset.
[0047] As a preferred technical solution, the feature extraction module is used to extract features from the dynamic sign language dataset, and extract single-finger features and double-finger features, specifically including:
[0048] Through the coordinate transformation formula, transform the vector description in the Leap Motion device center coordinate system L into the vector description in the palm coordinate system H;
[0049] The single-finger gesture features include: the distance between the finger tip and the palm center, the angle formed by the line connecting the finger tip and the palm center and the palm plane, and the distance from the finger tip to the palm plane;
[0050] The double-finger gesture features include the distance between the adjacent two finger tips and the angle between the line connecting the adjacent two finger tips and the palm center.
[0051] As a preferred technical solution, the distance between the fingertip and the palm center is expressed as:
[0052] D i =|| L F i - L C|| / D max
[0053] where D i represents the Euclidean distance from the fingertip L F i to the palm center position L C in the central coordinate system L of the Leap Motion device. D max represents the maximum distance value from all fingertips to the palm center position, and the distance value between the fingertip and the palm center is normalized to the interval [0, 1];
[0054] The angle formed by the line connecting the fingertip and the palm center and the palm plane is expressed as:
[0055] A i =π / 2 - ∠( H F i - H C, H N)
[0056] where A i represents the angle formed by the line connecting the fingertip H F i and the palm center H C and the palm plane H N in the palm coordinate system H. ∠( H F i - H C, H N) represents the angle between the line connecting the fingertip H F i and the palm center H C and the normal vector, i = 1,..., 5;
[0057] The distance from the fingertip to the palm plane is expressed as:
[0058] H i =D i ·D max cos(∠( H F i - H C, H N))
[0059] where H i represents the distance from the fingertip to the palm plane, D i ·D maxIndicates calculating the distance from the fingertip to the center of the palm.
[0060] As a preferred technical solution, the distance between the fingertips of two adjacent fingers is expressed as:
[0061] DD i =|| H F i - H F i+1 || / DD max
[0062] Where DD i represents the Euclidean distance between the fingertips of two adjacent fingers, H F i+1 represents the adjacent fingertip of the fingertip under the palm coordinate system H H F i and DD max represents the maximum distance between the fingertips of two adjacent fingers;
[0063] The included angle between the line connecting the fingertips of two adjacent fingers and the center of the palm is expressed as:
[0064] AA i =∠( H F i - H C, H F i+1 - H C)
[0065] Where AA i represents the included angle between the line connecting the fingertips of two adjacent fingers and the center of the palm, H C represents the center of the palm under the palm coordinate system H.
[0066] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0067] (1) The present invention reduces the calculation cost and improves the recognition accuracy through the key frame acquisition and supplementary frame technology.
[0068] (2) The present invention uses a 23-dimensional gesture feature vector composed of five gesture features to represent different gesture features, further improving the recognition efficiency;
[0069] (3) The deep learning model constructed by the present invention based on the two-layer bidirectional long short-term memory network can capture the front and back time series features of the gesture actions simultaneously, improving the recognition accuracy and efficiency.
[0070] (4) The present invention realizes the remote operation application through real-time dynamic sign language recognition, enhancing the convenience and flexibility of human-computer interaction. Brief Description of the Drawings
[0071] Figure 1 Schematic flowchart of the sign language recognition method based on Leap Motion sensor and deep learning according to the present invention;
[0072] Figure 2 Schematic diagram of dynamic sign language data composed of key frames and supplementary frames according to the present invention;
[0073] Figure 3 Schematic diagram of the key point vector in the hand according to the present invention;
[0074] Figure 4 Schematic diagram of the coordinate transformation principle according to the present invention;
[0075] Figure 5 Schematic diagram of the structure of the dynamic sign language recognition model based on long short - term memory network and Sofmax activation function according to the present invention. Detailed implementation manners
[0076] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0077] Embodiment 1
[0078] As Figure 1 shown, this embodiment provides a sign language recognition method based on Leap Motion sensor and deep learning, including the following steps:
[0079] S1: Data acquisition: Use the Leap Motion sensor to collect a dynamic sign language data set;
[0080] S11: Use the Leap Motion sensor to collect dynamic sign language data, collect a total of 10 kinds of dynamic sign language data from Chinese numbers 0 to 9, 50 copies of each kind of dynamic sign language are collected, and a total of 500 dynamic sign language samples are collected. Each dynamic sign language sample consists of 30 frames of data, and each frame of data includes the palm normal vector L N, the palm center position L C, the finger tip positions L F i , i = 1,..., 5;
[0081] S12: As Figure 2As shown, when collecting each dynamic sign language sample, set the thresholds for the speeds of the fingertips of the five fingers to collect the key frame data of 6 consecutive frames and the data of the frame before these 6 key frames. Copy the data of the frame before the key frame a times as the first a frames of the n + a + b frames of data, and copy the last frame of the consecutive n key frames b times as the last b frames of the n + a + b frames of data. In this embodiment, the stability and integrity of the data are ensured through key frame extraction and frame filling techniques;
[0082] Specifically, in this embodiment, a is preferably 12, b is preferably 12, and n is preferably 6;
[0083] S13: Add labels to the collected dynamic sign language samples to form a dynamic sign language data set;
[0084] S2: Extract features from the data after frame filling, as Figure 3 shown, based on the palm normal vector L N, the palm center position L C, and the finger tip positions L F i (i = 1,..., 5) collected by Leap Motion, 3 single - finger features and 2 double - finger features are extracted, which are the distance D i between the finger tip and the palm center, the angle A i formed by the line connecting the finger tip and the palm center and the palm plane, the distance H i from the finger tip to the palm plane, the distance DD i between adjacent finger tips, and the angle AA i formed by the line connecting adjacent finger tips and the palm center;
[0085] S21: As Figure 4 shown, through the coordinate transformation formula, transform the vector description L P in the Leap Motion device center coordinate system L into the vector description H P in the palm coordinate system H:
[0086]
[0087] Among them, L P represents the coordinate value in the Leap Motion coordinate system L, H P represents the coordinate value in the palm coordinate system H, T represents the transformation matrix, and the palm normal vector H N, the palm center position H C, and the finger tip positions H F i in the palm coordinate system H, i = 1,..., 5;
[0088] S22: The single-finger gesture features extracted include: the distance D between the fingertip and the palm center i , the angle A formed by the line connecting the fingertip and the palm center and the palm plane i , the distance H from the fingertip to the palm plane i , where the distance D between the fingertip and the palm center i is obtained through the following formula:
[0089] D i =|| L F i - L C|| / D max
[0090] where, D i represents the Euclidean distance from the fingertip to the palm center position, and D max represents the maximum distance value from all fingertips to the palm center position. The distance value between the fingertip and the palm center is normalized to the interval [0, 1];
[0091] The angle A formed by the line connecting the fingertip and the palm center and the palm plane i is obtained through the following formula:
[0092] A i =π / 2 - ∠( H F i - H C, H N)(i = 1,..., 5)
[0093] where, ∠( H F i - H C, H N) represents the angle between the line connecting the fingertip and the palm center and the normal vector;
[0094] The distance H from the fingertip to the palm plane i is obtained through the following formula:
[0095] H i =D i ·D max cos(∠( H F i - H C, H N))(i = 1,..., 5)
[0096] where, D i ·D max represents calculating the distance from the fingertip to the palm center position;
[0097] S23: The extracted two-finger gesture features include: the distance DD between the fingertips of two adjacent fingers i , and the angle AA between the lines connecting the fingertips of two adjacent fingers and the center of the palm i , where the distance DD between the fingertips of two adjacent fingers i is obtained through the following formula:
[0098] DD i = || H F i - H F i+1 || / DD max (i = 1,..., 4)
[0099] where H F i+1 represents the adjacent finger fingertips of H F i , DD i represents the Euclidean distance between the fingertips of two adjacent fingers, and DD max represents the maximum distance between the fingertips of two adjacent fingers, which is used to normalize DD i to the interval [0, 1];
[0100] The angle between the lines connecting the fingertips of two adjacent fingers and the center of the palm is obtained through the following formula:
[0101] AA i = ∠( H F i - H C, H F i+1 - H C)(i = 1,..., 4)
[0102] First, find the vectors of the positions of the two adjacent fingertips and the center of the palm, and then find the angle between the two fingertip vectors, that is, find the angle between the lines connecting the two adjacent finger fingertips and the center of the palm;
[0103] S24: The gesture features extracted from each frame of the dynamic sign language sample form a 23-dimensional feature vector. The extracted gesture features are used to form a dynamic sign language feature dataset, and the dynamic sign language feature dataset is divided into a training dataset and a test dataset according to a ratio of 8:2;
[0104] S3: As Figure 5 shown, build a dynamic sign language recognition model based on a two-layer bidirectional long short-term memory network, train it with the training dataset, and test it with the test dataset;
[0105] S31: Configure the input layer, bidirectional LSTM layer, fully connected layer, and output layer of the model to ensure that the dynamic sign language recognition model can correctly process and model the time series characteristics of dynamic sign language;
[0106] Specifically, in this embodiment, the output layer dimension of the two-layer bidirectional long short-term memory network is set to 23, the output dimension of the two-layer bidirectional long short-term memory network and the output dimension of the fully connected layer are set to 10, the activation function of the fully connected layer is set to the Softmax function, and a random deactivation function is added after each layer of the bidirectional neural network, and the random deactivation rate is set to 0.5;
[0107] S32: training the model using the training data set, optimizing the training process by adjusting the training parameters, evaluating the effect of the model using the test data set, saving the dynamic sign language recognition model with the best accuracy in the test data set as the optimal model, and saving the trained model parameters for use in real-time dynamic sign language recognition;
[0108] Specifically, in this embodiment, the hyperparameter batch_size is set to 30 and the epoch is set to 500 to ensure that the model is fully trained and avoid overfitting;
[0109] S4: Apply the optimal model to the elevator system with two-digit floors;
[0110] S41: Using the Leap Motion sensor to collect dynamic sign language in real time: The first digit in the dynamic sign language represents the digit in the tens place, and the second digit in the dynamic sign language represents the digit in the ones place. This is in line with the public's habit of writing multi-digit numbers from high to low.
[0111] S42: The dynamic sign language data collected in real time each time is processed by key frame extraction and frame supplementation technology to generate 30 frames of data;
[0112] S43: input the processed 30 frames of data into the optimal model for real-time dynamic sign language recognition, and output the recognized sign language result as the number of elevator floors;
[0113] S44: Output the corresponding sign language category according to the recognition result, and the stairs reach the designated floor.
[0114] Experimental results show that the present invention can adapt to the hand shapes, hand speeds and reaction times of different users, and can also achieve accurate real-time sign language recognition in practical applications such as elevator systems, with high adaptability, real-time performance and accuracy.
[0115] Example 2
[0116] This embodiment provides a sign language recognition system based on Leap Motion sensor and deep learning, including: a dynamic sign language data set acquisition module, a feature extraction module, a data set division module, a dynamic sign language recognition model construction module, a dynamic sign language recognition model training module, and a dynamic sign language recognition module;
[0117] In this embodiment, the dynamic sign language data set acquisition module is used to acquire a dynamic sign language data set based on a Leap Motion sensor;
[0118] In this embodiment, the feature extraction module is used to extract features from the dynamic sign language data set, extract single-finger features and double-finger features, and form a dynamic sign language feature data set;
[0119] In this embodiment, the data set division module is used to divide the dynamic sign language feature data set into a training data set and a test data set;
[0120] In this embodiment, the dynamic sign language recognition model construction module is used to construct a dynamic sign language recognition model based on a two-layer bidirectional long short-term memory network;
[0121] In this embodiment, the dynamic sign language recognition model training module is used to train the dynamic sign language recognition model based on the training data set, test the dynamic sign language recognition model based on the test data set, and use the dynamic sign language recognition model with the optimal accuracy as the trained dynamic sign language recognition model;
[0122] In this embodiment, the dynamic sign language recognition module is used to perform dynamic sign language recognition based on the trained dynamic sign language recognition model and output the corresponding sign language category.
[0123] In this embodiment, the dynamic sign language data set acquisition module is used to acquire a dynamic sign language data set based on a Leap Motion sensor, specifically including:
[0124] Each frame of data includes the palm normal vector L N, the palm center position L C, the finger tip position L F i ;
[0125] When collecting each dynamic sign language sample, copy a frames of data before the key frame as the first a frames of the n + a + b frames of data, and copy the last frame of the consecutive n frames of key frames into b copies as the last b frames of the n + a + b frames of data;
[0126] Add labels to the collected dynamic sign language samples to form a dynamic sign language data set.
[0127] In this embodiment, the feature extraction module is used to extract features from the dynamic sign language data set, extract single-finger features and double-finger features, specifically including:
[0128] Through the coordinate transformation formula, transform the vector description in the Leap Motion device center coordinate system L into the vector description in the palm coordinate system H;
[0129] In this embodiment, the single - finger gesture features include: the distance between the fingertip and the palm center, the angle formed by the line connecting the fingertip and the palm center and the palm plane, and the distance from the fingertip to the palm plane;
[0130] In this embodiment, the two - finger gesture features include the distance between the adjacent two fingertips and the angle between the lines connecting the adjacent two fingertips and the palm center.
[0131] In this embodiment, the distance between the fingertip and the palm center is expressed as:
[0132] D i =|| L F i - L C|| / D max
[0133] Where D i represents the Euclidean distance from the fingertip L F i to the palm center position L C in the central coordinate system L of the Leap Motion device. D max represents the maximum distance value from all fingertips to the palm center position, and normalizes the distance value between the fingertip and the palm center to the interval [0, 1];
[0134] In this embodiment, the angle formed by the line connecting the fingertip and the palm center and the palm plane is expressed as:
[0135] A i =π / 2 - ∠( H F i - H C, H N)
[0136] Where A i represents the angle formed by the line connecting the fingertip H F i and the palm center H C and the palm plane H N in the palm coordinate system H. ∠( H F i - H C, H N) represents the angle between the line connecting the fingertip H F i and the palm center H C and the normal vector, i = 1,..., 5;
[0137] In this embodiment, the distance from the fingertip to the palm plane is expressed as:
[0138] H i =D i ·Dmax cos(∠( H F i - H C, H N))
[0139] wherein, H i represents the distance from the fingertip to the palm plane, and D i ·D max represents the calculated distance from the fingertip to the center position of the palm.
[0140] In this embodiment, the distance between the fingertips of two adjacent fingers is expressed as:
[0141] dD i =|| H F i - H F i+1 || / DD max
[0142] wherein, DD i represents the Euclidean distance between the fingertips of two adjacent fingers, H F i+1 represents the adjacent fingertip of the fingertip in the palm coordinate system H H F i , and DD max represents the maximum distance between the fingertips of two adjacent fingers;
[0143] In this embodiment, the included angle between the connection lines of the fingertips of two adjacent fingers and the palm center is expressed as:
[0144] AA i =∠( H F i - H C, H F i+1 - H C)
[0145] wherein, AA i represents the included angle between the connection lines of the fingertips of two adjacent fingers and the palm center, H and C represents the palm center in the palm coordinate system H.
[0146] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A sign language recognition method based on Leap Motion sensor and deep learning, characterized in that, Including the following steps: Collecting a dynamic sign language data set based on a Leap Motion sensor; Performing feature extraction on the dynamic sign language data set, extracting single-finger features and double-finger features to form a dynamic sign language feature data set; Dividing the dynamic sign language feature data set into a training data set and a test data set; Constructing a dynamic sign language recognition model based on a two-layer bidirectional long short-term memory network; Training the dynamic sign language recognition model based on the training data set, testing the dynamic sign language recognition model based on the test data set, and using the dynamic sign language recognition model with the optimal accuracy as the trained dynamic sign language recognition model; Performing dynamic sign language recognition based on the trained dynamic sign language recognition model and outputting the corresponding sign language category.
2. The sign language recognition method based on Leap Motion sensor and deep learning according to claim 1, characterized in that, Collecting a dynamic sign language data set based on a Leap Motion sensor, specifically including: Each frame of data includes the palm normal vector L N, the palm center position L C, the finger tip positions L F i ; When collecting each dynamic sign language sample, copying the data of one frame before the key frame a times as the first a frames of the n + a + b frames of data, and copying the last frame of the consecutive n key frames b times as the last b frames of the n + a + b frames of data; Adding labels to the collected dynamic sign language samples to form a dynamic sign language data set.
3. The sign language recognition method based on Leap Motion sensor and deep learning according to claim 1, characterized in that Performing feature extraction on the dynamic sign language data set, extracting single-finger features and double-finger features, specifically including: Transforming the vector description in the Leap Motion device center coordinate system L into a vector description in the palm coordinate system H through a coordinate transformation formula; The single-finger gesture features include: the distance between the finger tip and the palm center, the angle formed by the line connecting the finger tip and the palm center and the palm plane, and the distance from the finger tip to the palm plane; The double-finger gesture features include the distance between the fingertips of two adjacent fingers and the angle between the line connecting the fingertips of two adjacent fingers and the palm center.
4. The sign language recognition method based on Leap Motion sensor and deep learning according to claim 3, wherein, The distance between the finger tip and the palm center is expressed as: D i = || L F i - L C || / D max Among them, D i represents the Euclidean distance from the fingertip L F i to the palm center position L C in the central coordinate system L of the Leap Motion device. D max represents the maximum distance value from all fingertip positions to the palm center position, and normalizes the distance value between the fingertip and the palm center to the interval [0, 1]; The angle formed by the line connecting the finger tip and the palm center and the palm plane is expressed as: A i = π / 2 - ∠( H F i - H C, H N) Among them, A i represents the included angle formed by the line connecting the fingertip H F i and the palm center H C and the palm plane H N, ∠( H F i - H C, H N) represents the included angle between the line connecting the fingertip H F i and the palm center H C and the normal vector, i = 1,..., 5; The distance from the finger tip to the palm plane is expressed as: H i = D i · D max cos(∠( H F i - H C, H N)) Among them, H i represents the distance from the fingertip to the palm plane, and D i ·D max represents the calculated distance from the fingertip to the center position of the palm.
5. The sign language recognition method based on Leap Motion sensor and deep learning according to claim 3, characterized in that, The distance between the fingertips of two adjacent fingers is expressed as: DD i = || H F i - H F i+1 || / DD max Among them, DD i represents the Euclidean distance between the fingertips of two adjacent fingers, H F i+1 represents the adjacent fingertip of the fingertip under the palm coordinate system H H F i ; DD max represents the maximum distance between the fingertips of two adjacent fingers; The angle between the line connecting the fingertips of two adjacent fingers and the palm center is expressed as: AA i = ∠( H F i - H C, H F i+1 - H C) Among them, AA i represents the angle between the lines connecting the fingertips of two adjacent fingers and the center of the palm, H C represents the center of the palm in the palm coordinate system H.
6. A sign language recognition system based on Leap Motion sensor and deep learning, characterized in that, Including: A dynamic sign language data set collection module, a feature extraction module, a data set division module, a dynamic sign language recognition model construction module, a dynamic sign language recognition model training module, and a dynamic sign language recognition module; The dynamic sign language data set collection module is used to collect a dynamic sign language data set based on a Leap Motion sensor; The feature extraction module is used to perform feature extraction on the dynamic sign language data set, extract single-finger features and double-finger features to form a dynamic sign language feature data set; The data set division module is used to divide the dynamic sign language feature data set into a training data set and a test data set; The dynamic sign language recognition model construction module is used to construct a dynamic sign language recognition model based on a two-layer bidirectional long short-term memory network; The dynamic sign language recognition model training module is used to train the dynamic sign language recognition model based on the training data set, test the dynamic sign language recognition model based on the test data set, and use the dynamic sign language recognition model with the optimal accuracy as the trained dynamic sign language recognition model; The dynamic sign language recognition module is used to perform dynamic sign language recognition based on the trained dynamic sign language recognition model and output the corresponding sign language categories.
7. The sign language recognition system based on Leap Motion sensor and deep learning according to claim 6, characterized in that, The dynamic sign language data set acquisition module is used to acquire a dynamic sign language data set based on the Leap Motion sensor, specifically including: Each frame of data includes the palm normal vector L N, the palm center position L C, the finger tip positions L F i ; When collecting each dynamic sign language sample, copy a frames of data before the key frame as the first a frames of the n + a + b frames of data, and copy the last frame of the consecutive n key frames into b copies as the last b frames of the n + a + b frames of data; Add labels to the collected dynamic sign language samples to form a dynamic sign language data set.
8. The sign language recognition system based on Leap Motion sensor and deep learning according to claim 6, characterized in that, The feature extraction module is used to extract features from the dynamic sign language data set, and extract single-finger features and double-finger features, specifically including: Through the coordinate transformation formula, transform the vector description in the center coordinate system L of the Leap Motion device into the vector description in the palm coordinate system H; The single-finger gesture features include: the distance between the fingertip and the palm center, the angle formed by the line connecting the fingertip and the palm center and the palm plane, and the distance from the fingertip to the palm plane; The double-finger gesture features include the distance between the fingertips of two adjacent fingers and the angle between the line connecting the fingertips of two adjacent fingers and the palm center.
9. The sign language recognition system based on Leap Motion sensor and deep learning according to claim 8, wherein, The distance between the fingertip and the palm center is expressed as: D i = || L F i - L C || / D max Among them, D i represents the Euclidean distance from the fingertip L F i to the palm center position L C in the central coordinate system L of the Leap Motion device. D max represents the maximum distance value from all fingertip positions to the palm center position, and normalizes the distance value between the fingertip and the palm center to the interval [0, 1]; The angle formed by the line connecting the fingertip and the palm center and the palm plane is expressed as: A i = π / 2 - ∠( H F i - H C, H N) Among them, A i represents the angle formed by the line connecting the fingertip H F i and the palm center H C and the palm plane H N, ∠( H F i - H C, H N) represents the angle between the line connecting the fingertip H F i and the palm center H C and the normal vector, i = 1,..., 5; The distance from the fingertip to the palm plane is expressed as: H i = D i · D max cos(∠( H F i - H C, H N)) Among them, H i represents the distance from the fingertip to the palm plane, and D i ·D max represents the calculated distance from the fingertip to the center position of the palm.
10. The sign language recognition system based on Leap Motion sensor and deep learning according to claim 8, characterized in that, The distance between the fingertips of two adjacent fingers is expressed as: DD i = || H F i - H F i+1 || / DD max Among them, DD i represents the Euclidean distance between the fingertips of two adjacent fingers, H F i+1 represents the adjacent finger fingertips of the finger fingertips in the palm coordinate system H H F i and DD max represents the maximum distance between the fingertips of two adjacent fingers; The angle between the line connecting the fingertips of two adjacent fingers and the palm center is expressed as: AA i = ∠( H F i - H C, H F i+1 - H C) Among them, AA i represents the angle between the lines connecting the fingertips of two adjacent fingers and the center of the palm, H C represents the center of the palm in the palm coordinate system H.