A 3D gesture recognition method based on Leap motion combined with BPN neural network
Through Leap motion combined with the three-dimensional gesture recognition method of BPN neural network, the problem of low accuracy and poor convenience of gesture sign language recognition system in the prior art is solved, and efficient and simple communication between deaf and mute people and ordinary people is achieved.
Patent Information
- Application Number
- CN202110629446.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-06-07
AI Technical Summary
The existing neural network gesture sign language recognition system has low accuracy, high latency, and the convenience and cost of wearing peripherals need to be optimized.
Leap motion combined with BPN neural network three-dimensional gesture recognition method is adopted to realize gesture recognition through user-defined gesture image acquisition, long sample cutting, sample encoding and decoding and BPN neural network training.
It improves the accuracy of gesture recognition, simplifies the operation process, reduces the computational complexity, and is suitable for communication between deaf and mute people.
Smart Images

Figure CN115512432B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of somatosensory control human-computer interaction, is based on a Leap motion somatosensory sensor, and is particularly suitable for the fields of user-defined gesture recognition and sign language learning for the deaf-mute. Background Art
[0002] Since the 1990s, artificial intelligence, as an interdisciplinary and comprehensive subject involving many fields, has ushered in its own spring of development. Neural networks and data dimensionality reduction have gradually become hot topics. Neural networks abstract the human brain neuron network from the perspective of information processing, establish a simple model, and form different networks according to different connection methods. Data dimensionality reduction is to reduce high-dimensional data to low dimensions through a certain algorithm, which plays a role in reducing the amount of calculation. Among them, gesture and sign language recognition, as a key link in human-computer interaction and communication for the disabled, has also ushered in more possibilities with the east wind of artificial intelligence. As a compact somatosensory controller, Leap motion has high recognition accuracy for hand features and is easy to use.
[0003] Ovur Salih Ertug et al. proposed an autonomous learning framework for gesture recognition based on sEMG [1] Jin Biao et al. proposed a millimeter wave radar dynamic gesture recognition method based on a serial one-dimensional neural network (1D-ScNN) [2] ; Yu Xiaoyang et al. proposed a method for virtual reality gesture recognition using a depth camera of an HTC VIVE device [3] ; Chen Yingrou et al. proposed a static gesture recognition method based on multi-feature weighted fusion [4] ; Vanita Jain et al. proposed a method for recognizing American Sign Language using support vector machine (SVM) and convolutional neural network (CNN) [5] ; Roy Partha Pratim et al. proposed a novel end-to-end sign language recognition system [6] , using Camshift tracker to extract hand motion trajectory, and using sequence classification based on Hidden Markov Model (HMM) to recognize gestures; Zhong Jianmin et al. used data gloves combined with mobile phone apps to propose a low-cost, high-accuracy intelligent sign language translation system based on KNN-HMM [7] .
[0004] According to research and analysis, the existing neural network gesture and sign language recognition system has shortcomings such as low accuracy and high latency, which affects the efficiency of human-computer interaction and communication for people with disabilities. The recognition method of wearable peripherals still needs to be optimized in terms of convenience and cost.
[0005] [1] Ovur Salih Ertug, Zhou Xuanyi, Qi Wen, Zhang Longbin, Hu Yingbai, Su Hang, Ferrigno Giancarlo, De Momi Elena. A novel autonomous learning framework to enhance sEMG-based hand gesture recognition using depth information[J]. Biomedical Signal Processing and Control, 2021, 66.
[0006] [2] Jin Biao, et al. "Millimeter-wave radar dynamic gesture recognition method based on 1D-ScNN." Journal of Electronics and Information Technology. doi:10.11999 / JEIT200894.
[0007] [3] Yu Xiaoyang, Jiang Lin, wang Lijun. P-15.4: Virtual reality gesture recognition based on depth information[J]. SID Symposium Digest of Technical Papers, 2021, 52.
[0008] [4] Chen Yingrou, Tian Qiuhong, Yang Huimin, Liang Qinglong, Bao Jiaxin. Static gesture recognition based on multi-feature weighted fusion[J]. Computer Systems & Applications, 2021, 30(02): 20-27.
[0009] [5] Vanita Jain, Achin Jain, Abhinav Chauhan, Srinivasu Soma Kotla, Ashish Gautam. American Sign Language recognition using Support Vector Machine and Convolutional Neural Network[J]. International Journal of Information Technology, 2021(prepublish).
[0010] [6]Roy Partha Pratim, Kumar Pradeep, Kim Byung Gyu. An Efficient Sign Language Recognition(SLR)System Using Camshift Tracker and Hidden Markov Model(HMM)[J]. SN Computer Science, 2021, 2(2).
[0011] [7]Zhong Jianmin, Li Xiaodong, Li Jiajian, Lu Rengui, Chang Zijian. Intelligent Sign Language Translation System Based on KNN-HMM[J]. Science & Technology Vision, 2021(03): 43-46. Summary of the Invention
[0012] Technical Problem to be Solved: To overcome the problem of communication between the deaf-mute and ordinary people, the present invention proposes a three-dimensional gesture sign language recognition method based on Leap motion combined with BPN neural network with relatively high gesture recognition accuracy and simple and fast operation.
[0013] The technical solution adopted by the present invention is: A gesture sign language recognition method for hand joint information based on Leap motion and BPN neural network, and the gesture recognition method includes the following steps:
[0014] Step 1: User-specified special gesture image acquisition
[0015] Among sign languages, iconic gestures, anti-character gestures, phonetic gestures, ideographic gestures, demonstrative gestures, and comprehensive gestures are the most common classifications of sign languages, which can almost cover most finger languages. The present invention selects a part of gestures from these six gesture classifications as representatives for gesture recognition, which can ensure the functional integrity of the invention. Provide a group of gestures by the user, use this group of gestures as standard gestures, collect the information of this group of gestures, and record it as the standard group; construct its corresponding name label according to the general meaning of the standard group gestures;
[0016] Step 2: Gesture long sample acquisition
[0017] Use the Leap motion sensor to sample each defined gesture in Step 1, and continuously sample 1280 frames for each gesture;
[0018] Step 3: Gesture long sample cutting
[0019] Cut the long sample into segments of 128 frames, and each segment is a short sample; all short samples form a training set, which has the following characteristics and requirements: The training set consists of several samples. Each sample includes a short sample segment and a related gesture (the gesture defined in step 1). The short sample segments of each sample are of equal length, and the number of samples of the gestures defined in the recognition set is equal in the sample set;
[0020] Step 4: Sample encoding and decoding
[0021] Each sample segment contains 128 frames of data, and each frame of data contains 3D information (x, y, z three coordinate axes) of each finger joint; The input of the BPN neural network is a vector, and each component is expressed in float type. The encoding process is a calculation process from the sample segment data to the float vector; The output of the BPN neural network is also a float vector, and the gesture definition recognition result is a set of text information (text). The decoding process is a calculation process from the output of the BPN neural network to the recognition result;
[0022] Step 5: BPN neural network design
[0023] The input of the BPN neural network is a float component with a length of 54, and the output is a float vector with a length of 10 (the number of components is closely related to the 10 recognition sets defined in step 1). The BPN neural network is designed as a standard network with 5 layers, and the output component dimensions of each layer are: 54 -> 128 -> 64 -> 32 -> 10. Each layer contains a bias, and the activation function uses the hyperbolic tangent function (tanh);
[0024] Step 6: BPN neural network training
[0025] For the training sample set output in step 3, perform encoding and decoding through step 4, train the BPN neural network designed in step 5, iterate for several generations, and finally output a trained weight file;
[0026] Step 7: Online recognition
[0027] The system loads the trained weight file, online and in real-time recognizes the incoming hand information of the Leap motion sensor, and quickly prints the recognition result.
[0028] Furthermore, in the above step 4, the dimensionality reduction operation of the three-dimensional skeletal information of the sample encoding and decoding as the input of the BPN neural network includes:
[0029] In the first step, use formula (1) to substitute the x, y, and z axis coordinates respectively to calculate the average value of the palm positions, n is taken as 128, and calculate 128 average positions of the palm positions
[0030]
[0031] In the second step, the average position obtained in 2.1 is regarded as the coordinate origin, and the average value (x1, y1, z1) of the palm normal vectors of each of the 128 frames obtained from the drive is calculated again through formula (1).
[0032] In the third step, the standard deviation of the palm normal vector is calculated using formula (2) to obtain (x2, y2, z2).
[0033]
[0034] In the fourth step, the average position and standard deviation of the fingertips of 5 fingers are calculated using formulas (1) and (2) to obtain 5×3×2 = 30 components.
[0035] In the fifth step, for the multi-frame data included in the short sample, the data type of each frame includes ① palm normal vector, ② palm speed, ③ hand shape angle, ④ distances from 5 fingers to the palm center, and ⑤ direction coordinates of 5 fingers. The average value and standard deviation of these 5 types of data are calculated respectively. The average value is used to measure the most probable positions of all frames in the short sample, and the standard deviation is used to measure the fluctuation of the dynamic information of all frames in the short sample, that is, the changing gesture actions. For the palm normal vector, calculate the mean and standard deviation of the palm normal vectors of all frames in the short sample to obtain 2 vectors, containing 6 components (two for each of the x, y, and z coordinate axes); for the palm speed, calculate the mean and standard deviation of the palm speeds of all frames in the short sample to obtain 2 vectors, containing 6 components; for the hand shape angle, this parameter is the angle between the thumb and the 4 fingers, calculate the mean and standard deviation of the hand shape angles of all frames in the short sample to obtain 2 scalars, containing 2 components; for the distances of 5 fingers from the palm center, calculate the mean and standard deviation of all frames in the short sample, and each finger obtains 2 components; for the direction coordinates of 5 fingers, calculate the mean and standard deviation of the finger directions of 128 frames, and each finger obtains 2 vectors, containing 6 components. Through the above operations, the complex four-dimensional data of multi-frame three-dimensional space plus one-dimensional time is reduced to multi-frame one-dimensional data, which successfully meets the input format requirements of the BPN neural network.
[0036] Furthermore, perturbations of training samples are added to the BPN neural network. When training the BPN neural network in the present invention, in order to enhance the robustness of the samples, for the already cut short samples, operations such as rotation and translation are performed. Random rotation and translation operations are performed on each frame of data of the samples in the 3D space (the x, y, and z coordinate axes), which is equivalent to observing the same gesture from different positions and angles. This greatly expands the coverage of the samples and makes the sample data universal.
[0037] Furthermore, in the real-time online gesture recognition, a buffer queue with a length of 1280 frames is designed. The Leapmotion driver outputs one frame of data each time it is called back, which is pushed to the end of the queue, and a frame counter is designed to accumulate 1; if and only if the frame counter is an integer multiple of 64, the recognition engine is started for recognition once, and 128 frames are obtained from the end of the buffer queue each time as a data buffer. After encoding and dimensionality reduction in step 2, a 54-dimensional vector is obtained and input to the BPN neural network; after the BPN neural network outputs a 10-dimensional vector, it returns the confidence through decoding (decoding calculates the Euclidean distance between the standard output and the actual output, and returns the result with the closest distance). When the Euclidean distance is less than 0.3, the recognition is judged to be successful, otherwise it returns failure; when the recognition is successful, the buffer queue is cleared; otherwise, the buffer queue cannot be cleared.
[0038] The beneficial effects of the present invention are as follows: a three-dimensional gesture and sign language recognition method based on Leap motion combined with a neural network, especially a gesture and sign language recognition method combined with a common sign language behavior analysis model for the deaf-mute, embodies humanized considerations for different individual users. The input and output of the classic framework of the BPN neural network can only be one-dimensional. This solution is particularly aimed at the limitation of the dimensionality of the input data of the BPN neural network, and adopts a dimensionality reduction method to calculate and reorganize the four-dimensional gesture coordinate information extracted in real time by the hardware driver into one-dimensional information, incorporating the idea of dimensionality reduction. In the specific dimensionality reduction operation, the x, y, and z coordinate information of 128 frames of data are averaged, and then the standard deviation of the important information is calculated, so that the dynamic information is completely accommodated in one data, which greatly facilitates the judgment of the subsequent BPN neural network training to recognize sign language actions. The risk brought by the "dimensionality disaster" is greatly reduced, further reflecting the importance of data dimensionality reduction. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 Schematic diagram of long sample cutting
[0040] Figure 2 This is a sample encoding and decoding modeling diagram
[0041] Figure 3 is the total picture of all data components
[0042] Figure 4 Define gesture output as a 10-dimensional vector graph
[0043] Figure 5 It is a flow chart of all steps of the present invention
[0044] Figure 6 This is the output dimension diagram of each layer of the BPN neural network design
[0045] Figure 7 This is the BPN neural network training flow chart
[0046] Figure 8 is the online recognition flow chart
[0047] Figure 9 are the schematic diagrams of 10 custom gestures Detailed implementation manners
[0048] The following combines the accompanying drawings and the detailed implementation manners to further explain the detailed implementation manners of the present invention. The following examples or drawings are used to explain the present invention, but are not limited to explaining the scope of the present invention.
[0049] The implementation process of the present invention is divided into 3 stages, including a total of 7 steps, and some steps are common to multiple processes. These processes, steps, and operation flows are shown in Figure 5 . The details of each step will be explained in detail below.
[0050] Step 1: Definition of the recognition set. The first step of the system must define a finite number of gestures. As an example, the system defines the following 10 gestures:
[0051] Gesture 1 = "OK"
[0052] Gesture 2 = "No"
[0053] Gesture 3 = "Bye"
[0054] Gesture 4 = "Victory"
[0055] Gesture 5 = "Fist"
[0056] Gesture 6 = "Strong"
[0057] Gesture 7 = "Weak"
[0058] Gesture 8 = "Come"
[0059] Gesture 9 = "Love you"
[0060] Gesture 10 = "Pistol"
[0061] The 10 custom gestures are shown in Figure 9 .
[0062] Step 2: Long sample collection. Using the Leap motion sensor, sample each defined gesture in Step 1. Each gesture is sampled continuously for 1280 frames (about 10 seconds). Save it as a file, so that this file is associated with the specific gesture name. Each gesture sample is collected once, and a total of 10 files are formed to constitute a complete set of long samples. When training samples, the complete set of samples must be used as input for cutting to ensure that the number of training samples for each gesture is the same.
[0063] Step 3: Long sample cutting. Cut the long sample into segments of 128 frames, and each segment is a short sample. Each long sample has a length of 1024 + 256 frames, each short sample has a length of 128 frames, and the cutting step size is 8 frames. Each long sample will be cut into 128 short samples. 1024 / 8 = 128. See Figure 1 . All short samples form a training set, and the number of samples for each gesture in the training set is equal.
[0064] Step 4: Sample encoding. Each training sample = 128-frame gesture segment + gesture name. The process of converting the 128-frame gesture segment into a 54-dimensional vector is called input encoding. The process of converting the gesture name into a 10-dimensional vector is called output encoding. The 54-dimensional vector generated by input encoding meets the input requirements of the BPN neural network; the 10-dimensional vector generated by output encoding meets the output dimension of the BPN neural network.
[0065] Input encoding uses a dimensionality reduction algorithm. According to the selection of driving data in the modeling process, 5 data items of the driving data are selected and calculated using mathematical statistics methods to generate one-dimensional data.
[0066] Each short sample contains 128 frames of data. The data type of each frame includes ① palm normal vector, ② palm speed, ③ hand shape angle, ④ distances from 5 fingers to the palm center, and ⑤ 5-finger direction coordinates. Calculate the mean and standard deviation of these 5 types of data respectively. The mean is used to measure the most likely position of the 128 frames of the short sample, and the standard deviation is used to measure the fluctuation of the dynamic information of the 128 frames of the short sample, that is, the changing gesture actions;
[0067] (1) Palm normal vector, calculate the mean and standard deviation of the palm normal vectors of 128 frames to obtain 2 vectors, containing 6 components.
[0068] (2) Palm speed, calculate the mean and standard deviation of the palm speeds of 128 frames to obtain 2 vectors, containing 6 components.
[0069] (3) Hand shape angle, this parameter is the angle between the thumb and the 4 fingers, and it is a float scalar. Calculate the mean and standard deviation of the hand shape angles of 128 frames to obtain 2 scalars, containing 2 components.
[0070] (4) Distances of 5 fingers from the palm center, these are float scalars. Calculate the mean and standard deviation of 128 frames, and each finger gets 2 components.
[0071] (5) 5-finger directions, calculate the mean and standard deviation of the finger directions of 128 frames, and each finger gets 2 vectors, containing 6 components.
[0072] According to the above method, the four-dimensional sample (one time dimension plus three spatial dimensions of x, y, and z) is reduced to a one-dimensional vector. The meaning of each component is shown in Figure 3 . According to the definition of the recognition set in Step 1 for the output encoding (taking 10 gestures as an example), the output encoding of the BPN neural network is defined as a 10-dimensional vector, as shown in Figure 4 . The output of the BPN neural network is defined as 10 points in a 10-dimensional vector space, and the distance between any two points is equal, all being , meeting the characteristics of discreteness and equidistance.
[0073] Step 5: Design of the BPN neural network. The BPN neural network is designed as a standard 5-layer network. The output component dimensions of each layer are: 54, 128, 64, 32, and 10 respectively. Each layer contains a bias, and the hyperbolic tangent function (tanh) is used as the activation function. The output dimensions of each layer are shown in Figure 6 .
[0074] Step 6: Training of the BPN neural network. For the training sample set output in Step 3, after encoding and decoding through Step 4, the BPN neural network is trained for several generations, and finally a trained weight file is output. The training error of the BPN neural network is carried out according to the standard backpropagation. The network training process is shown in Figure 7 .
[0075] Step 7: Online recognition. The system loads the weight file trained in Step 6, online and in real time recognizes the incoming hand information of the Leapmotion sensor, and quickly prints the recognition result.
[0076] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A three-dimensional gesture recognition method based on Leap motion combined with a BPN neural network, characterized in that, The gesture recognition method includes the following stages: Stage 1: User-specified special gesture image acquisition A set of gestures is provided by the user. This set of gestures is used as the standard gestures, and the information of this set of gestures is collected and recorded as the standard group; According to the meanings of the standard group of gestures, the corresponding name tags are constructed; Training samples are constructed from the original samples. The original samples are used as the input source of the training samples, and perturbation processing operations are performed on them, and the original samples are cut into several short samples, which are recorded as the training sample group, in order to achieve the purpose of enhancing robustness and diversifying samples; Stage 2: Dimensionality reduction operation on the three-dimensional skeleton information used as the input of the BPN neural network The real-time three-dimensional skeleton information of the hand extracted from the Leap motion driver is a set of multi-frame dynamic multi-dimensional information. Each frame contains x, y, and z axis information, plus the time axis, forming four-dimensional data. This step reduces the above-obtained four-dimensional data to one-dimensional and uses it as the input vector of the BPN neural network; Stage 3: Gesture recognition by the BPN neural network Design the preliminary structure model of the BPN neural network, train and test and adjust the BPN neural network model with the sample data after dimensionality reduction in Stage 2, iteratively train the BPN neural network, and finally output the trained weight file. Immediately, the system loads this weight file to perform real-time recognition on the three-dimensional skeleton information transmitted by the Leap motion sensor and quickly print the recognition result; The specific dimensionality reduction operation method in Stage 2 is as follows: The short samples contain multiple frames of data. The data types of each frame include the palm normal vector, palm speed, hand shape angle, the distance from each finger to the palm center, and the direction coordinates of each finger. Calculate the mean and standard deviation of these 5 types of data in turn. The mean is used to measure the most likely position of all frames of the short sample, and the standard deviation is used to measure the fluctuation of the dynamic information of all frames of the short sample, that is, the changing gesture actions; for the palm normal vector, calculate the mean and standard deviation of the palm normal vectors of all frames of the short sample to obtain 2 vectors, containing 6 components, which are two for each of the x, y, and z coordinate axes; for the palm speed, calculate the mean and standard deviation of the palm speeds of all frames of the short sample to obtain 2 vectors, containing 6 components; for the hand shape angle, this parameter is the angle between the thumb and the 4 fingers. Calculate the mean and standard deviation of the hand shape angles of all frames of the short sample to obtain 2 scalars, containing 2 components; for the distance from each finger to the palm center, calculate the mean and standard deviation of all frames of the short sample, and each finger gets 2 components; for the direction coordinates of each finger, calculate the mean and standard deviation of the finger directions of all frames of the short sample, and each finger gets 2 vectors, containing 6 components; through the above operations, the complex multi-frame three-dimensional space data plus one-dimensional time data of multiple data is reduced to multi-frame one-dimensional data, successfully meeting the input format requirements of the BPN neural network.
2. The 3D gesture recognition method based on Leap motion combined with BPN neural network according to claim 1, characterized in that In Stage 3: (1) BPN neural network structure The input of the BPN neural network is a 54-dimensional vector, and the output is a 10-dimensional vector; the BPN neural network is designed as a standard network with 5 layers, and the output component dimensions of each layer are: 54 -> 128 -> 64 -> 32 -> 10; each layer contains a bias, and the hyperbolic tangent function is used as the activation function; (2) Decision criteria for online gesture recognition A buffer queue with a length of 1280 frames is designed. Each time the Leap motion driver makes a callback, one frame of data is output and pushed to the tail of the queue, and the frame counter is incremented by 1; when and only when the frame counter is an integer multiple of 64, the recognition engine is started for recognition once, and each time 128 frames are obtained from the tail of the buffer queue as the data buffer. After encoding and dimensionality reduction in step 2, a 54-dimensional vector is obtained and input to the BPN neural network; after the BPN neural network outputs a 10-dimensional vector, the confidence level is returned via decoding. If the Euclidean distance is less than 0.3, the recognition is determined to be successful, otherwise it returns failure; when the recognition is successful, the buffer queue is cleared; otherwise, the buffer queue cannot be cleared.
Citation Information
Patent Citations
Dynamic gesture recognition method and system based on deep neural network
CN108932500A
Sign language word recognition method
CN111913575A