A gesture segmentation and recognition method, electronic device, and storage medium

Through intelligent data gloves, gesture action data is collected, and the sliding window and similar feature judgment method is used to realize the precise segmentation and recognition of gesture actions, solving the problem of insufficient personalized adaptation and transition gesture detection in the prior art, and significantly improving the accuracy and efficiency of gesture recognition.

CN119646605BActive Publication Date: 2025-06-13HARBIN INST OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411672462.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-06-13
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

The existing gesture recognition technology is difficult to achieve personalized adaptation to different users, and the detection and recognition of transition gesture actions are insufficient, resulting in inaccuracy and efficiency of gesture recognition.

Method used

The gesture action data is collected through intelligent data gloves, the data is divided using a fixed sliding window, and by calculating the similar features inside the window and the dissimilar features of the adjacent window, the gesture window is judged as a basic gesture action window or a transition gesture action window, and precise segmentation and identification are performed.

Benefits of technology

The accurate segmentation of continuous gesture movements is achieved, and the accuracy and efficiency of gesture recognition is improved. The accuracy rate of gesture segmentation is as high as 94.8%, the accuracy rate of basic gesture recognition reaches 96.3%, and the rate of transition gesture recognition reaches 93.2%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119646605B_ABST
    Figure CN119646605B_ABST
Patent Text Reader

Abstract

A gesture segmentation and recognition method, electronic device, and storage medium, belonging to the technical field of gesture recognition. To solve the problem of fast and accurate gesture recognition, the present invention includes using an intelligent data glove to collect gesture action data, dividing the collected gesture action data into n consecutive window data using a fixed sliding window, further dividing each window data into three sub-windows of equal size, calculating the internal similarity features between the sub-windows of each window using a distance calculation function, establishing an adjacent window similarity feature judgment method for judging whether a gesture window is a basic gesture action window or a transitional gesture action window, and respectively obtaining the basic gesture action window and the transitional gesture action window; performing upsampling processing on the transitional gesture action window, and respectively inputting the obtained basic gesture action window and the processed transitional gesture action window into a gesture action recognition model for recognition to obtain the basic gesture action and the transitional gesture action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of gesture recognition, and particularly relates to a gesture segmentation and recognition method, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of wearable technology and artificial intelligence technology, gesture recognition technology has been widely applied in various fields. Due to its advantages such as unrestricted usage scenarios and obvious interaction features, gesture recognition technology based on a configured multi-sensor data glove has also achieved numerous applications in the field of human-computer interaction. As a preprocessing technology for gesture recognition, the segmentation effect of gesture segmentation directly determines the accuracy of gesture recognition. The existing fixed-threshold segmentation algorithm has extremely poor segmentation effects for new users with different operation habits, and it is difficult to achieve personalized adaptation for different users.

[0003] The technical solution disclosed in the invention patent with the application number 202311327282.1 and the invention name "A High-Precision Gesture Data Recognition Method, Electronic Device, and Storage Medium" is as follows: "Collect new target gesture data, construct a new target gesture data set, obtain source domain gesture data, and construct a source domain gesture data set; construct a useless gesture filtering model based on the mPUL algorithm and the TSC algorithm; use the constructed useless gesture filtering model to filter the new target gesture data to obtain target domain gesture data, and construct a target domain gesture data set; construct a cross-domain gesture recognition model based on transfer learning; input the collected source domain gesture data set and target domain gesture data set into the constructed cross-domain gesture recognition model based on transfer learning to perform gesture recognition from the source domain to the target domain, and obtain a high-precision gesture data recognition result." However, the above technology does not consider the detection and recognition of transitional gesture actions, and uses a deep learning algorithm to filter useless gestures, which requires high computing resources and has a slow inference speed. Summary of the Invention

[0004] The problem to be solved by the present invention is to quickly and accurately recognize gestures, and a gesture segmentation and recognition method, an electronic device, and a storage medium are proposed.

[0005] To achieve the above object, the present invention is realized through the following technical solutions:

[0006] A gesture segmentation and recognition method includes the following steps:

[0007] S1. Use an intelligent data glove to collect gesture action data, and divide the collected gesture action data into n consecutive window data using a fixed sliding window to obtain a continuous window data set X = {x 1 , x 2 , …, x n}, where the i-th window data is marked as x i ;

[0008] S2. Divide each window data obtained in step S1 into three sub - windows of equal size. Then the i - th window data is divided into three sub - window data, which are respectively

[0009] S3. For the sub - window data of each window obtained in step S2, use the distance calculation function to calculate the internal similarity features between the sub - windows of each window;

[0010] S4. Based on the internal similarity features between the sub - windows of each window obtained in step S4, calculate the absolute difference between the minimum value of the internal similarity features in the previous window and the maximum value of the internal similarity features in the current window, and the absolute difference between the minimum value of the internal similarity features in the current window and the maximum value of the internal similarity features in the previous window. Establish a method for judging the similarity features of adjacent windows to judge whether the gesture window is a basic gesture action window or a transitional gesture action window, and obtain the basic gesture action window and the transitional gesture action window respectively;

[0011] S5. Upsample the transitional gesture action window obtained in step S4;

[0012] S6. Input the obtained basic gesture action window and the processed transitional gesture action window into the gesture action recognition model for recognition respectively to obtain the basic gesture action and the transitional gesture action.

[0013] Further, the sensors on the intelligent data glove in step S1 are a three - axis accelerometer and a three - axis gyroscope, and the i - th window data x i ∈R w×C , where w is the window size and C is the number of channels of the sensor.

[0014] Further, the specific implementation method of step S3 includes the following steps:

[0015] S3.1. For the i - th window x i , the internal similarity feature of the i - th window is calculated by the following formula:

[0016]

[0017] where, is the internal similarity feature of the first sub - window of the i - th window, is the internal similarity feature of the second sub - window of the i - th window, is the internal similarity feature of the third sub - window of the i - th window, and d is the Manhattan distance calculation function; the Manhattan distance calculation function is as follows

[0018] d(sx p ,sxq ) = ∑|sx p - sx q |

[0019] where d(sx p , sx q ) refers to the distance between sub - window p and sub - window q;

[0020] S3.2. Calculate the Manhattan distance calculation function using accelerometer and gyroscope data, that is, replace sx p , sx q in the formula with accelerometer and gyroscope data, and obtain the internal similarity feature f pq for the p - th sub - window and the q - th sub - window, and the calculation formula is:

[0021]

[0022] where a is the acceleration data, g is the gyroscope data, and by dividing the window into three sub - windows, represents the average acceleration value of the j - axis in the q - th sub - window, represents the average angular velocity value of the j - axis in the q - th sub - window, the j - axis is the x - axis or y - axis or z - axis, p = 1 or 2 or 3, q = 1 or 2 or 3, and p and q are not equal.

[0023] Furthermore, the specific implementation method of step S4 includes the following steps:

[0024] S4.1. Set the (i - 1) - th window as the previous window and the i - th window as the current window, and calculate the absolute value b between the minimum value of the internal similarity feature in the previous window and the maximum value of the internal similarity feature in the current window i-1 , and the calculation formula is:

[0025]

[0026] where min is the minimum value function and max is the maximum value function;

[0027] S4.2. Calculate the absolute difference e between the minimum value of the internal similarity feature in the current window and the maximum value of the internal similarity feature in the previous window i , and the calculation formula is:

[0028]

[0029] S4.3. Establish a method for judging the similarity feature of adjacent windows:

[0030] S4.3.1. Set the detection flag indicating the transitional gesture action window as Tag t , and the default situation is Tagt Set as a non-transition gesture, the start threshold of the transition gesture action is TH ts , and the end threshold of the transition gesture is TH te ;

[0031] S4.3.2. In the case where Tag t is a non-transition gesture, compare e i with the start threshold TH of the transition gesture action ts . If e i is greater than TH ts , the previous window is regarded as the basic gesture action window, and the current window is regarded as the start of the transition gesture action window, so the current window is added to the transition gesture action window X tans ; otherwise, the previous window is regarded as the transition gesture action window.

[0032] Furthermore, the specific implementation method of step S5 includes the following steps:

[0033] S5.1. In the case where Tag t is the transition gesture action window, start aggregating the transition window by adding the current window to the buffer X tans ; the aggregation process is to compare b i-1 with the end threshold TH of the transition gesture te . If b i-1 is higher than the threshold TH te , it is considered that the end of the transition gesture action window will be reached, and the previous window is aggregated as the end of the basic gesture action;

[0034] All detected transition windows are aggregated into X tans . Then, perform upsampling on X tans to adjust its size to a fixed window size to match the input size of the recognition model; otherwise, the window aggregation will continue in the next round;

[0035] S5.2. When the process of aggregating the window exceeds the maximum length of the transition gesture action, it is processed as a static gesture in the basic gesture action;

[0036] S5.3. When there are mixed samples of transition gestures and basic gestures in both windows, the value e of two adjacent windows is higher than the threshold transition gesture start threshold TH ts . In this case, the latest window is considered as the start of the transition gesture action window.

[0037] Furthermore, the specific implementation method of step S6 includes the following steps:

[0038] In S6.1, when constructing two gesture action recognition models, each gesture action recognition model is a CNN-LSTM network model architecture;

[0039] In S6.2, input the obtained basic gesture action window and the processed transitional gesture action window into the two gesture action recognition models constructed in step S6.1 for recognition, to obtain the basic gesture action and the transitional gesture action.

[0040] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the gesture segmentation and recognition method are implemented.

[0041] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the gesture segmentation and recognition method is implemented.

[0042] Advantages of the present invention:

[0043] For the gesture segmentation and recognition method of the present invention, by calculating the similar features inside the window and the dissimilar features of adjacent windows, and based on the differences in motion patterns between transitional gesture actions and basic gesture actions, the accurate segmentation of continuous gesture actions is achieved, and the basic gesture actions and transitional gesture actions in continuous gestures can be clearly distinguished. In addition, the gesture segmentation and recognition method based on a fixed sliding window is improved, and the large window is divided into smaller windows to reduce the confusion of multiple gesture information caused by an overly large window. And the incomplete data segmented is aggregated to ensure the integrity of the gesture data in the data window, thereby significantly improving the accuracy of continuous gesture segmentation, and the gesture segmentation accuracy rate is as high as 94.8%.

[0044] For the gesture segmentation and recognition method of the present invention, compared with traditional gesture segmentation methods, the transitional gesture is usually included in the segmented gesture data window, resulting in a decline in gesture recognition effect. However, in the process of continuous gesture segmentation of the present invention, the influence of the transitional gesture is fully considered. By separating the basic gesture action from the transitional gesture action, the interference of the transitional gesture (invalid information) on the recognition performance of the basic gesture is avoided, ensuring that only valid information is included in the basic gesture data window, thereby significantly improving the recognition accuracy of the basic gesture. And in the present invention, two independent gesture recognition models are adopted to independently train and recognize the segmented transitional gesture and basic gesture respectively, so that each model focuses on its own gesture action set, thereby improving the recognition accuracy. The experimental results show that the recognition accuracy of the basic gesture reaches 96.3%, and the recognition rate of the transitional gesture reaches 93.2%.

[0045] A gesture segmentation and recognition method according to the present invention has the advantages of low computational resource requirements and high efficiency, and supports a minimum gesture interval time of 0.2 seconds in continuous gesture segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a flowchart of a gesture segmentation and recognition method according to the present invention;

[0047] Figure 2 is a framework diagram of a gesture segmentation and recognition method according to the present invention;

[0048] Figure 3 is a differential pattern diagram of transitional gestures and basic gesture signals of the present invention;

[0049] Figure 4 is a dissimilarity feature map between two cases of two consecutive windows of the present invention, where (a) is a case where a basic gesture action before a transitional gesture action results in a lower feature value in the current window, while the feature value in the previous window is medium-low or high, and (b) is a case where a transitional gesture action before a basic gesture action results in a medium-low or high feature value in the current window, while the feature value in the previous window is lower. DETAILED DESCRIPTION OF THE INVENTION

[0050] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only a part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention usually described and shown in the accompanying drawings here can be arranged and designed in various different configurations, and the present invention can also have other embodiments.

[0051] Therefore, the detailed description of the specific embodiments of the present invention provided in the accompanying drawings below is not intended to limit the scope of the claimed invention, but merely represents the selected specific embodiments of the present invention. All other specific embodiments obtained by those skilled in the art based on the specific embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0052] To further understand the content, features and effects of the present invention, the following specific embodiments are exemplified and combined with the attached Figure 1 - Attached Figure 4 The details are as follows:

[0053] Example 1:

[0054] A gesture segmentation and recognition method includes the following steps:

[0055] S1. Collect gesture action data using an intelligent data glove, and segment the collected gesture action data into n consecutive window data using a fixed sliding window to obtain a continuous window data set X = {x 1 , x 2 , …, x n}, where the i-th window data is marked as x i ;

[0056] Furthermore, the sensors on the intelligent data glove in step S1 are a three-axis accelerometer and a three-axis gyroscope, and the obtained i-th window data x i ∈ R w×C , w is the window size, and C is the number of channels of the sensor;

[0057] S2. Divide each window data obtained in step S1 into three sub-windows of equal size. Then the i-th window data is divided into three sub-window data, which are respectively

[0058] S3. For the sub-window data of each window obtained in step S2, use a distance calculation function to calculate the internal similarity features between the sub-windows of each window;

[0059] Furthermore, the specific implementation method of step S3 includes the following steps:

[0060] S3.1. For the i-th window x i , the calculation formula for the internal similarity feature of the i-th window is as follows:

[0061]

[0062] Where is the internal similarity feature of the first sub-window of the i-th window, is the internal similarity feature of the second sub-window of the i-th window, is the internal similarity feature of the third sub-window of the i-th window, and d is the Manhattan distance calculation function; the Manhattan distance calculation function is as follows

[0063] d(sx p , sx q ) = ∑|sx p - sx q |

[0064] Where d(sx p , sx q ) refers to the distance between sub-window p and sub-window q;

[0065] S3.2. Calculate the Manhattan distance calculation function using the accelerometer and gyroscope data, that is, substitute sx in the formula p , sx q with the accelerometer and gyroscope data, and obtain the internal similarity feature f for the p-th sub-window and the q-th sub-window pq . The calculation formula is as follows:

[0066]

[0067] where a is the acceleration data and g is the gyroscope data. By dividing the window into three sub-windows, represents the average acceleration value of the j-axis in the q-th sub-window, represents the average angular velocity value of the j-axis in the q-th sub-window, and the j-axis is the x-axis or y-axis or z-axis. p = 1 or 2 or 3, q = 1 or 2 or 3, and p and q are not equal.

[0068] Furthermore, d(sx m , sx n ) is the distance function for calculating the internal similarity. The Manhattan distance is used as the distance calculation function between two vectors. The cumulative result of the Manhattan distance can reflect the total acceleration and gyroscope changes. For example, a large change in the Manhattan distance may indicate rapid movement or sudden stop, while a small change may indicate a relatively stable motion state. This is very useful for distinguishing different gesture action patterns. Its definition is as follows:

[0069] d(sx m , sx n ) = ∑|sx m - sx n |

[0070] where d(sx m , sx n ) refers to the distance between two sub-windows sx m , sx n . The Manhattan distance is the distance between two vectors,

[0071] Furthermore, the value of the internal similarity feature is affected by the gesture action rate within the window. The more intense the gesture action within the window, the greater the difference between the sub-windows. On the contrary, the smoother the hand movement within the window, the smaller the difference between the sub-windows and the smaller the feature value. Since in continuous gestures, the basic gesture actions are more intense, the feature values of the basic gesture actions are higher than those of the transitional gesture actions.

[0072] S4. Based on the internal similarity features between sub - windows of each window obtained in step S4, calculate the absolute difference between the minimum value of the internal similarity features in the previous window and the maximum value of the internal similarity features in the current window, and the absolute difference between the minimum value of the internal similarity features in the current window and the maximum value of the internal similarity features in the previous window. Establish a method for judging the similarity features of adjacent windows to determine whether the gesture window is a basic gesture action window or a transitional gesture action window, and obtain the basic gesture action window and the transitional gesture action window respectively;

[0073] Furthermore, the signal characteristics of transitional gesture actions and basic gesture actions are different (regardless of whether the activity is static or dynamic). The signals of static gesture actions are usually low - frequency and low - amplitude, while the signals of dynamic activities are composed of high - frequency and high - amplitude. On the other hand, the signal frequency of transitional gesture actions is lower, but relatively shorter compared to basic gesture actions. Therefore, the characteristics of transitional gesture actions and basic gesture actions are different. A high difference between the characteristics of two consecutive windows indicates the presence of a transitional gesture action window at the beginning or end of the window. Therefore, the segmentation method proposed in the present invention distinguishes transitional gesture actions and basic gesture actions by calculating two dissimilarity values between two adjacent windows, that is, calculating the degree of dissimilarity between the internal similarity features of two consecutive windows.

[0074] Furthermore, the specific implementation method of step S4 includes the following steps:

[0075] S4.1. Set the (i - 1)th window as the previous window and the ith window as the current window, and calculate the absolute value b between the minimum value of the internal similarity features in the previous window and the maximum value of the internal similarity features in the current window i-1 , and the calculation formula is:

[0076]

[0077] where min is the minimum value function and max is the maximum value function;

[0078] Furthermore, b is intended to determine whether the current window is the start of a transitional gesture action. This value is measured by calculating the absolute difference between the minimum value of the internal similarity features in the previous window and the maximum value of the internal similarity features in the current window. If the current window is the start of a transitional gesture action and the previous window is the end of a basic gesture action, this value is low;

[0079] S4.2. Calculate the absolute difference e between the minimum value of the internal similarity features in the current window and the maximum value of the internal similarity features in the previous window i , and the calculation formula is:

[0080]

[0081] Further, e is intended to determine whether the current window is the end of a transitional gesture action. The value is measured by calculating the absolute difference between the minimum of the similar features inside the current window and the maximum of the similar features inside the previous window. The ratio of this difference is calculated based on the lowest feature value of the current window. Given an internal similarity feature vector fi- 1 , f i and respectively for the previous window and the current window;

[0082] Further, as Figure 4 (a) shows, the current window (x i ) is a transitional gesture action, while the previous window (x i-1 ) is a basic gesture action. In the present invention, since the starting points of each gesture action are similar, the amplitude of the transitional gesture action is relatively small. Therefore, the maximum and minimum feature values of the current window indicated by and are low values because the distance between sub-windows is small and the pattern of signal data is similar. For the previous window, the maximum feature value is a high value, and the minimum feature value is a low value or a medium value because the basic action changes from a stable state to a state of violent fluctuation, the fluctuation value of signal data is large, and the distance between sub-windows is large. According to the minimum and maximum feature values, b i-1 and e i are determined. Therefore, in this case, the values of b i-1 and e i will be small values and large values respectively.

[0083] Conversely, as Figure 4 (b) shows, when the current window (x i ) is a basic gesture action and the previous window (x i-1 ) is a transitional gesture action, indicates that the maximum feature value of the current window is a high value, and the minimum feature value is a low value or a medium value; while for the previous window, and are both low values. Then, through calculation, the values of b i-1 and e i will be large values and small values respectively.

[0084] Therefore, it can be concluded that in all cases, a large value of b can determine the start of the basic gesture action, and a large value of e is used to determine the end of the basic gesture action. The large value, medium value, and small value used here are only used to describe the feature values b and e.

[0085] S4.3. Establish a method for judging the similar features of adjacent windows:

[0086] S4.3.1. Set the detection flag indicating the transitional gesture action window to Tag t , with the default being Tag t Set it as a non - transitional gesture, with the transitional gesture action start threshold being TH ts , and the transitional gesture end threshold being TH te ;

[0087] S4.3.2. When Tag t is a non - transitional gesture, compare e i with the transitional gesture action start threshold TH ts . If e i is greater than TH ts , then the previous window is regarded as the basic gesture action window, and the current window is regarded as the start of the transitional gesture action window, thus adding the current window to the transitional gesture action window X tans ; otherwise, the previous window is regarded as the transitional gesture action window.

[0088] S5. Upsample the transitional gesture action windows obtained in step S4;

[0089] Furthermore, the specific implementation method of step S5 includes the following steps:

[0090] S5.1. When Tag t is the transitional gesture action window, start aggregating the transitional windows by adding the current window to the buffer X tans ; the aggregation process is to compare b i-1 with the transitional gesture end threshold TH te . If b i-1 is higher than the threshold TH te , it is considered that the end of the transitional gesture action window will be reached, and the previous window is aggregated as the end of the basic gesture action;

[0091] All detected transitional windows are aggregated into X tans , then, perform upsampling on X tans to adjust its size to a fixed window size to match the input size of the recognition model; otherwise, the window aggregation will continue in the next round;

[0092] S5.2. When the process of aggregating windows exceeds the maximum length of the transitional gesture action, it is processed as a static gesture in the basic gesture action;

[0093] S5.3. When there are mixed samples of transitional gestures and basic gestures in both windows, and the value e of two adjacent windows is higher than the threshold transitional gesture start threshold TH ts , then the latest window is considered as the start of the transitional gesture action window.

[0094] S6. Input the obtained basic gesture action window and the processed transitional gesture action window into the gesture action recognition model respectively for recognition, to obtain the basic gesture action and the transitional gesture action;

[0095] Furthermore, the specific implementation method of step S6 includes the following steps:

[0096] S6.1. Build 2 gesture action recognition models, and each gesture action recognition model is a CNN-LSTM network model architecture;

[0097] Furthermore, the used CNN-LSTM network model architecture is that two two-dimensional convolutional layers are used to extract the spatial features of multi-dimensional sensor data, and then connected to the long short-term memory network layer (LSTM) to model the temporal changes. Its specific model architecture is that first the data is input into a two-dimensional convolutional layer (Conv2D), and dropout is used to prevent overfitting in network training. Then it passes through another two-dimensional convolutional layer (Conv2D), and through average pooling (AveragePooling), batch normalization (BatchNormalization), dropout and flattening, the data features are extracted and processed, and dimensionality reduction is performed. Then the data is input into the long short-term memory network layer (LSTM) to model the temporal changes of the data, and the learned features are fused through the fully connected layer and finally used for the classification task. The specific configuration parameters are shown in Table 1;

[0098] Table 1 CNN-LSTM gesture recognition model parameters

[0099]

[0100] S6.2. Input the obtained basic gesture action window and the processed transitional gesture action window into the 2 gesture action recognition models built in step S6.1 respectively for recognition, to obtain the basic gesture action and the transitional gesture action.

[0101] Furthermore, use two gesture action recognition models based on the CNN-LSTM network with the same architecture to train and recognize the transitional gesture action and the basic gesture action, divide the recognition task into two models, so that each model can focus on learning its own gesture action set. After completing the gesture segmentation based on the similarity difference, the transitional gesture action window and the basic gesture action window will be input into the above model for recognition. The window size is initially set to 60 pieces of data (1.2 seconds), and will be adjusted according to different gesture recognition requirements in the later stage to adapt to different gesture recognition task requirements.

[0102] In this embodiment, first, the data signal is divided into multiple consecutive windows. Then, these windows are passed into a continuous gesture segmentation model based on data similarity, and two types of data feature types (intra-window similarity features and inter-window dissimilarity features) are extracted simultaneously to distinguish basic gesture action windows and transitional gesture action windows. Intra-window similarity is achieved by dividing the window into three sub-windows of equal size and then measuring the similarity of the signals between these sub-windows. Inter-window dissimilarity is measured by calculating the difference between the features of the current window and the previous window in a fixed sliding window. To perform gesture segmentation, the current window is the last window read from the sensor. As Figure 2 shown, the current window is labeled as x i , and the window before the current window is labeled as x i-1 . If the difference degree between two adjacent windows is very small, it indicates that there is no transitional gesture action window in either of the two windows. However, if the difference degree is high, it indicates that there is a transitional gesture action window between the two windows. Based on the dissimilarity degree, the transitional gesture action window and the basic gesture action window can be distinguished. Then, interpolation and scaling processing are performed on the data windows. Oversized windows are downsampled until the window is reduced to the length required by the gesture recognition model; undersized windows are upsampled until the window is enlarged to the length required by the gesture recognition model. Finally, the distinguished and processed gesture action windows are sent to the gesture action recognizer, where the basic gesture action windows are sent to the basic gesture action recognition model, and the transitional gesture action windows are sent to the transitional gesture action recognition model for gesture action recognition. Each gesture recognition model is implemented using LSTM-CNN deep learning.

[0103] The present invention establishes a continuous gesture dataset based on a data glove to verify the effectiveness of the model. The dataset contains 1500 continuous gesture sequences, each sequence consisting of 10 - 15 gestures, and the starting and ending points of each basic gesture and transitional gesture are manually labeled. To verify the effectiveness of transitional gesture segmentation, two advanced gesture segmentation algorithms are selected for comparison: the gesture segmentation model based on a fixed sliding window and the adaptive differential threshold gesture segmentation model.

[0104] The accuracy rate of gesture segmentation refers to the overlap rate between each gesture segmented manually (true label) and the gesture segmented automatically by the algorithm. The specific results are shown in Table 2:

[0105] Table 2 Comparison of segmentation accuracy rates of different gesture segmentation models

[0106]

[0107]

[0108] The experimental results show that the accuracy rate of the gesture segmentation model based on similarity difference proposed by the present invention is as high as 94.8%, which is better than the other two segmentation models. Although the segmentation accuracy rate of the adaptive differential threshold gesture segmentation model is very close to that of the model of the present invention (only 0.3% lower), it does not consider transitional gestures during segmentation, resulting in the segmented gestures containing both transitional gestures and basic gesture actions, which affects the subsequent recognition effect. To verify this, we input the segmented gestures into a unified recognition model for verification. The recognition models all adopt the CNN-LSTM structure proposed by the present invention to reduce the influence of other variables. The experimental results are shown in Table 3.

[0109] In addition, due to the low computational resource requirements and high efficiency of the gesture segmentation model based on similarity difference of the present invention, the minimum gesture interval time supported in continuous gesture segmentation is 0.2 seconds, and the segmentation time is the shortest compared with other gesture segmentation models.

[0110] Table 3 Comparison of recognition accuracy rates of gesture action windows segmented by different gesture segmentation models

[0111]

[0112]

[0113] As can be seen from Table 3, when using the gesture segmentation model based on a fixed sliding window, the gesture recognition effect is the worst, and the accuracy rate is 88.6%. This is because the gesture segments segmented by this model may have data missing or redundancy, resulting in poor recognition effect. Secondly, when using the adaptive differential threshold gesture segmentation model, the segmented gesture segments contain useless transitional gesture data, which affects the overall recognition effect, and the recognition accuracy rate is 92.5%. In contrast, the gesture segmentation method based on similarity difference proposed by the present invention, on the basis of considering the segmentation and recognition of transitional gestures, the gesture recognition effect is significantly better than other segmentation algorithms, the recognition accuracy rate reaches 96.3%, and the recognition accuracy rate for transitional gestures reaches 93.2%.

[0114] Embodiment 2:

[0115] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a gesture segmentation and recognition method described in Embodiment 1 are implemented.

[0116] The computer device of the present invention may be a device including a processor and a memory, such as a single-chip microcomputer including a central processing unit. And, when the processor is used to execute the computer program stored in the memory, the steps of the above-mentioned gesture segmentation and recognition method are implemented.

[0117] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0118] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.

[0119] Embodiment 3:

[0120] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the gesture segmentation and recognition method described in Embodiment 1.

[0121] The computer-readable storage medium of the present invention may be any form of storage medium readable by the processor of the computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. A computer program is stored on the computer-readable storage medium. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned gesture segmentation and recognition method can be implemented.

[0122] The computer program includes computer program code, which may be in the form of source code, object code, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROMs), random access memories (RAMs), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0123] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0124] Although the present application has been described above with reference to specific embodiments, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application can be combined with each other in any way, and the situations of these combinations are not exhaustively described in this specification only for the sake of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A gesture segmentation and recognition method, characterized in that: The steps include: S1. Use the smart data glove to collect gesture data, and use a fixed sliding window to divide the collected gesture data into n continuous window data to obtain a continuous window data set X = {x 1 ,x 2 ,…,x n }, Among them, the i-th window data is marked as x i ; S2. Divide each window data obtained in step S1 into three sub-windows of equal size. Then the i-th window data is divided into three sub-window data, namely S3. For the sub-window data of each window obtained in step S2, the internal similarity features between the sub-windows of each window are calculated using the distance calculation function; S4. Based on the internal similarity features between the sub-windows of each window obtained in step S4, the absolute difference between the minimum value of the internal similarity features in the previous window and the maximum value of the internal similarity features in the current window, and the absolute difference between the minimum value of the internal similarity features in the current window and the maximum value of the internal similarity features in the previous window are calculated, and a method for determining similarity features of adjacent windows is established to determine whether the gesture window is a basic gesture action window or a transition gesture action window, and obtain a basic gesture action window and a transition gesture action window respectively; S5. Upsampling the transition gesture action window obtained in step S4; S6. Input the obtained basic gesture action window and the processed transition gesture action window into the gesture action recognition model for recognition, respectively, to obtain the basic gesture action and the transition gesture action.

2. A gesture segmentation and recognition method according to claim 1, characterized in that: The sensors on the smart data glove in step S1 are a three-axis accelerometer and a three-axis gyroscope. The i-th window data x i ∈R w×C , w is the window size, and C is the number of channels of the sensor.

3. The method for hand gesture segmentation and recognition according to claim 2, characterized in that: The specific implementation method of step S3 includes the following steps: S3.

1. For the i-th window x i , the internal similarity feature of the i-th window The calculation formula is as follows: in, is the internal similarity feature of the first sub-window of the i-th window, is the internal similarity feature of the second sub-window of the i-th window, is the internal similarity feature of the third sub-window of the i-th window, and d is the Manhattan distance calculation function; the Manhattan distance calculation function is as follows: d(sx p ,sx q )=∑|sx p -sx q | Among them, d(sx p , sx q ) refers to the distance between sub-window p and sub-window q; S3.

2. The Manhattan distance calculation function is calculated using the accelerometer and gyroscope data, that is, sx in the Manhattan distance calculation function p , sx q Use the accelerometer and gyroscope data to replace and obtain the internal similarity features f for the p-th sub-window and the q-th sub-window pq The calculation formula is: Among them, a is the acceleration data, g is the gyroscope data, and the window is divided into three sub-windows. Represents the average acceleration value of the j-axis in the q-th sub-window, It represents the average angular velocity value of the j-axis in the q-th sub-window, the j-axis is the x-coordinate axis, the y-coordinate axis or the z-coordinate axis, p=1 or 2 or 3, q=1 or 2 or 3, and p and q are not equal.

4. The method for hand gesture segmentation and recognition according to claim 3, characterized in that: The specific implementation method of step S4 includes the following steps: S4.

1. Set the i-1th window as the previous window, the i-th window as the current window, and calculate the absolute value b between the minimum value of the internal similarity feature in the previous window and the maximum value of the internal similarity feature in the current window i-1 , the calculation formula is: Among them, min is the minimum function and max is the maximum function; S4.

2. Calculate the absolute difference e between the minimum value of the similarity feature within the current window and the maximum value of the similarity feature within the previous window i , the calculation formula is: S4.

3. Establish a method for determining similar features of adjacent windows: S4.3.

1. Set the detection flag representing the transition gesture action window to Tag t , the default is Tag t Set to non-transition gesture, the transition gesture action start threshold is TH ts , the transition gesture end threshold is TH te ; S4.3.

2. In Tag t For non-transition gestures, e i The transition gesture action starts at the threshold TH ts For comparison, if e i Greater than TH ts , the previous window is regarded as the basic gesture action window, and the current window is regarded as the beginning of the transition gesture action window, so the current window is added to the transition gesture action window X tans ; Otherwise, the previous window is considered the transition gesture action window.

5. A gesture segmentation and recognition method according to claim 4, characterized in that: The specific implementation method of step S5 includes the following steps: S5.

1. In Tag t In the case of a transition gesture action window, by adding the current window to the buffer X tans To start the aggregation transition window; the aggregation process is to convert b i-1 The transition gesture ends at the threshold TH te For comparison, if b i-1 Above the threshold TH te , it is considered that the end of the transition gesture action window will be reached, and the previous window will be aggregated as the end of the transition gesture action; All detected transition windows are aggregated into X tans , then, for X tans Upsampling is performed to resize it to a fixed window size to match the input size of the recognition model; otherwise, window aggregation continues in the next round; S5.

2. When the process of aggregating the windows exceeds the maximum length of the transition gesture action, it is processed as a static gesture in the basic gesture action; S5.

3. When both windows have mixed samples of transition gestures and basic gestures, the values ​​e of the two adjacent windows will be higher than the threshold transition gesture start threshold TH ts , the latest window is considered to be the beginning of the transition gesture action window.

6. A gesture segmentation and recognition method according to claim 5, characterized in that: The specific implementation method of step S6 includes the following steps: S6.

1. Construct two gesture recognition models, each of which is a CNN-LSTM network model architecture; S6.

2. The obtained basic gesture action window and the processed transition gesture action window are respectively input into the two gesture action recognition models constructed in step S6.1 for recognition to obtain the basic gesture action and the transition gesture action.

7. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a gesture segmentation and recognition method as described in any one of claims 1 to 6 when executing the computer program.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the gesture segmentation and recognition method described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • High-precision gesture data recognition method, electronic equipment and storage medium

    CN117292404A

  • Personalized gesture recognition system for multiple application scenes and gesture recognition method thereof

    CN115294658A

  • Method and apparatus for dividing gesture

    CN1249454A