Gesture recognition and classification method based on Kalman filtering in multi-hand environment
By using Kalman filtering technology to model and update the gesture state in a multi-hand environment, the missed and mis-checked gesture recognition in a multi-hand environment in the existing technology is solved, and high-precision gesture recognition and continuous tracking are achieved.
Patent Information
- Application Number
- CN202411919597.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-13
AI Technical Summary
Existing gesture recognition technology based on visual and deep neural networks is prone to missed detection and missed detection in a multi-handed environment, and it is difficult to achieve accurate estimation of gesture state.
The gesture recognition and classification method based on Kalman filtering in a multi-hand environment is adopted. The gesture state is modeled and updated by extended Kalman filtering, and combined with the prediction and update steps of hand key points, accurate estimation and recognition of gesture state is achieved.
It improves the accuracy of gesture recognition, reduces the occurrence of missed and missed detection, and realizes continuous tracking and dynamic classification of gesture status in a multi-handed environment.
Smart Images

Figure CN119992642A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gesture recognition, and in particular to a gesture recognition and classification method based on Kalman filtering in a multi-hand environment. Background Art
[0002] Gesture recognition has been widely used in the fields of smart cockpits, game control, human-computer interaction, etc. Early gesture recognition mostly relied on wearable devices, such as gloves equipped with inertial components or pressure sensors. These wearable devices recognize gestures by detecting hand movements. However, the cost of using wearable devices is high, and they are not an ideal solution for gesture recognition. At present, gesture recognition technology based on vision and deep neural networks has become a research hotspot due to its wide applicability and low deployment cost, and it has made significant progress in gesture recognition accuracy. However, due to factors such as lighting changes, hand occlusion, environmental interference, gesture posture diversity, and insufficient training data, the trained gesture state model may miss or misdetect the input hand motion video frames. Therefore, reducing or avoiding the missed detection and misdetection of gesture recognition in hand motion videos by the gesture state model is a technical problem to be solved by this application. Summary of the invention
[0003] The present invention aims to improve the accuracy of gesture recognition and enhance the human-computer interaction experience in a multi-hand environment, and provides a gesture recognition and classification method based on Kalman filtering in a multi-hand environment.
[0004] To achieve this object, the present invention adopts the following technical solutions:
[0005] A method for gesture recognition based on Kalman filtering in a multi-hand environment is provided, the steps comprising:
[0006] S1, determine the measurement data z obtained at the current time T T Is it valid data?
[0007] If yes, the gesture state is updated based on the extended Kalman filter to obtain the gesture state at the current moment. Then proceed to step S2;
[0008] If not, get the gesture status at the previous T-1 time The gesture state is predicted based on the extended Kalman filter, and the prediction result is used as the gesture state at the current moment. Then proceed to step S2;
[0009] S2, according to the measurement data z T Perform gesture recognition and output the recognition results.
[0010] Preferably, updating the gesture state based on the extended Kalman filter specifically includes the following steps:
[0011] S11, input the preset sampling time interval Δt and the gesture state posterior information of the previous moment Perform the prediction step in Kalman filtering to predict the prior state of the gesture at the current moment
[0012] in, Contains: pixel coordinates of several hand key points on the x-axis, pixel coordinates of several hand key points on the y-axis, velocity components of several hand key points on the x-axis, velocity components of several hand key points on the y-axis, acceleration components of several hand key points on the x-axis, acceleration components of several hand key points on the y-axis, the chirality of the hand, and the mean confidence value of the key points; the chirality of the left hand is 0, and the chirality of the right hand is 1; the mean confidence value of the key points represents the average value of the reliability of the detection results of all hand key points by the hand key point detector; The state variables contained in correspond;
[0013] S12, based on the gesture state prior at the current moment Combined with the observation information of gesture T , execute the Kalman filter update step and output the optimal estimate of the gesture state at the current moment If multiple hands are continuously detected, the prior gesture state at the current moment is matched with the measured gesture state to achieve continuous gesture tracking.
[0014] Preferably, in step S1, verify the measurement data z T The effectiveness method includes the steps of:
[0015] A1, calculate the measurement data z T The centroid c of the hand key points in T , and obtain the measurement data c of the previous moment T-1 ;
[0016] A2, calculate c T With c T-1 The distance d;
[0017] A3, judging whether the distance d exceeds a threshold value e,
[0018] If so, then determine c T Deviation from prediction, z T Invalid input;
[0019] If not, then determine z T is a valid input.
[0020] Preferably, in step A1, the centroid c is calculated by the following formula (13): T
[0021]
[0022] In formula (13), n represents z T The number of hand key points in ;
[0023] x i ,y i Respectively represent the coordinate values of the i-th hand key point on the x-axis and y-axis.
[0024] Preferably, in step A2, c T (x, y) and c T-1 The distance d of (x′, y′) is calculated by the following formula (14):
[0025]
[0026] In formula (14), x and y represent the centroid c T The coordinates of x′ and y′ represent the center of mass c at the previous moment. T-1 The coordinates of .
[0027] The present invention also provides a gesture classification method, which performs the gesture recognition method based on Kalman filtering in a multi-hand environment to realize the state recognition of the gesture in the measurement data corresponding to at least one frame in the continuous N frames, and then performs the following steps to realize the classification of static gestures:
[0028] L1, build and train a static gesture classification model based on a feedforward neural network;
[0029] L2, extract the normalized coordinate landmarks of the key points of the hand from the recognized gesture state, and perform local normalization processing to obtain local normalized coordinate landmarks L , expressed as follows:
[0030] landmarks L = landmarks-(x0, y0)
[0031] =[0,0,x1-x0,y1-y0,…,x i -x0,y i -y0,…,x 20 -x0,y 20 -y0] (17)
[0032] In expression (17), x0 and y0 represent the normalized coordinates of the wrist key point on the x-axis and y-axis respectively;
[0033] x i ,y i Respectively represent the normalized coordinates of the i-th hand key point on the x-axis and y-axis; i = 1, 2, ..., 20;
[0034] L3, the local normalized coordinates landmarks L The data is input into the trained static gesture classification model, and the model outputs the classification result of the static gesture.
[0035] The present invention also provides a method for dynamic gesture classification. After static gesture classification is implemented, dynamic gesture classification is implemented through the following steps:
[0036] M1, at the current time T, updates the history queue deque'=[g1, g2, ..., g N ],g N is the static gesture classification result of the Nth frame collected at the previous moment before the current moment T;
[0037] M2, determine whether the length of deque' is greater than the preset length threshold,
[0038] If yes, go to step M3;
[0039] If not, return to step M1;
[0040] M3, determining whether the occurrence frequency of the gesture of the corresponding type with the highest occurrence frequency in the deque′ exceeds a preset frequency threshold,
[0041] If not, it is determined that there is no valid dynamic gesture in deque′;
[0042] If yes, go to step M4;
[0043] M4, remove static gestures irrelevant to dynamic gesture classification from deque′;
[0044] M5, merge and remove the static gestures of the same type and adjacent to each other in the remaining deque′;
[0045] M6, traverse the merged deque′ to see if there is a predefined gesture combination,
[0046] If it exists, the category ID corresponding to the gesture combination of the corresponding continuous gesture is recorded as the classification result of the dynamic gesture and output;
[0047] If not present, return to step M1.
[0048] The present invention also provides a method for continuous tracking of gestures in a multi-hand continuous detection scenario. When the gesture state recognition is realized by the above scheme, in the multi-hand continuous detection scenario, the prior information of each two gesture states of at most two gestures detected at the current time T is defined. Measurement information Then follow these steps:
[0049] N1, Establishment and Z T The matching state vector of each element in The corresponding matching state vector is expressed as Z T The corresponding matching state vector is expressed as in, They are The corresponding matching state vector; They are The corresponding matching state vector;
[0050] N2, calculation and Z T The similarity of and Z T The matching relationship I=[IOU 11 , IOU 12 , IOU 21 , IOU 22 ];
[0051] N3, in IOU 11 and IOU 21 Select the first element with the largest IOU, and 12 and IOU 22 The second element with the largest IOU is selected. The first element determines the matching relationship between the prior and the measurement of the first gesture, and the second element determines the matching relationship between the prior and the measurement of the second gesture. Since the prior is predicted based on the gesture state at the previous moment, the matching relationship between the observation at the current moment and the gesture state at the previous moment is confirmed, thereby achieving continuous tracking of the same gesture in multi-hand scenarios.
[0052] Preferably, in step N1, the matching state vector M=[c x , c y The calculation method of [, w, h] is expressed by the following formulas (18)-(21):
[0053]
[0054] w=max(x1,x2,…,x n )-min(x1, x2, …, x n ) (20)
[0055] h=max(y1,y2,...,y n )-min(y1,y2,...,y n ) (twenty one)
[0056] In formulas (18)-(21), c x 、c y Respectively represent the coordinates of the center of mass of the key points of the hand on the x and y axes, w and h are the width and length of the minimum circumscribed rectangle of the hand; x i ,y i Respectively represent the coordinates of the i-th hand key point on the x-axis and y-axis.
[0057] Preferably, in step N2, calculate and Z T The similarity method is: through the following formulas (22)-(24) Calculate the intersection-over-union (IOU) ratio between the two pairs, and get I = [IOU 11 , IOU 12 , IOU 21 , IOU 22 ] as and Z T Matching relationship; IOU 11 express and The intersection-over-union ratio, IOU 12 express and The intersection-over-union ratio, IOU 21 express and The intersection-over-union ratio, IOU 22 express and The intersection and union ratio of
[0058]
[0059] A inter =w inter *h inter (twenty three)
[0060] A union =A p +A z -A inter (twenty four)
[0061] In formulas (22)-(24), A inter represents the area intersection area, A union A represents the area union region; p express or The rectangular box area, A z express or The rectangular box area of w inter 、h inter Indicates the width and height of the intersection area.
[0062] Preferably, after realizing the gesture state recognition of multiple hands through the above scheme, steps N1-N3 are executed.
[0063] The present invention has the following beneficial effects:
[0064] 1. The gesture state is modeled based on the extended Kalman filter, thereby achieving effective and accurate estimation of the gesture state. The extended Kalman filter and acceleration motion model are used to model the hand state, fully considering the nonlinear problem in the motion process, making the estimation of the gesture state more accurate. The hand's chirality information and confidence are also integrated into the gesture state estimation, which improves the accuracy of gesture state estimation. A correction mechanism for missed frames and false detection is designed. When the hand key point detection fails to output effectively or outputs abnormally, the estimated hand key point coordinates are used as the hand monitoring results at the current moment to solve the common missed detection and false detection problems in existing gesture recognition methods. Combined with prior information, an unbiased optimal estimation of hand key point detection is achieved.
[0065] 2. The coordinates of the hand key points are locally normalized with the coordinates of the wrist key points as the origin, eliminating the influence of the absolute position between the hand key points on the accuracy of gesture recognition, ensuring that the model can more effectively learn the relative spatial position characteristics between the hand key points. A static gesture classifier based on a fully connected network model is designed to achieve effective recognition of 7 types of static gestures. By analyzing the static gesture transformation of continuous frames and designing a judgment strategy, effective recognition of 6 types of dynamic gestures is achieved.
[0066] 3. By extracting the minimum bounding rectangle from the gesture state and calculating the IOU, the similarity of the multi-hand states to be matched in the multi-hand scenario is evaluated. Based on the IOU indicator, the posterior state of the gesture at the previous moment and the prior state at the current moment are matched, and the matching relationship between consecutive frames corresponding to the same hand is determined, thereby achieving continuous tracking of the same gesture at different times in the multi-hand scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0068] Figure 1 This is a flow chart of implementing gesture recognition, classification and control transfer based on Kalman filtering in a multi-hand environment provided by an embodiment of the present invention. ;
[0069] Figure 2 This is an example of the existing gesture recognition technology causing misdetection, missed detection, and not supporting multi-hand gesture recognition;
[0070] Figure 3 It is a schematic diagram of the 21 key points of the hand;
[0071] Figure 4 is a schematic diagram of various categories of static gestures;
[0072] Figure 5 is a schematic diagram of various categories of dynamic gestures;
[0073] Figure 6 It is a schematic diagram of the transfer of control. DETAILED DESCRIPTION
[0074] The technical solution of the present invention is further described below with reference to the accompanying drawings and through specific implementation methods.
[0075] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on this patent; in order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0076] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "inner", "outer", etc. appear, the orientation or position relationship indicated is based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as a limitation on this patent. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0077] In the description of the present invention, unless otherwise clearly specified and limited, if the term "connection" or the like appears to indicate the connection relationship between components, the term should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two components or the interaction relationship between two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0078] Existing gesture recognition methods based on vision and deep neural networks are prone to Figure 2 The error detection shown in Figure b (the position recognition error of the hand key point, Figure 2 Figure a in the figure is the accurate position of the key points of the hand, and the position of the key points of the hand identified in figure b is inconsistent with that in figure a), missed detection shown in figure c (the rectangular frame does not completely cover the hand), and gesture recognition that does not support multiple hands shown in figure d (when multiple hands appear in the same image, the detection action of the key points of the hand is not performed). In order to solve the above problems, this embodiment provides a gesture state estimation method based on Kalman filtering in a multi-hand environment, which includes the following technical links:
[0079] 1. Get the gesture observation information of each hand in the hand area image of the current frame
[0080] The gesture observation information of each hand in the hand region image of the current frame is obtained by the gesture key point detector. In order to ensure the precision and accuracy of gesture recognition, preferably, the gesture key point detector is set to obtain at most two groups of hand data of two hands each time (one hand corresponds to one group of hand data), and each group of hand data includes 21 hand key points in the corresponding hand (such as Figure 3 The two-dimensional coordinates of the hand (shown in ), the chirality information (the chirality includes the left hand and the right hand. In this embodiment, the chirality of the left hand is assigned to "0" and the chirality of the right hand is assigned to "1") and the confidence value. It should be noted here that the gesture observation information of each hand can be obtained through the existing gesture key point detector.
[0081] 2. Combine the acquired gesture observation information and perform gesture estimation based on the extended Kalman filter
[0082] In this embodiment, the gesture estimation subsystem based on the extended Kalman filter includes: a gesture state tracker based on the extended Kalman filter and a gesture state estimation module. The construction of the gesture state tracker based on the extended Kalman filter and the method for gesture tracking include the following steps:
[0083] 1) Model the system state of the gesture. The model is based on the gesture state vector x 手势The gesture observation state is modeled by the following formula (1): 观测 And expressed by the following formula (2); construct the state transfer matrix F of the system state of the gesture expressed by the following formula (3); construct the observation state vector x expressed by the following formula (5) 观测 The observation matrix H is:
[0084] x 手势
[0085] =[x1,y1,x2,y2,…,x n ,y n , v x1 , v y1 , v x2 , v y2 , …, v xn , v yn , a x1 , a y1 , a x2 , a y2 , …, a xn , a yn , lr, s] (1)
[0086] x 观测 =[x1,y1,x2,y2,…,x n ,y n , lr, s] (2)
[0087]
[0088] In expressions (1)-(2), x n ,y n Respectively represent the coordinate values of the nth hand key point on the hand on the x-axis and y-axis;
[0089] v xn , v yn Respectively represent the movement speed of the nth hand key point on the x-axis and y-axis;
[0090] a xn , a yn Respectively represent the movement acceleration of the nth hand key point on the x-axis and y-axis;
[0091] lr represents the handedness, where lr = 0 for the left hand and lr = 1 for the right hand. The handedness is identified by the gesture keypoint detector;
[0092] s is the average value of the reliability of the detection results of the hand key point detector for all hand key points; n=1, 2, ..., 21.
[0093] Gesture state vector x 手势The total dimension is 128. The observation state vector x 观测 It includes the coordinates of 21 hand key points in the x-axis and y-axis directions, as well as the chirality and confidence values, totaling 44 dimensions.
[0094] In expression (3), F pk Represents the state transition submatrix of the kth hand key point;
[0095] 0 represents a zero matrix whose dimension depends on the dimension of the submatrix in the row or column; 1 represents a 1*1 matrix whose only element is 1;
[0096] F pk It is expressed as follows:
[0097]
[0098] In expression (4), Δt represents the time interval between consecutive adjacent states; 1 represents direct transfer or unchanged relationship between corresponding variables; 0 represents no direct relationship between corresponding variables; here, consecutive adjacent states refer to the states of gestures at two consecutive time points, and direct transfer between corresponding variables means that within the Δt time interval, the corresponding variables in the first gesture state are unchangedly transferred to the corresponding variables in the second gesture state. No direct relationship between corresponding states means that within the Δt time interval, the corresponding variables in the first gesture state have no effect on the corresponding variables in the second gesture state.
[0099]
[0100] In expression (5), I 21 is a 21×21 identity matrix, representing a 21-dimensional mapping of the keypoint states of the hand to the observations; this means that the keypoints in each state correspond one-to-one to the keypoints in the observations, without transformations such as scaling or rotation. The identity matrix ensures this direct correspondence. The meanings of 0 and 1 are the same as in equation (4), where 0 represents a zero matrix whose dimension depends on the dimension of the submatrix in the row or column; 1 represents a 1*1 matrix whose only element is 1.
[0101] 2) Establishing the state estimation problem
[0102] At time T, the posterior states of the motion equation and observation equation at the previous time T-1 are Linearize the gesture prior state at time T and the observed state z T is represented as:
[0103]
[0104] remember Then we have:
[0105]
[0106] In expressions (6)-(9), Represents the posterior state of the gesture at time T-1;
[0107] z T Represents the observed state of the gesture at time T;
[0108] Represents the state transfer function, which is used to convert the estimated state of the previous moment Switch to the predicted state at the current moment;
[0109] Represents the error between the posterior state and the estimated state at the previous moment;
[0110] Represents the Jacobian matrix of the state transfer function f, in Calculation is performed at
[0111] w T represents the process noise at the current time T, w T The uncertainty in gesture prediction introduced by model errors or external influences is simulated, and Q T represents the process noise covariance matrix, where Q T Set to I 128 *0.1,I 128 represents the identity matrix with a dimension of 128*128, Q T The dimension of corresponds to the dimension of the state transfer matrix F; represents a normal probability density distribution with a mean of 0 and a variance of QT;
[0112] Represents the observation function, applied to the prior of the current state
[0113] Represents the prior of the current state, that is, the estimated state;
[0114] v T Indicates z T The corresponding process noise, and R T represents the observation noise covariance matrix, R T Set to I 44 *0.1,I 44 Represents the identity matrix with a dimension of 44*44, R T The dimension of corresponds to the dimension of the observation matrix H;
[0115] x0 represents the given initial state;
[0116] u 1:T Represents the control input from the first time step to the current moment;
[0117] z 0:T-1 represents the observation from the initial moment to the previous moment;
[0118] x k |x0,u 1:T , z 0:T-1 Represents the state prior given the initial state, control input, and previous observations;
[0119] P(x T |x0,u 1:T , z 0:T-1 ) represents the probability of the state prior given the initial state, control input, and previous observations;
[0120] u T Represents the control input at the current moment;
[0121] represents the state transfer function, which is applied to the estimated state at the previous moment and the current control input;
[0122] Represents the state covariance matrix at the previous moment;
[0123] F T represents the transpose of the state transfer matrix F;
[0124] Represents the prior state at the current moment The probability that the mean is The variance is Normal distribution of
[0125] P(z T |x T ) represents the observation probability of the current state;
[0126] represents the observation z at the current time T The probability that the mean is Variance is R T The normal distribution of .
[0127] 3) Establish the prediction and update process of the extended Kalman filter, including:
[0128] (1) Status prediction: Represents the prediction result of the gesture state at time T;
[0129] (2) Error covariance matrix prediction: Represents the error covariance matrix predicted at the current time T;
[0130] (3) Kalman gain calculation: H T represents the transpose of the observation matrix H;
[0131] (4) Prediction status update:
[0132] (5) Covariance update: Represents the updated covariance at the current time T; Represents the error covariance matrix predicted at the current time T.
[0133] The model prediction and update process of gesture recognition based on Kalman filtering in this embodiment is: combining the prior information of the previous moment (T-1 moment) and the observation of the current moment (T moment) to achieve an unbiased optimal estimation of the gesture state at the current moment T. In this embodiment, the technical principle of gesture state estimation based on the gesture state estimation module of the extended Kalman filter is: after the program runs, first determine whether the gesture key point detector outputs the hand key point detection result for the current frame. If there is a detection result, then combine the a posteriori state information of the previous frame The measurement data z at the current time T T The validity of the measurement data z is verified. T If the data is valid, the gesture state model is updated at the current time T, and then the gesture state is estimated with the updated model. Otherwise, the update step is skipped and the gesture state is estimated directly with the model used at time T-1.
[0134] For the measured data z T The method for verifying the effectiveness of includes the following steps:
[0135] (1) Calculate the measurement data z using the following formula (13): T The centroid c of the hand key points in T ,
[0136]
[0137] In formula (13), n represents z T The number of hand key points in , n = 21;
[0138] x i ,y i Respectively represent the coordinate values of the i-th hand key point on the x-axis and y-axis.
[0139] (2) Obtain the measurement data c at the previous moment T-1 (3) Calculate c T (x, y) and c T-1 The distance is calculated as follows:
[0140]
[0141] In formula (14), x and y represent the centroid c T The coordinates of x′ and y′ represent the center of mass c at the previous moment. T-1 The coordinates of . The representation is determined by the following method:. (4) Determine whether the distance d exceeds a threshold value e (preferably e = 50),
[0142] If so, then determine c T Deviation from prediction, z T Invalid input;
[0143] If not, then determine z T is a valid input.
[0144] (5) By calculating the Kalman gain K T , and update the posterior estimate of the gesture state Specifically:
[0145]
[0146] In formula (10), Represents the updated covariance at the current time T;
[0147] K T Represents the Kalman gain at the current time T;
[0148] H represents the observation matrix;
[0149] Represents the error covariance matrix predicted at the current time T;
[0150]
[0151] In addition, the error covariance matrix predicted at the current time T is It is calculated by the following expression (11):
[0152]
[0153] Kalman gain K T It is calculated by the following expression (12):
[0154]
[0155] H T represents the transpose of the observation matrix H;
[0156]
[0157] If the measurement data z T If it is not valid data, the above update step is skipped at time T and the following prediction process is directly executed:
[0158] (1) Based on the posterior at the previous time T-1 The prior state at the current time T The estimation method is expressed as:
[0159]
[0160] (2) Keep the covariance unchanged and convert the prior error covariance matrix at the current time T Let the posterior error covariance matrix at time T-1 be
[0161] (3) If measurement information is added at a future time, the above model update process is entered; otherwise, the above prediction processes (1) to (2) are repeated at times T+1, T+1, etc.
[0162] (4) A counter is set to count the number of consecutive frames in which no valid measurement value is received. When the counter count exceeds "15", it is considered that the Kalman filter estimation is no longer valid, and the prediction is stopped at this time, and it is considered that there is no valid hand area in the input image.
[0163] (5) When performing real-time detection, if missed detection occurs in multiple consecutive frames (no detection results are output), the gesture state estimation module will continue to estimate the state of the gesture based on previous information within the time counted by the counter. In this case, due to the lack of effective measurement information, the uncertainty of the system's state estimation will accumulate rapidly. When the uncertainty exceeds an acceptable level, stop using the model for gesture state estimation. The percentage of missed detections in the gesture state estimation module is generally less than 3 frames. In a short period of time, the uncertainty of the gesture state estimation is within a relatively small range. Therefore, the gesture state can be effectively estimated and recursively inferred based on prior information, thereby eliminating the impact of missed detections.
[0164] After identifying the gesture in each frame of measurement data in N consecutive frames, this embodiment further provides a gesture classification method to combine the gesture state recognition results of the N consecutive frames to output corresponding classification results for various gestures.
[0165] In this embodiment, static gestures are classified based on a fully connected network model, and dynamic gestures are classified by formulating a dynamic gesture judgment strategy and combining the static gesture classification results.
[0166] The static gesture classification method and the dynamic gesture classification method provided in this embodiment are specifically described below:
[0167] 1. Static gesture classification method
[0168] In this embodiment, static gesture classification is implemented by a fully connected gesture classification model. The input data of the model is the coordinates of the key points of the hand in the gesture state. When the target hand is detected as the left hand, the coordinates of the key points of the hand are flipped with the central axis of the image as the axis of symmetry, and then the classification process continues. The specific implementation method is as follows:
[0169] 1) Extract the normalized coordinate landmarks of the key points of the hand from the gesture state obtained by the above gesture recognition method,
[0170] landmarks=[x0,y0,x1,y1,…,x i ,y i , …, x 20 ,y 20 ]
[0171] x0, y0 represent the 0th hand key point, i.e. the normalized coordinates of the wrist key point on the x-axis and y-axis respectively;
[0172] x i ,y i Respectively represent the normalized coordinates of the i-th hand key point on the x-axis and y-axis; i = 1, 2, ..., 20;
[0173] The original coordinates x of the i-th hand key point on the x-axis and y-axis i ,y i The normalization method is expressed by the following expression:
[0174] For the absolute pixel coordinates of the i-th hand key point on the x-axis and y-axis (u i , v i ), let the length and width of the image be h and w respectively, then the normalized coordinate (x i ,y i )for:
[0175] x i =u i / w
[0176] y i =v i / h
[0177] In fact, gesture recognition methods directly output normalized coordinates rather than absolute pixel coordinates. The normalized output representation can reduce errors in numerical calculations and improve the stability of the model.
[0178] The landmarks data here are derived from the output of the gesture state model. Using normalized coordinates can make the gesture classification model more versatile when processing images of different resolutions and sizes, without having to learn how to cope with different image size changes.
[0179] 2) For the normalized coordinate landmarks of the hand key points, the local coordinates are normalized with the wrist key point coordinates as the origin to obtain the local normalized coordinate landmarks L , which is expressed as follows:
[0180] landmarks L = landmarks-(x0, y0)
[0181] =[0,0,x1-x0,y1-y0,…,x i -x0,y i -y0,…,x 20 -x0,y 20 -y0] (17)
[0182] In expression (17), x0 and y0 represent the normalized coordinates of the wrist key point on the x-axis and y-axis respectively;
[0183] x i ,y i Respectively represent the normalized coordinates of the i-th hand key point on the x-axis and y-axis; i = 1, 2, …, 20.
[0184] The purpose of local normalization is to focus on the gesture category rather than the specific position of the gesture in the image. In other words, we hope that the gesture classification model will focus on the relative position relationship between the key points of the hand rather than the absolute position relationship. By using local normalization, the model can more effectively learn the relative spatial position characteristics between the key points.
[0185] 3) Training static gesture classification model
[0186] In this embodiment, a static gesture classification model is trained by a feedforward neural network, which includes a 42-dimensional input layer for receiving 42-dimensional local normalized coordinate data, and then connected to the input layer is a first Dropout layer with a dropout rate of 20% to reduce overfitting during training; then connected to the first Dropout layer is a first fully connected layer containing 64 neurons, which uses a ReLU activation function to increase the nonlinear ability of the network; then connected to the first fully connected layer is a second Dropout layer with a dropout rate increased to 40%, which further helps the generalization of the model; then connected to the second Dropout layer is a second fully connected layer containing 32 neurons, which also uses a ReLU activation function; then connected to the second fully connected layer is a third Dropout layer with a dropout rate of 40%; finally, connected to the third Dropout layer is an output layer containing N neurons, and the number of neurons in the output layer corresponds to the type of preset static gestures; finally, using a softmax function, the output is converted into a probability q for predicting each gesture category and the cross entropy loss is calculated.
[0187] The following is an explanation of the input and output data of each network layer in the feedforward neural network used above:
[0188] The input data of the first Dropout is 42-dimensional local normalized coordinate data, and the output data is a 42-dimensional feature vector, but 20% of the neurons are randomly discarded during training;
[0189] The output of the first Dropout layer is used as the input of the first fully connected layer. The output data of the first fully connected layer is a 64-dimensional feature vector.
[0190] The output of the first fully connected layer is used as the input of the second Dropout layer. The output data of the second Dropout is a 64-dimensional feature vector, but 40% of the neurons are randomly discarded during training;
[0191] The output of the second Dropout layer is used as the input of the second fully connected layer, and the output data of the second fully connected layer is 32;
[0192] The output of the second fully connected layer is used as the input of the third Dropout layer. The output data is 32-dimensional, and 40% of the neurons are randomly discarded during training. The output of the third Dropout layer is used as the input of the output layer, and a 10-dimensional vector is output.
[0193] In this embodiment, there are 7 types of static gestures, namely, upward finger (corresponding to Figure 4 "0-UP" in the Figure 4 "1-close"), extend your index finger (corresponding to Figure 4"2-Pointer") in Figure 4 "3-Hold" in the Figure 4 "4-Forward" in the Figure 4 "5-Down"), left finger (corresponding to Figure 4 "6-Left") in the Figure 4 in the “7-Right” section).
[0194] 4) Local normalized coordinate landmarks obtained through gesture state recognition L The gesture is input into the trained static gesture classification model, and the model outputs the classification result of the static gesture (preferably outputting the gesture classification ID, for example, the ID corresponding to the static gesture category "right finger" is "7").
[0195] Based on the above steps, static gestures are classified based on the coordinate data of the hand key points obtained through gesture state recognition. However, in actual application scenarios, gestures are usually dynamically changing, so a dynamic gesture classification method with high accuracy is needed to ensure the recognition accuracy of dynamic gesture control instructions.
[0196] The dynamic gesture classification method provided in this embodiment is implemented by analyzing the static gesture changes of consecutive frames, specifically:
[0197] In this embodiment, dynamic gestures include 6 categories, namely: Figure 5 The classification of dynamic gestures depends on the classification results of static gestures in the past N moments (N frames). In this embodiment, a history queue deque′=[g1, g2, ..., g N ] is used to store the static gesture classification results of the most recent N consecutive moments. The specific method of its implementation is as follows:
[0198] In real-time detection, the static gesture classification result of the N+1th frame at the current time T is defined as g N+ 1, and then update the history queue deque'. Then, determine whether the length of deque' is greater than the preset length threshold. If so, delete the tail element, such as deleting "g1" in the queue and replacing g N+1 Add to the head of deque′, that is, add to g N Then execute the following dynamic gesture classification process:
[0199] 1) Check whether the frequency of the gesture of the corresponding type with the highest frequency in deque′ exceeds the preset frequency threshold. If so, it is determined that there is no valid dynamic gesture in deque′, and the dynamic gesture classification process is terminated. Otherwise, continue to execute the following process 2);
[0200] 2) Remove static gestures irrelevant to the dynamic gesture classification from deque′; in this embodiment, irrelevant gestures are Figure 5 "3-Hold" and "4-forward" as shown in;
[0201] 3) Merge and remove the static gestures of the same type and adjacent to each other in the remaining deque′ to simplify the subsequent analysis;
[0202] 4) traverse the merged deque′ to see if there is a predefined gesture combination, and record the category id corresponding to the gesture combination of the corresponding continuous gesture;
[0203] Here is a brief description of the method for determining whether a predefined gesture combination exists in the deque′:
[0204] like Figure 5 The defined, predefined gesture combinations include:
[0205] (6-Left, 7-Right), corresponding to "right swipe";
[0206] (7-Right, 6-Left), corresponding to "left swipe";
[0207] (0-Up, 5-Down), corresponding to "underline";
[0208] (5-Down, 0-Up), corresponding to "swipe up";
[0209] (2-Pointer, 1-Close), corresponding to "click";
[0210] (0-Up, 1-Close), corresponding to "Close".
[0211] For example, when "(6-Left, 7-Right)" is recognized, it is determined that there is a predefined gesture combination.
[0212] In addition, this embodiment also provides a gesture control right management method in a multi-hand continuous detection scenario, which is used to continuously track the gestures of multiple hands in a multi-hand gesture control scenario and manage the control rights, thereby realizing the transfer of control rights between different hands. The gesture control right management method includes the following technical links:
[0213] 1. Gesture tracking
[0214] When the hand is moving, the state of the gesture, including the coordinates and confidence of the key points of the hand, is changing dynamically. And in a scene with multiple hands, there may be a situation where the handedness of the two hands is the same. Therefore, simply relying on the state data of the gesture is difficult to be an accurate basis for anchoring the hand area. In this embodiment, through the above-mentioned gesture recognition and amplification based on Kalman filtering, effective tracking of gestures is first achieved. In the input image, the hand area is first identified using the above-mentioned gesture state model, and each detection area is accompanied by a confidence score, indicating the degree of credibility of the area as a hand. Preferably, in a multi-hand scenario, the model selects the two hand areas with the top two confidence levels from high to low, and the area with the highest confidence level is regarded as the "hand with control".
[0215] Assume that at the current time T, the priors of the two gesture states detected are At the current time T, the gesture state model detects and outputs the measurement information of the two gesture states In this embodiment, the maximum weight matching method is used to With Z T The elements of are matched to confirm the matching relationship between the two gesture objects observed at the current time T and the gesture object at the previous time T-1. The specific implementation steps are as follows:
[0216] 1) Status extraction
[0217] In order to avoid tedious calculations, this embodiment establishes the matching state vector M=[c x , c y ,w,h], where c x 、c y Respectively represent the coordinates of the center of mass of the key points of the hand on the x and y axes, w and h are the width and length of the minimum circumscribed rectangle of the hand. The specific calculation method is as follows:
[0218]
[0219] w=max(x1,x2,…,x n )-min(x1, x2, …, x n ) (20)
[0220] h=max(y1,y2,...,y n )-min(y1,y2,...,y n ) (twenty one)
[0221] According to formulas (18)-(21), we can calculate The corresponding matching state vectors And calculate separately The corresponding matching state vectors
[0222] 2) Calculate the Z of the gesture object observed at the current time T T , and the prior of the gesture state of the gesture object The similarity is calculated as:
[0223] right Calculate the intersection-and-union ratio (IOU) two by two, and get and Z T The matching relationship;
[0224]
[0225] A inter =w inter *h inter (twenty three)
[0226] A union =A p +A z -A inter (twenty four)
[0227] In formulas (22)-(24), A inter represents the area intersection area, A union A represents the area union region; p express or The rectangular box area, A z express or The rectangular box area of w inter 、h inter represents the width and height of the intersection area. And let I = [IOU 11 , IOU 12 , IOU 21 , IOU 22 ] as and Z T The matching relationship, where IOU 11 express and The intersection-over-union ratio, IOU 12 express and The intersection-over-union ratio, IOU 21 express and The intersection-over-union ratio, IOU 22 express and The intersection and ratio of .
[0228] 3) In IOU 11 and IOU 21Select the first element with the largest IOU, and 12 and IOU 22 The second element with the largest IOU is selected. The first element determines the matching relationship between the prior and the measurement of the first gesture, and the second element determines the matching relationship between the prior and the measurement of the second gesture. Since the prior is predicted based on the gesture state at the previous moment, the matching relationship between the observation at the current moment and the gesture state at the previous moment is confirmed, thereby realizing continuous tracking and detection of the same gesture at different time points in multi-hand scenarios, allowing the system to continuously track the same gesture, rather than just identifying static instantaneous states.
[0229] 2. Methods of Transferring Control
[0230] In the initial state, if there are two hand areas in the scene, the gesture object with higher confidence is the object with control by default. In the subsequent gesture recognition process, according to the control anchoring strategy, the control will always be kept in this hand. When the hand makes a specific gesture, the control transfer is triggered, and the other object tracked by the gesture is regarded as the hand with control. In this embodiment, the specific gesture includes any one of the following: pointing up, making a fist, extending the index finger, closing, pointing forward, pointing down, left finger, and right finger. It is preferred Figure 4 or Figure 6 The "3-Hold" gesture in is the "close" gesture.
[0231] In summary, the present invention has the following beneficial effects:
[0232] 1. The gesture state is modeled based on the extended Kalman filter, thereby achieving effective and accurate estimation of the gesture state. The extended Kalman filter and acceleration motion model are used to model the hand state, fully considering the nonlinear problem in the motion process, making the estimation of the gesture state more accurate. The hand's chirality information and confidence are also integrated into the gesture state estimation, which improves the accuracy of gesture state estimation. A correction mechanism for missed frames and false detection is designed. When the hand key point detection fails to output effectively or outputs abnormally, the estimated hand key point coordinates are used as the hand monitoring results at the current moment to solve the common missed detection and false detection problems in existing gesture recognition methods. Combined with prior information, an unbiased optimal estimation of hand key point detection is achieved.
[0233] 2. The coordinates of the hand key points are locally normalized with the coordinates of the wrist key points as the origin, eliminating the influence of the absolute position between the hand key points on the accuracy of gesture recognition, ensuring that the model can more effectively learn the relative spatial position characteristics between the hand key points. A static gesture classifier based on a fully connected network model is designed to achieve effective recognition of 7 types of static gestures. By analyzing the static gesture transformation of continuous frames and designing a judgment strategy, effective recognition of 6 types of dynamic gestures is achieved.
[0234] 3. By extracting the minimum bounding rectangle from the gesture state and calculating the IOU, the similarity of the multi-hand states to be matched in the multi-hand scenario is evaluated. Based on the IOU indicator, the posterior state of the gesture at the previous moment and the prior state at the current moment are matched, and the matching relationship between consecutive frames corresponding to the same hand is determined, thereby achieving continuous tracking of the same gesture at different times in the multi-hand scenario. In addition, by setting an anchoring strategy for the gesture control right and setting a specific control right transfer trigger gesture, the accurate transfer of gesture control right in the multi-hand scenario is achieved.
[0235] It should be noted that the above specific implementations are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art should understand that various modifications, equivalent substitutions, changes, etc. can be made to the present invention. However, as long as these changes do not deviate from the spirit of the present invention, they should be within the scope of protection of the present invention. In addition, some terms used in the specification and claims of this application are not restrictive, but are only for the convenience of description.
Claims
1. A gesture recognition method based on Kalman filtering in a multi-hand environment, characterized in that the steps include: S1, determine the measurement data z obtained at the current time T T Is it valid data? If yes, the gesture state is updated based on the extended Kalman filter to obtain the gesture state at the current moment. Then proceed to step S2; If not, get the gesture status at the previous T-1 time The gesture state is predicted based on the extended Kalman filter, and the prediction result is used as the gesture state at the current moment. Then proceed to step S2; S2, according to the gesture state Perform gesture recognition and output the recognition results.
2. The gesture recognition method based on Kalman filtering in a multi-hand environment according to claim 1, characterized in that: The specific steps of updating the gesture state based on the extended Kalman filter are as follows: S11, input the preset sampling time interval Δt and the gesture state posterior information of the previous moment Perform the prediction step in Kalman filtering to predict the prior state of the gesture at the current moment in, Contains: pixel coordinates of several hand key points on the x-axis, pixel coordinates of several hand key points on the y-axis, velocity components of several hand key points on the x-axis, velocity components of several hand key points on the y-axis, acceleration components of several hand key points on the x-axis, acceleration components of several hand key points on the y-axis, hand chirality, and key point confidence mean; the chirality of the left hand is 0, and the chirality of the right hand is 1; the key point confidence mean represents the average value of the reliability of the detection results of all hand key points by the hand key point detector; The state variables contained in correspond; S12, based on the gesture state prior at the current moment Combined with the observation information of gesture T , execute the Kalman filter update step and output the optimal estimate of the gesture state at the current moment In a multi-hand continuous detection environment, the prior gesture state and the measured gesture state are matched at the current moment to achieve continuous gesture tracking.
3. The gesture recognition method based on Kalman filtering in a multi-hand environment according to claim 1, characterized in that: In step S1, verify the measurement data z T The effectiveness method includes the steps of: A1, calculate the measurement data z T The centroid c of the hand key points in T , and obtain the measurement data c of the previous moment T-1 ;; A2, calculate c T With c T-1 The distance d; A3, judging whether the distance d exceeds a threshold value e, If so, then determine c T Deviation from prediction, z T Invalid input; If not, then determine z T is a valid input.
4. The gesture recognition method based on Kalman filtering in a multi-hand environment according to claim 3 is characterized in that: In step A1, the centroid c is calculated by the following formula (13): T : In formula (13), n represents z T The number of hand key points in ; x i ,y i Respectively represent the coordinate values of the i-th hand key point on the x-axis and y-axis.
5. The gesture recognition method based on Kalman filtering in a multi-hand environment according to claim 3 is characterized in that: In step A2, c T (x,y) and c T-1 The distance d of (x′, y′) is calculated by the following formula (14): In formula (14), x and y represent the centroid c T The coordinates of x′ and y′ represent the center of mass c at the previous moment. T-1 The coordinates of .
6. A gesture classification method, characterized in that: After executing the gesture recognition method based on Kalman filtering in a multi-hand environment as described in any one of claims 1 to 5 to realize the state recognition of the gesture in the measurement data corresponding to at least one frame in the continuous N frames, the following steps are performed to realize the classification of static gestures: L1, build and train a static gesture classification model based on a feedforward neural network; L2, extract the normalized coordinate landmarks of the key points of the hand from the recognized gesture state, and perform local normalization processing to obtain local normalized coordinate landmarks L , expressed as follows: landmarks L =landmarks-(x0,y0) =[0,0,x1-x0,y1-y0,...,x i -x0,y i -y0,...,x 20 -x0,y 20 -y0] (17) In expression (17), x0 and y0 represent the normalized coordinates of the wrist key point on the x-axis and y-axis respectively; x i ,y i Respectively represent the normalized coordinates of the i-th hand key point on the x-axis and y-axis; i = 1, 2, ..., 20; L3, the local normalized coordinates landmarks L The data is input into the trained static gesture classification model, and the model outputs the classification result of the static gesture.
7. A gesture dynamic classification method, characterized in that: After executing claim 6 to implement static gesture classification, the following steps are performed to implement dynamic gesture classification: M1, at the current time T, updates the history queue deque′=[g1,g2,...,g N ],g N is the static gesture classification result of the Nth frame collected at the previous moment before the current moment T; M2, determine whether the length of deque' is greater than the preset length threshold, If yes, go to step M3; If not, return to step M1; M3, determining whether the occurrence frequency of the gesture of the corresponding type with the highest occurrence frequency in the deque′ exceeds a preset frequency threshold, If not, it is determined that there is no valid dynamic gesture in deque′; If yes, go to step M4; M4, remove static gestures irrelevant to dynamic gesture classification from deque′; M5, merge and remove the static gestures of the same type and adjacent to each other in the remaining deque′; M6, traverse the merged deque′ to see if there is a predefined gesture combination, If it exists, the category ID corresponding to the gesture combination of the corresponding continuous gesture is recorded as the classification result of the dynamic gesture and output; If not present, return to step M1.
8. The gesture recognition method based on Kalman filtering in a multi-hand environment according to claim 2, characterized in that: In step S12, the method for implementing gesture tracking when performing gesture state recognition in a multi-hand continuous detection scenario is to define prior information of at most two gesture states detected at the current time T Measurement information Then follow these steps: N1, Establishment and Z T The matching state vector of each element in The corresponding matching state vector is expressed as Z T The corresponding matching state vector is expressed as in, They are The corresponding matching state vector; They are The corresponding matching state vector; N2, calculation and Z T The similarity of and Z T The matching relationship I=[IOU 11 ,IOU 12 ,IOU 21 ,IOU 22 ]; N3, in IOU 11 and IOU 21 Select the first element with the largest IOU, and 12 and IOU 22 The second element with the largest IOU is selected. The first element determines the matching relationship between the prior and the measurement of the first gesture, and the second element determines the matching relationship between the prior and the measurement of the second gesture. Since the prior is predicted based on the gesture state at the previous moment, the matching relationship between the observation at the current moment and the gesture state at the previous moment is confirmed, thereby realizing continuous tracking of the same gesture state in multi-hand scenarios.
9. The gesture recognition method based on Kalman filtering in a multi-hand environment according to claim 8, characterized in that: In step N1, the matching state vector M = [c x ,c y ,w,h] is calculated by the following formulas (18)-(21): w=max(x1,x2,…,x n )-min(x1,x2,…,x n ) (20) h=max(y1,y2,…,y n )-min(y1,y2,…,y n ) (21) In formulas (18)-(21), c x 、c y Respectively represent the coordinates of the center of mass of the key points of the hand on the x and y axes, w and h are the width and length of the minimum circumscribed rectangle of the hand; x i ,y i Respectively represent the coordinates of the i-th hand key point on the x-axis and y-axis.
10. The gesture recognition method based on Kalman filtering in a multi-hand environment according to claim 8, characterized in that: In step N2, calculate and Z T The similarity method is: through the following formulas (22)-(24) Calculate the intersection-over-union (IOU) ratio between the two pairs, and get I = [IOU 11 ,IOU 12 ,IOU 21 ,IOU 22 ] as and Z T Matching relationship; IOU 11 express and The intersection-over-union ratio, IOU 12 express and The intersection-over-union ratio, IOU 21 express and The intersection-over-union ratio, IOU 22 express and The intersection and union ratio of A inter =w inter *h inter (23) A union =A p +A z -A inter (24) In formulas (22)-(24), A inter represents the area intersection area, A union A represents the area union region; p express or The rectangular box area, A z express or The rectangular box area of w inter 、h inter Indicates the width and height of the intersection area.
11. The gesture recognition method based on Kalman filtering in a multi-hand environment according to claim 8, characterized in that: In a multi-hand continuous detection scenario, when the gesture state recognition of multiple hands is realized by the solution provided in any one of claims 1-5, steps N1-N3 are executed.
Citation Information
Patent Citations
Isomerous data fusion based coordinated gesture recognition method and system of sensor
CN102945362A
Intelligent wheelchair dynamic gesture recognition method based on Kinect depth information
CN103390168A
Hand gesture recognition method based on switching Kalman filtering model
CN104050488A
Gesture remote control system
CN104333793A
Gesture recognition method, gesture recognition device and electronic device
CN109190559A