A gesture recognition method based on integrating horizontal and vertical information in the YOLO-V3 framework
Through data expansion, the construction of the Decay BN layer and the correction loss function, the shortcomings of the traditional gesture recognition model in small-size gesture recognition are solved, and the gesture recognition accuracy and generalization ability of the model are improved.
Patent Information
- Application Number
- CN202111559476.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-20
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-12-20
AI Technical Summary
The traditional gesture recognition model has shortcomings in recognition accuracy, especially the inaccurate recognition of small-sized gestures, and the yolo-v3 framework has not been targetedly improved.
By augmenting the gesture images, building the Decay BN layer, increasing horizontal and vertical convolutions, and correcting the loss function, the generalization ability and small-objective prediction ability of the model are improved.
It significantly improves the accuracy of gesture recognition, can effectively identify various types of gestures, and enhances the generalization ability of the model and the prediction ability of small goals.
Smart Images

Figure CN114445908B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gesture recognition, and particularly to a gesture recognition method for integrating horizontal and vertical information based on the yolo-v3 framework. Background Art
[0002] With the increasing popularity of human-computer interaction, gesture recognition has become an important method of human-computer interaction. Gesture recognition is required in various fields to support business. For example, in live detection, it can be determined whether a photo is being impersonated by changing gestures; in deaf-mute education, gestures can be used to translate the intended meaning; in home entertainment, gesture recognition can be used to enhance the game experience, etc.
[0003] However, the recognition accuracy of traditional gesture recognition models is still not satisfactory, for the following reasons: 1. There are many types of gestures, and many gestures are similar. Gestures that are similar from different angles are likely to cause misjudgment; 2. The size and scale of gestures are inconsistent in different scenarios, and smaller gestures appear less frequently in pictures. Existing loss functions focus on larger-sized gestures, so smaller-sized gestures will be ignored during the training of the gesture recognition model, and smaller-sized gestures do not receive due attention, resulting in inaccurate detection of smaller-sized gestures by the gesture recognition model; 3. The backbone network of the gesture recognition model uses the general yolo-v3 framework and has not been improved based on the specific problem of gesture recognition.
[0004] Therefore, how to provide a gesture recognition method for integrating horizontal and vertical information based on the yolo-v3 framework to improve the accuracy of gesture recognition has become an urgent technical problem to be solved. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a gesture recognition method for integrating horizontal and vertical information based on the yolo-v3 framework to improve the accuracy of gesture recognition.
[0006] The present invention is implemented as follows: A gesture recognition method for integrating horizontal and vertical information based on the yolo-v3 framework includes the following steps:
[0007] Step S10: Obtain a large number of gesture pictures, and perform data augmentation on each of the gesture pictures to obtain a picture set;
[0008] Step S20: Use the moving average algorithm to construct a Decay BN layer;
[0009] Step S30: Construct the backbone network of yolo-v3 through a number of horizontal convolutions and vertical convolutions;
[0010] Step S40: Construct a loss function based on the small target loss;
[0011] Step S50: Construct a gesture recognition model based on the Decay BN layer, the backbone network, and the loss function;
[0012] Step S60: Use the picture set to train the gesture recognition model;
[0013] Step S70: Use the trained gesture recognition model to perform gesture recognition.
[0014] Further, the specific steps of step S10 include:
[0015] Step S11: Obtain a large number of gesture pictures;
[0016] Step S12: Use the similarity function to perform k-means clustering on each of the gesture pictures to obtain N categories;
[0017] Step S13: Select one gesture picture from each of the N categories in turn, crop the gesture area of each gesture picture to obtain sub-pictures, and correct the size of the sub-pictures to 1 / N of the gesture pictures;
[0018] Step S14: Stitch each of the sub-pictures into a first picture, stitch each of the sub-pictures after randomly rotating different angles into a second picture, splice the sub-pictures after randomly repeating sampling N times into a third picture, and form a batch with the sub-pictures, the first picture, the second picture, and the third picture;
[0019] Step S15: Based on each batch, form a data-augmented picture set.
[0020] Further, in step S12, the formula of the similarity function is:
[0021] S whole = α1 * (S 11 + S 12 ) + α2 * S2;
[0022]
[0023]
[0024]
[0025] Among them, S whole represents the similarity between picture A and picture B; anchor A represents the gesture target in picture A; anchor B represents the gesture target in picture B; S 11 represents anchor A the mean value of the abscissas of all points in, and anchorB The absolute difference of the mean values of the abscissas of all the points in; S 12 Denote anchor A The absolute difference between the mean value of the ordinates of all the points in and the ordinate of anchor B The absolute difference of the mean values of the ordinates of all the points in; S2 denotes the mean value of the absolute differences of the abscissas and ordinates of Picture A and Picture B; α1 and α2 denote similarity coefficients; m A Denote anchor A The number of pixel points in; m B Denote anchor B The number of pixel points in; m denotes the number of all pixel points of Picture A and Picture B; (x i , y i ) denote the coordinate points in; (x A , y j ) denote the coordinate points in; x j ) denote the coordinate points in; x B , x Ai , x Bj , y Ai and y Bj all denote gesture target adjustment coefficients.
[0026] Further, the step S20 specifically includes:
[0027] Step S21: Select a Decay BN layer, and calculate the mean value μ i and variance σ i of the i-th batch in the Decay BN layer. Based on i, calculate the mean value μ i+j and variance σ i+j of the (i + j)-th batch;
[0028] Step S22: Calculate the similarity of the data of the (i + j)-th batch and the i-th batch;
[0029] Step S23: Update the mean value μ i+j and variance σ i+j of the Decay BN layer based on the similarity, and complete the construction of the Decay BN layer.
[0030] Further, in the step S22, the calculation formula of the similarity is:
[0031] s i-(i+j) = ∑p c ∈batch i ∑p d ∈batch i+j s(p c , p d );
[0032]
[0033] Among them, s i-(i+j) represents the sum of the similarities of each pair of elements in batch i and batch i+j ; batch i represents the set of the i-th batch data; batch i+j represents the set of the (i + j)-th batch data; p c represents a data in batch i ; p d represents a data in batch i+j ; s(p c , p d ) represents the similarity between p c and p d ;
[0034] represents the mean value of the ratio of the absolute difference of the corresponding elements of the data vectors to the sum of the corresponding elements; represents the cosine value of the data vector.
[0035] Furthermore, in the step S23, the formulas for updating the mean μ i+j and the variance σ i+j are as follows:
[0036]
[0037]
[0038] Among them, is the updated value of μ i+j , which is a linear combination of the mean μ i+j of the (i + j)-th batch and the mean μ i+k of the (i + k)-th batch; represents the adjustment coefficient of the (i + k)-th batch; β1 and β2 represent the linear adjustment coefficients.
[0039] Furthermore, in the step S40, the formula for the loss function is:
[0040] loss whole = l box + loss small ;
[0041]
[0042]
[0043]
[0044]
[0045]
[0046]
[0047] Among them, loss whole represents the overall loss; l box represents the loss of the prediction box; loss small represents the small target loss; (2 - w i h i ) represents the small target loss correction coefficient; S 2 represents dividing the picture into S 2 grids; B represents the number of preset prediction boxes in the picture; Takes the value of 1 when there is a target in the prediction box and 0 when there is no target in the prediction box, and is used to calculate the loss when there is a target in the prediction box; loss x represents the correction loss of the abscissa of the prediction box center point; loss y represents the correction loss of the ordinate of the prediction box center point; loss w represents the correction loss of the width of the prediction box; loss h represents the correction loss of the height of the prediction box; x i represents the actual value of the abscissa of the prediction box, represents the predicted value of the abscissa of the prediction box; y i represents the actual value of the ordinate of the prediction box, represents the predicted value of the ordinate of the prediction box; w i represents the actual value of the width of the prediction box, represents the predicted value of the width of the prediction box; h i represents the actual value of the height of the prediction box, represents the predicted value of the height of the prediction box; area i represents the area of the i-th prediction box, area mean represents the average value of the areas of all prediction boxes.
[0048] The advantages of the present invention are as follows:
[0049] By clustering each gesture image, then cropping the gesture area of the gesture image based on the clustering result to obtain sub-images, and performing non-rotated stitching, rotated stitching, and probability-based stitching on the sub-images, the scale and diversity of the image set are expanded, enabling the gesture recognition model trained by the image set to effectively recognize various types of gestures; constructing a Decay BN layer through the moving average algorithm, that is, determining the mean and standard deviation of the current batch by calculating the similarity of different batches, avoiding falling into the local information trap, and greatly improving the generalization ability of the gesture recognition model; constructing the backbone network of yolo-v3 through several horizontal convolutions and vertical convolutions, greatly improving the information integration ability of the gesture recognition model, and thus greatly improving the generalization ability of the gesture recognition model; constructing a loss function through small target loss, that is, correcting the traditional loss function, improving the prediction ability of the gesture recognition model for small targets, and ultimately greatly improving the accuracy of gesture recognition. Brief Description of the Drawings
[0050] The present invention will be further described below with reference to the accompanying drawings in conjunction with the embodiments.
[0051] Figure 1 It is a flowchart of a gesture recognition method based on integrating horizontal and vertical information in the yolo-v3 framework of the present invention.
[0052] Figure 2 It is a schematic structural diagram of the backbone network of yolo-v3 of the present invention.
[0053] Figure 3 It is a schematic structural diagram of the backbone network of traditional yolo-v3. Detailed Embodiments
[0054] The overall idea of the technical solution in the embodiments of the present application is as follows: By clustering each gesture image, then cropping the gesture area of the gesture image based on the clustering result to obtain sub-images, and performing non-rotated stitching, rotated stitching, and probability-based stitching on the sub-images, the scale and diversity of the image set are expanded to effectively recognize various types of gestures; determining the mean and standard deviation of the current batch by calculating the similarity of different batches to avoid falling into the local information trap for generalization ability; constructing the backbone network of yolo-v3 through several horizontal convolutions and vertical convolutions to improve the information integration ability; constructing a loss function through small target loss to improve the prediction ability of small targets, and ultimately achieving an improvement in the accuracy of gesture recognition.
[0055] Please refer to Figures 1 to 3 As shown, a preferred embodiment of a gesture recognition method based on integrating horizontal and vertical information in the yolo-v3 framework of the present invention includes the following steps:
[0056] Step S10: Obtain a large number of gesture pictures, and perform data augmentation on each of the gesture pictures to obtain a picture set;
[0057] Step S20: Construct a Decay BN layer using the moving average algorithm; Since the Decay BN layer can normalize data to ensure the consistency of the distribution of training data, but the data of each batch is only a very small sample of the overall data, and only focuses on the mean and standard deviation of this batch, some data may fall into the dilemma of excessive proportion of local information. Therefore, it is necessary to introduce the data information of other batches to enhance its generality, that is, by calculating the similarity between the current batch data and the previous batch data, and then based on the similarity, use the moving average algorithm to update the mean and standard deviation of the Decay BN layer;
[0058] Step S30: Construct the backbone network of yolo-v3 through several horizontal convolutions and vertical convolutions; That is, on the basis of the original yolo-v3 framework, add 1x1 convolutions from two angles, horizontal and vertical, to integrate channel information, and then expand the information perception of the model through the splicing of the channel dimension, so as to integrate more effective information to contribute to the model;
[0059] Step S40: Construct a loss function based on the small target loss; That is, modify the formula for the prediction box loss in the loss function, add the size loss function part, and optimize the loss function, so as to improve the generalization performance of the model;
[0060] Step S50: Construct a gesture recognition model based on the Decay BN layer, the backbone network, and the loss function;
[0061] Step S60: Use the picture set to train the gesture recognition model;
[0062] Step S70: Use the trained gesture recognition model to perform gesture recognition.
[0063] That is, through data augmentation, redesigning the Decay BN layer, modifying the backbone network, and modifying the loss function, to improve the recognition accuracy of the gesture recognition model for gestures.
[0064] The specific steps of step S10 include:
[0065] Step S11: Obtain a large number of gesture pictures;
[0066] Step S12: Use the similarity function to perform k-means clustering on each of the gesture pictures to obtain N categories;
[0067] Step S13: Select one of the gesture pictures from each of the N categories in turn, crop the gesture area of each gesture picture to obtain sub-pictures, and correct the size of the sub-pictures to 1 / N of the gesture picture;
[0068] Step S14: Stitch the sub-pictures into a first picture, randomly rotate the sub-pictures by different angles and stitch them into a second picture, perform random repeated sampling on each sub-picture N times and then stitch them into a third picture, and form a batch with the sub-pictures, the first picture, the second picture, and the third picture; that is, perform non-rotated stitching, rotated stitching, and probability-based stitching on the sub-pictures;
[0069] Step S15: Based on each batch, form an augmented picture set.
[0070] In the above step S12, the formula of the similarity function is:
[0071] S whole = α1 * (S 11 + S 12 ) + α2 * S2;
[0072]
[0073]
[0074]
[0075] Among them, S whole represents the similarity between picture A and picture B; anchor A represents the gesture target in picture A; anchor B represents the gesture target in picture B; S 11 represents the absolute difference between the mean value of the abscissas of all points in anchor A and the mean value of the abscissas of all points in anchor B ; S 12 represents the absolute difference between the mean value of the ordinates of all points in anchor A and the mean value of the ordinates of all points in anchor B ; S2 represents the mean value of the absolute differences of the abscissas and ordinates of picture A and picture B; α1 and α2 represent similarity coefficients; m A represents the number of pixel points in anchor A ; m B represents the number of pixel points in anchor B ; m represents the number of all pixel points of picture A and picture B; (x i , y i ) represents anchorA The coordinate points in; (x j , y j ) represent the coordinate points in anchor B ; x Ai , x Bj , y Ai and y Bj all represent the gesture target adjustment coefficients. That is, the similarity of the picture is characterized by the linear combination of S 11 , S 12 and S2.
[0076] The specific steps of S20 include:
[0077] Step S21: Select a Decay BN layer (Decay Batch Normalization), and calculate the mean μ i and variance σ i of the i-th batch in the Decay BN layer. Based on i, calculate the mean μ i+j and variance σ i+j of the (i + j)-th batch;
[0078] Step S22: Calculate the similarity between the data of the (i + j)-th batch and the i-th batch;
[0079] Step S23: Update the mean μ i+j and variance σ i+j of the Decay BN layer based on the similarity, and complete the construction of the Decay BN layer.
[0080] In step S22, the calculation formula for the similarity is:
[0081] s i-(i+j) = ∑p c ∈batch i ∑p d ∈batch i+j s(p c , p d );
[0082]
[0083] where s i-(i+j) represents the sum of the similarities of each pair of elements in batch i and batch i+j ; batch i represents the set of data of the i-th batch; batch i+j represents the set of data of the (i + j)-th batch; p c represents batchi one of the data in; p d represents batch i+j one of the data in; s(p c , p d ) represents p c and p d similarity;
[0084] represents the mean of the ratio of the absolute difference of the corresponding elements of the data vectors to the sum of the corresponding elements; represents the cosine value of the data vector.
[0085] In the said step S23, the updated mean μ i+j and variance σ i+j The formula is:
[0086]
[0087]
[0088] where is the updated value of μ i+j is the mean μ of i + j batches, i+j and the mean μ of i + k batches i+k linear combination; represents the adjustment coefficient of the (i + k)-th batch; β1 and β2 represent linear adjustment coefficients.
[0089] In the said step S30, since the backbone network of yolo-v3 is DarkNet-53, and the most important one is the Residual block framework. By using 1x1 convolution and 3x3 convolution and then summing with the original layer, it is ensured that the final fitting of the model will not fall into the dilemma of model degradation. However, the summation fits the channel information into one dimension completely, which may cause the loss of data information. Therefore, the present invention performs data summation and channel splicing in both horizontal and vertical directions, and integrates information from two perspectives of aggregation and dispersion to ensure that the data information will not be lost.
[0090] The said backbone network is as Figure 2 described, and the steps are as follows:
[0091] Step S31: Perform 1x1 convolution and 3x3 convolution longitudinally on layer A1 to generate an output of the first column of 208*208*128;
[0092] Step S32: Perform 1x1 convolution and 3x3 convolution longitudinally on layer A2 to generate an output of the second column of 208*208*64;
[0093] Step S33: Perform a 1x1 vertical convolution on layer A3 to generate an output of 208*208*64 for the third column;
[0094] Step S34: Perform a concat operation on the data from Step S32 and Step S33 to generate an output of 208*208*128;
[0095] Step S35: Perform an ADD operation by adding the results of Step S31 and Step S34, and then perform a 1x1 convolution to generate an output of 208*208*64;
[0096] Step S36: Perform a concat operation on the output of Step S35 and the original output to generate an output of 208*208*128, and then perform a 1x1 convolution to finally generate an output of 208*208*64; The final output contains not only the channel information of the horizontal convolution but also the channel information of the vertical convolution, and the information of the data is retained to the greatest extent.
[0097] In the said Step S40, the formula of the loss function is:
[0098] loss whole = l box + loss small ;
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105] Among them, loss whole represents the overall loss; l box represents the loss of the prediction box; loss small represents the small target loss, which is used to quantify the gap between the overall size of the prediction box and the actual box size, so as to correct the prediction deviation of the small target. The formula means that if the mean ratio of all prediction boxes is less than 0.1, the loss is three times the ratio; if the mean ratio of all prediction boxes is greater than 0.1, the loss value is 0.3;
[0106] (2 - w i h i ) represents the small target loss correction coefficient; S 2It means dividing the picture evenly into S 2 grids; B represents the number of preset prediction boxes in the picture; It takes the value of 1 when there is a target in the prediction box and 0 when there is no target in the prediction box, and is used to calculate the loss when there is a target in the prediction box; loss x represents the correction loss of the abscissa of the center point of the prediction box; loss y represents the correction loss of the ordinate of the center point of the prediction box; loss w represents the correction loss of the width of the prediction box; loss h represents the correction loss of the height of the prediction box; that is, through loss x 、loss y 、loss w and loss h correct the coordinates of the prediction box, change the absolute index into a relative index, and eliminate the influence of the absolute index on the prediction box; x i represents the actual value of the abscissa of the prediction box, represents the predicted value of the abscissa of the prediction box; y i represents the actual value of the ordinate of the prediction box, represents the predicted value of the ordinate of the prediction box; w i represents the actual value of the width of the prediction box, represents the predicted value of the width of the prediction box; h i represents the actual value of the height of the prediction box, represents the predicted value of the height of the prediction box; area i represents the area of the i-th prediction box, area mean represents the average value of the areas of all prediction boxes. That is, a size penalty loss is added to the loss to increase the weight of small-size gestures.
[0107] Since the original loss function of yolo-v3 has two drawbacks. The first is for the prediction of the prediction box, the absolute value of the predicted value and the actual value is calculated, rather than the relative value, resulting in that even if a small target is predicted to be larger, the numerical difference from the smaller loss of a large target may not be significant, and such a loss is not conducive to the detection of small targets; the second is that in the loss function, although the coefficient of the prediction box loss is used for small target punishment, there is no explicit addition of the punishment for small target prediction errors, and there is still a problem of inaccurate small target prediction in the actual prediction process. Therefore, the present invention corrects the original loss function based on the existing problems.
[0108] To sum up, the advantages of the present invention are as follows:
[0109] By clustering each gesture picture, then cropping the gesture area of the gesture picture based on the clustering result to obtain sub-pictures, and performing non-rotated splicing, rotated splicing and probability-based splicing on the sub-pictures, the scale and diversity of the picture set are expanded, so that the gesture recognition model trained by the picture set can effectively recognize various types of gestures; by constructing a Decay BN layer through the moving average algorithm, that is, by calculating the similarity of different batches to determine the mean and standard deviation of the current batch, avoiding falling into the local information trap, and greatly improving the generalization ability of the gesture recognition model; by constructing the backbone network of yolo-v3 through several horizontal convolutions and vertical convolutions, the information integration ability of the gesture recognition model is greatly improved, and then the generalization ability of the gesture recognition model is greatly improved; by constructing a loss function through small target loss, that is, by correcting the traditional loss function, the prediction ability of the gesture recognition model for small targets is improved, and finally the accuracy of gesture recognition is greatly improved.
[0110] Although the specific implementation manners of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope protected by the claims of the present invention.
Claims
1. A gesture recognition method for integrating horizontal and vertical information based on the yolo-v3 framework, characterized in that: It includes the following steps: Step S10: Obtain a large number of gesture pictures, and perform data augmentation on each of the gesture pictures to obtain a picture set; Step S20: Use the moving average algorithm to construct a Decay BN layer; Step S30: Construct the backbone network of yolo-v3 through several horizontal convolutions and vertical convolutions; Step S40: Construct a loss function based on the small target loss; Step S50: Construct a gesture recognition model based on the Decay BN layer, the backbone network, and the loss function; Step S60: Use the picture set to train the gesture recognition model; Step S70: Use the trained gesture recognition model to perform gesture recognition; The specific content of step S10 includes: Step S11: Obtain a large number of gesture pictures; Step S12: Use the similarity function to perform k-means clustering on each of the gesture pictures to obtain N categories; Step S13: Select one gesture picture from each of the N categories in sequence, crop the gesture area of each gesture picture to obtain sub-pictures, and correct the size of the sub-pictures to 1 / N of the gesture pictures; Step S14: Stitch each of the sub-pictures into a first picture, randomly rotate each of the sub-pictures by different angles and stitch them into a second picture, randomly resample each of the sub-pictures N times and then stitch them into a third picture, and form a batch with the sub-pictures, the first picture, the second picture, and the third picture; Step S15: Based on each of the batches, form a picture set after data augmentation; In step S40, the formula of the loss function is: loss whole = l box + loss small ; Among them, loss whole represents the overall loss; l box represents the loss of the prediction box; loss small represents the loss of small targets; (2 - w i h i ) represents the correction coefficient of the small target loss; S 2 represents dividing the picture into S 2 grids; B represents the number of preset prediction boxes in the picture; takes the value of 1 when there is a target in the prediction box and 0 when there is no target in the prediction box, and is used to calculate the loss when there is a target in the prediction box; loss x represents the correction loss of the abscissa of the center point of the prediction box; loss y represents the correction loss of the ordinate of the center point of the prediction box; loss w represents the correction loss of the width of the prediction box; loss h represents the correction loss of the height of the prediction box; x i represents the actual value of the abscissa of the prediction box, represents the predicted value of the abscissa of the prediction box; y i represents the actual value of the ordinate of the prediction box, represents the predicted value of the ordinate of the prediction box; w i represents the actual value of the width of the prediction box, represents the predicted value of the width of the prediction box; h i represents the actual value of the height of the prediction box, represents the predicted value of the height of the prediction box; area i represents the area of the i-th prediction box, area mean represents the average value of the areas of all prediction boxes.
2. The gesture recognition method for integrating horizontal and vertical information based on the YOLO-V3 framework according to claim 1, wherein: In step S12, the formula of the similarity function is: S whole = α1 * (S 11 + S 12 ) + α2 * S2; Among them, S whole represents the similarity between Picture A and Picture B; anchor A represents the gesture target in Picture A; anchor B represents the gesture target in Picture B; S 11 represents anchor A the absolute difference between the mean value of the abscissas of all points in B and the mean value of the abscissas of all points in anchor 12 represents anchor A the absolute difference between the mean value of the ordinates of all points in B and the mean value of the ordinates of all points in anchor A represents anchor A the number of pixel points in B represents anchor B the number of pixel points in i , y i represents the coordinate point in anchor A ; (x j , y j ) represents the coordinate point in anchor B ; x Ai , x Bj , y Ai and y Bj all represent gesture target adjustment coefficients.
3. The gesture recognition method for integrating horizontal and vertical information based on the yolo-v3 framework according to claim 1, wherein: The specific content of step S20 includes: Step S21: Select a Decay BN layer, and calculate the mean μ of the i-th batch in the Decay BN layer i and the variance σ i . Taking i as the reference, calculate the mean μ i+j and the variance σ i+j of the (i + j)-th batch; Step S22: Calculate the similarity between the data of the (i + j)-th batch and the data of the i-th batch; Step S23: Update the mean μ of the Decay BN layer based on the similarity i+j and the variance σ i+j , and complete the construction of the Decay BN layer.
4. The gesture recognition method for integrating horizontal and vertical information based on the YOLO-V3 framework according to claim 3, wherein: In step S22, the calculation formula of the similarity is: s i-(i+j) = ∑p c ∈ batch i ∑p d ∈ batch i+j s(p c , p d ); Among them, s i-(i+j) represents the sum of the similarities of each pair of elements in batch i and batch i+j ; batch i represents the set of the i-th batch of data; batch i+j represents the set of the (i + j)-th batch of data; p c represents a data in batch i ; p d represents a data in batch i+j ; s(p c , p d ) represents the similarity between p c and p d ; represents the mean value of the ratio of the absolute difference of the corresponding elements of the data vector to the sum of the corresponding elements; represents the cosine value of the data vector.
5. The gesture recognition method for integrating horizontal and vertical information based on the yolo-v3 framework according to claim 3, characterized in that: In the step S23, the updated mean μ i+j and variance σ i+j are calculated using the following formulas: Among them, is the updated value of μ i+j which is the mean value of μ for i + j batches i+j and the mean value of μ for i + k batches i+k as a linear combination; represents the adjustment coefficient for the (i + k)-th batch; β1 and β2 represent the linear adjustment coefficients.
Citation Information
Patent Citations
Gesture tracking and recognition method based on deep learning
CN111709310A