Gesture recognition method based on attention-yolov5 network of CBC classifier

By introducing an attention mechanism and an improved CBC classifier into the YOLOv5 network, the problem of insensitivity to gesture features is solved, and more efficient gesture recognition results are achieved.

CN115862127BActive Publication Date: 2025-11-21SHENYANG JIANZHU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211034219.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-26
Publication Date
2025-11-21
Estimated Expiration
2042-08-26

AI Technical Summary

Technical Problem

Existing gesture recognition technologies suffer from low recognition efficiency and accuracy due to insensitivity to gesture features.

Method used

We employ an Attention-YOLOv5 network gesture recognition method based on a CBC classifier. By adding a CBMA module to the YOLOv5 network and improving the CBC classifier, combined with the R-vine Copula model, we enhance feature processing capabilities and preserve the correlation between attributes.

Benefits of technology

It significantly improves the accuracy, average precision, and recall of gesture recognition, thereby enhancing the overall performance of gesture recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862127B_ABST
    Figure CN115862127B_ABST
Patent Text Reader

Abstract

The application discloses an Attention-YOLOv5 network gesture recognition method based on a CBC classifier, and comprises the following steps: obtaining a YOLOv5 network model by training based on a gesture image dataset and labels; adding a CBMA module in the YOLOv5 network model so that important target features account for a larger processing proportion, and obtaining an Attention-YOLOv5 network model by training; establishing an improved CBC classifier based on R-vineCopula, and connecting the CBC classifier to an output end of the YOLOv5; obtaining a gesture image to be recognized, and detecting and recognizing based on the YOLOv5 network model and the improved CBC classifier, and outputting a result. By introducing an attention mechanism, the application improves the problem that a feature difference of a YOLOv5 target detection network is not sensitive. A R-vineCopula model is used to improve a naive Bayes classifier, the CBC classifier retains the correlation between attributes, and the problem of missing picture classification accuracy is solved. Compared with the previous YOLOv5 network, the accuracy, average accuracy and recall rate of the application are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the field of target detection and recognition, and in particular to an Attention-YOLOv5 network gesture recognition method based on a CBC classifier. BACKGROUND

[0002] In the past few decades, the rapid development of computer technology has made people inseparable from computers, and the information interaction between people and computers is an indispensable step. The most common mode of human-computer interaction is to rely on simple mechanical devices, namely keyboard and mouse, and other devices such as touch screen. Although the interaction mode widely exists in people's daily life and is skillfully used, one of the mainstream trends is gesture recognition technology, which is a more human-computer interaction mode in line with human habits. Gesture recognition has been widely used in AR, human-computer interaction, gesture control, auxiliary driving, sign language recognition, machine control and other fields, among which visual-based gesture recognition, data glove-based gesture recognition, special marker-based gesture recognition and acceleration sensor-based gesture recognition are also increasingly perfect. In addition, gesture recognition can also be applied to real-time monitoring of the behavior of prisoners in prison, realizing the prediction and early warning of dangerous behavior of prisoners in prison. Gesture is a basic feature of human behavior and an indispensable part of interpersonal communication. The development of gesture recognition technology makes it possible for people to interact with machines or other devices. According to the differences of gestures in time and space, gestures can be divided into static gestures and dynamic gestures. The study of static gestures mainly considers the position information of gestures, while the study of dynamic gestures needs to consider not only the spatial position change of gestures but also the change rule of gestures in time sequence. At present, there are problems in the field of gesture interaction due to the insensitivity of gesture features, resulting in low recognition efficiency and accuracy. SUMMARY

[0003] In view of the problems existing in the prior art, the application provides an Attention-YOLOv5 network gesture recognition method based on a CBC classifier, which can improve the recognition effect and recognition accuracy of gesture features.

[0004] The application discloses an Attention-YOLOv5 network gesture recognition method based on a CBC classifier, which comprises the following steps:

[0005] S1: training a YOLOv5 network model based on a gesture image dataset and labels prepared;

[0006] S2: adding a CBMA module in the YOLOv5 network model so that important target features account for a larger processing proportion, and training an Attention-YOLOv5 network model;

[0007] S3: Establish a CBC classifier based on the improved R-vine Copula, which is connected to the output end of YOLOv5;

[0008] S4: Obtain the gesture image to be recognized, and detect and recognize based on the YOLOv5 network model and the improved CBC classifier, and output the result.

[0009] Further, the step S1 comprises:

[0010] The gesture image data set is divided into several gesture pictures, and the gesture pictures are labeled.

[0011] The obtained gesture feature pictures and the labeled labels are respectively placed in the imgs and labels modules of the to-be-trained network.

[0012] The YOLOv5 network model is trained by using the labeled gesture pictures.

[0013] Further, the step S1 further comprises: testing the trained YOLOv5 network model, and if qualified, continuing.

[0014] Further, the step S2 comprises:

[0015] For the YOLOv5 network model, modify yolov5s.yaml, modify the C3 module of Backbone to CBAMC3 module after adding attention mechanism, and the parameters remain unchanged.

[0016] Add CBAMC3 module in common.py.

[0017] Modify yolo.py and add additional judgment statements.

[0018] Call the modified yolov5s.yaml during training, which is used to verify the effectiveness of the attention mechanism on the YOLOv5 network model.

[0019] Further, the step S2 further comprises:

[0020] Test the running result of the Attention-YOLOv5 network model through the test image.

[0021] Further, the step S3 comprises:

[0022] The CBC classifier is improved based on the R-vine Copula method, so that the CBC classifier fits the Pair Copula function according to the data characteristics, so as to preserve the correlation between the data.

[0023] Further, the step S3 specifically comprises:

[0024] The Kendall correlation coefficient matrix between attributes under each category is calculated, the structure of each tree is constructed, the absolute value of each tree tau is maximized, and thus the R-vine Copula model is established;

[0025] Each Pair Copula model is fitted with data, the type of each Pair Copula function is selected through the AIC criterion, and the parameters are estimated by the maximum likelihood estimation method;

[0026] The edge probability density function is estimated by using the Gaussian kernel density function, and the joint probability density of the attributes under each category is obtained in combination with the results of the Pair Copula function calculated in the foregoing.

[0027] For unknown sample X, the class to which the sample X belongs is determined by comparing the size of P(Ci|X).

[0028] The application further discloses a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the gesture recognition method based on the CBC classifier Attention-YOLOv5 network.

[0029] The application further discloses a computer device, which comprises a memory, a processor and a computer program stored in the processor and capable of running on the processor, and the processor realizes the gesture recognition method based on the CBC classifier Attention-YOLOv5 network when executing the computer program.

[0030] The application has at least the following beneficial effects:

[0031] The application improves the problem of unsensitivity of feature difference of the YOLOv5 target detection network by introducing the attention mechanism. The R-vine Copula model is used to improve the naive Bayes classifier, the CBC classifier retains the correlation between attributes, and the problem of missing picture classification precision is solved. Compared with the previous YOLOv5 network, the accuracy, average precision and recall rate of the application are significantly improved.

[0032] Other beneficial effects of the application will be described in detail in the specific embodiment part. BRIEF DESCRIPTION OF DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings described below only show some embodiments of the present application, and all other drawings obtained by those of ordinary skill in the art without creative effort based on these drawings belong to the scope of protection of the present application.

[0034] Figure 1 is a flowchart of the gesture recognition method of the Attention-YOLOv5 network based on the CBC classifier according to the preferred embodiment of the present application.

[0035] Figure 2 is a structural diagram of the CBAM module according to the preferred embodiment of the present application.

[0036] Figure 3 is a structural diagram of the Attention-YOLOv5 network model according to the preferred embodiment of the present application.

[0037] Figure 4 is Figure 3 is a specific structural diagram of the CBAM module in the embodiment.

[0038] Figure 5 is a structural diagram of the R-vine according to the preferred embodiment of the present application.

[0039] Figure 6 is a flowchart of the method for improving the CBC classifier according to the preferred embodiment of the present application.

[0040] Figure 7 is an implementation state diagram of the trained YOLOv5 network model according to the preferred embodiment of the present application.

[0041] Figure 8 is an implementation state diagram of the input gesture image according to the preferred embodiment of the present application.

[0042] Figure 9 is an effect diagram of the running process according to the preferred embodiment of the present application.

[0043] Figure 10 is an implementation state diagram of the modified CBMA3 module according to the preferred embodiment of the present application. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described in detail below. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort belong to the scope of protection of the present application.

[0045] As Figures 1 to 6 shown, the application discloses an Attention-YOLOv5 network gesture recognition method based on a CBC classifier, which includes a training process and a detection process, wherein the training process includes: making a training set img, making a label label, adding an attention mechanism module CBAM, improving a CBC classifier based on R-vine Copula theory, and training an improved YOLOv5 network model to obtain an Attention-YOLOv5 network model. The detection process includes: inputting a gesture image into the Attention-YOLOv5 network, recognizing the image detect, classifying the image by the CBC classifier, and obtaining a gesture category.

[0046] The network model running in the application adopts YOLOv5. In each embodiment of the application, each module of the used YOLOv5 is not a single network model, but includes multiple different versions such as YOLOv5s, YOLOv5m, YOLOv5l, YOLOv5x and YOLOv5x+TTA. The YOLOv5 network model mainly includes the following parts:

[0047] (1) Input end: the input end of YOLOv5 still adopts the Mosaic data enhancement method of YOLOv4, and its splicing method adopts random scaling, arrangement and reduction, which obviously improves the detection accuracy of small targets. Adaptive anchor frame calculation is used to calculate and compare the initial prediction framework with the real frame, and the difference obtained by calculation is used for back propagation.

[0048] (2) Backbone: the Focus structure is an innovation point of the YOLOv5 network structure, and the key point is that it contains a slicing operation. YOLOv5 network designs two CSP structures, and in the YOLOv5 network, CSP1_X structure is applied to the backbone network Backbone, and CSP2_X structure is applied to the Neck.

[0049] (3) Neck: YOLOv5 network adopts a similar FPN+PAN structure as YOLOv4. The difference is that the Neck structure of YOLOv4 adopts a normal convolution operation, while YOLOv5 further strengthens the feature fusion capability of YOLOv5 network by referring to the CSP2 structure in CSPNet.

[0050] (4) Output end: YOLOv5 network adopts the loss function of Bounding box and GIOU_Loss, and in the post-processing target detection network process, the nms weighted operation method is adopted to screen and suppress non-maximum value for many target frames.

[0051] In this embodiment, after combining the attention mechanism with the YOLOv5 network, the operation steps are as follows: obtaining a gesture picture data set, performing data set feature division and labeling; placing the obtained gesture feature picture and the labeled label under the imgs and labels folders of the training network respectively; training the YOLOv5 gesture recognition network by using the labeled gesture picture; building a YOLOv5 network model with good precision by using the trained network; detecting the input gesture picture and video by using the built YOLOv5 network model; adding an attention mechanism module to the original YOLOv5 target detection network, mainly including the following three aspects: adding an attention module in commmon.py; adding a judgment condition in yolo.py; adding a corresponding module in the yaml file; according to the actual training effect, selecting to appropriately add an attention module (CBAM module) after the C3 module in the configuration file, and attention should be paid to the change of the channel number and the network layer number in the implementation process.

[0052] The CBAM mixed domain attention mechanism is adopted in the application, and the channel attention and the spatial attention are simultaneously included. The CBAM includes two sub-modules: a Channel Attention Module (CAM) and a Spartial Attention Module (SAM), which respectively realize the attention of the channel and the space. In the spatial attention module, the input feature map is compressed and reduced in dimension, and the difference between the background and the target in the spatial domain is highlighted.

[0053] CBAM is the full name of Convolutional Block Attention Module, which is one of the representative works of attention mechanism published on ECCV2018. The application relates to the attention in the network architecture. The attention not only tells us where to focus on, but also improves the representation of the focus point. The goal is to increase the expressiveness by using the attention mechanism, focus on important features and suppress unnecessary features. In order to emphasize meaningful features in the two dimensions of space and channel, the channel and spatial attention modules are applied in sequence to learn what to focus on and where to focus on in the channel and spatial dimensions respectively. In addition, by understanding the information to be emphasized or suppressed, the information flow in the network is also helpful.

[0054] For the main network architecture of the CBAM module, one is the channel attention module, and the other is the spatial attention module. The CBAM integrates the channel attention module and the spatial attention module in sequence.

[0055] The present application aims at the problem that the existing YOLOv5 feature extraction network feature difference is not sensitive, and introduces an attention mechanism in the YOLOv5 network model. By adding the attention module CBAM, the feature vector is weighted, so that the important target feature occupies a larger processing proportion. Including making gesture image dataset and label, training yolov5 network model, using the trained model to test gesture picture and video. Adding attention module CBAM, improving YOLOv5 network model, weighting feature vector, aiming at the problem that feature difference is not sensitive, so that the important target feature occupies a larger processing proportion.

[0056] The processed data set is sent into the Attention-YOLOv5 model for training, and the running result of the model is detected through the test image.

[0057] The existing YOLOv5 model and the YOLOv5 model added with the method proposed in the present application are compared, the test effects of the two networks are compared, and the results are shown in Table 1.

[0058]

[0059] Table 1 comparison of experimental results

[0060] The R-vine Copula model decomposes the joint probability density function between attributes into a product of a series of Pair Copula functions and edge probability density functions, as long as the corresponding Pair Copula function type is selected and its parameters are estimated, and the edge probability density function is combined, the joint probability density of each category under the attribute can be optimized.

[0061] The specific implementation method of the improved classifier is to apply the R-vine Copula theory to the Bayesian classifier, and the obtained CBC classifier fits the Pair Copula function according to the data characteristics, and the correlation between the data is retained.

[0062] As shown in Figure 6 The present application improves the naive Bayesian classifier based on R-vine Copula, and the specific steps are:

[0063] The Kendall correlation coefficient matrix between attributes under each category is calculated, the structure of each tree is constructed, the absolute value of each tree tau is maximized, and thus the R-vine Copula model is established. Each Pair Copula model is fitted with data, the type of each Pair Copula function is selected through the AIC criterion, and the parameters are estimated by maximum likelihood estimation. The edge probability density function is estimated by using the Gaussian kernel density function, and the joint probability density of the attributes under each category can be obtained by combining the results of the Pair Copula function calculated above. According to the Bayesian classification principle, the unknown sample X belongs to the category with the maximum posterior probability, that is, by comparing the size of P(Ci|X), it is determined which category the sample belongs to.

[0064] The R-vine Copula improved Naive Bayes classifier disclosed in the application is compared with the traditional classifier, and the following table 2 is obtained:

[0065] Class / Correct % Rate CBC NBC PCA-NBC 0 92.7 85.4 91.4 1 90.4 80.3 86.2 2 85.3 83.8 82.7 3 84.1 72.9 73.7 4 86.2 80.2 76.4 5 93.8 83.3 82.7

[0066] Table 2 Classification results under three classifiers

[0067] As can be seen from the table, compared with the existing YOLOv5 network model, the accuracy of the Attention-YOLOv5 network model is improved by 3.1%, the average accuracy is improved by 3.2%, and the recall rate is improved by 5.9%. The improved network model of the application has better detection effect and better recognition and detection accuracy.

[0068] The CBC classifier uses R-vine structure to form the Copula function of the correlation between attributes, and retains the correlation between gesture features. The classification accuracy of 6 data sets under 3 classifiers is shown in table 2. The accuracy of the CBC classifier is higher than that of the NBC classifier and the PCA-NBC classifier, and the PCA-NBC classifier only improves the classification accuracy on some data sets.

[0069] Average precision (Average Precision): The PR curve is a two-dimensional curve drawn with Recall and Precision as the horizontal and vertical coordinate axes. The average precision is the average value of the precision corresponding to all values from 0 to 1 of the recall rate, and is equal to the area under the PR curve. In practice, the value of the integral is approximately equal to the sum of the precision multiplied by the change value of the recall rate at each possible threshold value.

[0070] Mean Average Precision: mAP is the average of AP of each class. The reasons affecting mAP mainly include poor training data, uncertainty of data, insufficient training data, and inaccurate frame labeling in the graph. Sometimes, increasing training data does not necessarily improve mAP much. Using a better training network, mAP will be higher. Recall R: = the number of correctly predicted positive samples / the total number of positive samples.

[0071] In view of the problem that the YOLOv5 target detection network is not sensitive to gesture features, the traditional naive Bayes classifier discards the independence between attributes, causing the problem of missing classification accuracy, and the improved YOLOv5 network and CBC classifier gesture recognition algorithm provided by the application comprises the following steps. Make gesture image dataset and label, train YOLOv5 network model, test gesture picture and video using trained model. Add attention module CBAM, improve YOLOv5 feature extraction network, weight feature vector screening, and solve the problem of feature difference insensitivity, so that important target features occupy a larger processing proportion. Use CBC classifier based on R-vineCopula theory to classify gesture pictures and solve classification bias. Improve the existing framework to further improve the classification accuracy and speed.

[0072] Embodiment

[0073] The application discloses an embodiment of an Attention-YOLOv5 network gesture recognition method based on a CBC classifier, which is specifically as follows:

[0074] Train the training set picture of the training train input: refer to Figure 7 , modify classes_path, run train.py to start training, and after training multiple epochs, the weight will be generated in the logs folder.

[0075] Detect the input gesture picture of the detection: refer to Figure 8 and Figure 9 , run the picture path to store the input picture before running, and run detect.py for detection after modification.

[0076] By making gesture image dataset and label, the yolov5 network model is trained, and the gesture picture and video are tested using the trained model.

[0077] Add attention module CBAM (refer to Figure 2 ), improve YOLOv5 network model, weight feature vector screening, and solve the problem of feature difference insensitivity, so that important target features occupy a larger processing proportion.

[0078] Add CBAM module in YOLOv5 network model, the specific steps are as follows:

[0079] ①. In the structure of yolov5s, modify yolov5s.yaml, modify C3 module into CBAMC3 module after adding attention mechanism, the parameters remain unchanged;

[0080] ②. Add CBAMC3 module in common.py;

[0081] ③. Modify yolo.py, add additional judgment statement, and the modified CBAMC3 module is as shown in Figure 10

[0082] At this point, when training the model, the modified yolov5s.yaml is called, so as to verify the effectiveness of attention mechanism on YOLOv5 network model.

[0083] The CBC classifier is improved as follows:

[0084] The Kendall correlation coefficient matrix between attributes in each category is calculated, the structure of each tree is constructed, the absolute value of each tree tau is maximized, and the R-vine Copula model is established. Figure 5 .

[0085] Each Pair Copula model is fitted with data, the type of each Pair Copula function is selected through AIC criterion, and the parameters are estimated by maximum likelihood estimation.

[0086] The edge probability density function is estimated by using Gaussian kernel density function, and the joint probability density of the attributes in each category can be obtained by combining the results of the Pair Copula function calculated before.

[0087] According to the Bayesian classification principle, the unknown sample X belongs to the category with the maximum posterior probability, that is, by comparing the size of P(Ci|X), it is determined which category the sample belongs to.

[0088] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any skilled person in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.​

Claims

1. A gesture recognition method based on an Attention-YOLOv5 network of a CBC classifier, characterized in that, The method comprises: S1: training a YOLOv5 network model based on a gesture image dataset and labels; S2: adding a CBMA module to the YOLOv5 network model to make important target features occupy a larger processing proportion, and training an Attention-YOLOv5 network model; S3: Establishing a classification based on An improved CBC classifier is connected to the output end of YOLOv5. The step S3 comprises: Based on The method improves the Bayesian classifier to obtain a CBC classifier, so that the CBC classifier fits a function according to characteristics of data, to preserve the correlation between data; The step S3 specifically comprises: Kendall correlation coefficient matrix between attributes under each category is calculated, the structure of each tree is constructed so that the absolute value of each tree is maximum, thereby establishing the model; Fit each of the models with the data, select the type of each function by AIC criterion, and estimate the parameters by maximum likelihood estimation method. Fit each of the models with the data, select the type of each function by AIC criterion, and estimate the parameters by maximum likelihood estimation method. Fit each of the models with the data, select the type of each function by AIC criterion, and estimate the parameters by maximum likelihood estimation The edge probability density function is estimated by using a Gaussian kernel density function, and the joint probability density of the attributes under each category is obtained by combining the results of the functions. function. For unknown sample X, by comparing the size of the distance, it is determined which class the sample X belongs to; S4: obtaining a gesture image to be recognized, and detecting and recognizing based on the YOLOv5 network model and the improved CBC classifier to output a result.

2. The Attention-YOLOv5 network gesture recognition method based on the CBC classifier according to claim 1, characterized in that, The step S1 comprises: performing dataset feature division and labeling on a plurality of gesture pictures in the gesture image dataset; placing the obtained gesture feature pictures and the labeled labels in imgs and labels modules of a network to be trained, respectively; training the YOLOv5 network model by using the labeled gesture pictures.

3. The Attention-YOLOv5 network gesture recognition method based on the CBC classifier according to claim 1, characterized in that, The step S1 further comprises: testing the trained YOLOv5 network model, and if qualified, continuing.

4. The Attention-YOLOv5 network gesture recognition method based on the CBC classifier according to claim 1, characterized in that, The step S2 comprises: For the YOLOv5 network model, modify the C3 module of Backbone to CBAMC3 module after adding attention mechanism, and the parameters remain unchanged; In add CBAMC3 module; modifications , adding additional conditional statements; The modified , for verifying the effectiveness of the attention mechanism on the YOLOv5 network model.

5. The Attention-YOLOv5 network gesture recognition method based on the CBC classifier according to claim 1, characterized in that, The step S2 further comprises: testing the running result of the Attention-YOLOv5 network model by using test images.

6. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 5. The computer program is executed by the processor to implement the gesture recognition method based on the CBC classifier and the Attention-YOLOv5 network as claimed in any one of claims 1 to 5.

7. A computer device comprising: The memory, the processor, and the computer program stored in the processor and executable on the processor, wherein the processor executes the computer program to implement the gesture recognition method based on the CBC classifier and the Attention-YOLOv5 network as claimed in any one of claims 1 to 5.