A Fine-Grained Behavior Recognition Method Based on Progressive Hierarchical Weighted Attention Network
The progressive hierarchical attention network enhances fine-grained human action classification in static images by integrating multi-layer features through self-attention and weighted fusion, addressing the limitations of existing models in similar action differentiation.
Patent Information
- Application Number
- CN202111481340.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-12-06
AI Technical Summary
Existing models perform poorly in similar behavior classifications, making it difficult to achieve high-accurate fine-grained behavior recognition.
A method based on a progressive hierarchical weighted attention network is adopted, and the Resnet50 backbone network and YOLO v5 human detection is combined, and loss value weighted fusion is performed, and feature extraction and classification is used by the self-attention mechanism.
It improves the accuracy of fine-grained behavior recognition, can better distinguish similar behaviors, simulate the process of human recognition behavior, and improve classification effect.
Smart Images

Figure CN114360051B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to a fine-grained behavior recognition method based on a progressive hierarchical weighted attention network. Background Technique
[0002] Human behavior recognition is a research hotspot in the current field of pattern recognition. Human behavior recognition mainly involves photographing a human target and analyzing the acquired data to finally identify the behavior type of the current target. Currently, most methods used in human behavior recognition usually simplify the behavior recognition problem to a video classification or image classification problem, that is: given a cropped video segment or image, the model is required to return a predefined action label, such as: playing football, skydiving, playing basketball.
[0003] Behavior recognition is divided into video and static image behaviors. The main difference between the two is that video has temporal information while an image can only complete the behavior through a single image, with a greater difficulty coefficient. Currently, the focus of most methods is mainly on video, but in real life, human vision can often convey the information of the behavior occurring in the image through a single image, such as reading a book, playing the guitar, etc. Thus, it can be seen that behavior recognition can be completed without temporal information. Therefore, it is feasible to recognize human behavior through static images, and images have advantages that videos do not have, such as small data size, relatively simple annotation, fast speed, and wide application.
[0004] Most existing models are for some relatively common datasets and aim to perform coarse-grained action understanding (such as football, skydiving, etc.). In this case, the background context usually provides discriminative signals rather than the action itself, and this information helps the neural network classify the image. For example, in the UCF101 dataset, the algorithm can rely on the background color to classify human activities. When faced with a dataset with a single background and similar actions, the existing models show much lower accuracy in action discrimination than that which can be achieved by humans. Summary of the Invention
[0005] To solve the problem that existing models do not perform well in the division of similar behaviors, the present invention provides a fine-grained behavior recognition method based on a progressive hierarchical weighted attention network, including the following steps:
[0006] The collected images are preprocessed and then input into a pre-defined neural network model for training. The training model is composed of a 4-layer progressive network with Resnet50 as the backbone, where:
[0007] The first layer of the progressive network trains the original image, calculates its parent class and subclass losses, and jointly performs backpropagation to update the parameters of the first-stage model;
[0008] The second layer of the progressive network uses YOLO v5 for human detection. The detected results are cropped and interpolated, and then fused with the original image as the input for training. The loss calculation and model parameter update methods are the same as those in the first stage;
[0009] The third layer of the progressive network trains the original image, calculates the losses of its parent class and subclasses, and jointly performs backpropagation to update the parameters of the first-stage model;
[0010] The fourth layer of the progressive network further extracts fine-grained features and crops the linearly interpolated image, fuses it with the original image for training, and when dividing its subclasses, introduces the hidden vectors after calculating the self-attention of the third, fourth, and fifth layers of Resnet50. The subclass loss is calculated by weighted fusion according to the loss values of each layer, and the subclass loss and the parent class loss are jointly backpropagated to update the parameters of the fourth-stage model;
[0011] Input the real-time data into the trained network for recognition.
[0012] Furthermore, the process of progressive training of the model includes:
[0013] For the first layer of the progressive network, the original image is used as the input, features are extracted using the Resnet50 network, the feature map of the third-to-last layer of the Resnet50 network is subjected to average pooling of a fixed size. After pooling, it is input into the self-attention network to capture global dependencies, and then input into the classifier for subclass division. The calculated loss is denoted as L1;
[0014] At the same time, the Resnet50 classification layer is used for parent class prediction, and the loss of this parent class prediction is denoted as L r1 , and L1 and L r1 are jointly used for the first backpropagation of the model, and the parameters are updated;
[0015] For the second layer of the progressive network, the original input image and the target image obtained by bilinear interpolation after cropping the human body are fused as the input of this layer. The cropping range of the human body comes from the results of human detection of the original image by YOLOv5;
[0016] After fusing the original input and the target image, use the Resnet50 network to extract features, take the feature map of the second-to-last layer of the Resnet50 network for average pooling of a fixed size. After pooling, connect to the self-attention mechanism, input into the classifier for subclass division, and calculate the loss denoted as L2;
[0017] At the same time, the Resnet50 network classification layer is used for parent class prediction, and the loss of this parent class prediction is denoted as L r2 ; L2 and the loss value L of the parent class r2Perform the second backpropagation and update of the model jointly.
[0018] For the third layer of the progressive network, take the original image as the input, use the Resnet50 network to extract features, perform average pooling of a fixed size on the feature map of the last layer of the Resnet50 network, input the pooled result into the self-attention network to capture global dependencies after pooling, and then input it into the classifier for subclass division. Calculate the loss denoted as L3;
[0019] Meanwhile, use the classification layer of the Resnet50 network for parent class prediction, and denote the loss of this parent class prediction as L r3 , and jointly perform the third backpropagation of the model with L3 and L r3 and update the parameters;
[0020] In the fourth layer of the progressive network, take the original input image and the randomly cropped image after fine-grained feature extraction as the input. After fusing the two images, connect them to the self-attention module. Before subclass division, introduce a hierarchical weighting mechanism, which uses the hidden vectors after attention calculation in the third, fourth, and fifth layers of the Resnet50 network, and weight-fuses the hidden vectors according to the loss values of the first three layers of the progressive network. After the fusion is completed, calculate the subclass loss value denoted as L4;
[0021] Meanwhile, use the classification layer of the Resnet50 network for parent class prediction, and denote the loss of this parent class prediction as L r4 , and jointly perform the final backpropagation and parameter update with L4 and the parent class loss L r4 .
[0022] Furthermore, the process of fusing the original input image and the target image cropped from the human body includes:
[0023] For a given normalized original image X ∈ R (c,h,w) , obtain the approximate image of the human body and extract the image X′, including the following process:
[0024] x center, y center, w, h = YOLO(X)
[0025] lefttopx = int(x center - w / 2.0)
[0026] lefttopy = int(y center - h / 2.0)
[0027] X′ = X[:, lefttpoy + 1:lefttopy + h + 3, lefttpox + 1:lefttopx + w + 1]
[0028] Bilinear interpolation is performed on the obtained rough image extraction image X' to obtain the target image X'';
[0029] The obtained target image X'' is fused with the input image to obtain the input image of the second layer of the progressive network, expressed as:
[0030]
[0031] where x center represents the x value of the center point of the detection box in the output of YOLOv5 for human detection, y center represents the y value of the center point of the detection box in the output of YOLOv5 for human detection, w represents the width of the detection box in the output of YOLOv5 for human detection, and h represents the height of the detection box in the output of YOLOv5 for human detection; int represents the rounding operation; represents the input image of the second layer of the progressive network.
[0032] Furthermore, the process of cropping the image after fine-grained feature extraction in the fourth layer of the progressive network includes:
[0033] For the input image of the second layer of the progressive network Convolution operation is performed using a depthwise separable convolution module, that is, fine-grained feature extraction is performed;
[0034] The result after performing depthwise separable convolution is denoted as X c ∈R (c,h,w) , if h, w are the sizes of the original image, the predefined length and width of the cropped target image are h c <h, w c <w, the initial cropping center coordinates are (i, j), the trainable cropping parameters α, β are initialized, and the process of cropping the image is expressed as:
[0035]
[0036]
[0037] X' c = X[:, leftdy+1:leftdy+h c , leftdx+1:leftdx+w c
[0038] X' c = ZerPad(X' c )
[0039] where X' c represents the image obtained after cropping; ZerPad represents padding the cropped image X' c Perform zero-padding to the same size as the original image, then perform image fusion and progressive network training. The cropping area in this stage is updated along with the model training until the accurate small-range target is cropped. When the classification accuracy is the highest during the model training process, it is determined that the accurate small-range target is cropped, and the model convergence is completed.
[0040] Further, in the fourth layer of the progressive network, the weight calculation formula includes:
[0041]
[0042] Among them, hidden weight [i] represents the weight when splicing different layers of resnet after passing through the attention network during the training of the fourth layer of the progressive network, i ∈ {1, 2, 3}; LY i * represents the loss value after training in the i-th stage of the progressive network.
[0043] Further, the gradient calculation formulas for the first, second, and third layers of the progressive network during backpropagation of the model include:
[0044]
[0045] Among them, F Yi represents the gradient calculation formula during backpropagation of the model during the training of the i-th layer of the progressive network, i ∈ {3, 4, 5}; Y resnet_i represents the result after the first classification of the model during the training of the i-th layer of the progressive network; Y i represents the result after the second classification of the model during the training of the i-th layer of the progressive network; L Yresne_i represents the loss function of the first classification of the model during the training of the i-th layer of the progressive network; L Yi represents the loss function of the second classification of the model during the training of the i-th layer of the progressive network.
[0046] Further, the gradient calculation formula for the fourth layer of the progressive network during backpropagation of the model includes:
[0047]
[0048] Among them, F Yconcat represents the gradient calculation formula during backpropagation of the model during the training of the fourth layer of the progressive network; Y resnet_4 represents the result after the first classification during the training of the fourth layer of the progressive network; Y concate represents the result after the second classification during the training of the fourth layer of the progressive network; L Yresnet_4 represents the loss function of the first classification during the training of the fourth layer of the progressive network; LYconcate It represents the loss function of the second classification during the training process of the fourth layer of the progressive network.
[0049] Furthermore, when inputting real-time data into the trained network for recognition, the input data is classified into large categories through the Resnet50 network. The output results of the third, fourth, and fifth layers in the Resnet50 network are respectively input into the attention network to extract features. The results are concatenated together and input into the classifier for sub-category classification, and this classification result is used as the final classification result.
[0050] Furthermore, before inputting the real-time data into the trained network, preprocess the real-time data, that is, unify the data input into the network to a fixed size.
[0051] The present invention mimics the activity process of human recognition behavior, performs two-stage classification, and improves the classification accuracy by fusing the progressive and data augmentation modules, changing the current situation that the prior art cannot well solve the classification of similar behaviors. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is the overall flowchart of a fine-grained behavior recognition method based on a progressive hierarchical weighted attention network of the present invention;
[0053] Figure 2 It is the system framework diagram of the model of the present invention;
[0054] Figure 3 It is the flowchart for prediction of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0056] The present invention proposes a fine-grained behavior recognition method based on a progressive hierarchical weighted attention network, as Figure 1 , including the following steps:
[0057] Preprocess the collected images and then input them into a predefined neural network model for training. The training model is composed of a 4-layer progressive network with Resnet50 as the backbone, where:
[0058] The first layer of the progressive network trains the original image, calculates its parent class and sub-class losses, and jointly updates the parameters of the first-stage model through backpropagation;
[0059] The second layer of the progressive network uses YOLO v5 for human detection. The detected results are cropped and interpolated, and then fused with the original image as the input for training. The loss calculation and model parameter update methods are the same as those in the first stage;
[0060] The third layer of the progressive network trains the original image, calculates the losses of its parent class and subclasses, and jointly performs backpropagation to update the parameters of the first-stage model;
[0061] The fourth layer of the progressive network further extracts fine-grained features and crops the linearly interpolated image, fuses it with the original image for training, and when dividing its subclasses, introduces the hidden vectors after calculating self-attention in the third, fourth, and fifth layers of Resnet50. Weighted fusion is performed according to the loss values of each layer to calculate the subclass loss, and joint backpropagation with the parent class loss is used to update the parameters of the fourth-stage model;
[0062] Input the real-time data into the trained network for recognition.
[0063] In this embodiment, the collected images are preprocessed and then input into a pre-defined neural network model for training. The training model consists of a 4-layer progressive network with Resnet50 as the backbone. Among them, the first layer of the progressive network trains the original image, calculates the losses of its parent class and subclasses, and jointly performs backpropagation to update the parameters of the first-stage model; the second layer of the progressive network uses YOLO v5 for human detection. The detected results are cropped and interpolated, and then fused with the original image as the input for training. The loss calculation and model parameter update methods are the same as those in the first stage; the training of the third layer of the progressive network is the same as that of the first layer; the fourth layer of the progressive network further extracts fine-grained features and crops the linearly interpolated image, fuses it with the original image for training, but when dividing its subclasses, it introduces the hidden vectors after calculating self-attention in the third, fourth, and fifth layers of Resnet50. Weighted fusion is performed according to the loss values of each layer to calculate the subclass loss, and joint backpropagation with the parent class loss is used to update the parameters of the fourth-stage model.
[0064] In this embodiment, the input images are preprocessed, and the preprocessing includes the following steps:
[0065] Adjust the size of the image to the set unified size of 260×260, crop the image to 256×256, and pad the cropped image with a padding of 8.
[0066] Embodiment 1
[0067] This embodiment further describes the process of progressive training of the model, mainly explaining the acquisition of the dataset input to each layer of the network during the training process. The process of progressive training of the model includes:
[0068] For the first layer of the progressive network, the original image is used as the input. The Resnet50 network is utilized to extract features. The feature map of the third-to-last layer of the Resnet50 network is taken for average pooling of a fixed size. After the pooling is completed, it is input into the self-attention network to capture global dependencies, and then input into the classifier for subclass division. The calculated loss is denoted as L1;
[0069] Meanwhile, the classification layer of Resnet50 is used for parent class prediction, and the loss of this parent class prediction is denoted as L r1 . L1 and L r1 are jointly used for the first backpropagation of the model, and the parameters are updated;
[0070] For the second layer of the progressive network, the original input image and the target image obtained by bilinear interpolation after human body cropping are fused and used as the input of this layer. The cropping range of the human body comes from the result of human body detection on the original image by YOLOv5;
[0071] After fusing the original input and the target image, the Resnet50 network is used for feature extraction. The feature map of the second-to-last layer of the Resnet50 network is taken for average pooling of a fixed size. After the pooling is completed, the self-attention mechanism is connected, and it is input into the classifier for subclass division. The calculated loss is denoted as L2;
[0072] Meanwhile, the classification layer of the Resnet50 network is used for parent class prediction, and the loss of this parent class prediction is denoted as L r2 ; L2 and the loss value L of the parent class r2 are jointly used for the second backpropagation and update of the model.
[0073] For the third layer of the progressive network, the original image is used as the input. The Resnet50 network is utilized to extract features. The feature map of the last layer of the Resnet50 network is taken for average pooling of a fixed size. After the pooling is completed, it is input into the self-attention network to capture global dependencies, and input into the classifier for subclass division. The calculated loss is denoted as L3;
[0074] Meanwhile, the classification layer of the Resnet50 network is used for parent class prediction, and the loss of this parent class prediction is denoted as L r3 , and L3 and L r3 are jointly used for the third backpropagation of the model, and the parameters are updated;
[0075] In the fourth layer of the progressive network, the original input image and the randomly cropped image after fine-grained feature extraction are used as inputs. After fusing the two images, they are fed into the self-attention module. Before subclass division, a hierarchical weighting mechanism is introduced, which utilizes the hidden vectors after attention calculation in the third, fourth, and fifth layers of the Resnet50 network, and weights and fuses the hidden vectors according to the loss values of the first three layers of the progressive network. After the fusion is completed, the subclass loss value is calculated and denoted as L4;
[0076] Meanwhile, the classification layer of the Resnet50 network is used for parent class prediction, and the loss of this parent class prediction is denoted as L r4 , and L4 and the parent class loss L r4 are combined for the final backpropagation and parameter update.
[0077] The process of fusing the original input image and the target image obtained by bilinear interpolation after human body cropping in the second layer of the progressive network includes:
[0078] For a given normalized original image X ∈ R (c,h,w) , to obtain the approximate image of the human body and extract the image X′, the following process is included:
[0079] x center, y center, w, h = YOLO(X)
[0080] lefttopx = int(x center - w / 2.0)
[0081] lefttopy = int(y center - h / 2.0)
[0082] X′ = X[:, lefttpoy + 1:lefttopy + h + 3, lefttpox + 1:lefttopx + w + 1]
[0083] Perform bilinear interpolation on the obtained approximate image extraction image X′ to obtain the target image X″;
[0084] Fuse the obtained target image X″ with the input image to obtain the input image of the second layer of the progressive network, expressed as:
[0085]
[0086] where, x center represents the x value of the center point of the detection box in the output of YOLOv5 for human body detection, and y centerLet y represent the y - value of the center point of the detection box in the output of YOLOv5 for human detection, w represent the width of the detection box in the output of YOLOv5 for human detection, and h represent the height of the detection box in the output of YOLOv5 for human detection; int represents the rounding operation; Represents the input image of the second layer of the progressive network.
[0087] The process of cropping the image after fine - grained feature extraction in the fourth layer of the progressive network includes:
[0088] For the input image of the second layer of the progressive network Perform convolution operation using the depth - separable convolution module, that is, perform fine - grained feature extraction;
[0089] Denote the result after performing the depth - separable convolution as X c ∈R (c,h,w) , if h and w are the sizes of the original image, pre - define the length and width of the cropped target image as h c <h, w c <w, the initial cropping center coordinates are (i, j), initialize the trainable cropping parameters α and β, and the process of cropping the image is expressed as:
[0090]
[0091]
[0092] X′ c =X[:, leftdy + 1:leftdy + h c , leftdx + 1:leftdx + w c
[0093] X′ c =ZerPad(X′ c )
[0094] Among them, X′ c represents the image obtained after cropping; ZerPad means padding the cropped image X′ c with 0 to the same size as the original image, performing image fusion and then training the progressive network, and the cropping area is updated along with the model training.
[0095] Example 2
[0096] In this example, progressive training is performed on the model, as Figure 2 , specifically including:
[0097] Step1. The training process of the first layer of the progressive network
[0098] For the first layer of the progressive network, the original image is input, and the Resnet50 network is used to extract features. This network consists of 5 layers, namely layer1 to layer5. The output result of the Resnet50 network is input into the classifier for the first classification, and the loss of this classification result is calculated and denoted as L Yresnet_1 ;
[0099] Here, the feature map of the third-to-last layer of the Resnet50 network is taken for average pooling of a fixed size. After pooling, the self-attention mechanism is connected, and then it is input into the classifier for subclass division. The calculated loss is denoted as L Y1 , and at this time, the total loss function is denoted as L Y1 *, expressed as:
[0100] L Y1 * = L Y1 +L Yresnet_1 ;
[0101] The gradients and parameters of the model are updated using the above total loss to complete the training.
[0102] Step2. Training process of the second layer of the progressive network After Step1 is completed, the second training is carried out.
[0103] For the second layer of the progressive network, the input here includes two parts. One is the original input image, and the other is the target image cropped from the human body. The cropping area comes from the result of human body detection of the original image by YOLOv5. After fusing the original input and the target image, Resnet50 is used for feature extraction. The output result of the Resnet50 network is input into the classifier for the first classification, and the loss of this classification result is calculated and denoted as L Yresnet_2 ;
[0104] Here, the feature map of the second-to-last layer of the Resnet50 network is taken for average pooling of a fixed size. After pooling, the self-attention mechanism is connected, and then it is input into the classifier for subclass division. The loss of this classification is calculated as L Y2 , and at this time, the total loss function is denoted as L Y2 *, expressed as:
[0105] L Y2 * = L Y2 +L Yresnet_2 ;
[0106] The gradients and parameters of the model are updated using the above total loss to complete the training.
[0107] Step3. Training process of the third layer of the progressive network
[0108] After Step2 ends, the third training is carried out. For the first layer of the progressive network, the original image is input, and the Resnet50 network is used to extract features. The output result of the Resnet50 network is input into the classifier for the first classification, and the loss of this classification result is calculated and denoted as L Yresnet_3 ;
[0109] Here, the feature map of the last layer of the Resnet50 network is taken for average pooling of a fixed size. After pooling, the self-attention mechanism is connected, and then it is input into the classifier for subclass division, and the calculated loss is denoted as L Y3 , and at this time the total loss function is denoted as L Y3 *, which is expressed as:
[0110] L Y3 * = L Y3 +L Yresnet_3 ;
[0111] Use the above total loss to update the gradients and parameters of the model to complete the training.
[0112] Step4. Training process of the fourth layer of the progressive network
[0113] After Step3 ends, the fourth training is carried out. In the fourth layer of the progressive network, the input contains two parts, one is the original input image, and the other is the randomly cropped image after fine-grained feature extraction. Similarly, after fusing the original input and the target image, the Resnet50 network is used to extract features. The output result of the Resnet50 network is input into the classifier for the first classification, and the loss of this classification result is calculated and denoted as L Yresnet_4 ;
[0114] Here, the feature maps of the last three layers of the Resnet50 network are taken for average pooling of a fixed size. After pooling, the self-attention mechanism is connected. For the subclass division of this layer, a hierarchical weighting mechanism is introduced, and the hidden vectors after the attention calculation of the last three layers of the Resnet50 are used. The hidden vectors are weighted and fused according to the loss values L Y1 *、L Y2 *、L Y3 * of the first three layers of the progressive network. After the fusion is completed, the subclass loss value is calculated and denoted as L Yconcat .
[0115] Among them, the weight calculation formula includes:
[0116]
[0117] Then it is input into the classifier for the second classification, and at this time the total loss function is denoted as L Yconcat , and at this time the total loss function is denoted as LYconcat *, expressed as:
[0118] L Yconcat * = L Yconcat +L Yresnet_4 ;
[0119] Update the gradients and parameters of the model using the above total loss to complete the training.
[0120] In this embodiment, the process of model prediction, such as Figure 3 , includes the following steps:
[0121] Preprocess the image to be recognized, directly input the preprocessed data into the fourth layer of the progressive network, directly use the Resnet50 network to perform large-category classification on the input data, then input the output results of the third, fourth, and fifth layers in the Resnet50 network into the attention network to extract features, splice the results together, input them into the classifier for sub-category classification, and use the classification result as the final classification result.
[0122] In the process of backpropagation in the above steps Step1 to Step3, that is, in the backpropagation of the model, the gradient calculation formula includes:
[0123]
[0124] Among them, F Yi represents the gradient calculation formula for the model during backpropagation when training the i-th layer of the progressive network, i ∈ {3, 4, 5}; Y resnet_i represents the result after the first classification of the model when training with the i-th layer of the progressive network; Y i represents the result after the second classification of the model when training with the i-th layer of the progressive network; L Yresne_i represents the loss function of the first classification of the model when training with the i-th layer of the progressive network; L Yi represents the loss function of the second classification of the model when training with the i-th layer of the progressive network.
[0125] In the process of backpropagation in the above step Step4, that is, in the backpropagation of the model, the gradient calculation formula includes:
[0126]
[0127] Among them, F Yconcat represents the gradient calculation formula for the model during backpropagation when training the fourth layer of the progressive network; Y resnet_4 represents the result after the first classification during the training of the fourth layer of the progressive network; Y concate represents the result after the second classification during the training of the fourth layer of the progressive network; LYresnet_4 It represents the loss function of the first classification during the training process of the fourth layer of the progressive network; L Yconcate It represents the loss function of the second classification during the training process of the fourth layer of the progressive network.
[0128] The model of the present invention consists of a 4-layer progressive network with Resnet50 as the backbone. Among them, the self-attention mechanism is introduced into each layer of the progressive network to assist in modeling, and subclass division is performed, while the Resnet50 network is mainly responsible for parent class division and feature extraction. Specifically, during the training of the second layer of the progressive network, YOLO v5 is used for human detection, and the detected results are cropped and interpolated, and then fused with the original image as the input. The purpose is to assist the progressive network in capturing from rough behaviors to subtle behaviors; at the fourth layer of the progressive network, the linearly interpolated image is further subjected to fine-grained feature extraction, cropped and padded, and then fused with the original image as the input. The purpose is to enable the model to learn more subtle actions in human behaviors and better help the model distinguish similar human actions. In addition, the loss values of the parent class and subclass are calculated at each stage of the model, and the parameters are backpropagated and updated. However, when subclassifying at the fourth layer of the progressive network, the hidden vectors after calculating the self-attention of the third, fourth, and fifth layers of Resnet50 are introduced, and weighted fusion is performed according to the loss values of each layer. The main purpose is to enable the model to spontaneously select which pixel points and which pixel units contribute the most to classification, and improve the classification effect of the model in fine-grained scenarios. Starting from the perspective of fine-grained classification, the present invention encourages cross-level information exchange, fuses the two-stage losses of the parent class and subclass, combines the progressive idea, and introduces a human detection model, improving the accuracy of fine-grained human behavior classification and changing the current situation that the prior art cannot well solve the classification of similar behaviors.
[0129] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A fine-grained behavior recognition method based on a progressive hierarchical weighted attention network, characterized in that It includes the following steps: The collected images are preprocessed and then input into a pre-defined neural network model for training. The training model consists of a 4-layer progressive network with Resnet50 as the backbone, where: The first layer of the progressive network trains the original image, calculates its parent class and subclass losses, and jointly performs backpropagation to update the parameters of the model in the first stage; The second layer of the progressive network uses YOLO v5 for human detection, crops and interpolates the detected results, and fuses them with the original image as the input for training. The loss calculation and model parameter update methods are the same as those in the first stage; The third layer of the progressive network trains the original image, calculates its parent class and subclass losses, and jointly performs backpropagation to update the parameters of the model in the first stage; The fourth layer of the progressive network further extracts fine-grained features and crops the linearly interpolated image, fuses it with the original image for training, and when dividing subclasses, introduces the hidden vectors after calculating self-attention by the third, fourth, and fifth layers of Resnet50. According to the loss values of each layer, weighted fusion is performed to calculate the subclass loss, and joint backpropagation with the parent class loss is used to update the parameters of the model in the fourth stage; Input the real-time data into the trained network for recognition; The process of progressive training of the model includes: In the first layer of the progressive network, the original image is used as the input, features are extracted using the Resnet50 network, the feature map of the third-to-last layer of the Resnet50 network is taken for average pooling of a fixed size. After pooling, it is input into the self-attention network to capture global dependencies, and then input into the classifier for subclass division. The calculated loss is denoted as L1; Meanwhile, the Resnet50 classification layer is used for parent class prediction, and the loss of this parent class prediction is denoted as L r1 , and L1 and L r1 are jointly used for the first backpropagation of the model to update the parameters; In the second layer of the progressive network, the original input image and the target image obtained by bilinear interpolation after cropping the human body are fused as the input of this layer. The cropping range of the human body comes from the results of human detection of the original image by YOLOv5; After fusing the original input and the target image, features are extracted using the Resnet50 network, the feature map of the second-to-last layer of the Resnet50 network is taken for average pooling of a fixed size. After pooling, the self-attention mechanism is connected, and it is input into the classifier for subclass division. The calculated loss is denoted as L2; Meanwhile, the classification layer of the Resnet50 network is used for the prediction of the parent class, and the loss of this parent class prediction is denoted as L r2 ; L2 and the loss value L of the parent class r2 Jointly perform the second backpropagation and update of the model; In the third layer of the progressive network, the original image is used as the input, features are extracted using the Resnet50 network, the feature map of the last layer of the Resnet50 network is taken for average pooling of a fixed size. After pooling, it is input into the self-attention network to capture global dependencies, and input into the classifier for subclass division. The calculated loss is denoted as L3; Meanwhile, the classification layer of the Resnet50 network is used for parent class prediction, and the loss of this parent class prediction is denoted as L r3 , combine L3 and L r3 to jointly perform the third backpropagation of the model and update the parameters; In the fourth layer of the progressive network, the original input image and the randomly cropped image after fine-grained feature extraction are used as the input. After fusing the two images, they are connected to the self-attention module. Before dividing subclasses, a hierarchical weighting mechanism is introduced, which uses the hidden vectors after attention calculation by the third, fourth, and fifth layers of the Resnet50 network, and the hidden vectors are weighted and fused according to the loss values of the first three layers of the progressive network. After fusion, the subclass loss value is calculated and denoted as L4; Meanwhile, the classification layer of the Resnet50 network is used for parent class prediction, and the loss of this parent class prediction is denoted as L r4 , combine L4 with the parent class loss L r4 to jointly perform the final backpropagation and parameter update.
2. The fine-grained behavior recognition method based on a progressive hierarchical weighted attention network according to claim 1, wherein The process of fusing the original input image and the target image obtained by bilinear interpolation after human body cropping in the second layer of the progressive network includes: For a given original image X ∈ R after standardization (c , h ,w) , a rough image of the human body is obtained to extract the image X′, including the following process: x center , y center , w, h = YOLO(X) lefttopx = int(x center - w / 2.0) lefttopy = int(t center - h / 2.0) X′ = X[:, lefttpoy+1:lefttopy+h+3, lefttopx+1:lefttopx+w+1] Perform bilinear interpolation on the obtained rough image to extract the image X′ to obtain the target image X″; Fuse the obtained target image X″ with the input image to obtain the input image of the second layer of the progressive network, denoted as: Among them, x center represents the x-value of the center point of the detection box in the output of YOLOv5 for human detection, y center represents the y-value of the center point of the detection box in the output of YOLOv5 for human detection, w represents the width of the detection box in the output of YOLOv5 for human detection, and h represents the height of the detection box in the output of YOLOv5 for human detection; int represents the rounding operation; represents the input image of the second layer of the progressive network.
3. A fine-grained behavior recognition method based on a progressive hierarchical weighted attention network according to claim 1, characterized in that The process of cropping the image after fine-grained feature extraction in the fourth layer of the progressive network includes: The input image to the second layer of the progressive network Perform convolution operations using a depthwise separable convolution module, that is, perform fine-grained feature extraction; Denote the result after performing depthwise separable convolution as X c ∈R (c,h,w) , if h and w are the sizes of the original image, pre-define the length and width of the cropped target image as h c <h, w c <w, the initial cropping center coordinates are (i, j), initialize the trainable cropping parameters α and β, and the process of cropping the image is expressed as: X′ c = X[:, leftdy + 1:leftdy + h c , leftdx + 1:leftdx + w c X′ c = ZerPad(X′ c ) Among them, X' c represents the image obtained after cropping; ZerPad means padding the cropped image X' c with 0s until it has the same size as the original image, performing image fusion and then progressive network training, and the cropping area is updated along with the model training.
4. A fine-grained behavior recognition method based on a progressive hierarchical weighted attention network according to claim 1, characterized in that In the fourth layer of the progressive network, the weight calculation formula includes: Among them, hidden weight [i] represents the weight when splicing different layers of resnet after passing through the attention network during the training process of the fourth layer of the progressive network, i ∈ {1, 2, 3}; LY i * represents the loss value after training in the i-th stage of the progressive network.
5. A fine-grained behavior recognition method based on a progressive hierarchical weighted attention network according to claim 1, characterized in that The gradient calculation formula for backpropagation of the model in the first, second, and third layers of the progressive network includes: Among them, F Yi represents the calculation formula of the gradient in the backpropagation of the model during the training of the i-th layer of the progressive network, where i ∈ {3, 4, 5}; Y resnet_i represents the result after the first classification of the model when training with the i-th layer of the progressive network; Y i represents the result after the second classification of the model when training with the i-th layer of the progressive network; L Yresne_i represents the loss function of the first classification of the model when training with the i-th layer of the progressive network; L Yi represents the loss function of the second classification of the model when training with the i-th layer of the progressive network.
6. The fine-grained behavior recognition method based on a progressive hierarchical weighted attention network according to claim 1, wherein The gradient calculation formula for backpropagation of the model in the fourth layer of the progressive network includes: Among them, F Yconcat represents the calculation formula of the gradient in the backpropagation of the model during the training of the fourth layer of the progressive network; Y resnet_4 represents the result after the first classification during the training process of the fourth layer of the progressive network; Y concate represents the result after the second classification during the training process of the fourth layer of the progressive network; L Yresnet_4 represents the loss function of the first classification during the training process of the fourth layer of the progressive network; L Yconcate represents the loss function of the second classification during the training process of the fourth layer of the progressive network.
7. A fine-grained behavior recognition method based on a progressive hierarchical weighted attention network according to claim 1, characterized in that When inputting real-time data into the trained network for recognition, the input data is classified into large categories through the Resnet50 network. The output results of the third, fourth, and fifth layers in the Resnet50 network are respectively input into the attention network to extract features. The results are concatenated together and input into the classifier for sub-category classification, and this classification result is used as the final classification result.
8. A fine-grained behavior recognition method based on a progressive hierarchical weighted attention network according to claim 7, characterized in that Preprocess the real-time data before inputting it into the trained network, that is, unify the data input into the network to a fixed size.