Multi-person Pose Detection Method in Dense Scenes
By building a parallel branch network for positioning classification tasks, the problem of insufficient robustness and accuracy of multi-person pose detection in dense scenarios is solved, and more efficient target detection effect is achieved.
Patent Information
- Application Number
- CN202310649852.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-02
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-06-02
AI Technical Summary
The existing multi-person pose detection method has low detection robustness and insufficient classification accuracy in dense scenarios, especially in terms of target occlusion and category imbalance.
Build a parallel branch network for positioning classification tasks, including a shared backbone network, shallow CNN module and feature fusion module, define target quantity loss, dynamic category weights and dynamic difficulty weight functions, and optimize feature learning to alleviate imbalance problems.
The robustness and classification accuracy of object detection are improved, and the robustness and accuracy of multi-person pose detection in dense scenarios is solved.
Smart Images

Figure CN116682178B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a multi-person pose detection method, which can be used for target recognition in dense scenes. Background Art
[0002] The current research methods for behavior recognition can be divided into single-person pose detection and multi-person pose detection. Good results have been achieved in the research of single-person pose detection, but there are still some obstacles in the research of multi-person pose detection, such as: small body size, target occlusion, similar postures, etc.
[0003] Multi-person pose detection is to locate and recognize the postures of all people in an image, and it is a basic research topic for human action recognition and human-computer interaction. With the successful application of deep learning in the field of computer vision, human pose detection has become increasingly important in the field of computer vision. The purpose of human pose detection is to capture the behavior targets in a video or image sequence and judge their behavior categories. Human pose detection methods can be divided into top-down structures and bottom-up structures. In the top-down method, first, object detection methods such as YOLO, SSD, and FPN are used to locate the target, and then the target is extracted for pose detection. The top-down method is more like the process of human eye vision's perception of target behavior from coarse to fine, and has good recognition performance. The bottom-up method is to detect all body joints in the entire image and group and match the joints to complete behavior categories.
[0004] The existing multi-person pose detection methods are mainly divided into two categories: the first category is the pose detection method based on the analysis of human recognition key points and point motion laws, and the second category is the pose detection method based on target location and image classification. Among them:
[0005] In the first category of methods, the top-down method needs to perform single-target pose detection on each detected target, which will cause a sharp increase in computational cost when the number of image targets increases. In contrast, the bottom-up method is more attractive, which only depends on the context information and the relationship between the internal joints of the body to recognize the target behavior. However, when the object of interest is in a dense scene, such as students in a crowded classroom, conference participants in a hall, or spectators in a large stadium, if this technology is directly applied, it is very difficult to detect severely occluded joints, and the incomplete pose extraction will hinder further behavior analysis. Therefore, the bottom-up method is not feasible in dense scenes.
[0006] The second category of methods does not directly target the integrity of the pose like the first category of methods, but relies on more reliable target positions to classify the target behavior. Most of the existing technical solutions are based on this category of methods to study multi-person pose detection methods.
[0007] The University of Electronic Science and Technology of China discloses a "system and method for dense human pose estimation based on Mask-RCNN" in its patent document with the application number CN201910289577.1, which mainly includes an object detection module, a semantic segmentation module, and an instance segmentation module. Among them, the object detection module is used to obtain accurate object detection frames; the semantic segmentation module defines a semantic segmentation loss function, which supervises the network by treating all people in the picture as foregrounds during training to obtain a semantic segmentation mask; the instance segmentation module further classifies the foreground and background based on the obtained semantic segmentation mask to generate an instance segmentation mask. The system makes the segmentation result more precise through the combination of the semantic segmentation module and the instance segmentation module. This method solves the problem in the traditional technology that it is impossible to accurately perform dense human pose estimation due to multiple objects in the object detection frame during instance segmentation, improving the overall accuracy of detection in dense scenes, but the detection accuracy of small objects with few pixels, low resolution, and weak feature expression ability is low.
[0008] Tang et al. proposed a region of interest (ROI) pooling module in the paper "Pose detection in complex classroom environment based on improved Faster R-CNN" published in IET Image Processing. This module combines semantic features and high-resolution features from the last two layers of the convolutional feature map, making the combined features more expressive than single-layer features. In addition, this method retains local features in the last fully connected layer of feature extraction, making the target features belonging to the same class closer in the feature space, thus having stronger classification ability. However, this method hinders the improvement of its detection accuracy due to the imbalance between foreground and background classes not being considered.
[0009] Gao et al. proposed a pose detection method based on a single-stage object detector in the paper "Multi-scale single-stage pose detection with adaptive sample training in the classroom scene" published in Knowledge-Based Systems. They proposed a multi-scale feature enhancement branch to obtain balanced and robust features; adopted an adaptive fusion mechanism to learn complementary spatial features to make the feature extractor more discriminative; and adopted an adaptive positive sample training strategy to make full use of high-quality predicted positive samples during training to obtain better positive samples. Although this method solves the imbalance problem between positive and negative samples, there is still an imbalance problem in the class distribution of positive samples, resulting in low detection accuracy for some classes.
[0010] When performing object detection using these existing technologies, the network architectures they use mainly consist of a backbone network, a task feature network, and a task head network. Among them, the backbone part is responsible for extracting general features such as color, shape, and texture. The task feature network enhances the general features extracted by the backbone and extracts features specific to the task. The task head network outputs results in different forms for the extracted features to complete the two tasks of object localization and classification. These network architectures can still be improved in terms of inference cost and detection accuracy. The newly proposed YOLOv5 network architecture in recent years has achieved state-of-the-art object detection results compared to traditional network architectures. The paper "Classroom Behavior Detection Based on Improved YOLOv5 Algorithm Combining Multi-Scale Feature Fusion and Attention Mechanism" published by Tang et al. applied YOLOv5 to object detection in dense scenes, proposed a spatial and channel convolutional attention mechanism to extract deep semantic features, and significantly improved the detection accuracy. However, due to the fact that objects are dense and categories are similar in practice, this method still has the problem of low detection robustness.
[0011] Due to the neglect of the contradiction between the localization task and the fine-grained classification task and the small difference between categories in fine-grained classification in the above-mentioned existing technologies, it will lead to the constraint of the localization task on the classification task and reduce the object classification accuracy. Summary of the Invention
[0012] The purpose of the present invention is to propose a multi-person pose detection method in dense scenes aiming at the deficiencies of the above-mentioned existing technologies, so as to avoid the constraint of the localization task on the classification task, alleviate the foreground-background class imbalance problem and the foreground category imbalance problem, and improve the robustness of object detection and the object classification accuracy.
[0013] The technical solutions for achieving the purpose of the present invention are as follows:
[0014] (1) Select an image set of multi-person poses in dense scenes in public, and divide it into a training set, a validation set, and a test set according to the ratio of 8:1:1;
[0015] (2) Construct a parallel branch network for the localization and classification tasks:
[0016] (2a) Construct a common backbone network identical to the backbone network of the existing object detection network YOLOv5;
[0017] (2b) Construct two shallow CNN modules cascaded in sequence by two traditional convolutional layers, two dilated convolutional layers, and one traditional convolutional layer;
[0018] (2c) Construct two feature fusion modules composed of deconvolution layers and upsampling layers, namely the first feature fusion module and the second feature module, respectively, for further extracting the features of the two shallow CNN modules and the shared backbone network, obtaining the N-layer feature maps of the shallow CNN modules and the N-layer feature maps of the shared backbone network after feature extraction, where N ≥ 3;
[0019] (2d) Match the N-layer feature maps of the two shallow CNN modules after feature extraction with the N-layer feature maps of the shared backbone network layer by layer, and perform element-wise multiplication on each layer to obtain the N-layer feature maps for localization and classification;
[0020] (2e) Add a task feature network with the same structure as the existing one to the existing task feature network to form parallel localization and classification branches, and input the N-layer feature maps for localization and classification into the corresponding branches respectively to generate a localization feature pyramid and a classification feature pyramid;
[0021] (2f) Construct a localization task head and a classification task head that are the same as the localization task head and the classification task head of the existing object detection network YOLOv5;
[0022] (2g) Cascade the shared backbone network, one shallow CNN module in (2b), the first feature fusion module, the localization branch, and the localization task head in sequence, and cascade the shared backbone network, the other shallow CNN module in (2b), the second feature fusion module, the classification branch, and the classification task head in sequence to form a parallel branch network for localization and classification tasks;
[0023] (3) Define the functions required in the parallel branch network for localization and classification tasks:
[0024] (3a) Define the target number loss function L num as:
[0025] L num = L MSE (n p , n′ p ),
[0026] where n p is the number of predicted targets in the p-th image, n′ p is the number of true targets in the p-th image, and L MSE (n p , n′ p ) represents calculating the mean square error between n p and n′ p ;
[0027] (3b) Define the dynamic class weight function ω i and the dynamic difficulty weight function d p as:
[0028]
[0029] Among them, c i represents the predicted quantity of the i-th category in the p-th image;
[0030] represents the average target prediction quantity of t images;
[0031] (4) Set the corresponding number of training generations according to the selected public image set, and input the images in the training set into the parallel branch network of the localization and classification task for training:
[0032] (4a) Use the N-layer feature map of the shared backbone network in (2c) to perform regression prediction on the target quantity to obtain the predicted target number n p , and then calculate the target quantity loss value L according to the existing true target number n' p in the training set and the function defined in (3a); num ;
[0033] (4b) Input the localization feature pyramid and the classification feature pyramid obtained in (2e) into the localization task head and the classification task head constructed in (2f) respectively for target localization and classification, obtain the prediction results of localization and classification, and calculate the dynamic difficulty weight d p and the dynamic category weight ω i of each image according to the function defined in (3b) to optimize the feature learning in the task feature network;
[0034] (4c) Calculate the regression loss value L reg and the classification loss value L cls respectively by the existing regression loss function and classification loss function, and calculate the total loss value L according to L cls , L reg , the target quantity loss value L num and the existing total loss calculation function, and update the weights in the parallel branch network of the localization and classification task generation by generation through backpropagation until the set number of training generations is reached to obtain the preliminarily trained parallel branch network of the localization and classification task;
[0035] (4d) Input the validation set in the image set into the preliminarily trained parallel branch network of the localization and classification task, adjust its hyperparameters the same as those defined in the existing target detection network, and then train again,
[0036] (4e) Repeat (4d) until the best detection effect is achieved to obtain the finally trained parallel branch network of the localization and classification task;
[0037] (5) Input the test set into the finally trained parallel branch network of the localization and classification task to obtain the detection results on this data set.
[0038] Compared with the prior art, the present invention has the following advantages:
[0039] First, since two shallow CNN modules and a feature fusion module are provided in the parallel branch network for the localization and classification tasks, and the two independent shallow CNN modules use convolution and dilated convolution to extract the feature information of the image, and then fuse it with the feature information extracted by the backbone network respectively, the semantic information in the extracted features is enriched, and the robustness of object detection is improved;
[0040] Second, since the present invention constructs the object localization task and the object classification task into two parallel branch structures, and sets these two branches to have the same structure but not share parameters, the constraint of the localization task on the classification task is avoided;
[0041] Third, by defining the calculation formulas for the object number loss, the dynamic difficulty weight, and the dynamic classification weight, the present invention calculates the object number loss to backpropagate and train the parallel branch network structure for the localization and classification tasks. At the same time, the dynamic class weight and the dynamic difficulty weight are calculated respectively according to the predicted object number to optimize the feature learning process, alleviating the foreground-background class imbalance problem and the foreground class imbalance problem, and improving the object classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is the implementation flowchart of the present invention;
[0043] Figure 2 is the structure diagram of the parallel branch network for the localization and classification tasks constructed in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The embodiments of the present invention will be further described in detail with reference to the accompanying drawings.
[0045] This example is based on the detection of multiple human postures in a dense scene, and the detection targets mainly include students in a crowded classroom, conference participants in a hall, or audiences in a large stadium.
[0046] Refer to Figure 1 , the implementation steps of this example are as follows:
[0047] Step 1, select a public image set and divide it.
[0048] Select an image set from the publicly available database of multiple human posture images in a dense scene, and divide it into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0049] Step 2, construct a parallel branch network for the localization and classification tasks.
[0050] Refer to Figure 2 , the implementation of this step is as follows:
[0051] 2.1) Construct a shared backbone network identical to the backbone network of the existing object detection network YOLOv5;
[0052] The YOLOv5 backbone network is composed of a cascaded first cross-stage module CSP1, a second cross-stage module CSP2, a third cross-stage module CSP3, and an adaptive pooling module SSP. Among them:
[0053] The first cross-stage module CSP1 is used to extract small-scale features;
[0054] The second cross-stage module CSP2 is used to extract medium-scale features;
[0055] The third cross-stage module CSP3 is used to extract large-scale features;
[0056] The adaptive pooling module SSP is used to perform feature pooling operations on the features extracted by the first cross-stage module CSP1, the second cross-stage module CSP2, and the third cross-stage module CSP3.
[0057] 2.2) Construct two shallow CNN modules cascaded in sequence by two traditional convolutional layers, two dilated convolutional layers, and one traditional convolutional layer;
[0058] For the first two traditional convolutional layers, the kernel sizes are 3×3 and 1×1 respectively, and the strides are both 1;
[0059] For the two dilated convolutional layers, the dilation rates are 3 and 2 respectively, and the strides are both 1;
[0060] For the last traditional convolutional layer, the kernel size is 1×1 and the stride is 1;
[0061] 2.3) Construct two feature fusion modules, namely the first feature fusion module and the second feature module, each composed of a transposed convolutional layer with a kernel size of 3×3 and a stride of 1 and an upsampling layer with an upsampling factor of 2, for further extracting the features of the two shallow CNN modules and the shared backbone network, obtaining the N-layer feature maps of the shallow CNN modules after feature extraction and the N-layer feature maps of the shared backbone network, where N≥3;
[0062] 2.4) Match the N-layer feature maps of the two shallow CNN modules after feature extraction with the N-layer feature maps of the backbone network after feature extraction layer by layer, and perform element-wise multiplication on each layer to obtain the N-layer feature maps for localization and classification;
[0063] 2.5) Add a task feature network with the same structure as the existing one to the existing task feature network to form parallel localization and classification branches;
[0064] 2.6) Input the located N - layer feature maps in step 2.4) into the location branch to generate a location feature pyramid:
[0065] 2.6.1) Set N = 3, then the input located N - layer feature maps consist of three - layer feature maps;
[0066] 2.6.2) Generate a location feature pyramid:
[0067] Suppose the three - layer located feature maps are respectively the small - size feature map S1, the medium - size feature map S2, and the large - size feature map S3. Input S1 into the up - sampling layer with a multiple of 2 for up - sampling, and then input the up - sampled feature map into the first traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the first feature map S 1→2 , and then add it to the medium - size feature map S2 to obtain the second feature map S′2;
[0068] Input S2 into the up - sampling layer with a multiple of 2 for up - sampling, and then input the up - sampled feature map into the second traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the third feature map S 2→3 , and then add it to the large - size feature map S3 to obtain the large - scale feature map S′3 of the pyramid;
[0069] Input S′3 into the down - sampling layer with a multiple of 2 for down - sampling, and then input the down - sampled feature map into the third traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the fourth feature map S′ 3→2 , and then add it to the second feature map S′2 to obtain the medium - scale feature map of the pyramid
[0070] Input into the down - sampling layer with a multiple of 2 for down - sampling, and then input the down - sampled feature map into the fourth traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the fifth feature map Then add it to the small - size feature map S1 to obtain the small - scale feature map of the pyramid
[0071] Arrange the small - scale feature map of the pyramid the medium - scale feature map of the pyramid the large - scale feature map S′3 of the pyramid from top to bottom to form a location feature pyramid;
[0072] 2.7) Input the classified N - layer feature maps in step 2.4) into the classification branch to generate a classification feature pyramid:
[0073] 2.7.1) Set N = 3, then the input classified N - layer feature maps consist of three - layer feature maps;
[0074] 2.7.2) Generate classification feature pyramid:
[0075] Divide the three - layer classification feature maps into small - size feature maps S a , medium - size feature maps S b , and large - size feature maps S c . Input S a into an up - sampling layer with a multiple of 2 for up - sampling. Then input the up - sampled feature map into the I - th traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the first feature map S a→b . Then add it to the medium - size feature map S b to get the second feature map S′ b ;
[0076] Input S b into an up - sampling layer with a multiple of 2 for up - sampling. Then input the up - sampled feature map into the II - th traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the third feature map S b→c . Then add it to the large - size feature map S c to get the pyramid large - size feature map S′ c ;
[0077] Input S′ c into a down - sampling layer with a multiple of 2 for down - sampling. Then input the down - sampled feature map into the III - th traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the fourth feature map S′ c→b . Then add it to the second feature map S′ b to get the pyramid medium - size feature map
[0078] Input into a down - sampling layer with a multiple of 2 for down - sampling. Then input the down - sampled feature map into the IV - th traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the fifth feature map . Then add it to the small - size feature map S a to get the pyramid small - size feature map
[0079] Arrange the pyramid small - size feature map the pyramid medium - size feature map the pyramid large - size feature map S′ c from top to bottom to form a classification feature pyramid;
[0080] 2.8) Construct a localization task head and a classification task head that are the same as those of the existing object detection network YOLOv5's localization task head and classification task head;
[0081] The positioning task head is composed of two cascaded fully connected layers;
[0082] The classification task head is composed of a first fully connected layer, a second fully connected layer, and a softmax activation layer cascaded in sequence;
[0083] 2.9) Cascade the shared backbone network, one shallow CNN module in 2.2), the first feature fusion module, the positioning branch, and the positioning task head in sequence, and cascade the shared backbone network, the other shallow CNN module in 2.2), the second feature fusion module, the classification branch, and the classification task head in sequence to form a positioning and classification task parallel branch network.
[0084] Step 3, define the functions required in the positioning and classification task parallel branch network.
[0085] 3.1) Define the target number loss function L num as:
[0086] L num = L MSE (n p , n' p ),
[0087] where n p is the number of predicted targets in the p-th image, n' p is the number of true targets in the p-th image, and L MSE (n p , n' p ) represents calculating the mean squared error between n p and n' p ;
[0088] 3.2) Define the dynamic class weight function ω i and the dynamic difficulty weight function d p respectively as:
[0089]
[0090]
[0091] where c i represents the predicted quantity of the i-th class in the p-th image;
[0092] represents the average target prediction quantity of t images.
[0093] Step 4, train the positioning and classification task parallel branch network.
[0094] 4.1) Input the N - layer feature maps of the shared backbone network in step 2.3) into a traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain N - layer feature maps with 256 feature channels. Then, input the N - layer feature maps with 256 feature channels into two cascaded traditional convolutional layers with a stride of 1 and a convolutional kernel size of 3×3 respectively to obtain the number of predicted targets n p , and then calculate the target quantity loss value L according to the existing true target number n′ in the training set p and the function defined in step 3.1) num ;
[0095] 4.2) Input the localization feature pyramid and classification feature pyramid obtained in steps 2.6) and 2.7) into the localization task head and classification task head constructed in 2.8) for target localization and classification respectively. Obtain the prediction results of localization and classification from the localization task head and classification task head respectively, and calculate the dynamic difficulty weight d p and dynamic class weight ω i for each image to optimize the feature learning in the task feature network, where:
[0096] 4.3) Calculate the regression loss value L reg and classification loss value L cls respectively by the existing regression loss function and classification loss function. According to L cls , L reg , target quantity loss value L num and the existing total loss calculation function, calculate the total loss value L, and update the weights in the localization - classification task parallel - branch network generation by generation through backpropagation until the set number of training generations is reached, obtaining a preliminarily well - trained localization - classification task parallel - branch network;
[0097] The calculation formulas of the regression loss value L reg and classification loss value L cls are as follows:
[0098] L reg =L MAE (b i , b′ i ),
[0099] L cls =L focal (c i , c′ i ),
[0100] where b i is the predicted bounding box of the i - th target, b′ i is the true bounding box of the i - th target, and L MAE (b i , b′i ) represents the calculation of b i and b' i the absolute error between them; c i is the predicted category of the i-th target, c' i is the true category of the i-th target, L focal (c i , c' i ) represents the calculation of c i and c' i the focal loss between them;
[0101] In this example, the number of training epochs is set to 200;
[0102] 4.4) Input the validation set in the image set into the parallel branch network of the preliminary trained localization and classification task. After adjusting the hyperparameters that are the same as those defined in the existing object detection network, train again:
[0103] 4.4.1) The hyperparameters to be adjusted in the total loss value L of the parallel branch network of the localization and classification task are: the hyperparameter r cls of the classification loss L cls , the hyperparameter r reg of the regression loss L reg and the hyperparameter r num of the quantity loss L num . Initialize the initial values of these three hyperparameters to 1, and calculate the total loss value L of the parallel branch network of the localization and classification task:
[0104] L = r cls ·L cls + r reg ·L reg + r num ·L num ;
[0105] 4.4.2) Adjust the magnitudes of the three hyperparameters r cls , r reg and r num to obtain the optimal hyperparameters and
[0106] 4.4.3) Set the initial value of the learning rate of the parallel branch network of the localization and classification task to 0.001, and then adjust it:
[0107] When the oscillation amplitude of the total loss value curve obtained when the network is preliminarily trained exceeds 0.3, adjust the learning rate to half of the original learning rate and then train the parallel branch network of the localization and classification task again;
[0108] When the total loss value curve obtained when the network is initially trained does not converge, adjust the learning rate to twice the original learning rate and then train the parallel branch network for the localization and classification task again;
[0109] 4.5) Repeat step 4.4) until the parallel branch network for the localization and classification task achieves the best detection effect, and obtain the finally trained parallel branch network for the localization and classification task.
[0110] Step 5, detect the poses of multiple people in the test set.
[0111] Input the test set into the finally trained parallel branch network for the localization and classification task to obtain the detection results on this data set.
[0112] The above description is only a specific example of the present invention and does not constitute any limitation to the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these corrections and changes based on the idea of the present invention are still within the scope of protection of the claims of the present invention.
Claims
1. A method for multi-person pose detection in a dense scene, characterized in that, It includes the following steps: (1) Select an image set of multiple human postures in a dense scene that is publicly available, and divide it into a training set, a validation set, and a test set according to a ratio of 8:1:1; (2) Construct a parallel branch network for the localization and classification tasks: (2a) Construct a common backbone network that is the same as the backbone network of the existing object detection network YOLOv5; (2b) Construct two shallow CNN modules that are sequentially cascaded by two traditional convolutional layers, two dilated convolutional layers, and one traditional convolutional layer; (2c) Construct two feature fusion modules composed of a transposed convolutional layer and an upsampling layer, namely the first feature fusion module and the second feature module, respectively, for further extracting the features of the two shallow CNN modules and the common backbone network, obtaining the N-layer feature maps of the shallow CNN modules after feature extraction and the N-layer feature maps of the common backbone network, where N≥3; (2d) Match the N-layer feature maps of the two shallow CNN modules after feature extraction with the N-layer feature maps of the common backbone network layer by layer, and perform element-wise multiplication on each layer to obtain the N-layer feature maps for localization and classification; (2e) Add a task feature network with the same structure as the existing one to the existing task feature network to form parallel localization and classification branches, and input the N-layer feature maps for localization and classification into the corresponding branches respectively to generate a localization feature pyramid and a classification feature pyramid; (2f) Construct a localization task head and a classification task head that are the same as the localization task head and the classification task head of the existing object detection network YOLOv5; (2g) Cascade the common backbone network, one shallow CNN module in (2b), the first feature fusion module, the localization branch, and the localization task head in sequence, and cascade the common backbone network, the other shallow CNN module in (2b), the second feature fusion module, the classification branch, and the classification task head in sequence to form a parallel branch network for the localization and classification tasks; (3) Define the functions required in the parallel branch network for the localization and classification tasks: (3a) Define the target quantity loss function L num as follows: L num = L MSE (n p , n' p ), where n p is the number of predicted targets in the p-th image, and n' p is the number of true targets in the p-th image, and L MSE (n p , n' p ) represents calculating the mean squared error between n p and n'; p (3b) Define the dynamic class weight function ω i and the dynamic difficulty weight function d p as follows: Among them, c i represents the predicted quantity of the i-th category in the p-th image; Indicates the average number of target predictions for t images; (4) Set the corresponding number of training epochs according to the selected publicly available image set, and input the images in the training set into the parallel branch network for the localization and classification tasks to train it: (4a) Use the N-layer feature maps of the shared backbone network in (2c) to perform regression prediction on the target quantity, and obtain the predicted target number n p , and then calculate the target quantity loss value L according to the existing true target number n' in the training set p and the function defined in (3a) num ; (4b) Input the localization feature pyramid and classification feature pyramid obtained in (2e) into the localization task head and classification task head constructed in (2f) respectively for object localization and classification, obtain the prediction results of localization and classification, and calculate the dynamic difficulty weight d of each image according to the function defined in (4b) p and the dynamic class weight ω i , to optimize the feature learning in the task feature network; (4c) Calculate the regression loss value L and the classification loss value L respectively by using the existing regression loss function and classification loss function. reg And the classification loss value L cls , According to L cls , L reg , The target quantity loss value L num And the existing total loss calculation function, calculate the total loss value L, and update the weights in the parallel branch network of the localization classification task generation by generation through backpropagation until the set number of training generations is reached, so as to obtain a preliminarily trained parallel branch network of the localization classification task; (4d) Input the validation set in the image set into the preliminarily trained parallel branch network for the localization and classification tasks, adjust its hyperparameters that are the same as those defined in the existing object detection network, and then train it again, (4e) Repeat (4d) until the best detection effect is achieved to obtain the finally trained parallel branch network for the localization and classification tasks; (5) Input the test set into the finally trained parallel branch network for the localization and classification tasks to obtain the detection results on this data set.
2. The method according to claim 1, wherein In the shallow CNN module established in step (2b); the parameter settings of each layer are as follows: For the first two traditional convolutional layers, the convolutional kernel sizes are 3×3 and 1×1 respectively, and the strides are both 1; For the two dilated convolutional layers, the dilation rates are 3 and 2 respectively, and the strides are both 1; For the last traditional convolutional layer, the stride is 1 and the convolutional kernel size is 1×1.
3. The method according to claim 1, wherein In the feature fusion module established in step (2c); the parameter settings of each layer are as follows: For the transposed convolutional layer, the convolutional kernel size is 3×3 and the stride is 1; The upsampling multiple of the upsampling layer is 2.
4. The method according to claim 1, wherein In step (2e), the localization N-layer feature map is input into the localization branch to generate a localization feature pyramid, which is implemented as follows: (2e1) Set N = 3, then the input localization N-layer feature map consists of three layers of feature maps; (2e2) Generate a localization feature pyramid: Suppose the three - layer localization feature maps are the small - size feature map S1, the medium - size feature map S2, and the large - size feature map S3. Input S1 into the up - sampling layer with a multiple of 2, and then input the feature map after the up - sampling layer into the first traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the first feature map S 1→2 , and then add it to the medium - size feature map S2 to obtain the second feature map S'2; Input S2 into the upsampling layer with a multiple of 2, and then input the feature map after the upsampling layer into the second traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the third feature map S 2→3 , and then add it to the large-size feature map S3 to obtain the pyramid large-size feature map S'3; Input S′3 into the downsampling layer with a multiple of 2, and then input the feature map after the downsampling layer into the third traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the fourth feature map S′ 3→2 , and then add it to the second feature map S′2 to obtain the medium-scale feature map in the pyramid Input it into the downsampling layer with a multiple of 2, and then input the feature map after the downsampling layer into the fourth traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the fifth feature map Then add it to the small-size feature map S1 to obtain the pyramid small-size feature map Store the feature map of the small pyramid ruler Store the feature map of the medium pyramid ruler The feature maps S′3 of the large pyramid ruler stored from top to bottom form a positioning feature pyramid.
5. The method according to claim 1, characterized in that, In step (2e), the classification N-layer feature map is input into the classification branch to generate a classification feature pyramid, which is implemented as follows: (2ea) Set N = 3, then the input classification N-layer feature map consists of three layers of feature maps; (2eb) Generate a classification feature pyramid: Three layers The classification feature maps are respectively the small-size feature map S a , the medium-size feature map S b , and the large-size feature map S c . Input S a into the upsampling layer with a multiple of 2, and then input the feature map after passing through the upsampling layer into the I-th traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the first feature map S a→b . Then add it to the medium-size feature map S b to obtain the second feature map S' b ; Input S b into the upsampling layer with a multiple of 2, and then input the feature map after the upsampling layer into the second traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the third feature map S b→c , and then add it to the large-size feature map S c to obtain the pyramid large-size feature map S′ c ; Input S′ c into the downsampling layer with a multiple of 2, and then input the feature map after the downsampling layer into the third traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the fourth feature map S′ c→b , and then add it to the second feature map S′ b to obtain the medium-scale feature map in the pyramid Input it into the downsampling layer with a multiple of 2, and then input the feature map after passing through the downsampling layer into the IVth traditional convolutional layer with a stride of 1 and a convolutional kernel size of 1×1 to obtain the fifth feature map Then add it to the small-size feature map S a to obtain the pyramid small-size feature map Store the pyramid small-scale feature map Store the pyramid medium-scale feature map Store the pyramid large-scale feature map S′ c Arrange them from top to bottom to form a classification feature pyramid.
6. The method according to claim 1, characterized in that In step (4a), the N-layer feature map of the shared backbone network in (2c) is used to perform regression prediction on the number of targets, which is implemented as follows: (4a1) Input the N-layer feature map of the shared backbone network into a traditional convolutional layer with a stride of 1 and a kernel size of 1×1 to obtain an N-layer feature map with 256 feature channels; (4a2) Input the N-layer feature map with 256 feature channels into two cascaded traditional convolutional layers with a stride of 1 and a kernel size of 3×3 respectively to obtain the prediction result of the number of targets.
7. The method according to claim 1, characterized in that Calculate the regression loss value L in step (4c) reg and the classification loss value L cls , and the formulas are as follows: L reg = L MAE (b i , b' i ) L cls = L focal (c i , c' i ) where b i is the predicted bounding box of the i-th target, and b′ i is the ground truth bounding box of the i-th target. L MAE (b i , b′ i ) represents calculating the absolute error between b i and b′ i . c i is the predicted class of the i-th target, and c′ i is the ground truth class of the i-th target. L focal (c i , c′ i ) represents calculating the focal loss between c i and c′ i .
8. The method according to claim 1, characterized in that In step (4d), the hyperparameters in the localization and classification task parallel branch network that are the same as those defined in the existing object detection network are adjusted, which is implemented as follows: (4d1) Set the hyperparameter to be adjusted in the total loss of the parallel branch network for the positioning and classification task to be: classification loss L cls The hyperparameter r cls , regression loss L reg The hyperparameter r reg and quantity loss L num The hyperparameter r num , and initialize the initial values of these three hyperparameters to 1, and calculate the total loss value L of the parallel branch network for the positioning and classification task: L = r cls ·L cls +r reg ·L reg +r num ·L num ; (4d2) Adjust the hyperparameter r cls , r reg and r num The magnitudes of these three hyperparameters, namely r and 1 (4d3) Adjust the learning rate of the localization and classification task parallel branch network: When the oscillation amplitude of the total loss value curve obtained when the network is initially trained exceeds 0.3, adjust the learning rate to half of the original learning rate; When the total loss value curve obtained when the network is initially trained does not converge, adjust the learning rate to twice the original learning rate.
Citation Information
Patent Citations
System and method for dense human body posture estimation based on mask-RCNN
CN110008915A
Target detection method based on shallow spatial feature fusion and adaptive channel screening
CN113128558A
A lightweight object detection method based on multiple receptive fields and attention feature pyramids
CN114937151A