Human body detection method based on improved FCOS

Through the improved FCOS model, the head and body detection branches are constructed in the detection head, and combined with the attention mechanism and the keystroke algorithm, the problem of low detection accuracy of human bodies under complex backgrounds and varied postures is solved, achieving higher detection accuracy and robustness.

CN119942086APending Publication Date: 2025-05-06CHINA ACAD OF SAFETY SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510117229.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing human body detection algorithm has low accuracy in human body detection under complex backgrounds and varied postures.

Method used

Using the improved FCOS model, the head detection branch and body detection branch are simultaneously constructed in the detection head, and the positive samples of the detection box are screened based on the interleaving and comparison of the detection box during the training process, and their weight values ​​are calculated for weighted coordinate regression loss. Use the convolutional block attention module to replace the convolutional block detection head of the FCOS model, and use the key algorithm to solve the best match between the body prediction box and the body detection box.

Benefits of technology

It improves the accuracy and robustness of human detection, enhances the accuracy of the model's recognition of head features, reduces ID switching of the same violator, and improves the reliability of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942086A_ABST
    Figure CN119942086A_ABST
Patent Text Reader

Abstract

The invention relates to a human body detection method based on an improved FCOS, belongs to the technical field of target detection, and solves the problem of low human body detection precision of a human body detection algorithm under complex backgrounds and variable postures in the prior art. The method comprises the following steps: acquiring an image comprising a to-be-detected target; inputting the image into a trained human body detection model to obtain a human body detection result in the image; wherein the human body detection model is an improved FCOS model, a head detection branch and a body detection branch are constructed in each detection head at the same time, and head and body information is bound in the head detection branch; and when the human body detection model is trained, screening positive samples of the detection frame and calculating weight values of the positive samples based on the intersection-to-union ratio of the detection frame with the head labeling frame and the body labeling frame so as to carry out weighted coordinate regression loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a human body detection method based on improved FCOS. Background Art

[0002] Human body detection is a key task in the field of computer vision and is widely used in various scenarios, such as video surveillance, intelligent security, human-computer interaction, autonomous driving, etc. In these applications, accurately detecting and locating human bodies in images or videos is crucial for subsequent analysis and processing.

[0003] Traditional human detection methods are mostly based on manually designed features and classifiers, but these methods have limited effectiveness in human detection under complex backgrounds and changing postures. With the development of deep learning technology, human detection methods based on convolutional neural networks (CNN) have gradually become mainstream. Among them, anchor-based detection methods achieve accurate prediction of human positions by presetting a series of anchor boxes and performing classification and regression on these anchor boxes. However, the design and use of anchor boxes also bring problems such as large amount of calculation and imbalance of positive and negative samples.

[0004] FCOS (Fully Convolutional One-Stage Object Detection) is an anchor-free detection method that simplifies the detection process and improves efficiency by predicting four distance values ​​from the target boundary and the probability of the target category at each position in the feature map. However, the human body has variable postures, large scale changes, and differences in visual features between the head and body, making it difficult to accurately detect and distinguish the head and body. Summary of the invention

[0005] In view of the above analysis, an embodiment of the present invention aims to provide a human body detection method based on improved FCOS, so as to solve the problem that the human body detection accuracy of the existing human body detection algorithm is low under complex backgrounds and variable postures.

[0006] The purpose of the present invention is mainly achieved through the following technical solutions:

[0007] The present invention provides a human body detection method based on improved FCOS, comprising the following steps:

[0008] Acquire an image including a target to be detected;

[0009] The image is input into a trained human body detection model to obtain a human body detection result in the image; wherein the human body detection model is an improved FCOS model, a head detection branch and a body detection branch are simultaneously constructed in each detection head, and the head and body information are bound in the detection head; when training the human body detection model, based on the intersection and union ratio of the detection box with the head annotation box and the body annotation box, positive samples of the detection box are screened and their weight values ​​are calculated to perform weighted coordinate regression loss.

[0010] Furthermore, the head detection branch is used to obtain a head detection frame in the image and a body prediction frame corresponding to the head detection frame;

[0011] The body detection branch is used to obtain a body detection frame in the image;

[0012] The body prediction frame and the body detection frame are matched, and the matched body detection frame and the head detection frame corresponding to the body prediction frame are used together as a human body detection result.

[0013] Furthermore, matching the body prediction frame with the body detection frame, and using the matching result and the head detection frame corresponding to the body prediction frame as the human body detection result, includes:

[0014] Based on the matching degree between the body prediction frame and the body detection frame, constructing a matching matrix;

[0015] The matching matrix is ​​solved using the Hungarian algorithm to obtain the best matching body detection frame for the head detection frame, which constitutes a human body detection result together with the head detection frame;

[0016] For the unmatched head detection frame, the corresponding body prediction frame is used as the human detection result.

[0017] Furthermore, the head detection branch uses a convolutional block attention module to replace the convolution module of the FCOS model detection head.

[0018] Furthermore, the detection frame positive samples include head detection frame positive samples, body detection frame positive samples and multi-category detection frame positive samples;

[0019] The screening of positive samples of the detection box and calculating the weight value thereof includes:

[0020] Based on the intersection-and-union ratio of the detection frame and the head annotation frame, a positive sample of the head detection frame is screened, and based on the distance between the positive sample of the head detection frame and the center of the head annotation frame, a head weight value of the positive sample of the head detection frame is calculated;

[0021] Based on the intersection-and-union ratio of the detection frame and the body annotation frame, the body detection frame positive sample is screened, and based on the distance between the body detection frame positive sample and the center of the body annotation frame, the body weight value of the body detection frame positive sample is calculated; wherein, when the head detection frame positive sample and the body detection frame positive sample are the same detection frame, the detection frame is set as a multi-category detection frame positive sample, and its body weight value is updated to be the sum of the corresponding head weight value and the body weight value;

[0022] The head weight value and the body weight value of the positive sample of the detection frame are added together to obtain the weight value of the positive sample of the detection frame.

[0023] Furthermore, the positive samples of the head detection box are screened and their head weight values ​​are calculated, including:

[0024] Calculate the distance between the center point of each detection frame and the center point of the head annotation frame, select the first k1 detection frames with the closest distance as candidate head detection frame positive samples and calculate the intersection-over-union ratio between them and the head annotation frame;

[0025] Based on the intersection-and-union ratios between the positive samples of the candidate head detection frames and the head annotation frame, the mean and variance of the intersection-and-union ratios of the positive samples of the candidate head detection frames are obtained;

[0026] Taking the mean and the sum of the variance of the intersection-over-union ratio of the candidate head detection frame positive samples as a threshold, further screening the head detection frame positive samples to obtain the final head detection frame positive samples;

[0027] Set the category value of the detection box that is a positive sample of the head detection box to the head and record its index value;

[0028] Based on the center point position of the head detection frame positive sample, the weight of the head detection frame positive sample is calculated as the head weight value of the detection frame positive sample.

[0029] Furthermore, the positive samples of the body detection frame are screened and their body weight values ​​are calculated, including:

[0030] Calculate the distance between the center point of each detection frame and the center point of the body annotation frame, select the first k2 detection frames with the closest distance as candidate body detection frame positive samples and calculate the intersection-union ratio between them and the body annotation frame;

[0031] Based on the intersection-and-union ratios between the positive samples of the candidate body detection frames and the body annotation frames, the mean and variance of the intersection-and-union ratios of the positive samples of the candidate body detection frames are obtained;

[0032] Taking the mean and the sum of the variance of the intersection-over-union ratio of the candidate body detection frame positive samples as a threshold, further screening the body detection frame positive samples to obtain the final body detection boundary frame positive samples;

[0033] Set the category value of the detection box that is a positive sample of the body detection box to body and record its index value;

[0034] Calculating a body weight value of the positive sample of the body detection frame based on the center point position of the positive sample of the body detection frame;

[0035] When the index value of the body detection frame positive sample is the same as the index value of the head detection frame positive sample, it is determined to be the same detection frame.

[0036] Furthermore, the head weight value is calculated using the following formula:

[0037]

[0038] Among them, W head,i Indicates the head weight of the i-th detection box; dist i,j Represents the distance between the i-th detection box and the j-th head annotation box; (GT w,j ,GT h,j ) represents the width and height of the jth head annotation box;

[0039] The body weight value is calculated using the following formula:

[0040]

[0041] Among them, W body,i represents the body weight of the i-th detection box; dist i,m Represents the distance between the i-th detection bounding box and the m-th body annotation box; (GT w,m ,GT h,m ) represents the width and height of the mth body annotation box.

[0042] Furthermore, the human body detection model is trained using the following method:

[0043] Constructing a training data set for the human body detection model; wherein the training data set includes a set of images with human bodies, category labels corresponding to the annotated heads and bodies in each image, bounding boxes, and center point positions;

[0044] The training data set is loaded, and the human body detection model is trained using a total loss function consisting of a classification loss function, a center point loss function, and a weighted coordinate regression loss function. The model parameters are updated using gradient back propagation, and the training is terminated when the loss function value converges to obtain a trained human body detection model; wherein, based on the weight value of the positive sample of the detection box and the coordinate loss function, the weighted coordinate regression loss of each positive sample of the detection box is obtained.

[0045] Furthermore, in the head detection branch, when there is no body annotation box that matches the head annotation box, its coordinate regression loss function is:

[0046] L offsethead =α(|lh * -lh|+|th * -th|+|rh * -rh|+|bh * -bh|)

[0047] Where α represents a hyperparameter; (lh * , th * , rh * , bh * ) represents the coordinates of the corner points of the head detection box; (lh, th, rh, bh) represents the coordinates of the corner points of the head annotation box;

[0048] When there is a body annotation box that matches the head annotation box, its coordinate regression loss function is:

[0049] L offsethead =α(|lh * -lh|+|th * -th|+|rh * -rh|+|bh * -bh|)+β(|lp * -lp|+|tp * -tp|+|rp * -rp|+|bp * -bp|)

[0050] Among them, β represents a hyperparameter; (lp * ,tp * ,rp * ,bp * ) represents the corner coordinates of the predicted body box; (lp, tp, rp, bp) represents the corner coordinates of the body annotation box;

[0051] In the body detection branch, its coordinate regression loss function is:

[0052]

[0053] in, Represents the corner point coordinates of the body detection box; (lp, tp, rp, bp) represents the corner point coordinates of the body annotation box.

[0054] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:

[0055] 1. The technical solution of the present invention uses an improved FCOS model to simultaneously construct a head detection branch and a body detection branch in each detection head, and binds the head and body information in the head detection branch, so that the detection model can more accurately identify and associate the head and body parts of the human body. Compared with the traditional separate detection method, the detection accuracy is higher.

[0056] 2. The technical solution of the present invention, by introducing an attention mechanism in the head detection branch, can focus more attention on capturing features of key parts such as the head, discard noise and irrelevant information, further improve the network model's recognition accuracy of head features, enhance the robustness of the model, and make the detection results more reliable.

[0057] 3. The technical solution of the present invention achieves the best matching between the body prediction frame and the body detection frame by constructing a matching matrix and solving it using the Hungarian algorithm, ensuring that the head detection frame and the matched body detection frame are used as human detection results together, and can effectively reduce the ID switching of the same violator even when the person's body is obscured.

[0058] 4. According to the technical solution of the present invention, in the process of matching positive samples, the size of the human body is often larger than the human head, so the number of anchor frames for human body matching is much higher than that of the human head, which reduces the accuracy of head detection. Therefore, in the process of matching anchor frames with real frames, a layered positive sample matching mechanism is proposed to adaptively adjust the matching of anchor frames with real frames. Through this mechanism, higher quality positive samples can be obtained and the weights of the anchor frames can be updated, thereby highlighting the importance of the overlapping area between the head and the body, which can improve the convergence speed and target positioning accuracy.

[0059] In the present invention, the above-mentioned technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can become obvious from the description, or can be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The drawings are only for the purpose of illustrating particular embodiments and are not to be considered limiting of the present invention. Like reference symbols denote like components throughout the drawings.

[0061] Figure 1 Schematic diagram of a human body detection method based on improved FCOS in an embodiment of the present invention;

[0062] Figure 2 A schematic diagram of the structure of a human body detection model in an embodiment of the present invention;

[0063] Figure 3Schematic diagram of the structure of the convolutional block attention module of the head detection branch in an embodiment of the present invention. DETAILED DESCRIPTION

[0064] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.

[0065] A specific embodiment of the present invention discloses a human body detection method based on improved FCOS, such as Figure 1 As shown, the steps S1 to S2 are included:

[0066] Step S1, acquiring an image including a target to be detected.

[0067] The target to be detected refers to the person to be identified. Specifically, an image including the person to be identified taken by the on-site monitoring equipment is obtained, and the image is preprocessed including but not limited to noise removal, image enhancement, etc., to facilitate subsequent image recognition.

[0068] Step S2, inputting the image into a trained human body detection model to obtain a human body detection result in the image; wherein the human body detection model is an improved FCOS model, constructing a head detection branch and a body detection branch in each detection head at the same time, and binding the head and body information in the detection head; when training the human body detection model, based on the intersection and union ratio of the detection box with the head annotation box and the body annotation box, respectively, the detection box positive samples are screened and their weight values ​​are calculated to perform weighted coordinate regression loss.

[0069] Specifically, Figure 2 As shown, the human body detection model of this embodiment is an improved FCOS model, which constructs a multi-branch detection head to extract head and body features respectively, and makes predictions to ensure that both body and head information are included, while establishing the correlation between them.

[0070] More specifically, the FCOS model is a fully convolutional single-stage anchor-free detection model. Different from traditional anchor-based methods (such as YOLO), FCOS does not use predefined anchor boxes, but directly predicts the position and category of objects at each position in the feature map. FCOS can process input images of any size and output feature maps of corresponding sizes.

[0071] The human body detection model includes a backbone network, a neck network and a head network.

[0072] The backbone network is used to extract features of the image and downsample the image to generate multiple feature maps of different sizes. The backbone network in this embodiment adopts a pre-trained classification network. Exemplarily, a ResNet50 structure can be used; wherein the backbone network includes an initial convolution module and three convolution modules connected in sequence to generate feature maps of different sizes.

[0073] Specifically, in the initial convolution module of the backbone network, the image is sequentially passed through a convolution layer with a convolution kernel size of 7x7 and a step size of 2, a batch normalization layer, a ReLU activation function, and a pooling layer to perform preliminary feature extraction and downsampling on the image to reduce the resolution of the feature map.

[0074] In the first convolution module of the backbone network, the output feature map of the initial convolution module is sequentially passed through a number of convolution blocks to gradually extract detailed features of the image, and each convolution block includes a convolution layer with a convolution kernel of 3x3, a batch normalization layer and a ReLU activation function; after each convolution block, a residual block structure is used to add the input directly to the output through a jump connection to alleviate the gradient vanishing and gradient exploding problems, and finally down-sampled through a convolution layer with a step size of 2 to obtain the output feature map C3 of the first convolution module.

[0075] In the second convolution module of the backbone network, the output feature map C3 of the first convolution module gradually extracts the detailed features of the image through several convolution blocks in sequence, and each convolution block includes a convolution layer with a convolution kernel of 3x3, a batch normalization layer and a ReLU activation function; after each convolution block, a residual block structure is used to add the input directly to the output through a jump connection to alleviate the gradient vanishing and gradient exploding problems, and finally downsampled through a convolution layer with a step size of 2 to obtain the output feature map C4 of the second convolution module.

[0076] In the third convolution module of the backbone network, the output feature map C4 of the second convolution module gradually extracts the detailed features of the image through several convolution blocks in sequence, and each convolution block includes a convolution layer with a convolution kernel of 3x3, a batch normalization layer and a ReLU activation function; after each convolution block, a residual block structure is used to add the input directly to the output through a jump connection to alleviate the gradient vanishing and gradient exploding problems, and finally down-sampled through a convolution layer with a step size of 2 to obtain the output feature map C5 of the third convolution module.

[0077] Through the three convolution modules of the above backbone network, three output feature maps are obtained, whose resolution decreases successively, and the semantic information is gradually enriched, retaining rich details and semantic information at different scales, providing high-quality features for the subsequent neck network and head network.

[0078] The neck network is used to further fuse the multi-scale feature maps obtained by the backbone network to enhance the expressiveness of the features. Exemplarily, the neck network adopts a feature pyramid network (FPN) structure to transfer high-level semantic information to low-level feature maps through top-down paths and lateral connections to enhance the semantic information of low-level feature maps.

[0079] Furthermore, in the neck network, 1x1 convolution operations are performed on the output feature map C3 of the first convolution module, the output feature map C4 of the second convolution module, and the output feature map C5 of the third convolution module, respectively, to adjust the number of channels thereof to make them consistent.

[0080] Specifically, in the neck network, the output feature map C5 of the third convolution module is subjected to a 1x1 convolution operation to obtain a fused feature map P5; after upsampling the fused feature map P5, it is element-wise added to the feature map after the 1x1 convolution operation on the output feature map C4 of the second convolution module, and then passes through a convolution layer with a convolution kernel size of 3x3, a batch normalization layer, and a ReLU activation function in sequence to obtain the fused feature map P4.

[0081] After upsampling the fused feature map P4, it is element-wise added to the feature map after the 1x1 convolution operation on the output feature map C3 of the first convolution module, and then passes through a convolution layer with a convolution kernel size of 3x3, a batch normalization layer and a ReLU activation function in sequence to obtain the fused feature map P3.

[0082] The fused feature map P5 is downsampled in turn through a convolution layer with a convolution kernel size of 3x3 and a step size of 2, a batch normalization layer, and a ReLU activation function to obtain a fused feature map P6.

[0083] The fused feature map P6 is downsampled in turn through a convolution layer with a convolution kernel size of 3x3 and a step size of 2, a batch normalization layer, and a ReLU activation function to obtain a fused feature map P7.

[0084] It should be noted that the neck network effectively integrates feature information of different scales through the above processing flow, and generates a multi-scale feature pyramid through upsampling and downsampling operations, providing rich feature support for subsequent target detection.

[0085] The head network includes 5 detection heads, which are used to detect the fused feature map P3, fused feature map P4, fused feature map P5, fused feature map P6 and fused feature map P7 generated by the neck network respectively, obtain the human body detection results in each fused feature map and merge them to obtain the final human body detection result.

[0086] Furthermore, a head detection branch and a body detection branch are simultaneously constructed on each detection head. The detection results include category information, coordinate information and center point confidence of the head detection frame and the body detection frame.

[0087] The head detection branch is used to obtain a head detection frame in the image and a body prediction frame corresponding to the head detection frame.

[0088] The body detection branch is used to obtain a body detection frame in the image.

[0089] The body prediction frame and the body detection frame are matched, and the matched body detection frame and the head detection frame corresponding to the body prediction frame are used together as a human body detection result.

[0090] Furthermore, the head detection branch uses a convolutional block attention module to replace the convolution module of the FCOS model detection head.

[0091] like Figure 3 As shown, the convolutional block attention module (CBAM) is a module that combines channel attention and spatial attention. By connecting the channel attention module and the spatial attention module in series, the attention map of the feature map is calculated from the two dimensions of channel and space, and then multiplied with the input feature map to achieve adaptive learning of features.

[0092] The head detection branch includes a head shared attention convolutional layer, a head regression sub-branch, a head classification sub-branch and a head center sub-branch.

[0093] Each of the fused feature maps passes through the four convolution block attention modules of the head shared attention convolution layer in turn to enhance the features of the feature map.

[0094] Specifically, in each convolutional block attention module, the feature map enhances the features of important channels and suppresses unimportant channel features through the channel attention module. In the channel attention module, the feature map is first processed by average pooling and maximum pooling, and then the channel attention weight is obtained by a multi-layer perceptron (MLP), and the normalized attention weight is obtained by the Sigmoid function. The channel attention weight is multiplied with the original input feature map to enhance the important channel.

[0095] The output feature map of the channel attention module is enhanced by the spatial attention module to enhance the features of important spatial positions and suppress the features of unimportant spatial positions. In the spatial attention module, the output feature map of the channel attention module is first processed by channel average pooling and channel maximum pooling, and the two pooling results are spliced ​​in the channel dimension. Then, after passing through a 7x7 convolution layer and an activation function, a spatial attention weight is generated, and the spatial attention weight is multiplied with the output feature map of the channel attention module to enhance the features of important spatial positions.

[0096] It should be noted that the use of the attention mechanism in the head detection branch can focus more attention on the key parts and discard errors in the processing of noise and irrelevant information, thereby improving the accuracy and robustness of the model and highlighting the accuracy of head detection.

[0097] Furthermore, the output feature map of the head shared attention convolution layer is passed through the 1x1 convolution layer and linear activation function of the head regression sub-branch to obtain 8 coordinate values ​​of each pixel in the feature map, where the first 4 coordinate values ​​represent the head detection frame, and the last 4 coordinate values ​​represent the body prediction frame corresponding to the head detection frame.

[0098] The output feature map of the head shared attention convolution layer is passed through the 1x1 convolution layer and Sigmoid activation function of the head classification sub-branch to obtain the probability value of each pixel in the feature map belonging to each category.

[0099] The output feature map of the head shared attention convolution layer is passed through the 1x1 convolution layer and Sigmoid activation function of the head center degree sub-branch to obtain the center point confidence of each pixel in the feature map.

[0100] More specifically, the body detection branch includes a body shared convolutional layer, a body regression sub-branch, a body classification sub-branch and a body centrality sub-branch.

[0101] The fused feature map is sequentially passed through four convolution layers with 3×3 convolution kernels of the body shared convolution layer, a batch normalization layer, extraction and ReLU activation to enhance the features of the feature map.

[0102] The output feature map of the body shared convolution layer is passed through the 1x1 convolution layer and linear activation function of the body regression sub-branch to obtain four coordinate values ​​of each pixel in the feature map, representing the body detection frame.

[0103] The output feature map of the body shared convolution layer is passed through the 1x1 convolution layer and the Sigmoid activation function of the body classification sub-branch to obtain the probability value of each pixel point in the feature map belonging to each category.

[0104] The output feature map of the body shared convolution layer is passed through the 1x1 convolution layer and Sigmoid activation function of the body centrality sub-branch to obtain the central point confidence of each pixel in the feature map.

[0105] Furthermore, matching the body prediction frame with the body detection frame, and using the matching result and the head detection frame corresponding to the body prediction frame as the human body detection result, includes steps S211 to S213:

[0106] Step S211: construct a matching matrix based on the matching degree between the body prediction frame corresponding to one detection head and the body detection frame.

[0107] Specifically, for each body prediction frame corresponding to a head detection frame detected by the head detection branch in a detection head, the intersection and union ratio between the body prediction frame and all body detection frames obtained by the body detection branch is calculated to construct a matrix. For example, the rows represent the body prediction frames, the columns represent the body detection frames, and each element in the matrix represents the intersection and union ratio between the body prediction frame and the body detection frame.

[0108] Step S212: Use the Hungarian algorithm to solve the matching matrix to obtain the best matching body detection frame of the head detection frame, which together with the head detection frame constitutes a human body detection result.

[0109] Specifically, the Hungarian algorithm is an algorithm for finding the minimum weight match in a weighted bipartite graph. The Hungarian algorithm seeks the minimum value, while this embodiment hopes to maximize the intersection-over-union ratio. Therefore, in this embodiment, the negative value of the intersection-over-union ratio is used as the weight, and the Hungarian algorithm is applied to find the optimal match. According to the result of the Hungarian algorithm, a body detection frame is assigned to each head detection frame. If there are multiple head detection frames with similar IoU values ​​to the same body detection frame, the pair with the highest intersection-over-union ratio is selected.

[0110] Step S213: For the unmatched head detection frame, use it and the corresponding body prediction frame as the human body detection result.

[0111] Specifically, for each successfully matched head detection frame, it is combined with the corresponding body detection frame to form a complete human detection result. Since the body parts corresponding to some head detection frames are blocked, missing or failed to be detected in the image, for the head detection frame that is not matched, the corresponding body prediction frame is selected as the human detection result.

[0112] More specifically, for the human body detection result obtained in step S212, the categories of the head detection frame and the body detection frame are both marked as "detection frame" and are assigned the same person identifier;

[0113] For the human body detection result obtained in step S213, the category of the head detection frame is set to "detection frame", the category of the body prediction frame is set to "prediction frame", and the same person identifier is also assigned.

[0114] It should be noted that, for the body detection frame that is not matched to the head detection frame, it is treated separately as part of the human detection result, its category is marked as "detection frame", and a unique person identifier is given.

[0115] Furthermore, the human body detection model is trained using the following method:

[0116] Construct a training data set for the human body detection model; wherein the training data set includes a set of images with human bodies, category labels corresponding to the annotated heads and bodies in each image, bounding boxes, and center point positions.

[0117] Specifically, image data is obtained from multiple sources, including public data sets, self-taken photos or online pictures; the head and body in each image data are annotated, the corresponding bounding boxes are drawn, and the corresponding category labels are assigned to each bounding box; illustratively, the head and body in the image data are annotated separately, and the center position of each bounding box is annotated. It should be noted that the diversity of training data is increased by data enhancement techniques including but not limited to random cropping, rotation, scaling, color transformation, adding noise, affine transformation, etc., so as to improve the robustness of the model.

[0118] The training data set is loaded, and the human body detection model is trained using a total loss function consisting of a classification loss function, a center point loss function, and a weighted coordinate regression loss function. The model parameters are updated using gradient back propagation, and the training is terminated when the loss function value converges to obtain a trained human body detection model; wherein, based on the weight value of the positive sample of the detection box and the coordinate loss function, the weighted coordinate regression loss of the positive sample of each detection box is obtained.

[0119] Furthermore, the detection frame positive samples include head detection frame positive samples, body detection frame positive samples and multi-category detection frame positive samples.

[0120] Specifically, in the process of matching positive samples, the size of the human body is often larger than the human head to a certain extent. Therefore, the number of anchor frames of traditional human body matching is much higher than that of the human head, thereby reducing the accuracy of human head detection. Therefore, this embodiment uses a layered positive sample matching mechanism to adaptively adjust the matching of the anchor frame and the real frame. Through this mechanism, higher quality positive samples can be obtained, and the weight of the anchor frame can be updated to highlight the importance of the overlapping area between the head and the body, which can improve the convergence speed and target positioning accuracy.

[0121] Furthermore, the screening of positive samples of the detection box and calculation of their weight values ​​include steps S221 to S223:

[0122] Step S221: based on the intersection-over-union ratio of the detection frame and the head annotation frame, filter the positive sample of the head detection frame, and calculate the head weight value of the positive sample of the head detection frame based on the distance between the positive sample of the head detection frame and the center of the head annotation frame.

[0123] Specifically, through intersection-over-union screening, it can ensure that the selected positive samples have a high degree of overlap with the actual head position, thereby improving the accuracy of detection, and removing negative samples in the detection frame to reduce computational costs and reduce the risk of overfitting of the model during training. At the same time, by calculating the head weight value through the distance between the positive sample of the head detection frame and the center of the head annotation frame, the detection results can be further refined, so that the detection frame closer to the actual head center obtains a higher weight, which is easier to be identified and retained in subsequent processing, helping to improve the overall detection performance.

[0124] Further, the positive samples of the head detection frame are screened and their head weight values ​​are calculated, including steps S2211 to S2215:

[0125] Step S2211, calculate the distance between the center point of each detection frame and the center point of the head annotation frame, select the first k1 detection frames with the closest distance as candidate head detection frame positive samples and calculate the intersection-over-union ratio between them and the head annotation frame.

[0126] Specifically, for each detection box, the coordinates of its center point C are calculated based on the coordinate values ​​of its four corner points. det,a ; For each head annotation box, the coordinates of its center point C are calculated based on the coordinate values ​​of its four corner points gt,j ; Use the Euclidean distance formula to calculate the distance d between the center point of each detection box and the center point of the head annotation box a,j Based on the distance d a,j , select the first k1 detection frames closest to the center point of the head annotation frame as the candidate head detection frame positive samples; k1 is a hyperparameter and can be adjusted according to the actual situation. By selecting the first k1 detection frames closest to the center point of the head annotation frame as the candidate head detection frame positive samples, the area that is unlikely to contain the target can be preliminarily screened, reducing the number of detection frames that need to be processed and reducing the subsequent calculation cost and time.

[0127] More specifically, for each candidate head detection box positive sample, the intersection-over-union ratio between it and each head annotation box is calculated to further evaluate the degree of matching.

[0128] Step S2212: based on the intersection-and-union ratios between the positive samples of the candidate head detection frames and the head annotation frame, obtain the mean and variance of the intersection-and-union ratios of the positive samples of the candidate head detection frames.

[0129] Specifically, the mean and variance of the intersection-over-union ratios of the candidate head detection frame positive samples that match each head annotation frame as a positive sample are calculated; wherein the variance reflects the discreteness of the intersection-over-union ratio, that is, the consistency of the matching degree between the candidate head detection frame positive sample and the head annotation frame; and a high mean value indicates that the candidate frame has a good average match.

[0130] Step S2213: Taking the mean and the sum of the variance of the intersection-over-union ratios of the candidate head detection frame positive samples as a threshold, further screening the head detection frame positive samples to obtain the final head detection frame positive samples.

[0131] Step S2214: Set the category value of the detection box that is a positive sample of the head detection box to head and record its index value.

[0132] Specifically, by setting the category value, the model can clearly know the target category corresponding to each detection box positive sample, which helps the model learn to distinguish targets of different categories. The index value of the detection box that is the head detection box positive sample is saved in the positive sample index set POS, which is used to quickly locate and reference these positive samples in subsequent calculations.

[0133] Step S2215: Based on the center point position of the head detection frame positive sample, calculate the weight of the head detection frame positive sample as the head weight value of the detection frame positive sample.

[0134] Furthermore, the head weight value is calculated using the following formula:

[0135]

[0136] Among them, W head,i Indicates the head weight of the i-th detection box; dist i,j Represents the distance between the i-th detection box and the j-th head annotation box; (GT w,j ,GT h,j ) represents the width and height of the jth head annotation box.

[0137] Specifically, by assigning higher weights to positive samples of detection boxes that are closer to the true annotation boxes, the model can pay more attention to these key samples, thereby improving learning efficiency and detection accuracy.

[0138] Step S222: based on the intersection-and-union ratio of the detection frame and the body annotation frame, filter the positive samples of the body detection frame, and calculate the body weight value of the positive samples of the body detection frame based on the distance between the positive samples of the body detection frame and the center of the body annotation frame; wherein, when the positive sample of the head detection frame and the positive sample of the body detection frame are the same detection frame, set the detection frame as a multi-category detection frame positive sample, and update its body weight value to the sum of the corresponding head weight value and body weight value.

[0139] Specifically, since the head and body may have size differences, relative position changes, and mutual occlusion in the image, the detection frames of the head and body may have different best matching anchor frames and different feature representations; therefore, in order to more accurately reflect the degree of matching between each detection frame and the true annotation frame, and to improve the model's recognition ability for different target parts, independent positive sample screening and weight calculation are performed on the head detection frame and the body detection frame, respectively, so that the model can optimize its learning process for different target parts, ultimately improving the overall detection performance and robustness.

[0140] Further, the positive samples of the body detection frame are screened and their body weight values ​​are calculated, including steps S2221 to S2226:

[0141] Step S2221, calculate the distance between the center point of each detection frame and the center point of the body annotation frame, select the first k2 detection frames with the closest distance as candidate body detection frame positive samples and calculate the intersection and union ratio between them and the body annotation frame.

[0142] Specifically, similar to the above head detection frame processing steps, for each detection frame, the coordinates of its center point C are calculated based on the coordinate values ​​of its four corner points. det,a ; For each body annotation box, the coordinates of its center point C are calculated based on the coordinate values ​​of its four corner points gt,m ; Use the Euclidean distance formula to calculate the distance d between the center point of each detection box and the center point of the body annotation box a,m Based on the distance d a,m ; Select the first k2 detection frames closest to the nearest distance as the candidate body detection frame positive samples, where k2 is a hyperparameter.

[0143] Step S2222: based on the intersection-and-union ratios between the positive samples of the candidate body detection frames and the body annotation frame, obtain the mean and variance of the intersection-and-union ratios of the positive samples of the candidate body detection frames.

[0144] Step S2223: Taking the mean and the sum of the variance of the intersection-over-union ratios of the candidate body detection frame positive samples as a threshold, further screening the body detection frame positive samples to obtain the final body detection boundary frame positive samples.

[0145] Step S2224: Set the category value of the detection box that is a positive sample of the body detection box to body and record its index value.

[0146] Similarly, the index values ​​of the detection boxes that are positive samples of the body detection boxes are saved in the positive sample index set POS, which is used to quickly locate and reference these positive samples in subsequent calculations.

[0147] Step S2225: Calculate the body weight value of the positive sample of the body detection frame based on the center point position of the positive sample of the body detection frame.

[0148] Furthermore, the body weight value is calculated using the following formula:

[0149]

[0150] Among them, W body,i represents the body weight of the i-th detection box; dist i,m Represents the distance between the i-th detection bounding box and the m-th body annotation box; (GT w,m ,GT h,m ) represents the width and height of the mth body annotation box.

[0151] Step S2226: When the index value of the positive sample of the body detection frame is the same as the index value of the positive sample of the head detection frame, it is determined that they are the same detection frame.

[0152] Specifically, in some cases, when the detection box is large enough, it not only covers the head area, but also covers part of the body area, so that the detection box is judged as a positive sample in both the head and body detection tasks.

[0153] When it is determined to be the same detection frame, the body weight value of the positive sample of the detection frame is updated to the sum of the corresponding head weight value and body weight value, and its category value is set to multiple categories.

[0154] Step S223: Add the head weight value and the body weight value of the positive sample of the detection frame to obtain the weight value of the positive sample of the detection frame.

[0155] Specifically, the calculated weight values ​​can highlight those detection boxes with high prediction accuracy and good matching with the true annotation boxes, so that the model pays more attention to these important samples during the training process. When calculating the loss, the weight value is used to weight the loss function to ensure that the model pays more attention to the contribution of high-quality positive samples during the optimization process, thereby improving the learning efficiency and accuracy of the model.

[0156] It should be noted that, when the category of the detection frame positive sample is head, only the head weight value is calculated. At this time, the body weight value of the current detection frame positive sample is 0, and the weight value of the detection frame positive sample is the head weight value; when the category of the detection frame positive sample is body, only the body weight value is calculated. At this time, the head weight value of the current detection frame positive sample is 0, and the weight value of the detection frame positive sample is the body weight value; when the category of the detection frame positive sample is multi-category, the head weight value and the body weight value are calculated at the same time. Therefore, the weight value of the multi-category detection frame positive sample is the sum of the head weight value and the body weight value.

[0157] Exemplarily, the weight value formula of the positive sample of the detection box is expressed as:

[0158] W i =W head,i +W body,i

[0159] Among them, W i Represents the weight value of the positive sample of the i-th detection box; W head,i W represents the head weight value of the positive sample of the i-th detection box; body,i Represents the body weight value of the positive sample of the i-th detection box.

[0160] Using higher weights for positive samples of multi-category detection boxes can improve the performance of the model when dealing with targets with complex relationships or overlapping areas, making the model pay more attention to samples that are difficult to distinguish but rich in information, thereby optimizing the decision boundary during the training process and improving the overall detection accuracy and generalization ability, especially when there is mutual occlusion or proximity between targets, so that different categories of targets can be more effectively identified and located.

[0161] Furthermore, in the head detection branch, when there is no body annotation box that matches the head annotation box, its coordinate regression loss function is:

[0162] L offsethead =α(|lh * -lh|+|th * -th|+|rh * -rh|+|bh * -bh|)

[0163] Among them, α represents a hyperparameter, which can be set according to the training situation; represents the corner coordinates of the head detection box; (lh, th, rh, bh) represents the corner coordinates of the head annotation box.

[0164] When there is a body annotation box that matches the head annotation box, its coordinate regression loss function is:

[0165] L offsethead=α(|lh * -lh|+|th * -th|+|rh * -rh|+|bh * -bh|)+β(|lp * -lp|+|tp * -tp|+|rp * -rp|+|bp * -bp|)

[0166] Among them, β represents a hyperparameter, which can be set according to the training situation; (lp * , tp * , rp * , bp * ) represents the corner coordinates of the predicted body box; (lp, tp, rp, bp) represents the corner coordinates of the body annotation box.

[0167] Specifically, the coordinate regression loss function is used to measure the difference between the detection box predicted by the model and the true annotation box. Due to occlusion and other reasons, when annotating the bounding box in the image, only the head bounding box may be annotated without annotating its body bounding box. In the head detection branch, the obtained detection box includes the head detection box and the predicted body box. Therefore, when the detection box positive sample is a head detection box positive sample, it means that there is no body annotation box corresponding to the head annotation box. The accuracy of the head detection box predicted by the model is measured by only calculating the sum of the absolute values ​​of the difference between the coordinates of the corner points of the head detection box and the coordinates of the corner points of the head annotation box; when the detection box positive sample is a multi-classification detection box positive sample and there is a body annotation box corresponding to the head annotation box in the image, the sum of the absolute values ​​of the difference between the coordinates of the corner points of the head detection box and the coordinates of the corner points of the head annotation box, as well as the sum of the absolute values ​​of the difference between the coordinates of the corner points of the predicted body box and the coordinates of the corner points of the body annotation box are calculated at the same time to measure the accuracy of the detection box predicted by the model.

[0168] When the coordinates of the predicted box differ greatly from those of the labeled box, the value of the loss function will increase. During training, the model reduces the value of this loss function through optimization algorithms such as gradient descent, so that the predicted detection box is as close to the actual labeled box as possible, thereby improving the detection accuracy.

[0169] It should be noted that α and β in the coordinate regression loss function are hyperparameters and can be adjusted according to the training situation.

[0170] In the body detection branch, its coordinate regression loss function is:

[0171]

[0172] in, Represents the corner point coordinates of the body detection box; (lp, tp, rp, bp) represents the corner point coordinates of the body annotation box.

[0173] Specifically, since the body detection branch only obtains the body detection frame, in the body detection branch, for the positive sample of the body detection frame, its coordinate regression loss function only calculates the distance between the body detection frame and the body annotation frame.

[0174] It should be noted that based on the weight value of the positive sample of the detection box and the coordinate loss function, the weighted coordinate regression loss of the positive sample of each detection box is obtained.

[0175] During the training process, the classification loss function used in this embodiment is a standard cross entropy loss function, which is used to measure the prediction ability of the model for the head and body. The center point loss function uses a binary cross entropy loss function to measure the prediction accuracy of the target center point.

[0176] The total loss function is the sum of the classification loss function, center point loss function and weighted coordinate regression loss function of the head detection branch and the classification loss function, center point loss function and weighted coordinate regression loss function of the body detection branch.

[0177] In summary, the human body detection method based on improved FCOS in the embodiment of the present invention has the following beneficial effects:

[0178] 1. The technical solution of the present invention uses an improved FCOS model to simultaneously construct a head detection branch and a body detection branch in each detection head, and binds the head and body information in the head detection branch, so that the detection model can more accurately identify and associate the head and body parts of the human body. Compared with the traditional separate detection method, the detection accuracy is higher.

[0179] 2. The technical solution of the present invention, by introducing an attention mechanism in the head detection branch, can focus more attention on capturing features of key parts such as the head, discard noise and irrelevant information, further improve the network model's recognition accuracy of head features, enhance the robustness of the model, and make the detection results more reliable.

[0180] 3. The technical solution of the present invention achieves the best matching between the body prediction frame and the body detection frame by constructing a matching matrix and using the Hungarian algorithm to solve it, ensuring that the head detection frame and the matched body detection frame are used as human detection results together, and can effectively reduce the ID switching of the same person even when the person's body is obscured.

[0181] 4. According to the technical solution of the present invention, in the process of matching positive samples, the size of the human body is often larger than the human head, so the number of anchor frames for human body matching is much higher than that of the human head, which reduces the accuracy of head detection. Therefore, in the process of matching anchor frames with real frames, a layered positive sample matching mechanism is proposed to adaptively adjust the matching of anchor frames with real frames. Through this mechanism, higher quality positive samples can be obtained and the weights of the anchor frames can be updated, thereby highlighting the importance of the overlapping area between the head and the body, which can improve the convergence speed and target positioning accuracy.

[0182] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A human body detection method based on improved FCOS, characterized in that: The steps include: Acquire an image including a target to be detected; The image is input into a trained human body detection model to obtain a human body detection result in the image; wherein the human body detection model is an improved FCOS model, a head detection branch and a body detection branch are simultaneously constructed in each detection head, and the head and body information are bound in the detection head; when training the human body detection model, based on the intersection and union ratio of the detection box with the head annotation box and the body annotation box, positive samples of the detection box are screened and their weight values ​​are calculated to perform weighted coordinate regression loss.

2. The method according to claim 1, characterized in that: The head detection branch is used to obtain a head detection frame in the image and a body prediction frame corresponding to the head detection frame; The body detection branch is used to obtain a body detection frame in the image; The body prediction frame and the body detection frame are matched, and the matched body detection frame and the head detection frame corresponding to the body prediction frame are used together as a human body detection result.

3. The method according to claim 2, characterized in that: The matching of the body prediction frame and the body detection frame, and using the matching result and the head detection frame corresponding to the body prediction frame as the human body detection result, includes: Based on the matching degree between the body prediction frame and the body detection frame, constructing a matching matrix; The matching matrix is ​​solved using the Hungarian algorithm to obtain the best matching body detection frame for the head detection frame, which constitutes a human body detection result together with the head detection frame; For the unmatched head detection frame, the corresponding body prediction frame is used as the human detection result.

4. The method according to any one of claims 1 to 3, characterized in that: The head detection branch uses a convolutional block attention module to replace the convolution module of the FCOS model detection head.

5. The method according to claim 1, characterized in that: The detection frame positive samples include head detection frame positive samples, body detection frame positive samples and multi-category detection frame positive samples; The screening of positive samples of the detection box and calculating the weight value thereof includes: Based on the intersection-and-union ratio of the detection frame and the head annotation frame, a positive sample of the head detection frame is screened, and based on the distance between the positive sample of the head detection frame and the center of the head annotation frame, a head weight value of the positive sample of the head detection frame is calculated; Based on the intersection-and-union ratio of the detection frame and the body annotation frame, the body detection frame positive sample is screened, and based on the distance between the body detection frame positive sample and the center of the body annotation frame, the body weight value of the body detection frame positive sample is calculated; wherein, when the head detection frame positive sample and the body detection frame positive sample are the same detection frame, the detection frame is set as a multi-category detection frame positive sample, and its body weight value is updated to be the sum of the corresponding head weight value and the body weight value; The head weight value and the body weight value of the positive sample of the detection frame are added together to obtain the weight value of the positive sample of the detection frame.

6. The method according to claim 5, characterized in that: Filter the positive samples of the head detection box and calculate its head weight value, including: Calculate the distance between the center point of each detection frame and the center point of the head annotation frame, select the first k1 detection frames with the closest distance as candidate head detection frame positive samples and calculate the intersection-over-union ratio between them and the head annotation frame; Based on the intersection-and-union ratios between the positive samples of the candidate head detection frames and the head annotation frame, the mean and variance of the intersection-and-union ratios of the positive samples of the candidate head detection frames are obtained; Taking the mean and the sum of the variance of the intersection-over-union ratio of the candidate head detection frame positive samples as a threshold, further screening the head detection frame positive samples to obtain the final head detection frame positive samples; Set the category value of the detection box that is a positive sample of the head detection box to the head and record its index value; Based on the center point position of the head detection frame positive sample, the weight of the head detection frame positive sample is calculated as the head weight value of the detection frame positive sample.

7. The method according to claim 6, characterized in that: Filter the positive samples of the body detection box and calculate its body weight value, including: Calculate the distance between the center point of each detection frame and the center point of the body annotation frame, select the first k2 detection frames with the closest distance as candidate body detection frame positive samples and calculate the intersection-union ratio between them and the body annotation frame; Based on the intersection-and-union ratios between the positive samples of the candidate body detection frames and the body annotation frames, the mean and variance of the intersection-and-union ratios of the positive samples of the candidate body detection frames are obtained; Taking the mean and the sum of the variance of the intersection-over-union ratio of the candidate body detection frame positive samples as a threshold, further screening the body detection frame positive samples to obtain the final body detection boundary frame positive samples; Set the category value of the detection box that is a positive sample of the body detection box to body and record its index value; Calculating a body weight value of the positive sample of the body detection frame based on the center point position of the positive sample of the body detection frame; When the index value of the body detection frame positive sample is the same as the index value of the head detection frame positive sample, it is determined to be the same detection frame.

8. The method according to claim 7, characterized in that: The head weight value is calculated using the following formula: Among them, W head,i Indicates the head weight of the i-th detection box; dist i,j Represents the distance between the i-th detection box and the j-th head annotation box; (GT w,j ,GT h,j ) represents the width and height of the jth head annotation box; The body weight value is calculated using the following formula: Among them, W body,i represents the body weight of the i-th detection box; dist i,m Represents the distance between the i-th detection bounding box and the m-th body annotation box; (GT w,m ,GT h,m ) represents the width and height of the mth body annotation box.

9. The method according to claim 8, characterized in that: The human body detection model is trained using the following method: Constructing a training data set for the human body detection model; wherein the training data set includes a set of images with human bodies, category labels corresponding to the annotated heads and bodies in each image, bounding boxes, and center point positions; The training data set is loaded, and the human body detection model is trained using a total loss function consisting of a classification loss function, a center point loss function, and a weighted coordinate regression loss function. The model parameters are updated using gradient back propagation, and the training is terminated when the loss function value converges to obtain a trained human body detection model; wherein, based on the weight value of the positive sample of the detection box and the coordinate loss function, the weighted coordinate regression loss of each positive sample of the detection box is obtained.

10. The method according to claim 9, characterized in that: In the head detection branch, when there is no body annotation box that matches the head annotation box, its coordinate regression loss function is: L offsethead =α(|lh * -lh|+|th * -th|+|rh * -rh|+|bh * -bh|) Where α represents a hyperparameter; (lh * ,th * ,rh * ,bh * ) represents the coordinates of the corner points of the head detection box; (lh, th, rh, bh) represents the coordinates of the corner points of the head annotation box; When there is a body annotation box that matches the head annotation box, its coordinate regression loss function is: L offsethead =α(|lh * -lh|+|th * -th|+|rh * -rh|+|bh * -bh|)+β(|lp * -lp|+|tp * -tp|+|rp * -rp|+|bp * -bp|) Where β represents a hyperparameter; (lp * ,tp * ,rp * ,bp * ) represents the coordinates of the corner points of the predicted body box; (lp, tp, rp, bp) represents the coordinates of the corner points of the body annotation box; In the body detection branch, its coordinate regression loss function is: in, Represents the corner point coordinates of the body detection box; (lp, tp, rp, bp) represents the corner point coordinates of the body annotation box.

Citation Information

Cited By

  • Method and system for training image human body detection model

    CN121190836A

  • A method and system for training an image human body detection model

    CN121190836B