A method for detecting small targets of urban pedestrian flow based on convolutional neural network

By introducing shallow information layer and BIAFPN feature fusion network in YOLOX, combined with data enhancement methods, the problem of insufficient detection accuracy of small targets in urban flows is solved, and more efficient detection results are achieved.

CN114821665BActive Publication Date: 2025-07-29ZHEJIANG UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210574388.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-24
Publication Date
2025-07-29
Estimated Expiration
2042-05-24

AI Technical Summary

Technical Problem

The existing target detection algorithms have shortcomings in small target detection accuracy, especially in urban flow detection, which is difficult to effectively improve the detection accuracy.

Method used

The shallow information layer was introduced in the YOLOX technical solution, and the improved feature fusion network BIAFPN was adopted, combined with Mosaic and MixUp data enhancement methods, multi-scale feature fusion and feature processing were performed, feature information was enhanced through the CBAM attention mechanism, and the prediction head without anchor frame was used for detection.

Benefits of technology

The accuracy of urban abortion target detection has been improved, the detection speed and accuracy have been improved, and the problem of unbalanced positive and negative samples has been overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821665B_ABST
    Figure CN114821665B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting small targets of urban pedestrian flow based on a convolutional neural network. The image training dataset with detection frames of small portrait targets marked is subjected to Mosaic data augmentation and MixUp data augmentation. The augmented image training dataset is adjusted to the input picture size and input into the backbone network to obtain four kinds of feature maps of different sizes of the backbone network. The four kinds of feature maps are input into the feature fusion network BIAFPN for feature processing. The fused feature maps are respectively transmitted to their corresponding prediction heads, and after convolution in the classification branch and the regression branch respectively, they are connected along the channel part. Then the connected feature maps are stretched into one dimension, and then the stretched feature maps are connected to obtain the final feature map. The loss is calculated and backpropagation is carried out to update the network parameters, completing the training of the network. The present invention introduces shallow fine-grained features and then uses a feature fusion network to detect targets, which can effectively improve the accuracy of detecting small targets of urban portraits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of deep learning image processing, and particularly relates to a method for detecting small targets of urban pedestrian flow based on a convolutional neural network. Background Art

[0002] Object detection is a basic problem in machine vision, supporting visual tasks such as instance segmentation, object tracking, and action recognition, and has a wide range of applications in fields such as automotive autonomous driving, satellite images, and surveillance. Most of the existing object detection algorithms use the method of anchor boxes, but this often causes the problem of unbalanced positive and negative samples, which is even more difficult for small target detection. How to improve the accuracy of small target detection is still a difficult problem in current detection.

[0003] Currently, the mainstream technical solutions for object detection are one-stage algorithms and two-stage algorithms. The mainstream two-stage algorithms such as the Faster R-CNN series first screen out a large number of candidate regions where objects may exist, and then detect the candidate regions. This algorithm has high accuracy but slow speed and poor real-time image detection effect. The mainstream one-stage algorithms such as the YOLO series directly complete end-to-end prediction, and the model detection speed is faster, but to a certain extent, it reduces the object detection accuracy. Summary of the Invention

[0004] The purpose of this application is to provide a method for detecting small targets of urban pedestrian flow based on a convolutional neural network, introducing a shallow information layer to make multi-scale in the original YOLOX technical solution, and at the same time adopting a better feature fusion method BIAFPN, which overcomes the problem of low accuracy in detecting small targets of urban human figures.

[0005] To achieve the above purpose, the technical solution of this application is as follows:

[0006] A method for detecting small targets of urban pedestrian flow based on a convolutional neural network, including:

[0007] Obtain an image training dataset with small target detection frames of human figures marked, and perform Mosaic data augmentation and MixUp data augmentation on the image training dataset;

[0008] Adjust the augmented image training dataset to the input picture size, and input it into the backbone network CSPDarknet-53 to obtain four sizes of feature maps F1, F2, F3, and F4 output by the dark2 unit, dark3 unit, dark4 unit, and dark5 unit in the backbone network CSPDarknet-53;

[0009] Input the four sizes of feature maps F1, F2, F3, and F4 into the feature fusion network BIAFPN for feature processing to obtain the fused feature map F 12, F 22 , F 32 , F 42 ;

[0010] Input the fused feature maps F 12 , F 22 , F 32 , F 42 into their respective corresponding prediction heads, perform convolution on the classification branch and the regression branch respectively, and then connect them along the channel dimension. Then stretch the connected feature maps into one-dimensional to obtain the stretched feature maps F 13 , F 23 , F 33 , F 43 . Then connect the stretched feature maps F 13 , F 23 , F 33 , F 43 to obtain the final feature map, calculate the loss and perform backpropagation to update the network parameters to complete the training of the network;

[0011] Input the image to be detected into the trained network to obtain the detection result.

[0012] Furthermore, the Mosaic data augmentation includes:

[0013] Take out 4 images and splice them together by means of random scaling, random cropping, and random arrangement;

[0014] The MixUp data augmentation includes: overlaying 2 images together.

[0015] Furthermore, inputting the four-sized feature maps F1, F2, F3, and F4 into the feature fusion network BIAFPN for feature processing to obtain the fused feature maps F 12 , F 22 , F 32 , F 42 , including:

[0016] Input the feature map F1 directly into the feature fusion network BIAFPN. First, from top to bottom, after 1×1 convolution and upsampling, perform adaptive feature fusion with the feature map F2 to obtain the feature map F 21 ; Then continue to perform 1×1 convolution on the feature map F 21 , and after upsampling, perform adaptive feature fusion with the feature map F3 to obtain the feature map F 31 ; Then perform 1×1 convolution and upsampling on the feature map F 31 and perform adaptive feature fusion with the feature map F4 to obtain the feature map F 41 ; Output the feature map F 41 directly to obtain the feature map F42 ; Then, perform bottom-up and cross-scale fusion, and combine F 42 After 1×1 convolution and downsampling, it is combined with the previous F3 and F 31 to obtain the feature map F 32 ; Combine F 32 After 1×1 convolution and downsampling, it is combined with the previous F2 and F 21 to obtain the feature map F 22 ; Combine F 22 After 1×1 convolution and downsampling, it is combined with the previous F1 and F 11 to obtain the feature map F 12 ; After each feature fusion, perform the CBAM attention mechanism to enhance spatial and channel information once.

[0017] Furthermore, the calculation of the loss includes: classification loss, bounding box loss, and object score loss. The classification loss and the object score loss are BCELoss, and the bounding box loss uses IOULoss.

[0018] A method for detecting small targets of urban pedestrian flow based on a convolutional neural network proposed in this application introduces a better feature layer of shallow fine-grained features into the existing YOLOX technical solution, and then uses a better BIAFPN to replace the original PANet to detect targets, which can effectively improve the accuracy of detecting small targets of urban human figures. Description of the Drawings

[0019] Figure 1 is the flowchart of the method for detecting small targets of urban pedestrian flow based on neural network in this application.

[0020] Figure 2 is the network model diagram of the method for detecting small targets of urban pedestrian flow based on neural network in this application. Detailed Embodiments

[0021] In order to make the purpose, technical solution and advantages of this application clearer, the following further details this application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain this application and are not used to limit this application.

[0022] The method for detecting small targets of urban pedestrian flow based on neural network in this application mainly includes the following steps: First, perform data augmentation on the image, and then start training in batches. In each batch, pass the image through a convolutional neural network to obtain the feature maps F1, F2, F3, F4, and then pass the obtained feature maps through BIAFPN for feature fusion to obtain F 12 , F 22 , F 32 , F 42, the feature map is put into the prediction head for classification and regression to obtain prediction values, which are compared with the ground truth values of the images to calculate the loss. After each batch of training, backpropagation is performed to reduce the loss, and at the same time, the network parameters are updated to complete the training of the network.

[0023] In one embodiment, as Figure 1 shown, a method for detecting small targets of urban pedestrian flow based on a neural network is proposed, including:

[0024] Step S1: Obtain an image training dataset with annotated small target detection frames for human figures, and perform Mosaic data augmentation and MixUp data augmentation on the image training dataset.

[0025] In this embodiment, Mosaic data augmentation and MixUp data augmentation are performed on the training dataset. Among them, Mosaic data augmentation means taking out 4 images and splicing them in a way of random scaling, random cropping, and random arrangement. The advantage is that it enriches the background and small targets of the detected objects, and when calculating, the data of 4 images will be calculated at one time, without much overhead, and a single GPU can achieve good results. MixUp data augmentation means superimposing 2 images together, which can reduce the memory of wrong labels to enhance robustness.

[0026] Step S2: Adjust the augmented image training dataset to the input image size and input it into the backbone network CSPDarknet-53 to obtain four sizes of feature maps F1, F2, F3, and F4 output by the dark2 unit, dark3 unit, dark4 unit, and dark5 unit in the backbone network CSPDarknet-53.

[0027] As Figure 2 shown, this application uses CSPDarknet-53 as the backbone network for feature extraction. The CSPDarknet-53 used is pre-added with pre-trained weights trained on COCO, and is trained in batches. During the training process, the batch size is 16 (that is, 16 images are processed in each batch), the learning rate starts from 0.0025, and the method of learning rate warm-up is not adopted, and the cosine annealing method is used to update the learning rate.

[0028] Since the original image is large, this application scales the original image, scales it proportionally according to the long side to 640×640, and fills the part where the short side is less than 640 with 0. The scaled image is input into the backbone network CSPDarknet-53. After a series of convolutional operations, four sizes of feature maps F1, F2, F3, and F4 of 20×20, 40×40, 80×80, and 160×160 are output successively.

[0029] The size of the feature map is determined by the backbone network CSPDarknet-53, which will not be elaborated here. It should be noted that in YOLOX, usually only the features output by dark3, dark4, and dark5 are used for multi-scale fusion operations. In this embodiment, the features output by dark2, dark3, dark4, and dark5 are used for multi-scale fusion operations, and feature maps of four sizes are output, which can incorporate shallow fine-grained information and is beneficial to small target detection. Finally, there is also one more prediction head correspondingly, achieving a better detection effect.

[0030] Step S3: Input the four-sized feature maps F1, F2, F3, and F4 into the feature fusion network BIAFPN for feature processing to obtain the fused feature map F 12 、F 22 、F 32 、F 42 。

[0031] Directly input the feature map F1 ( Figure 2 the feature map output by Dark5 in Figure 2 ) into the feature fusion network BIAFPN. First, from top to bottom, after a 1×1 convolution and upsampling, it performs adaptive feature fusion with the feature map F2 ( Figure 2 the feature map output by Dark4 in 21 , and so on) to obtain the feature map F 21 . Continue to perform a 1×1 convolution on the feature map F 21 , and after upsampling, perform adaptive feature fusion with the feature map F3 to obtain the feature map F 31 . Then, perform a 1×1 convolution and upsampling on the feature map F 31 and perform adaptive feature fusion with the feature map F4 to obtain the feature map F 41 . Directly output the feature map F 41 41 to obtain the feature map F 41 42 42 . Then perform bottom-up and cross-scale fusion. After performing a 1×1 convolution and downsampling on F 42 42 , fuse it with the previous F3 and F 42 31 to obtain the feature map F 31 32 32 . After performing a 1×1 convolution and downsampling on F 32 32 , fuse it with the previous F2 and F 32 21 to obtain the feature map F 21 22 22 . After performing a 1×1 convolution and downsampling on F 22 22 , fuse it with the previous F1 and F 22 11 to obtain the feature map F 11 12 12 . After each feature fusion, a CBAM attention mechanism is performed to enhance spatial and channel information.

[0032] Specifically, the 20×20 feature map F1 is directly input into the top-down feature pyramid network BIAFPN. First, it is convolved by 1×1 to the same channel, and then upsampled to a 40×40 feature map, which is then adaptively feature fused with the feature map F2 through SUM. Then, its spatial and channel information is enhanced through the CBAM attention mechanism to obtain F 21 (40×40). Then continue to process F 21 by convolving it with 1×1, upsampling it to an 80×80 feature map, and then adaptively feature fusing it with the feature map F3 through SUM. Then, its spatial and channel information is enhanced through the CBAM attention mechanism to obtain F 31 (80×80). Next, process F 31 by convolving it with 1×1, upsampling it to a 160×160 feature map, and then adaptively feature fusing it with the feature map F4 through SUM. Then, its spatial and channel information is enhanced through the CBAM attention mechanism to obtain F 41 (160×160). Then directly output F 41 to obtain the feature map F 42 (160×160). After that is the bottom-up and cross-scale fusion. First, convert the number of channels of F 41 to be the same through 1×1 convolution, and then downsample it to 80×80, and then adaptively feature fuse it with the feature map F3 and F 31 through SUM to obtain the feature map F 32 (80×80). Next, convert the number of channels of F 32 to be the same through 1×1 convolution, and then downsample it to 40×40, and then adaptively feature fuse it with the feature map F2 and F 21 through SUM to obtain the feature map F 22 (40×40). Then convert the number of channels of F 22 to be the same through 1×1 convolution, and then downsample it to 20×20, and then adaptively feature fuse it with the feature map F1 and F 11 through SUM to obtain the feature map F 12 (20×20). At this point, all four sizes of output feature maps are obtained: F 12 (20×20), F 22 (40×40), F 32 (80×80), and F 42 (160×160).

[0033] It should be noted that Figure 2There is also a DWSConv between the adaptive feature fusion SUM and the CBAM attention mechanism, that is, the DW convolution and the PW convolution, which will not be elaborated here. In this embodiment, BIAFPN is used to replace the original PANet. The BIAFPN feature fusion adds cross-scale fusion and the CBAM attention mechanism (enhancing the features in terms of space and channels) on the basis of the original bidirectional fusion, serving as a better method for fusing features.

[0034] Step S4: Input the fused feature maps F 12 、F 22 、F 32 、F 42 into their respective corresponding prediction heads, respectively perform convolution on the classification branch and the regression branch and then connect them along the channel part, and then stretch the connected feature map into one dimension to obtain the stretched feature maps F 13 、F 23 、F 33 、F 43 . Then connect the stretched feature maps F 13 、F 23 、F 33 、F 43 to obtain the final feature map, calculate the loss and perform backpropagation to update the network parameters to complete the training of the network.

[0035] In this embodiment, after convolution on the classification branch and the regression branch of the prediction head Prediction and then connecting them along the channel part, the generated new feature maps are respectively four tensors of size {W×H×[(cls + reg + obj)]×N}, where W×H is the feature map size, cls is the detection category, reg is the predicted bounding box, obj is the prediction of the objectness score, and N is the number of predicted anchor boxes. Then multiply W by H and stretch the spatial dimension into one dimension to obtain the feature maps F 13 、F 23 、F 33 、F 43 . Then connect F 13 、F 23 、F 33 、F 43 along W*H to obtain the final feature map F.

[0036] Finally, calculate the classification loss, the bounding box loss, and the object score loss, and perform backpropagation to reduce the loss and update the network parameters at the same time.

[0037] Specifically, F 12 、F 22 、F 32 、F 42After performing convolutions on the classification branch and the regression branch respectively, each feature map generates 3 new feature maps F cls ∈{N×W×H×cls}, F obj ∈{N×W×H×1}, F reh ∈{N×W×H×4}. First, connect them along the channel dimension. The resulting new feature maps are four tensors of size {N×W×H×[(cls + reg + obj)]}, where W, H ∈ {20, 40, 80, 160}. Then multiply W and H, and stretch the spatial dimension into one dimension to obtain four tensors of size {N×(cls + reg + obj)×(W×H)}. Then connect F 13 、F 23 、F 33 、F 43 along W*H to obtain the final feature map F ∈ {N×(cls + reg + obj)×34000}.

[0038] where cls is the number of classes in the dataset, reg predicts the bounding box including the predicted upper left corner point (x1, y1) and the lower right corner point (x2, y2), N is the preset number of anchor boxes, which is 1 in this embodiment.

[0039] In this embodiment, the prediction head adopts the decoupled head method, which separates the classification branch and the regression branch for convolution operations, so as to achieve better detection effects. And the prediction at each position is reduced from 3 to 1 through connection operations, and the anchor-free method is adopted to avoid the problem of unbalanced positive and negative samples.

[0040] Since the output feature values cannot be directly used for loss calculation, regression needs to be performed first to obtain the actual predicted values. Calculate the classification loss, bounding box loss, and object score loss for the feature map F according to the following formulas. The classification loss and the object score loss are BCELoss, and the bounding box loss adopts IOULoss. The specific formulas are as follows:

[0041] BCELoss = -(ylog(p(x))+(1 - y)log(1 - p(x)))

[0042]

[0043] It should be noted that the grid setting in this application is on the finally obtained feature map, which is an abstract concept. The purpose is to facilitate the calculation of bounding box regression. For feature maps of 20*20, 40*40, 80*80, and 160*160, they have 20×20, 40×40, 80×80, and 160×160 grids respectively. Dividing the feature map into multiple grids is a relatively mature technology in this field and will not be elaborated here. At the same time, the CBAM attention mechanism is also a relatively mature technology in this field and will not be elaborated here either.

[0044] 1. Calculate the classification loss and the target score loss, using the Binary CrossEntropy Loss function for calculation:

[0045] BCELoss = -(ylog(p(x))+(1 - y)log(1 - p(x)))

[0046] where y represents whether it is the target, with values 1 or 0, and p(x) is the predicted target score.

[0047] 2. Calculate the bounding box loss. Calculate the IOU (Intersection of Union) between the predicted box information and the ground truth box information calculated according to the label. The IOU is the intersection - over - union ratio of the predicted box and the ground truth box. After NMS post - processing, the predicted boxes with high IOU values are obtained:

[0048]

[0049] where is the ground truth, B is the predicted box, is the area of the intersection of the ground truth and the predicted box, is the area of the union of the ground truth and the predicted box. The lower the IOULoss value, the more accurate the prediction.

[0050] It should be noted that the calculations of the classification loss, the target score loss, and the bounding box loss are already relatively mature technologies in this field and will not be elaborated here.

[0051] Thus, the loss between the predicted value and the true value is obtained. Before the end of each batch, backpropagation is performed to reduce the loss. At the same time, the network parameters are updated, and the training of the next batch starts until all batches of training data have been trained. Finally, the trained weights are obtained, and all updated parameters will be saved in the Outputs weight file.

[0052] Step S5: Input the image to be detected into the trained network to obtain the detection result.

[0053] This application also scales the image to be detected to a size of 640×640 and inputs it into the network. Four - sized feature maps are output through the CSPDarknet - 53 backbone network. After regression of the feature values, the predicted values are obtained, including the class cls, the bounding box reg, and the target score obj, to obtain a final prediction result.

[0054] This application also uses the SimOTA positive and negative sample allocation strategy. First, the prediction box is initially screened, and only those prediction boxes whose center points are within the groundtruth and within a square with a side length of 5 are retained. After the initial screening is completed, the bounding box loss of the prediction box and the groundtruth is calculated, and the classification loss is calculated using binary cross entropy, and the cost matrix is calculated:

[0055]

[0056] Represents the cost relationship between each true box and each feature point. The top k predicted boxes with the smallest loss relative to the groundtruth are fixed as positive samples, and the rest are negative samples, thus avoiding additional hyperparameters.

[0057] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A method for detecting small targets of urban pedestrian flow based on a convolutional neural network, characterized in that, The small target detection method of urban pedestrian flow based on convolutional neural network includes: Obtain an image training dataset with labeled small target detection frames of human figures, and perform Mosaic data augmentation and MixUp data augmentation on the image training dataset; Adjust the augmented image training dataset to the input picture size, and input it into the backbone network CSPDarknet-53 to obtain four types of feature maps F1, F2, F3, and F4 output by the dark2 unit, dark3 unit, dark4 unit, and dark5 unit in the backbone network CSPDarknet-53; Input the feature maps F1, F2, F3, and F4 of four sizes into the feature fusion network BIAFPN for feature processing to obtain the fused feature maps F 12 , F 22 , F 32 , F 42 ; The fused feature map F 12 , F 22 , F 32 , F 42 are respectively input into their corresponding prediction heads, and after convolution in the classification branch and the regression branch, they are connected along the channel part, and then the connected feature maps are stretched into one dimension to obtain the stretched feature map F 13 , F 23 , F 33 , F 43 . Then, the stretched feature maps F 13 , F 23 , F 33 , F 43 are connected to obtain the final feature map, calculate the loss and perform backpropagation to update the network parameters to complete the training of the network; Input the image to be detected into the trained network to obtain the detection result; Among them, inputting the feature maps F1, F2, F3, and F4 of four sizes into the feature fusion network BIAFPN for feature processing to obtain the fused feature maps F 12 , F 22 , F 32 , F 42 , including: The feature map F1 is directly input into the feature fusion network BIAFPN. First, it is from top to bottom. After passing through a 1×1 convolution and upsampling, it performs adaptive feature fusion with the feature map F2 to obtain the feature map F 21 ; Then continue to process the feature map F 21 After passing through a 1×1 convolution and upsampling, it performs adaptive feature fusion with the feature map F3 to obtain the feature map F 31 ; Then again, the feature map F 31 After passing through a 1×1 convolution and upsampling, it performs adaptive feature fusion with the feature map F4 to obtain the feature map F 41 ; The feature map F 41 is directly output to obtain the feature map F 42 ; Then perform bottom-up and cross-scale fusion. After passing the F 42 through a 1×1 convolution and downsampling, it is fused with the previous F3 and F 31 to obtain the feature map F 32 ; After passing the F 32 through a 1×1 convolution and downsampling, it is fused with the previous F2 and F 21 to obtain the feature map F 22 ; After passing the F 22 through a 1×1 convolution and downsampling, it is fused with the previous F1 and F 11 to obtain the feature map F 12 ; After each feature fusion, a CBAM attention mechanism is performed to enhance spatial and channel information.

2. The method for detecting small targets of urban pedestrian flow based on convolutional neural network according to claim 1, characterized in that, The Mosaic data augmentation includes: Take out 4 images and splice them together by means of random scaling, random cropping, and random arrangement; The MixUp data augmentation includes: overlaying 2 images together.

3. The method for detecting small targets of urban pedestrian flow based on a convolutional neural network according to claim 1, characterized in that, The calculation of the loss includes: classification loss, bounding box loss, and target score loss. The classification loss and the target score loss are BCELoss, and the bounding box loss adopts IOULoss.

Citation Information

Patent Citations

  • Rotating target detection method and device based on convolutional neural network

    CN113298169A