A small target detection method based on a novel semi-fragmented FPN+PAN feature fusion network
By adopting a new semi-fortune FPN+PAN feature fusion network in the small object detection method, the shallow feature map scale confusion problem is solved, and the detection accuracy and model generalization are improved.
Patent Information
- Application Number
- CN202310591616.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-05-24
AI Technical Summary
When existing small object detection methods process images containing complex backgrounds and multi-scale objects, they are prone to scale confusion problems in shallow feature maps, resulting in a decrease in detection accuracy.
The small object detection method based on the new semi-fault FPN+PAN feature fusion network is adopted to enhance the semantic information of the small object by adding new paths, and cut off the information transmission process of the deeper FPN, weakening the scale confusion problem of shallow feature maps.
It effectively improves the detection accuracy of small targets, enhances the generalization of network models, and can better process images containing complex backgrounds and multi-scale targets.
Smart Images

Figure CN116543290B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of small target detection, and in particular to a small target detection method based on a novel semi-fault FPN+PAN feature fusion network. Background Art
[0002] Object detection has always been a key research direction in the field of computer vision. With the emergence of deep neural networks, object detection based on deep networks has made significant progress in detection efficiency and detection accuracy. However, there are still many areas that need to be improved for small target detection. Before the popularity of deep learning methods, for targets of different scales, it was common to construct image pyramids of different resolutions from the original image, and then use detectors with fixed input resolutions to detect targets at each layer of the pyramid in order to detect small targets at the bottom of the pyramid. However, for some images with complex backgrounds and multi-scale targets, feature map scale confusion problems will occur, which reduces the detection accuracy of small targets.
[0003] In recent years, many achievements have been made in target detection using deep learning methods. Among them, domestic and foreign scholars have proposed a series of solutions for the problem of small target detection. The general small target detection solutions mainly include: using feature pyramids and multi-scale sliding windows, such as FPN, PAN, FPN+PAN; using data enhancement methods, such as oversampling and copying and pasting small targets, Mosaic, and GAN. Among them, the FPN, PAN, and FPN+PAN feature fusion networks enhance the features of large targets, but there are scale confusion problems in the shallow feature maps of the network. Typical data enhancement methods such as oversampling, Mosaic, and GAN violently increase the features of small targets, but they are limited by the shortcomings of the small targets themselves with few pixels and cannot solve the scale confusion problem in the shallow feature maps. Summary of the invention
[0004] In view of the shortcomings of the prior art, the present invention provides a small target detection method based on a new semi-fault FPN+PAN feature fusion network. On the basis of the FPN+PAN feature fusion network, a new path is added to enhance the semantic information of small targets, and the information transmission process of the deeper FPN layer is cut off, which is called the semi-fault FPN+PAN feature fusion network. This method can effectively weaken the scale confusion problem of the shallow feature map, thereby improving the detection effect of small targets. At the same time, the present invention also designs a New CSP-Darknet53 backbone network with deformable convolution and C3-Res modules, which can adaptively extract features according to the target shape. The entire network consists of a backbone network (multi-level feature extraction network), a feature fusion network (a new semi-fault FPN+PAN feature fusion network) and a detection network. First, the backbone network is used to extract features from the input image to obtain feature maps of different sizes; the multi-size feature maps are subjected to a feature information propagation process that will be interrupted through the semi-fault FPN+PAN network; then the detection network is used for multi-scale prediction, and the K-means++ method is used to generate target proposal boxes for classification and regression tasks. The present invention invents a new feature fusion method, which can be directly applied to a detector using a feature pyramid structure and has better detection effect and robustness for small target detection.
[0005] A small target detection method based on a novel semi-fault FPN+PAN feature fusion network includes the following steps:
[0006] Step 1: Use the improved feature extraction network New CSP-Darknet53 with Yolov5 backbone network to extract multi-scale target images containing small targets;
[0007] The backbone network is composed of 5 groups of feature extraction modules connected in series in sequence; the first group of feature extraction modules is composed of Focus data enhancement module and C3_Res module; the second, third and fourth groups of feature extraction modules are all composed of Dconv module and C3_Res module; the fifth group of feature extraction modules is composed of Dconv module and SPP module;
[0008] The C3_Res module structure is a convolutional module in which the C3 module in Yolov5 is connected to a residual connection channel at the head and the end;
[0009] The Dconv module is a module composed of a deformable convolution module, a BN link, and a SiLU activation function. This module changes the shape of the convolution kernel by adding an offset to adaptively extract target features. The deformable convolution module convolves a picture of size C*H*W, reduces the picture size and increases the number of channels to obtain a feature map of size 2C*H / 2*W / 2, and its formula is:
[0010]
[0011] In the formula, w represents the corresponding weight of the sampling value; is a regular grid, p 0 is the convolution center, p n for Medium element, Δp n is the offset, {Δp n |n=1,....,N}, where The convolution sampling position depends on the irregular offset p n +Δp n .
[0012] The BN step uses the general batch normalization operation BatchNorm2d to make it have the statistical characteristics of describing global data, while ensuring that the network training process will not have the problems of gradient explosion and gradient disappearance;
[0013] The SiLU activation function outputs a feature map of 2C*H / 2*W / 2. The formula of the SiLU activation function is as follows:
[0014]
[0015] Step 2: Extract 5 groups of feature maps T1, T2, T3, T4, and T5 of different depths whose sizes are halved and whose channels are doubled in sequence, generated by the five feature extraction modules of the backbone network, and input them into the feature fusion network; the feature fusion network first fuses the features of the shallow feature maps T3, T2, and T1 in the form of top-down feature fusion in the FPN network, and obtains 3 groups of new feature maps L1, L2, and L3 from T1, T2, and T3 respectively; secondly, obtains 2 groups of new feature maps L4 and L5 from T4 and T5 respectively; finally, 5 groups of new feature maps L1, L2, L3, L4, and L5 of different depths whose sizes are halved and whose channels are doubled in sequence are obtained;
[0016] Step 2.1: Extract the feature map T5, input it into the C3 module and the convolution module Conv2d in sequence, and output the new feature map L5;
[0017] Step 2.2: Upsample the feature map L5 and perform a Concat operation with the feature map T4. The Concat operation is an operation that directly merges two different feature maps in the channel dimension. After merging the channels, they are sequentially input into the C3 module and the convolution module Conv2d to obtain a new feature map L4.
[0018] Step 2.3: Extract the feature maps T3 and T4 output by Dconv of the fourth feature extraction module of the backbone network, input them into the C3 module and the convolution module Conv2d operation in turn, and output feature maps L3 and L4; perform upsampling operation on the feature maps L3 and L4 and perform Concat operation with the T3 feature map, merge the channels, and obtain the new feature map L3.
[0019] Step 2.4: Perform C3 module and convolution module Conv2d operations on L3, perform upsampling operation and Concat operation with T2 feature map, merge channels, and obtain the new feature map L2.
[0020] Step 2.5: Perform C3 module and convolution module Conv2d operations on L2, perform upsampling operation and Concat operation with feature map T1, merge channels, and obtain a new feature map L1.
[0021] Step 3: Establish a feature transmission channel between the L1, L2, L3, L4, and L5 feature maps through the PAN feature pyramid structure to ensure the effective transmission of context information between different feature layers, and output new feature maps Z1, Z2, Z3, Z4, and Z5 for detection;
[0022] Step 3.1: Input feature map L1 into the C3 module with half the number of channels and unchanged size, and output feature map Z1;
[0023] Step 3.2: Input Z1 into the Conv2d module with the same number of channels and half the size, and perform Concat fusion with the feature map L2, and then perform the C3 module with the same number of channels and size, and output the feature map Z2;
[0024] Step 3.3: Perform a Conv2d module on Z2 with the same number of channels and half the size, and concat it with the feature map L3, and then perform a C3 module with the same number of channels and size, and output the feature map Z3;
[0025] Step 3.4: Perform a Conv2d module on Z3 with the same number of channels and half the size, and perform Concat fusion with the feature map L4, and then perform a C3 module with the same number of channels and size, and output the feature map Z4;
[0026] Step 3.5: Perform a Conv2d module on Z3 with the same number of channels and half the size, and concat it with the feature map L5. Then perform a C3 module with the same number of channels and size, and output the feature map Z5.
[0027] Step 4: Use the K-means++ algorithm to obtain the prior frame, cluster the target frame scales of the objects in the Citypersons dataset, and obtain the prior frames of 5 scales through the k-means clustering algorithm and the genetic mutation algorithm respectively;
[0028] Step 4.1: Perform k-mens clustering algorithm on the labeled information of the data set to obtain the initial prior frame;
[0029] Step 4.1.1: Randomly select 5 true boxes in all labels as the centers of the cluster, where the true box is the target box with the target position and size marked in the dataset label information;
[0030] Step 4.1.2: Calculate the distance 1-IOU between each true box and each cluster, where IOU is the ratio of the intersection area to the union area of two true boxes;
[0031] Step 4.1.3: Calculate the nearest cluster center for each ground-truth box and assign it to the cluster closest to it;
[0032] Step 4.1.4: Recalculate the cluster center based on the ground-truth box in each cluster;
[0033] Step 4.1.5: Repeat steps 4.1.3-4.1.4 until the center of each cluster no longer changes.
[0034] Step 4.2: The wh information, i.e., height and width information, of the real box in the initial prior box is continuously adjusted, evaluated, changed, re-evaluated, and improved, and M variations are screened to obtain the final prior box result that conforms to the Citypersons dataset;
[0035] Step 4.2.1: Read the wh of each image in the Citypersons dataset and the wh of all ground truth boxes, and scale the maximum value of wh in each image to the specified image size;
[0036] Step 4.2.2: Change the real frame from relative coordinates to absolute coordinates, that is, multiply it by the scaled wh;
[0037] Step 4.2.3: Filter the true boxes and retain the true boxes whose wh is greater than or equal to two pixels;
[0038] Step 4.2.4: Use k-means clustering method to get n anchors;
[0039] Step 4.2.5: Use genetic algorithm to randomly mutate the anchors’ wh. If the effect is better after mutation, assign the mutated result to the anchors. If the effect is worse after mutation, skip it. The default mutation is M times.
[0040] The effect after mutation becomes better means: the fitness value calculated by the anchor_fitness method is evaluated, and the larger the fitness value, the better;
[0041] Step 4.2.6: Sort the final mutated anchors by area and return them.
[0042] Step 5: Finally, the fused feature maps Z1, Z2, Z3, Z4, and Z5 are output respectively and used to detect the candidate box information and the probability that the candidate box belongs to a certain category. The detection results are screened using the non-maximum suppression method, that is, the candidate boxes in the output detection results are sorted according to the probability values of the candidate boxes belonging to the category in the detection results, and the candidate box with the largest probability is selected as the final result to complete the small target detection method.
[0043] The beneficial effects of adopting the above technical solution are:
[0044] The present invention provides a small target detection method based on a novel semi-fault FPN+PAN feature fusion network. Traditional small target detection methods are generally based on image pyramids FPN, PAN, and FPN+PAN. Such methods have serious feature confusion problems in some feature maps. Or typical data enhancement methods such as oversampling, Mosaic, GAN, etc. Such methods violently increase the features of small targets, but are limited by the disadvantage that the small targets themselves have few pixels, and cannot solve the scale confusion problem in shallow feature maps. The present invention invents a method based on a novel semi-fault FPN+PAN feature fusion network, which weakens the scale confusion problem of shallow feature maps in target detection involving multi-morphology, multi-scale targets and complex backgrounds, improves the accuracy of small target detection and improves the generalization of the network model. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a diagram of the feature extraction network structure in an example of the present invention.
[0046] Figure 2 This is a structural diagram of the new semi-fault FPN+PAN feature fusion network in an example of the present invention.
[0047] Figure 3 This is the overall structure diagram of the network model in the example of the present invention.
[0048] Figure 4 This is a structural diagram of the C3_Res module in an example of the present invention.
[0049] Figure 5 This is a structural diagram of the Dconv module in an example of the present invention.
[0050] Figure 6Schematic diagram of the Citypersons dataset used for pedestrian small target detection in the example of the present invention.
[0051] Figure 7 This is the detection effect of small targets in an image using the improved algorithm model in the example of the present invention.
[0052] Figure 8 The comparison results of detection accuracy between the present invention example and the Yolov5s model and the continuous FPN+PAN model. DETAILED DESCRIPTION
[0053] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0054] The existing target detection algorithms have mostly improved the feature fusion network by violently stacking the number of feature fusions to improve the introduction of small target features, but this can easily cause scale confusion. In addition, the existing target detection algorithms are not capable of extracting features. For irregularly shaped targets such as pedestrians, most of the extracted features are weak, which leads to a decrease in the accuracy of small target detection. Based on this, Figure 1 , Figure 2 , Figure 3 As shown, a small target detection method based on a novel semi-fault FPN+PAN feature fusion network includes the following steps:
[0055] A small target detection method based on a novel semi-fault FPN+PAN feature fusion network includes the following steps:
[0056] Step 1: Use the improved feature extraction network New CSP-Darknet53 with Yolov5 backbone network to extract multi-scale target images containing small targets;
[0057] The backbone network is composed of 5 groups of feature extraction modules connected in series in sequence; the first group of feature extraction modules is composed of Focus data enhancement module and C3_Res module; the second, third and fourth groups of feature extraction modules are all composed of Dconv module and C3_Res module; the fifth group of feature extraction modules is composed of Dconv module and SPP module;
[0058] In this embodiment, the Dconv module is composed of the convolution module Conv2d, that is, conv2d is a part of the Dconv module, the C3_Res module is obtained by improving the C3 module in Yolov5, and the Focus data enhancement module and the SPP module are the original basic modules in Yolov5;
[0059] The C3_Res module structure is as follows Figure 4As shown in the figure, the C3 module in Yolov5 is connected to a convolution module with a residual connection channel at the head and the end, which can improve the extension potential of the network and help the network to extract features more effectively. Among them, the C3 module is a simplified version of the Bottleneck CSP module. Except for the Bottleneck part, there are only 3 convolutions, which can greatly reduce parameters, high precision and small amount of calculation. Among them, the Bottleneck CSP module is the basic module in Yolov4. It is a feature extraction module based on the residual structure network idea and composed of 1*1 convolution and 3*3 convolution. The C3 module splits the original feature map into two parts. One part is continuously extracted and filtered by the Bottleneck module, and the other part is obtained by the Concat operation. A new feature map. Although the C3 module has a residual module inside. In order to improve the screening effect, this method adds a residual connection channel to the C3 module to retain the features of the original feature map and ensure the normal handover and learning of the deep learning network.
[0060] The Dconv module is as Figure 5 As shown in the figure, it is a module composed of a deformable convolution module, a BN link, and a SiLU activation function, which can adaptively extract features and improve feature extraction capabilities. This module changes the shape of the convolution kernel by adding an offset to adaptively extract target features; wherein the deformable convolution module convolves a picture of size C*H*W, reduces the picture size and increases the number of channels, and obtains a feature map of size 2C*H / 2*W / 2, and its formula is:
[0061]
[0062] In the formula, w represents the corresponding weight of the sampling value; is a regular grid, which determines the receptive field and expansion size; p 0 is the convolution center, p n for Medium element, Δp n is the offset, {Δp n |n=1,....,N}, where The convolution sampling position depends on the irregular offset p n +Δp n .
[0063] The BN step uses the general batch normalization operation BatchNorm2d to make it have the statistical characteristics of describing global data, while ensuring that the network training process will not have the problems of gradient explosion and gradient disappearance;
[0064] The SiLU activation function has a smoother curve when it is close to zero, and outputs a feature map of 2C*H / 2*W / 2. The formula of the SiLU activation function is as follows:
[0065]
[0066] Step 2: Extract 5 groups of feature maps T1, T2, T3, T4, and T5 of different depths whose sizes are halved and channels are doubled in sequence, generated by the five feature extraction modules of the backbone network, and input them into the feature fusion network; the feature fusion network first fuses the shallow feature maps T3, T2, and T1 in the form of top-down feature fusion in the FPN network without fusing T3 and T4, and obtains 3 groups of new feature maps L1, L2, and L3 from T1, T2, and T3 respectively; secondly, obtain 2 groups of new feature maps L4 and L5 from T4 and T5 respectively; finally, obtain 5 groups of new feature maps L1, L2, L3, L4, and L5 of different depths whose sizes are halved and channels are doubled in sequence;
[0067] Step 2.1: Extract the feature map T5, input it into the C3 module and the convolution module Conv2d in sequence, integrate the features, and output the new feature map L5;
[0068] Step 2.2: Upsample the feature map L5 and perform a Concat operation with the feature map T4. The Concat operation is an operation that directly merges two different feature maps in the channel dimension. After merging the channels, they are sequentially input into the C3 module and the convolution module Conv2d to obtain a new feature map L4.
[0069] Step 2.3: Extract the feature maps T3 and T4 output by Dconv of the fourth feature extraction module of the backbone network, and input them into the C3 module and the convolution module Conv2d operation in sequence to integrate the features and output feature maps L3 and L4; perform upsampling operations on the feature maps L3 and L4 and perform Concat operations with the T3 feature map to merge the channels and obtain a new feature map L3.
[0070] Step 2.4: Perform C3 module and convolution module Conv2d operations on L3 to integrate features, perform upsampling operation and Concat operation with T2 feature map, merge channels, and obtain a new feature map L2.
[0071] Step 2.5: Perform C3 module and convolution module Conv2d operations on L2 to integrate features, perform upsampling operation and Concat operation with feature map T1, merge channels, and obtain a new feature map L1.
[0072] Step 3: Establish a feature transmission channel between the L1, L2, L3, L4, and L5 feature maps through the PAN feature pyramid structure to ensure the effective transmission of context information between different feature layers, and output new feature maps Z1, Z2, Z3, Z4, and Z5 for detection;
[0073] Step 3.1: Input feature map L1 into the C3 module with half the number of channels and unchanged size, and output feature map Z1;
[0074] Step 3.2: The Conv2d module with the same number of channels and half the size is input to Z1 to integrate the features and concat them with the feature map L2. Then the C3 module with the same number of channels and size is input to output the feature map Z2.
[0075] Step 3.3: Perform a Conv2d module on Z2 with the same number of channels and half the size to integrate the features and perform Concat fusion with the feature map L3. Then perform a C3 module with the same number of channels and size to output the feature map Z3.
[0076] Step 3.4: Perform a Conv2d module on Z3 with the same number of channels and half the size to integrate the features and perform Concat fusion with the feature map L4. Then perform a C3 module with the same number of channels and size to output the feature map Z4.
[0077] Step 3.5: Perform a Conv2d module on Z3 with the same number of channels and half the size to integrate the features, and perform Concat fusion with the feature map L5. Then perform a C3 module with the same number of channels and size to output the feature map Z5.
[0078] Step 4: Use the K-means++ algorithm to obtain the prior frame, cluster the target frame scales of the objects in the Citypersons dataset, and obtain the prior frames of 5 scales through the k-means clustering algorithm and the genetic mutation algorithm respectively;
[0079] Step 4.1: Perform k-mens clustering algorithm on the labeled information of the data set to obtain the initial prior frame;
[0080] Step 4.1.1: Randomly select 5 true boxes in all labels as the centers of the cluster, where the true box is the target box with the target position and size marked in the dataset label information;
[0081] Step 4.1.2: Calculate the distance 1-IOU between each true box and each cluster, where IOU is the ratio of the intersection area to the union area of two true boxes;
[0082] Step 4.1.3: Calculate the nearest cluster center for each ground-truth box and assign it to the cluster closest to it;
[0083] Step 4.1.4: Recalculate the cluster center based on the true box in each cluster; in this embodiment, the cluster position is updated by calculating the median value from each true box to the cluster center;
[0084] Step 4.1.5: Repeat steps 4.1.3-4.1.4 until the center of each cluster no longer changes.
[0085] Step 4.2: The wh information, i.e., height and width information, of the real box in the initial prior box is continuously adjusted, evaluated, changed, re-evaluated, and improved, and 1,000 variations are screened to obtain the final prior box result that conforms to the Citypersons dataset;
[0086] Step 4.2.1: Read the wh of each image and the wh of all ground truth boxes in the Citypersons dataset, and scale the maximum value of wh in each image to the specified image size; since the read ground truth boxes are relative to the internal coordinates of the image, no changes are required
[0087] Step 4.2.2: Change the real frame from relative coordinates to absolute coordinates, that is, multiply it by the scaled wh;
[0088] Step 4.2.3: Filter the true boxes and retain the true boxes whose wh is greater than or equal to two pixels;
[0089] Step 4.2.4: Use k-means clustering method to get n anchors;
[0090] Step 4.2.5: Use genetic algorithm to randomly mutate the anchors’ wh. If the mutation effect is better, assign the mutation result to the anchors. If the mutation effect is worse, skip it. The default mutation is 1000 times.
[0091] The effect after mutation becomes better means: the fitness value calculated by the anchor_fitness method is evaluated, and the larger the fitness value, the better;
[0092] Step 4.2.6: Sort the final mutated anchors by area and return them.
[0093] Step 5: Finally, the fused feature maps Z1, Z2, Z3, Z4, and Z5 are output respectively and used to detect the candidate box information and the probability that the candidate box belongs to a certain category. The detection results are screened using the non-maximum suppression method, that is, the candidate boxes in the output detection results are sorted according to the probability values of the candidate boxes belonging to the category in the detection results, and the candidate box with the largest probability is selected as the final result to complete the small target detection method.
[0094] The present invention is based on the small target detection method of the novel semi-fault FPN+PAN feature fusion network, as shown in the following example. Figure 6As shown in the figure, the Citypersons dataset is used for implementation. Citypersons is very different from previous datasets. It contains photos of crowded people in 27 cities, 3 seasons, and multiple weather conditions. It contains 2,975 training set photos, 500 validation set photos, and 1,575 photos for testing. Figure 6 As shown in the figure, the background of the Citypersons dataset is complex and the object scales are diverse. Figure 7 As shown in the figure, the model proposed by the present invention can well realize the detection of small targets. Figure 8 As shown in the figure, the detection accuracy comparison experiment of the original model Yolov5s and the continuous layer model with the model of the present invention is carried out, wherein the accuracy indicators AP, AP 50 , A.P. 75 , A.P. S , A.P. M , A.P. L The first AP is AP 0.50:0.05:0.95 The AP value is the average of AP values under 10 IOU thresholds (with 0.05 as a breakpoint, from 0.50 to 0.95), which is the most challenging indicator. The subscripts 50 and 70 refer to the two thresholds of 0.50 and 0.75. S, M, and L indicate small objects (area < 32 2 ), mid target(32 2 <area<96 2 ) and large targets (area>96 2 ).like Figure 8 As shown in Figure 2, the network proposed in this paper improves 4.2AP and 6.1AP compared to the original Yolov5 network. 50 , 7.6AP 75 , 7.8AP S , 6.7AP M , 2.4AP L It is worth noting that AP S With 6.7AP M Compared with AP L The results show that the proposed network can improve the detection of small targets. Compared with the continuous layer network, the proposed network improves 1.5AP and 2.7AP respectively. 50 , 1.2AP 75 , 2.2AP S , 1.1AP M , 0.6AP L ,It can be found that the network proposed by the ,invention effectively improves the detection accuracy of small targets and ,verifies the effectiveness of the fault operation.
[0095] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form a technical solution.
Claims
1. A small target detection method based on a new semi-fragmented FPN+PAN feature fusion network. It is characterized in that The following steps are involved: Step 1: Use the improved feature extraction network New CSP-Darknet53 with Yolov5 backbone network to extract multi-scale target images containing small targets; The backbone network is composed of 5 groups of feature extraction modules connected in series in sequence; the first group of feature extraction modules is composed of Focus data enhancement module and C3_Res module; the second, third and fourth groups of feature extraction modules are all composed of Dconv module and C3_Res module; the fifth group of feature extraction modules is composed of Dconv module and SPP module; The C3_Res module structure is a convolutional module in which the C3 module in Yolov5 is connected to a residual connection channel at the head and the end; The Dconv module is a module composed of a deformable convolution module, a BN link, and a SiLU activation function. This module changes the shape of the convolution kernel by adding an offset to adaptively extract target features. The deformable convolution module convolves a picture of size C*H*W, reduces the picture size and increases the number of channels to obtain a feature map of size 2C*H / 2*W / 2, and its formula is: In the formula, w represents the corresponding weight of the sampling value; is a regular grid, p 0 is the convolution center, p n for Medium element, Δp n is the offset, {Δp n |n=1,....,N}, where The convolution sampling position depends on the irregular offset p n +Δp n ; The BN step uses the general batch normalization operation BatchNorm2d to make it have the statistical characteristics of describing global data, while ensuring that the network training process will not have the problems of gradient explosion and gradient disappearance; The SiLU activation function outputs a feature map of 2C*H / 2*W / 2; the formula of the SiLU activation function is as follows: Step 2: Extract 5 groups of feature maps T1, T2, T3, T4, and T5 of different depths whose sizes are halved and whose channels are doubled in sequence, generated by the five feature extraction modules of the backbone network, and input them into the feature fusion network; the feature fusion network first fuses the features of the shallow feature maps T3, T2, and T1 in the form of top-down feature fusion in the FPN network, and obtains 3 groups of new feature maps L1, L2, and L3 from T1, T2, and T3 respectively; secondly, obtains 2 groups of new feature maps L4 and L5 from T4 and T5 respectively; finally, 5 groups of new feature maps L1, L2, L3, L4, and L5 of different depths whose sizes are halved and whose channels are doubled in sequence are obtained; The step 2 specifically includes the following steps: Step 2.1: Extract the feature map T5, input it into the C3 module and the convolution module Conv2d in sequence, and output the new feature map L5; Step 2.2: Upsample the feature map L5 and perform a Concat operation with the feature map T4. The Concat operation is an operation that directly merges two different feature maps in the channel dimension. After the channels are merged, they are sequentially input into the C3 module and the convolution module Conv2d to obtain a new feature map L4. Step 2.3: Extract the feature maps T3 and T4 output by Dconv of the fourth feature extraction module of the backbone network, input them into the C3 module and the convolution module Conv2d operation in turn, and output feature maps L3 and L4; perform upsampling operation on the feature maps L3 and L4 and perform Concat operation on the T3 feature map, merge the channels, and obtain the new feature map L3; Step 2.4: Perform C3 module and convolution module Conv2d operation on L3, perform upsampling operation and Concat operation with T2 feature map, merge channels, and obtain new feature map L2; Step 2.5: Perform C3 module and convolution module Conv2d operation on L2, perform upsampling operation and Concat operation with feature map T1, merge channels, and obtain new feature map L1; Step 3: Establish a feature transmission channel between the L1, L2, L3, L4, and L5 feature maps through the PAN feature pyramid structure to ensure the effective transmission of context information between different feature layers, and output new feature maps Z1, Z2, Z3, Z4, and Z5 for detection; Step 4: Use the K-means++ algorithm to obtain the prior frame, cluster the target frame scales of the objects in the Citypersons dataset, and obtain the prior frames of 5 scales through the k-means clustering algorithm and the genetic mutation algorithm respectively; Step 5: Finally, the fused feature maps Z1, Z2, Z3, Z4, and Z5 are output respectively and used to detect the candidate box information and the probability that the candidate box belongs to a certain category. The detection results are screened using the non-maximum suppression method, that is, the candidate boxes in the output detection results are sorted according to the probability values of the candidate boxes belonging to the category in the detection results, and the candidate box with the largest probability is selected as the final result to complete the small target detection method.
2. According to claim 1, a small target detection method based on a novel semi-fault FPN+PAN feature fusion network, It is characterized in that The step 3 specifically includes the following steps: Step 3.1: Input feature map L1 into the C3 module with half the number of channels and unchanged size, and output feature map Z1; Step 3.2: Input Z1 into the Conv2d module with the same number of channels and half the size, and perform Concat fusion with the feature map L2, and then perform the C3 module with the same number of channels and size, and output the feature map Z2; Step 3.3: Perform a Conv2d module on Z2 with the same number of channels and half the size, and concat it with the feature map L3, and then perform a C3 module with the same number of channels and size, and output the feature map Z3; Step 3.4: Perform a Conv2d module on Z3 with the same number of channels and half the size, and perform Concat fusion with the feature map L4, and then perform a C3 module with the same number of channels and size, and output the feature map Z4; Step 3.5: Perform a Conv2d module on Z3 with the same number of channels and half the size, and concat it with the feature map L5. Then perform a C3 module with the same number of channels and size, and output the feature map Z5.
3. According to claim 1, a small target detection method based on a novel semi-fault FPN+PAN feature fusion network, It is characterized in that The step 4 specifically comprises the following steps: Step 4.1: Perform k-mens clustering algorithm on the labeled information of the data set to obtain the initial prior frame; Step 4.1.1: Randomly select 5 true boxes in all labels as the centers of the cluster, where the true box is the target box with the target position and size marked in the dataset label information; Step 4.1.2: Calculate the distance 1-IOU between each true box and each cluster, where IOU is the ratio of the intersection area to the union area of two true boxes; Step 4.1.3: Calculate the nearest cluster center for each ground-truth box and assign it to the cluster closest to it; Step 4.1.4: Recalculate the cluster center based on the ground-truth box in each cluster; Step 4.1.5: Repeat steps 4.1.3-4.1.4 until the center of each cluster no longer changes. Step 4.2: The wh information, i.e., height and width information, of the real box in the initial prior box is continuously adjusted, evaluated, changed, re-evaluated, and improved, and M variations are screened to obtain the final prior box result that conforms to the Citypersons dataset; Step 4.2.1: Read the wh of each image in the Citypersons dataset and the wh of all ground truth boxes, and scale the maximum value of wh in each image to the specified image size; Step 4.2.2: Change the real frame from relative coordinates to absolute coordinates, that is, multiply it by the scaled wh; Step 4.2.3: Filter the true boxes and retain the true boxes whose wh is greater than or equal to two pixels; Step 4.2.4: Use k-means clustering method to get n anchors; Step 4.2.5: Use genetic algorithm to randomly mutate the anchors’ wh. If the effect is better after mutation, assign the mutated result to the anchors. If the effect is worse after mutation, skip it. The default mutation is M times. The effect after mutation becomes better means: the fitness value calculated by the anchor_fitness method is evaluated, and the larger the fitness value, the better; Step 4.2.6: Sort the final mutated anchors by area and return them.
Citation Information
Patent Citations
A pedestrian and vehicle detection method and system based on improved YOLOv3
CN109815886A
Road pothole detection method based on YOLO v5 model
CN113902729A