Small target detection method based on multilayer feature fusion and position sensitive mechanism
By introducing technical means such as multi-layer feature fusion, attention mechanism and optimization of loss function in small object detection, the problems of identification difficulties, background interference and scale changes in small object detection are solved, and a more efficient and accurate small object detection effect is achieved.
Patent Information
- Application Number
- CN202510188023.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art has difficulties in identifying, easy to be disturbed by background and scale changes in small object detection, resulting in poor detection results.
A small object detection method based on multi-layer feature fusion and position-sensitive mechanism is adopted to improve the accuracy of small object detection by introducing attention mechanisms, custom activation functions and splicing structures, optimizing training loss functions, and increasing the output scale of small objects.
It improves the accuracy and efficiency of small object detection, enhances the model's ability to identify small objects, reduces false detection, and provides a more reliable detection solution.
Smart Images

Figure CN120107736A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a small target detection method based on multi-layer feature fusion and position-sensitive mechanism. Background Art
[0002] In recent years, autonomous driving has made breakthrough progress in the historical process of traffic development. As the core foundation of autonomous driving, perception plays a vital role in the system. The application of deep learning in autonomous driving perception mainly includes target detection and target tracking, among which target detection is the basis of tasks such as target tracking and scene understanding. With the rapid iteration of artificial intelligence technology and the continuous improvement of computing hardware performance, target detection technology is also constantly updated, especially in the field of medium-sized and above target detection. However, in many practical scenarios, small target detection is still difficult to meet the needs. Small target detection usually refers to the detection of objects that occupy a small area in the image. Compared with large targets, small targets have: low-resolution features. Small targets are small in size and have few pixels, which leads to blurred features in the image, making it difficult for the model to recognize; they are easily disturbed by the background. Small targets are often in complex backgrounds and have low contrast with the background. They are easily covered by background noise, which increases the difficulty of detection; and scale change problems. In multi-scale environments, the scale of small targets varies greatly. It is usually necessary for the model to be able to handle scale invariance to avoid misidentification in targets of different sizes. These factors often lead to unsatisfactory detection results and great challenges in algorithms. Summary of the invention
[0003] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a small target detection method based on multi-layer feature fusion and position sensitive mechanism to improve the accuracy of small target detection.
[0004] The object of the present invention is achieved in that:
[0005] A small target detection method based on multi-layer feature fusion and position-sensitive mechanism, comprising:
[0006] Construct a small target dataset;
[0007] The clustering algorithm K-means is used to obtain the prior frame of the data set. Each predicted feature map outputs two sets of prior frames and four predicted features. Figure 1 A total of eight sets of prior frames are generated;
[0008] The rectangular box loss function that integrates the gravity formula of the artificial potential field function and the IOU function is used for training until convergence is reached, and the optimal weight is obtained for reasoning. The reasoning steps are as follows:
[0009] S1, input image information is extracted through two layers of convolution and residual structure C3 module;
[0010] S2, extract features from the output of S1 through a layer of convolution and C3 module;
[0011] S3, extract features from the output of S2 through a layer of convolution and two C3 modules;
[0012] S4, extract features from the output of S3 through a layer of convolution and C3 module;
[0013] S5, pass the output of S4 sequentially through the C3 module, convolution, attention mechanism CA module, pyramid pooling SPPF structure and convolution to extract features;
[0014] S6, after upsampling the output of S5, calculate the ACB structure with the output of S3. The ACB structure is the CONCAT function that fuses the ADD, CONV, and BN layers by channel weights.
[0015] S7, pass the output of S6 through a C3 module and a layer of convolution to extract features;
[0016] S8, after upsampling, perform ACB structure calculation on the output of S7 and the output of S2;
[0017] S9, extract features from the output of S8 through a C3 module;
[0018] S10, after a layer of convolution and upsampling, the output of S9 is combined with the output of S1 to perform ACB structure calculation;
[0019] S11, extract features from the output of S10 through a C3 module, and obtain a small target detection layer of 160*160 size here;
[0020] S12, after a layer of convolution, performs ACB calculation on the output of S11, and the output of S2 and S9. The calculation structure extracts features through the C3 module to obtain a detection layer of 80*80 size;
[0021] S13, after a layer of convolution, the output of S12 is used to perform ACB calculation with the outputs of S3 and S7. The calculation structure extracts features through the C3 module to obtain a detection layer of size 40*40;
[0022] S14, after a layer of convolution, performs ACB calculation on the output of S13 and the outputs of S4 and S5. The calculation structure extracts features through the C3 module to obtain a detection layer of size 20*20.
[0023] Furthermore, in the attention mechanism CA module, feature encoding of features aggregated along different directions captures long-range dependencies along one spatial direction, retains precise position information along another spatial direction, and encodes the generated feature maps separately to form a pair of direction-aware and position-sensitive feature maps.
[0024] Furthermore, the feature maps are weighted fused to avoid feature redundancy and help retain the key information of small objects.
[0025] Furthermore, after the inference is completed, the NMS stage is used to reduce false detections.
[0026] Furthermore, in the activation function of each convolution operation, channel-by-channel convolution is combined with the BN layer to form a DCB structure, and the DCB structure is used to calculate the maximum value of the current position and the corresponding position of the output of this structure to implement the activation function.
[0027] Furthermore, the calculation formula of the rectangular box loss function is as follows:
[0028]
[0029] Where S is the loss function; a is the weight coefficient of the ciou part; K att is the gain factor of the gravitational potential field function; s is the gravitational influence radius, which is set to the maximum length or width of the input.
[0030] Furthermore, the small target dataset includes most of the target groups of size 32*32, and the rest are target data of other sizes to increase the generalization ability of the model.
[0031] Due to the adoption of the above technical solution, the present invention has the following beneficial effects:
[0032] The present invention improves the detection accuracy and efficiency of small targets by improving the network structure, adding a small target detection layer, optimizing the training loss function, and other measures, providing a reliable solution for practical applications.
[0033] The present invention aims at the small target detection category of target detection, mainly by introducing the attention mechanism to establish the direction-aware and position-sensitive feature map; optimizing the activation function, using the channel-by-channel convolution and BN layer to achieve it; optimizing the training loss function, using the artificial potential field function gravity formula and the IOU function to act together as the rectangular box loss function; optimizing the feature map splicing, using channel-weighted fusion; improving the network structure, increasing the small target output scale and other measures to solve the small target detection problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0035] See also Figure 1 , a small target detection method based on multi-layer feature fusion and position-sensitive mechanism, the improvements are as follows:
[0036] (1) Add the CA attention mechanism to aggregate feature encodings along different directions, capture long-range dependencies along one spatial direction, retain precise position information along another spatial direction, and encode the generated feature maps separately to form a pair of direction-aware and position-sensitive feature maps;
[0037] (2) Since small targets usually have less pixel information and are easily masked or interfered by background noise, traditional activation functions are not resistant to interference. Convolution can be used at the activation function to increase nonlinearity and enhance spatial position representation. Therefore, channel-by-channel convolution (DC) and BN layers, referred to as DCB structure, are used to calculate the maximum value of the current position and the corresponding position of the output of this structure to implement the activation function.
[0038] (3) Optimize the network structure, add a small target detection layer, add initial low-dimensional features to the high- and low-layer networks in the FPN structure to increase the small target detection layer, while retaining the detection layer of targets of other sizes, and reduce false detections through the NMS stage;
[0039] The core implementation of the FPN structure is reflected in the following steps:
[0040] Steps S5, S6, S8, S10 and subsequent detection layer generation (S11-S14)
[0041] S5 (Pyramid Pooling SPPF Structure):
[0042] The SPPF module is used to perform multi-scale pooling on high-level features to enhance feature expression capabilities and provide a basis for subsequent FPN feature fusion.
[0043] S6 (S5 output upsampled and fused with S3):
[0044] The high-level features (S5) are upsampled and then fused with the middle-level features (S3) using ACB to achieve top-down feature propagation and lateral connection, forming the first-level feature pyramid of FPN.
[0045] S8 (S7 output upsampled and fused with S2):
[0046] Further integrate higher-level features (S7) with lower-level features (S2) to expand the multi-scale expression capability of FPN.
[0047] S10 (S9 output upsampled and fused with S1):
[0048] The lowest level initial features (S1) are introduced to supplement the detail information, optimize the small target detection, and complete the underlying feature fusion of FPN.
[0049] S11-S14 (multi-scale detection layer generation):
[0050] S11 generates a 160×160 small target detection layer;
[0051] S12 generates an 80×80 detection layer;
[0052] S13 generates a 40×40 detection layer;
[0053] S14 generates a 20×20 detection layer.
[0054] These detection layers output prediction results of different scales through multi-level feature fusion of FPN, covering the detection needs from tiny targets to large targets.
[0055] (4) The traditional CONCAT concatenation of the network causes redundant information to be contained in the feature map. Since the representation information of small targets is originally less, the redundant features will cover up the key information of the small targets. The weighted fusion of feature maps avoids the generation of feature redundancy and helps to retain the key information of small targets. The commonly used CONCAT part uses channel-wise weighted fusion (ADD) + convolution (CONV) + BN layer, referred to as the ACB structure.
[0056] (5) During the network training stage, the closer the artificial potential field function is to the target position, the smaller the gravity is, which is consistent with the direction of the IOU loss function, and small targets are more sensitive to positional relationships. Therefore, the gravity formula of the artificial potential field function and the IOU function are used together as the rectangular box loss function. The calculation formula is as follows:
[0057]
[0058] Where S is the loss function, a is the weight coefficient of the ciou part, K att is the gain factor of the gravitational potential field function, s is the gravitational influence radius, which is set to the maximum length or width of the input to prevent the gravitational value from being too large.
[0059] The process of small target detection based on multi-layer feature fusion and position-sensitive mechanism is as follows:
[0060] (1) Construct a small target dataset, mainly a 32*32 target group, and also construct regular target size data to increase the generalization ability of the model;
[0061] (2) Use the K-means clustering algorithm to obtain the prior box of the data set. Each feature map outputs two sets of prior boxes and four prediction features. Figure 1 A total of 8 sets of prior frames are generated;
[0062] (3) Use the modified fusion artificial potential field loss function to train until convergence is reached and obtain the optimal weights for inference;
[0063] (4) The specific reasoning steps of the small target detection method based on multi-layer feature fusion and position-sensitive mechanism are as follows:
[0064] S1, input image information is extracted through two layers of convolution and residual structure C3 module;
[0065] S2, extract features from the output of S1 through a layer of convolution and C3 module;
[0066] S3, extract features from the output of S2 through a layer of convolution and two C3 modules;
[0067] S4, extract features from the output of S3 through a layer of convolution and C3 module;
[0068] S5, pass the output of S4 sequentially through the C3 module, convolution, attention mechanism CA module, pyramid pooling SPPF structure and convolution to extract features;
[0069] S6, after upsampling the output of S5, perform ACB structure calculation with the output of S3;
[0070] S7, output a C3 module and a layer of convolution to extract features from S6;
[0071] S8, after upsampling, perform ACB structure calculation on the output of S7 and the output of S2;
[0072] S9, extract features from the output of S8 through a C3 module;
[0073] S10, after a layer of convolution and upsampling, the output of S9 is combined with the output of S1 to perform ACB structure calculation;
[0074] S11, extract features from the output of S10 through a C3 module, and obtain a small target detection layer of 160*160 size here;
[0075] S12, after a layer of convolution, performs ACB calculation on the output of S11, and the output of S2 and S9. The calculation structure extracts features through the C3 module to obtain a detection layer of 80*80 size;
[0076] S13, after a layer of convolution, the output of S12 is used to perform ACB calculation with the outputs of S3 and S7. The calculation structure extracts features through the C3 module to obtain a detection layer of size 40*40;
[0077] S14, after a layer of convolution, the output of S13 is used to perform ACB calculation with the outputs of S4 and S5. The calculation structure extracts features through the C3 module to obtain a detection layer of size 20*20;
[0078] Invention point:
[0079] The problem of small target detection is solved by combining attention mechanism, custom DCB structure activation function, custom ACB structure splicing feature map, using the artificial potential field function gravity formula and IOU function as the rectangular box loss function, and increasing the output scale of small targets.
[0080] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made in form and details without departing from the scope defined by the claims of the present invention.
Claims
1. A small target detection method based on multi-layer feature fusion and position-sensitive mechanism, characterized in that: include: Construct a small target dataset; The K-means clustering algorithm is used to obtain the prior frames of the data set. Each predicted feature map outputs two sets of prior frames, and four predicted feature maps generate a total of eight sets of prior frames. The rectangular box loss function that integrates the gravity formula of the artificial potential field function and the IOU function is used for training until convergence is reached, and the optimal weight is obtained for reasoning. The reasoning steps are as follows: S1, input image information is extracted through two layers of convolution and residual structure C3 module; S2, extract features from the output of S1 through a layer of convolution and C3 module; S3, extract features from the output of S2 through a layer of convolution and two C3 modules; S4, extract features from the output of S3 through a layer of convolution and C3 module; S5, pass the output of S4 sequentially through the C3 module, convolution, attention mechanism CA module, pyramid pooling SPPF structure and convolution to extract features; S6, after upsampling the output of S5, calculate the ACB structure with the output of S3. The ACB structure is the CONCAT function that fuses the ADD, CONV, and BN layers by channel weights. S7, pass the output of S6 through a C3 module and a layer of convolution to extract features; S8, after upsampling, perform ACB structure calculation on the output of S7 and the output of S2; S9, extract features from the output of S8 through a C3 module; S10, after a layer of convolution and upsampling, the output of S9 is combined with the output of S1 to perform ACB structure calculation; S11, extract features from the output of S10 through a C3 module, and obtain a small target detection layer of 160*160 size here; S12, after a layer of convolution, performs ACB calculation on the output of S11, and the output of S2 and S9. The calculation structure extracts features through the C3 module to obtain a detection layer of 80*80 size; S13, after a layer of convolution, the output of S12 is used to perform ACB calculation with the outputs of S3 and S7. The calculation structure extracts features through the C3 module to obtain a detection layer of size 40*40; S14, after a layer of convolution, performs ACB calculation on the output of S13 and the outputs of S4 and S5. The calculation structure extracts features through the C3 module to obtain a detection layer of size 20*20.
2. The small target detection method based on multi-layer feature fusion and position sensitive mechanism according to claim 1 is characterized by: In the attention mechanism CA module, feature encoding of features aggregated along different directions captures long-range dependencies along one spatial direction, retains precise position information along another spatial direction, and encodes the generated feature maps separately to form a pair of direction-aware and position-sensitive feature maps.
3. The small target detection method based on multi-layer feature fusion and position sensitive mechanism according to claim 2 is characterized by: The weighted fusion of feature maps avoids the generation of feature redundancy and helps to retain the key information of small targets.
4. The small target detection method based on multi-layer feature fusion and position sensitive mechanism according to claim 1 is characterized by: After inference is completed, the NMS stage is used to reduce false positives.
5. The small target detection method based on multi-layer feature fusion and position sensitive mechanism according to claim 1 is characterized by: In the activation function of each convolution operation, channel-by-channel convolution is combined with the BN layer to form a DCB structure. The DCB structure is used to calculate the maximum value of the current position and the corresponding position of the output of this structure to implement the activation function.
6. The small target detection method based on multi-layer feature fusion and position sensitive mechanism according to claim 1, characterized in that: The calculation formula of the rectangular box loss function is as follows: Where S is the loss function; a is the weight coefficient of the ciou part; K att is the gain factor of the gravitational potential field function; s is the gravitational influence radius, which is set to the maximum length or width of the input.
7. The small target detection method based on multi-layer feature fusion and position sensitive mechanism according to claim 1 is characterized by: The small target dataset includes most of the target groups with a size of 32*32, and the rest are target data of other sizes to increase the generalization ability of the model.