Micro target detection method and system based on multi-scale feature extraction and multi-branch fusion
By adopting multi-scale feature extraction and multi-branch fusion methods in micro-object detection, combined with the context attention module, the problem of traditional methods performing poorly in micro-object detection is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510209651.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional object detection methods do not perform well in micro-object detection, especially in complex backgrounds and high noise environments, making it difficult to accurately identify micro-objects.
Using a micro-object detection method based on multi-scale feature extraction and multi-branch fusion, rich multi-scale features are extracted and the global relationship between the target and the background is constructed through a multi-scale multi-branch convolution structure and context attention module.
It significantly improves the detection accuracy of small targets, enhances the characterization ability of small targets, reduces the probability of false detection and missed detection, and improves the calculation efficiency and robustness of the model.
Smart Images

Figure CN119992067A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for detecting tiny targets based on multi-scale feature extraction and multi-branch fusion, belonging to the technical field of computer vision image processing. Background Art
[0002] With the rapid development of computer vision technology, the application of target detection in various fields has become increasingly widespread. However, traditional target detection methods often perform poorly when facing tiny targets, mainly because these targets are small in size, have unclear appearance features, and are easily obscured by complex backgrounds. Tiny target detection is of great significance in many practical applications, and has therefore become a key research topic in the field of computer vision image processing, especially when processing high-resolution images or complex scenes. Compared with traditional target detection, tiny target detection faces more challenges, such as small target size, high image noise interference, and low contrast between the target and the background. These factors cause tiny targets to be often not obvious in the image and are easily ignored or misjudged by the detection algorithm.
[0003] With the rapid development of remote sensing technology, unmanned driving, and intelligent monitoring, the demand for small target detection is increasing. For example, in remote sensing images, small targets may be ships at sea, buildings, or fire sources in forests; in unmanned driving systems, small targets may be pedestrians or obstacles. Accurate small target detection can significantly improve the safety and efficiency of these application scenarios, and also provide an important reference for decision-making.
[0004] To address these challenges, researchers have gradually turned to deep learning methods, especially convolutional neural networks (CNNs), to improve the performance of small object detection. Through multi-scale feature extraction, attention mechanisms, and improved loss function design, deep learning models have performed well in handling small object detection tasks, not only improving detection accuracy, but also making significant progress in real-time and robustness. These technological advances provide solid technical support for the practical application of small object detection.
[0005] Multi-scale feature extraction refers to capturing information about objects of different scales in an image by constructing and utilizing multiple feature layers. Since small objects usually occupy a small pixel area in an image, these objects may be ignored or difficult to distinguish from the background if only single-scale features are used for detection. Multi-scale feature extraction methods extract and integrate feature maps from different scales, allowing the model to recognize objects of various sizes, whether large objects in the entire image or small local objects. This method has been widely used in architectures such as feature pyramid networks, significantly improving the accuracy and robustness of detection.
[0006] Expanding the receptive field is achieved by changing the network structure so that each feature point can cover a larger image area, thereby obtaining more contextual information. The expansion of the receptive field enables the model to not only focus on the target itself when detecting tiny targets, but also capture the environmental information around the target, thereby improving detection accuracy. Common implementation methods include the use of dilated convolutions and specially designed network structures, such as HRNet. These technologies expand the receptive field without significantly increasing the network depth or computational cost, allowing the model to perform better when dealing with complex backgrounds and dense targets.
[0007] In recent years, in the field of visual target detection such as UAV and remote sensing, the research and application results based on the YOLO algorithm have been rich and developed rapidly. Many UAV and remote sensing visual target detection models are developed based on YOLO.
[0008] In order to meet the industry needs in target detection applications, the YOLO series of algorithms need to be optimized and improved. Although YOLO has appeared in multiple versions in just a few years, it still plays an important role in small target detection tasks. Therefore, researchers continue to innovate and expand the model to further improve its detection accuracy.
[0009] The YOLO series of models started with YOLOv1 and gradually evolved to YOLOv4. Each version has improved accuracy, speed, and model architecture. YOLOv5 continues the advantages of this series and further optimizes ease of use and performance.
[0010] Compared with the general target detection framework, YOLOv5 adopts the structure of backbone network-neck network-head network, such as Figure 1 As shown in Figure 2, the backbone network of YOLOv5 adopts the CSP structure, which consists of multiple convolutional blocks. Each convolutional block includes a convolutional layer, a batch normalization layer, and an activation layer. As the image is processed by each convolutional block, the number of its feature channels will gradually increase, while the corresponding size will decrease.
[0011] Other improved models based on YOLOv5 mainly optimize the internal structure of YOLOv5 when processing multi-scale target detection, thereby further improving the detection accuracy. For example, Scaled-YOLOv4 improves YOLOv5 in feature extraction. This method improves feature extraction efficiency by introducing a more lightweight CSP structure while reducing the number of model parameters; YOLOv5-P6 improves YOLOv5 in multi-scale detection. This method enhances the detection capability of small targets by adding additional scale detection heads to the network structure. However, this improvement is achieved on the basis of increasing the amount of calculation; YOLOv5-Lite optimizes YOLOv5 by simplifying the backbone network and reducing the model size to adapt to the application scenarios of mobile terminals and embedded devices. This method achieves faster reasoning speed by reducing the amount of calculation and memory usage.
[0012] Most current network models are improved through multi-scale methods, expanding the receptive field, or combining with other models. However, mainstream neural network models often ignore the problems of information confusion and poor feature expression when blindly expanding the spatial receptive field. Summary of the invention
[0013] In view of the shortcomings of the prior art, the present invention provides a small target detection method based on multi-scale feature extraction and multi-branch fusion;
[0014] The method of the present invention obtains richer context information through a multi-scale multi-branch convolution structure and an expansion of the receptive field. At the same time, a multi-scale multi-branch feature extraction module and a context attention module are designed, and multiple parallel convolution kernels are used to extract feature images under different receptive fields, learn richer multi-scale features, and ensure that the feature map can retain more context information.
[0015] The present invention adopts deep learning method as the technical route for small target detection. By developing a more accurate detection algorithm, high-quality recognition of small targets in images can be achieved, thereby improving the effect of wide-area detection.
[0016] The method of the present invention can significantly improve the detection accuracy of tiny targets in remote sensing images and drones, so that the target object can be accurately identified while maintaining its integrity and clarity. This not only provides great convenience for detection tasks on computer platforms, but also enables detection personnel to easily extract targets, thus providing strong support for various civilian applications.
[0017] The method of the present invention can be applied in working scenarios in the fields of intelligent transportation, smart medical care, pest and disease detection, national defense security, search and rescue, and aerial imaging.
[0018] The present invention also provides a tiny target detection system based on multi-scale feature extraction and multi-branch fusion.
[0019] The technical solution of the present invention is:
[0020] A small target detection method based on multi-scale feature extraction and multi-branch fusion, comprising:
[0021] Acquire images and perform preprocessing;
[0022] The preprocessed image is input into the trained small target detection model for target detection to obtain the target detection result;
[0023] Among them, the tiny target detection model includes a multi-scale multi-branch feature extraction module, a contextual attention module and a detection head. Through the multi-scale multi-branch feature extraction module, from the perspective of increasing feature representation, a multi-branch convolution structure is used to extract semantic information of different scales; the global relationship between small targets and backgrounds is constructed through the contextual attention module; and tiny target detection is realized through the detection head.
[0024] Preferably, according to the present invention, an image is acquired and preprocessed, including: image size standardization, data enhancement, normalization and color space conversion, denoising and background suppression, annotation and data set division, and feature map alignment.
[0025] Preferably, according to the present invention, the multi-scale multi-branch feature extraction module includes four feature extraction branches; the four feature extraction branches are initialized to perform feature extraction from multiple angles respectively;
[0026] The first feature extraction branch uses 1x1 depthwise convolution to extract local information from features;
[0027] The second feature extraction branch, the third feature extraction branch, and the fourth feature extraction branch all capture and enhance features from multiple spatial dimensions through multi-scale and deep convolution, so that the extracted feature maps can extract more contextual information;
[0028] Finally, the outputs of the first feature extraction branch, the second feature extraction branch, the third feature extraction branch, and the fourth feature extraction branch are concatenated and integrated into a unified feature representation through a 1x1 convolution operation to achieve consistency in feature space information and channel dimension.
[0029] Preferably, according to the present invention, from the perspective of increasing feature representation, a multi-branch convolution structure is used to extract semantic information of different scales; including:
[0030] S1, input feature image, obtain feature information of feature image, including local texture, edge information and high-level semantic representation of the target;
[0031] S2, initialize multiple branches and convolutional layers of the multi-scale multi-branch feature extraction module to extract features from different angles;
[0032] The first feature extraction branch uses 1x1 depth convolution to extract local information from the original features. The formula is as follows:
[0033] X1=DwConv 1x1 (f);
[0034] Among them, Conv 1x1 represents the convolution operation of 1×1 convolution kernel, X1 represents the output feature map of the first feature extraction branch, Dw represents the depth in the depth convolution, and f represents the input feature map;
[0035] The second feature extraction branch captures and enhances features from different spatial directions through multi-scale convolution and depth convolution. The formula is as follows:
[0036] X2=(DwConv 3x3 (Conv 3x1 (Conv 1x3 (Conv 1x1 (f)))));
[0037] Among them, X2 represents the output feature map of the second feature extraction branch, DwConv 3x3 Represents the convolution operation of a 3×3 depth convolution kernel, Conv 3x1 、Conv 1x3 、Conv 1x1 Respectively represent the convolution operations of 3×1, 1×3, and 1×1 depth convolution kernels, and f represents the input feature map;
[0038] The third feature extraction branch and the fourth feature extraction branch further enrich the extracted feature dimensions. The formula is as follows:
[0039] X3=(DwConv 5x5 (Conv 5x1 (Conv 1x5 (Conv 1x1 (f)))));
[0040] X4=(DwConv 7x7 (Conv 7x1 (Conv 1x7 (Conv 1x1 (f)))));
[0041] Among them, X3 and X4 represent the output feature maps of the third feature extraction branch and the fourth feature extraction branch, and DwConv 5x5 Represents the convolution operation of a 5×5 depth convolution kernel, Conv5x1 、Conv 1x5 、Conv 1x1 Respectively represent the convolution operations of 5×1, 1×5, and 1×1 depth convolution kernels, DwConv 7x7 Represents the convolution operation of the 7×7 depth convolution kernel, Conv 7x1 、Conv 1x7 、Conv 1x1 Respectively represent the convolution operations of 7×1, 1×7, and 1×1 depth convolution kernels, and f represents the input feature map;
[0042] S3, concatenates the outputs of the four feature extraction branches and further integrates them into a unified feature representation through 1x1 convolution;
[0043] S4, apply scaling factor, apply a scaling factor α to the fused feature map, Y scaled =αiY fuse , where Y fuse Represents the fused feature map, Y scaled Represents a graph of applied scaling factors, fused and residually connected with the input features processed by 1x1 convolution, and outputs the final feature map after activation function processing.
[0044] Preferably, according to the present invention, the context attention module includes a first branch and a second branch;
[0045] First, the first branch uses a 1×1 convolution operation to generate weights. After the [1×H×M] weights are obtained through the Sigmoid function, they are used for the subsequent context information weighting.
[0046] Then, the second branch uses a 3×3 convolution operation, adjusts the shape to [N, 1, HW, 1] through the Reshape function, and normalizes it through the Softmax function to represent context information in multiple dimensions;
[0047] Finally, the original features are concatenated with the weighted context information using residual connections, and the final aggregated features are obtained through matrix addition.
[0048] Preferably, according to the present invention, for the context attention module, starting from constructing the global relationship between the small target and the background, ensuring that the spatial information is retained while representing the semantic features; including:
[0049] S1, input feature image, which is the final feature map output by the multi-scale multi-branch feature extraction module;
[0050] S2, initializes the convolution layer of the context attention module for feature compression and reconstruction, and performs multiple convolution operations, including:
[0051] Use 1x1 convolution to generate weight map a, and the weight obtained by sigmoid activation function is used for subsequent context information weighting. The formula is as follows:
[0052] a(x)=Conv 1×1 (x);
[0053] Among them, Conv 1×1 (x) represents the convolution operation of a 1×1 convolution kernel, x represents the input feature, and a(x) represents the generated weight map;
[0054] Use 3x3 convolution to generate context feature k, and then adjust the shape of context feature k to [N, 1, HW, 1], where N represents the batch size; H and W represent the height and width of the input feature map respectively; and normalize it through the softmax function to represent the context information in the spatial dimension. The formula is as follows:
[0055] k(x)=Softmax(Reshape(Conv 3×3 (x)));
[0056] Among them, k(x) represents context features, Softmax represents the normalization function, Reshape represents the shape adjustment function, Conv 3×3 represents the convolution operation of 3×3 convolution kernel, and x represents the input feature;
[0057] Use 3x3 convolution to generate feature map v and adjust the shape to [N, 1, C, HW]. C represents the number of channels of feature map v, which reflects the type of feature pattern extracted and captures the feature pattern in the input image. The formula is as follows:
[0058] v(x) = Reshape(Conv 3×3 (x));
[0059] Among them, v(x) represents the output feature map, Reshape represents the shape adjustment function, Conv 3×3 Represents the convolution operation of a 3×3 convolution kernel;
[0060] Use 3x3 convolution m, m function to reconstruct the fused context information, restore and enhance the features of the final context information, and output weighted features;
[0061] S3, combines v and k through matrix multiplication to obtain the fused context feature h, and adjusts the importance of the feature through weighted mapping a. The formula is as follows:
[0062] h(x)=Conv 3×3 (v(x)·k(x));
[0063] y(x)=m(h)·a(x);
[0064] Among them, h(x) represents the fused context feature, x represents the input feature, Conv 3×3 represents the convolution operation of the 3×3 convolution kernel, v(x) and k(x) represent the feature map and the context features before fusion respectively, m function represents the function of reconstructing the context features, a(x) represents the weight mapping function, and y(x) represents the feature map of the final output;
[0065] S4, use the ReLU activation function to perform nonlinear activation processing on the weighted features;
[0066] S5, performs a residual connection between the original input and the weighted contextual features, and outputs the final aggregated features; the original input is: the feature image input by the contextual attention module.
[0067] Preferably, according to the present invention, the final aggregated features are further input into a detection head to realize tiny target detection, and the detection head is YOLOv5.
[0068] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the processor implements the steps of a small target detection method based on multi-scale feature extraction and multi-branch fusion.
[0069] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a small target detection method based on multi-scale feature extraction and multi-branch fusion.
[0070] A small target detection system based on multi-scale feature extraction and multi-branch fusion, comprising:
[0071] The image acquisition and preprocessing module is configured to: acquire images and perform preprocessing;
[0072] The small target detection module is configured as follows: the preprocessed image is input into the trained small target detection model for target detection to obtain the target detection result; wherein the small target detection model includes a multi-scale multi-branch feature extraction module, a contextual attention module and a detection head. Through the multi-scale multi-branch feature extraction module, from the perspective of increasing feature representation, a multi-branch convolution structure is used to extract semantic information of different scales; the global relationship between the small target and the background is constructed through the contextual attention module; and the small target detection is realized through the detection head.
[0073] The beneficial effects of the present invention are:
[0074] 1. Multi-scale information capture: Through the multi-branch structure, convolution kernels of different sizes and different expansion rates are used to capture target features from multiple scales, thereby significantly improving the ability to represent tiny targets.
[0075] 2. Global context fusion: The context attention module constructs the global relationship between the target and the background, retaining the spatial information while ensuring semantic expression, effectively reducing the probability of false detection and missed detection.
[0076] 3. Efficient model structure: The deep convolution and 1×1 convolution fusion strategy is adopted to effectively extract information while controlling the number of parameters, thereby improving the computational efficiency and robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 It is a schematic diagram of the structure of YOLOv5;
[0078] Figure 2 It is a schematic diagram of the basic principle of the network architecture of the present invention;
[0079] Figure 3 It is a schematic diagram of the basic principle of the multi-scale multi-branch feature extraction module of the present invention;
[0080] Figure 4 Schematic diagram of the basic principle of the context attention module of the present invention;
[0081] Figure 5(a) is the visualization result a;
[0082] Figure 5(b) is the visualization result b. DETAILED DESCRIPTION
[0083] The present invention will be further defined below in conjunction with the accompanying drawings and embodiments, but is not limited thereto.
[0084] Example 1
[0085] A small target detection method based on multi-scale feature extraction and multi-branch fusion, comprising:
[0086] Acquire images and perform preprocessing;
[0087] The preprocessed image is input into the trained small target detection model for target detection to obtain the target detection result;
[0088] Among them, the tiny target detection model includes a multi-scale multi-branch feature extraction module, a contextual attention module and a detection head. Through the multi-scale multi-branch feature extraction module, from the perspective of increasing feature representation, a multi-branch convolution structure is used to extract semantic information of different scales; the global relationship between small targets and backgrounds is constructed through the contextual attention module; and tiny target detection is realized through the detection head.
[0089] This paper designs a new multi-scale multi-branch feature extraction module, which uses multiple parallel convolution kernels to extract feature maps under different receptive fields. Through this design, the model can learn richer multi-scale features and context information, significantly improving the feature representation ability of small targets, such as Figure 3 At the same time, this module uses deep convolution to optimize the model structure, effectively reducing the number of parameters and improving the computational efficiency of the model.
[0090] Aiming at the complex background information of small targets, this paper designs a new contextual attention module to construct the global relationship between small targets and background, such as Figure 4 This module utilizes global context information to ensure that spatial information is retained while representing semantic features, thereby effectively reducing the problem of false detection and missed detection of small targets.
[0091] like Figure 2 As shown, the core of the present invention is to extract the features of small targets through a multi-scale multi-branch feature extraction module. Ideally, accurate small target feature extraction helps the context attention module provide more accurate target information and location information. However, if there is a prediction error in the context information, the wrong features may mislead subsequent processing. Therefore, the present invention proposes a new network architecture, a small target detection method (MMER) based on multi-scale multi-branch feature extraction and receptive field expansion, which can effectively utilize features and context information to help the model better learn small target features.
[0092] The framework of the present invention focuses on extracting the features of small objects through a multi-scale, multi-branch structure. Unlike traditional multi-scale models, the main innovation of the present invention is reflected in the multi-scale multi-branch feature extraction module and the context attention module. The working principles of these two modules will be explained in detail below.
[0093] Example 2
[0094] The difference between the small target detection method based on multi-scale feature extraction and multi-branch fusion described in Example 1 is that:
[0095] Acquire images and perform preprocessing; including: image size standardization, data enhancement, normalization and color space conversion, denoising and background suppression, annotation and dataset partitioning, and feature map alignment.
[0096] Original images refer to high-resolution digital images collected through drones, remote sensing images, aerial images, and intelligent monitoring images.
[0097] Through public datasets: use MS COCO, VisDrone, DOTA, xView and other annotated datasets containing tiny objects;
[0098] Provided by cooperative units: Cooperate with remote sensing institutions, hospitals or industrial manufacturers to obtain images in specific areas;
[0099] Autonomous collection: Use drones, monitoring equipment, or professional sensors to take photos on demand, ensuring that the resolution is ≥ 1024×1024 to preserve tiny target details (such as targets below 10×10 pixels).
[0100] The preprocessing process includes the following key steps to improve the model's sensitivity to small objects and reduce noise interference:
[0101] (1) Image size standardization
[0102] Input resizing: Scale the original image to a fixed size (e.g. 640×640) and use bilinear interpolation to maintain the ratio (add edge padding to avoid deformation).
[0103] Multi-scale input: For high-resolution images (such as 4000×3000), they are first cut into overlapping sub-images (such as 512×512, step size 256) to avoid losing small objects due to downsampling.
[0104] (2) Data enhancement
[0105] Geometric enhancement: random horizontal / vertical flip (probability 0.5), rotation (-15°~+15°), translation (±10% offset), and shear (±10°) to simulate target perspective changes.
[0106] Small target enhancement: For sparse small targets, a copy-and-paste strategy is used to randomly copy small targets in reasonable background areas to balance the ratio of positive and negative samples.
[0107] (3) Normalization and color space conversion
[0108] Pixel value normalization: linearly map pixel values from [0,255] to [0,1] or standardize by channel (mean = [0.485, 0.456, 0.406], standard deviation = [0.229, 0.224, 0.225]).
[0109] (4) Denoising and background suppression
[0110] High-frequency enhancement: Use the Laplacian operator or Unsharp Masking to sharpen tiny object edges.
[0111] (5) Labeling and Dataset Division
[0112] Labeling method: Use tools such as LabelImg and CVAT to label the target bounding box and category. For small targets, they need to be enlarged 2 to 3 times and then finely calibrated.
[0113] Dataset division: The training set, validation set, and test set are divided into 7:2:1 to ensure that each category (especially the small object category) is evenly distributed in the subset.
[0114] (6) Feature map alignment (training phase)
[0115] Multi-scale feature matching: For the different scale feature maps output by FPN (such as P3-P5), they are assigned to the corresponding levels according to the target size (such as targets <32×32 pixels are preferentially associated with high-level features).
[0116] The multi-scale multi-branch feature extraction module includes four feature extraction branches; the four feature extraction branches are initialized to extract features from multiple angles respectively;
[0117] The first feature extraction branch uses 1x1 deep convolution to extract local information from the features; wherein, deep convolution reduces redundant calculations and extracts spatial features more efficiently.
[0118] The second feature extraction branch, the third feature extraction branch, and the fourth feature extraction branch all use 1×1 convolution operations to make preliminary adjustments to the number of channels to be processed to ensure that the number of channels in subsequent convolution operations is consistent. The second feature extraction branch, the third feature extraction branch, and the fourth feature extraction branch all use multi-scale and deep convolutions to capture and enhance features from multiple spatial dimensions, so that the extracted feature maps can extract more contextual information; the third feature extraction branch and the fourth feature extraction branch are similar to the second feature extraction branch, but the convolution kernel sizes of the third and fourth feature extraction branches are inconsistent, which further enriches the extracted feature dimensions. Finally, the outputs of the first feature extraction branch, the second feature extraction branch, the third feature extraction branch, and the fourth feature extraction branch are spliced and integrated into a unified feature representation through 1x1 convolution operations to achieve consistency in feature space information and channel dimensions.
[0119] From the perspective of increasing feature representation, a multi-branch convolutional structure is used to extract semantic information of different scales; including:
[0120] S1, input feature image, input feature image generally refers to the intermediate feature map after image preprocessing and / or extracted by the backbone network (Backbone). This feature map already contains low-level edge and texture information as well as high-level semantic information, and is the basic representation for subsequent modules to detect small targets. Obtain feature information of the feature image, including local texture, edge information and high-level semantic representation of the target; this information is extracted through operations such as the previous convolutional layer and pooling layer. The multi-scale feature extraction module further supplements local and global information at different scales on this basis.
[0121] S2, initialize multiple branches and convolutional layers of the multi-scale multi-branch feature extraction module to extract features from different angles;
[0122] The first feature extraction branch uses 1x1 depth convolution to extract local information from the original features. The formula is as follows:
[0123] X1=DwConv 1x1 (f);
[0124] Among them, Conv 1x1 represents the convolution operation of 1×1 convolution kernel, X1 represents the output feature map of the first feature extraction branch, Dw represents the depth in the depth convolution, and f represents the input feature map;
[0125] The second feature extraction branch captures and enhances features from different spatial directions through multi-scale convolution and depth convolution. The formula is as follows:
[0126] X2=(DwConv 3x3 (Conv 3x1 (Conv 1x3 (Conv 1x1 (f)))));
[0127] Among them, X2 represents the output feature map of the second feature extraction branch, DwConv 3x3 Represents the convolution operation of a 3×3 depth convolution kernel, Conv 3x1 、Conv 1x3 、Conv 1x1 Respectively represent the convolution operations of 3×1, 1×3, and 1×1 depth convolution kernels, and f represents the input feature map;
[0128] The third and fourth feature extraction branches are similar to the first feature extraction branch, but their convolution kernel sizes are different, which further enriches the extracted feature dimensions. The formula is as follows:
[0129] X3=(DwConv 5x5 (Conv 5x1 (Conv 1x5 (Conv 1x1 (f)))));
[0130] X4=(DwConv 7x7 (Conv 7x1 (Conv 1x7 (Conv 1x1 (f)))));
[0131] Among them, X3 and X4 represent the output feature maps of the third feature extraction branch and the fourth feature extraction branch, and DwConv5x5 Represents the convolution operation of a 5×5 depth convolution kernel, Conv 5x1 、Conv 1x5 、Conv 1x1 Respectively represent the convolution operations of 5×1, 1×5, and 1×1 depth convolution kernels, DwConv 7x7 Represents the convolution operation of the 7×7 depth convolution kernel, Conv 7x1 、Conv 1x7 、Conv 1x1 Respectively represent the convolution operations of 7×1, 1×7, and 1×1 depth convolution kernels, and f represents the input feature map;
[0132] S3, concatenates the outputs of the four feature extraction branches and further integrates them into a unified feature representation through 1x1 convolution;
[0133] S4, apply scaling factor, apply a scaling factor α to the fused feature map (this factor can be set to a fixed value or as a learnable parameter) Y scaled =αiY fuse , where Y fuse Represents the fused feature map, Y scaled A graph representing the applied scaling factor, fused and residually connected with the input features processed by 1x1 convolution, and outputting the final feature map after activation function (ReLU).
[0134] The context attention module includes a first branch and a second branch;
[0135] First, the first branch uses a 1×1 convolution operation to generate weights. After the [1×H×M] weights are obtained through the Sigmoid function, they are used for the subsequent context information weighting.
[0136] Then, the second branch uses a 3×3 convolution operation, adjusts the shape to [N, 1, HW, 1] through the Reshape function, and normalizes it through the Softmax function to represent context information in multiple dimensions;
[0137] Finally, the original features are concatenated with the weighted context information using residual connections, and the final aggregated features are obtained through matrix addition.
[0138] For the context attention module, we start from building the global relationship between small objects and background, ensuring that spatial information is retained while representing semantic features; thus effectively reducing the problems of false detection and missed detection of small objects, including:
[0139] S1, input feature image, which is the final feature map output by the multi-scale multi-branch feature extraction module;
[0140] S2, initializes the convolution layer of the context attention module for feature compression and reconstruction, and performs multiple convolution operations, including:
[0141] Use 1x1 convolution to generate weight map a, and the weight obtained by sigmoid activation function is used for subsequent context information weighting. The formula is as follows:
[0142] a(x)=Conv 1×1 (x);
[0143] Among them, Conv 1×1 (x) represents the convolution operation of a 1×1 convolution kernel, x represents the input feature, and a(x) represents the generated weight map; it is used to subsequently weight the importance of contextual features.
[0144] Use 3x3 convolution to generate context feature k, and then adjust the shape of context feature k to [N, 1, HW, 1], where N represents the batch size; H and W represent the height and width of the input feature map respectively; and normalize it through the softmax function to represent the context information in the spatial dimension. The formula is as follows:
[0145] k(x)=Softmax(Reshape(Conv 3×3 (x)));
[0146] Among them, k(x) represents context features, Softmax represents the normalization function, Reshape represents the shape adjustment function, Conv 3×3 represents the convolution operation of 3×3 convolution kernel, and x represents the input feature;
[0147] Use 3x3 convolution to generate feature map v and adjust the shape to [N, 1, C, HW]. C represents the number of channels of feature map v, which reflects the type of feature pattern extracted and captures the feature pattern in the input image. The formula is as follows:
[0148] v(x) = Reshape(Conv 3×3 (x));
[0149] Among them, v(x) represents the output feature map, Reshape represents the shape adjustment function, Conv 3×3 Represents the convolution operation of a 3×3 convolution kernel;
[0150] Use 3x3 convolution m, where the enhanced features obtained after the convolution operation, the m function is used to reconstruct the fused context information to make it have better discrimination ability. Perform feature recovery and enhancement on the final context information, and output weighted features;
[0151] S3, combines v and k through matrix multiplication to obtain the fused context feature h, and adjusts the importance of the feature through weighted mapping a. The formula is as follows:
[0152] h(x)=Conv 3×3 (v(x)·k(x));
[0153] y(x)=m(h)·a(x);
[0154] Among them, h(x) represents the fused context feature, x represents the input feature, Conv 3×3 represents the convolution operation of the 3×3 convolution kernel, v(x) and k(x) represent the feature map and the context features before fusion respectively, m function represents the function of reconstructing the context features, a(x) represents the weight mapping function, and y(x) represents the feature map of the final output;
[0155] S4, use the ReLU activation function to perform nonlinear activation processing on the weighted features;
[0156] S5, performs residual connection between the original input and the weighted context features, and outputs the final aggregated features; the original input is the feature image input by the context attention module, which is also the final feature map output by the multi-scale multi-branch module.
[0157] The final aggregated features are further input into the detection head to achieve small target detection. The detection head is YOLOv5. It includes:
[0158] The structure and workflow of YOLOv5 are as follows.
[0159] The detection head of YOLOv5 consists of the following parts:
[0160] Convolutional layer: Perform convolution operation on the input feature map and output the detection result. It usually includes multiple convolutional layers, which gradually compress the feature map into a smaller output size.
[0161] Output channels: Each output channel corresponds to the number of target categories, the number of bounding box coordinates (4 coordinates), and the target confidence. Usually, these output channels are organized in the form of a fixed-size output grid, and each grid cell is responsible for predicting the target in a certain area.
[0162] Specifically, the data output by each grid unit of YOLOv5 includes:
[0163] 4 bounding box coordinates: Rectangular box position representing the target.
[0164] Object confidence: Indicates the probability that the bounding box contains the object.
[0165] Target category: predict the type of target (e.g. pedestrian, vehicle, animal, etc.).
[0166] There are multiple tiny objects in the input image, such as pedestrians, vehicles, and buildings. After the above steps, the result output by the detection head may be similar to:
[0167] Pedestrian: [x1, y1, w1, h1, class_id = 0, confidence = 0.85]
[0168] Vehicle: [x2, y2, w2, h2, class_id = 1, confidence = 0.92]
[0169] Building: [x3,y3,w3,h3,class_id=2,confidence=0.79]
[0170] Where: [x1, y1, w1, h1] are the bounding box coordinates of the pedestrian. class_id = 0 represents the pedestrian category, and confidence = 0.85 is the confidence of the model in the detection. Other targets (such as vehicles and buildings) will also have similar detection results.
[0171] Through these detection results, users can know the specific location and category of each target in the image, thereby completing the task of tiny target detection.
[0172] The hardware environment used in the present invention is NVIDIA RTX 4070, video memory 8GB, and the software environment is Python 3.8, Pytorch1.8.2. In the training phase of the MMER framework of the present invention, the training is pre-trained on small target images such as drones or remote sensing. Figure 5 (a) is the visualization result a; Figure 5 (b) is the visualization result b.
[0173] The quantitative comparison data of each category on VisDrone2019 and YOLOv10s is shown in Table 1:
[0174] Table 1
[0175]
[0176]
[0177] In Table 1, mAP50 (Mean Average Precision at IoU = 0.5) is the average precision of the target detection model when the IoU threshold is 0.5. The higher the value, the more accurate the model is. mAP50:90 (Mean Average Precision at IoU = 0.5:0.95) is the average precision of the target detection model in the range of IoU thresholds from 0.5 to 0.95. This indicator is stricter than mAP50 and can more comprehensively evaluate the performance of the model under different overlaps.
[0178] The analysis results of Table 1 are as follows:
[0179] 1. Pedestrian:
[0180] YOLOv10s: mAP50 is 38.6%, mAP50:90 is 17.7%.
[0181] Ours: mAP50 is 53.9%, mAP50:90 is 24.5%.
[0182] Analysis: The model of the present invention is significantly better than YOLOv10s in pedestrian detection, especially in the evaluation of mAP50:90, which is improved by about 6.8 percentage points.
[0183] 2. People:
[0184] YOLOv10s: mAP50 is 30.5%, mAP50:90 is 11.8%.
[0185] Ours: mAP50 is 42.1%, mAP50:90 is 16.2%.
[0186] Analysis: The model of the present invention also performs better than YOLOv10s in detecting crowds, especially in terms of mAP50 and mAP50:90.
[0187] 3. Bicycle:
[0188] YOLOv10s: mAP50 is 10.7%, mAP50:90 is 4.23%.
[0189] Ours: mAP50 is 19.8%, mAP50:90 is 8.34%.
[0190] Analysis: The model of the present invention has significantly improved the accuracy of bicycle detection, especially the mAP50:90 index has been improved by 4.1 percentage points.
[0191] 4. Car:
[0192] YOLOv10s: mAP50 is 78.5%, mAP50:90 is 55.8%.
[0193] Ours: mAP50 is 83.8%, mAP50:90 is 58.3%.
[0194] Analysis: Although YOLOv10s performs well in car detection, our model has some improvements in both mAP50 and mAP50:90, especially in mAP50, which has increased by 5.3 percentage points.
[0195] 5. Van:
[0196] YOLOv10s: mAP50 is 44.1%, mAP50:90 is 30.8%.
[0197] Ours: mAP50 is 47.0%, mAP50:90 is 33.0%.
[0198] Analysis: The model of the present invention performs slightly better than YOLOv10s in van detection, especially in the mAP50:90 indicator, which is improved by 2.2 percentage points.
[0199] 6. Truck:
[0200] YOLOv10s: mAP50 is 33.6%, mAP50:90 is 22.1%.
[0201] Ours: mAP50 is 39.9%, mAP50:90 is 25.6%.
[0202] Analysis: The proposed model also performs better in truck detection, especially in mAP50:90, which is improved by 3.5 percentage points.
[0203] 7. Tricycle:
[0204] YOLOv10s: mAP50 is 24.1%, mAP50:90 is 13.3%.
[0205] Ours: mAP50 is 32.0%, mAP50:90 is 16.5%.
[0206] Analysis: In terms of tricycle detection, the accuracy of Ours model is significantly better than YOLOv10s, with mAP50 increased by 7.9 percentage points and mAP50:90 increased by 3.2 percentage points.
[0207] 8. Awning-tricycle:
[0208] YOLOv10s: mAP50 is 15.7%, mAP50:90 is 10.0%.
[0209] Ours: mAP50 is 15.0%, mAP50:90 is 9.3%.
[0210] Analysis: For this category, the performance gap between the two models is small, Ours is slightly inferior to YOLOv10s, but the gap is not large.
[0211] 9. Bus:
[0212] YOLOv10s: mAP50 is 53.0%, mAP50:90 is 38.1%.
[0213] Ours: mAP50 is 60.1%, mAP50:90 is 39.8%.
[0214] Analysis: The model of the present invention performs better in bus detection, especially in mAP50, which is improved by 7.1 percentage points.
[0215] 10. Motorcycle:
[0216] YOLOv10s: mAP50 is 41.1%, mAP50:90 is 17.8%.
[0217] Ours: mAP50 is 50.4%, mAP50:90 is 22.1%.
[0218] Analysis: The model of the present invention also performs better in motorcycle detection, especially in terms of mAP50 and mAP50:90.
[0219] In summary, the model of the present invention outperforms YOLOv10s in most categories, especially in the mAP50 and mAP50:90 indicators. In categories such as cars, pedestrians, crowds, and motorcycles, the advantage of Ours model is more obvious. For the awning tricycle category, the performance difference between the two models is small, but YOLOv10s has a slight advantage. Overall, the model of the present invention has improved the detection accuracy of most categories, especially under the more stringent mAP50:90 evaluation, its advantages are particularly prominent.
[0220] Example 3
[0221] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a small target detection method based on multi-scale feature extraction and multi-branch fusion described in Example 1 or 2 are implemented.
[0222] Example 4
[0223] A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of a small target detection method based on multi-scale feature extraction and multi-branch fusion described in Example 1 or 2 are implemented.
[0224] Example 5
[0225] A small target detection system based on multi-scale feature extraction and multi-branch fusion, comprising:
[0226] The image acquisition and preprocessing module is configured to: acquire images and perform preprocessing;
[0227] The small target detection module is configured to: input the preprocessed image into the trained small target detection model for target detection to obtain the target detection result; wherein the small target detection model includes a multi-scale multi-branch feature extraction module and a context attention module. Through the multi-scale multi-branch feature extraction module, from the perspective of increasing feature representation, a multi-branch convolution structure is used to extract semantic information of different scales; and the global relationship between the small target and the background is constructed through the context attention module.
[0228] The steps in the method of the present invention can be adjusted in order, combined or deleted according to actual needs.
[0229] The units in the device of the present invention can be combined, divided and deleted according to actual needs.
[0230] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0231] In the description of this specification, the description of the terms "one embodiment", "some embodiments", "specific embodiments", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0232] The above description is only a preferred example of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A small target detection method based on multi-scale feature extraction and multi-branch fusion, characterized in that: include: Acquire images and perform preprocessing; The preprocessed image is input into the trained small target detection model for target detection to obtain the target detection result; Among them, the tiny target detection model includes a multi-scale multi-branch feature extraction module, a contextual attention module and a detection head. Through the multi-scale multi-branch feature extraction module, from the perspective of increasing feature representation, a multi-branch convolution structure is used to extract semantic information of different scales; the global relationship between small targets and backgrounds is constructed through the contextual attention module; and tiny target detection is realized through the detection head.
2. The method for detecting small targets based on multi-scale feature extraction and multi-branch fusion according to claim 1, characterized in that: Acquire images and perform preprocessing; including: image size standardization, data enhancement, normalization and color space conversion, denoising and background suppression, annotation and dataset partitioning, and feature map alignment.
3. The method for detecting small targets based on multi-scale feature extraction and multi-branch fusion according to claim 1, characterized in that: The multi-scale multi-branch feature extraction module includes four feature extraction branches; the four feature extraction branches are initialized to extract features from multiple angles respectively; The first feature extraction branch uses 1x1 depthwise convolution to extract local information from features; The second feature extraction branch, the third feature extraction branch, and the fourth feature extraction branch all capture and enhance features from multiple spatial dimensions through multi-scale and deep convolution, so that the extracted feature maps can extract more contextual information; Finally, the outputs of the first feature extraction branch, the second feature extraction branch, the third feature extraction branch, and the fourth feature extraction branch are concatenated and integrated into a unified feature representation through a 1x1 convolution operation to achieve consistency in feature space information and channel dimension.
4. The method for detecting small targets based on multi-scale feature extraction and multi-branch fusion according to claim 1, characterized in that: From the perspective of increasing feature representation, a multi-branch convolutional structure is used to extract semantic information of different scales; including: S1, input feature image, obtain feature information of feature image, including local texture, edge information and high-level semantic representation of target; S2, initialize multiple branches and convolutional layers of the multi-scale multi-branch feature extraction module to extract features from different angles; The first feature extraction branch uses 1x1 depth convolution to extract local information from the original features. The formula is as follows: X1=DwConv 1x1 (f); Among them, Conv 1x1 represents the convolution operation of 1×1 convolution kernel, X1 represents the output feature map of the first feature extraction branch, Dw represents the depth in the depth convolution, and f represents the input feature map; The second feature extraction branch captures and enhances features from different spatial directions through multi-scale convolution and depth convolution. The formula is as follows: X2=(DwConv 3x3 (Conv 3x1 (Conv 1x3 (Conv 1x1 (f))))); Among them, X2 represents the output feature map of the second feature extraction branch, DwConv 3x3 Represents the convolution operation of a 3×3 depth convolution kernel, Conv 3x1 、Conv 1x3 、Conv 1x1 Respectively represent the convolution operations of 3×1, 1×3, and 1×1 depth convolution kernels, and f represents the input feature map; The third feature extraction branch and the fourth feature extraction branch further enrich the extracted feature dimensions. The formula is as follows: X3=(DwConv 5x5 (Conv 5x1 (Conv 1x5 (Conv 1x1 (f))))); X4=(DwConv 7x7 (Conv 7x1 (Conv 1x7 (Conv 1x1 (f))))); Among them, X3 and X4 represent the output feature maps of the third feature extraction branch and the fourth feature extraction branch, and DwConv 5x5 Represents the convolution operation of a 5×5 depth convolution kernel, Conv 5x1 、Conv 1x5 、Conv 1x1 Respectively represent the convolution operations of 5×1, 1×5, and 1×1 depth convolution kernels, DwConv 7x7 Represents the convolution operation of the 7×7 depth convolution kernel, Conv 7x1 、Conv 1x7 、Conv 1x1 Respectively represent the convolution operations of 7×1, 1×7, and 1×1 depth convolution kernels, and f represents the input feature map; S3, concatenates the outputs of the four feature extraction branches and further integrates them into a unified feature representation through 1x1 convolution; S4, apply scaling factor, apply a scaling factor α to the fused feature map, Y scaled =αiY fuse , where Y fuse Represents the fused feature map, Y scaled Represents a graph of applied scaling factors, fused and residually connected with the input features processed by 1x1 convolution, and outputs the final feature map after activation function processing.
5. The method for detecting small targets based on multi-scale feature extraction and multi-branch fusion according to claim 1, characterized in that: The context attention module includes a first branch and a second branch; First, the first branch uses a 1×1 convolution operation to generate weights. After the [1×H×M] weights are obtained through the Sigmoid function, they are used for the subsequent context information weighting. Then, the second branch uses a 3×3 convolution operation, adjusts the shape to [N, 1, HW, 1] through the Reshape function, and normalizes it through the Softmax function to represent context information in multiple dimensions; Finally, the original features are concatenated with the weighted context information using residual connections, and the final aggregated features are obtained through matrix addition.
6. The method for detecting small targets based on multi-scale feature extraction and multi-branch fusion according to claim 1, characterized in that: For the context attention module, we start from building the global relationship between the small object and the background to ensure that the spatial information is preserved while representing the semantic features; including: S1, input feature image, which is the final feature map output by the multi-scale multi-branch feature extraction module; S2, initializes the convolution layer of the context attention module for feature compression and reconstruction, and performs multiple convolution operations, including: Use 1x1 convolution to generate weight map a, and the weight obtained by sigmoid activation function is used for subsequent context information weighting. The formula is as follows: a(x)=Conv 1×1 (x); Among them, Conv 1×1 (x) represents the convolution operation of a 1×1 convolution kernel, x represents the input feature, and a(x) represents the generated weight map; Use 3x3 convolution to generate context feature k, and then adjust the shape of context feature k to [N, 1, HW, 1], where N represents the batch size; H and W represent the height and width of the input feature map respectively; and normalize it through the softmax function to represent the context information in the spatial dimension. The formula is as follows: k(x)=Softmax(Reshape(Conv 3×3 (x))); Among them, k(x) represents context features, Softmax represents the normalization function, Reshape represents the shape adjustment function, Conv 3×3 represents the convolution operation of 3×3 convolution kernel, and x represents the input feature; Use 3x3 convolution to generate feature map v and adjust the shape to [N, 1, C, HW]. C represents the number of channels of feature map v, which reflects the type of feature pattern extracted and captures the feature pattern in the input image. The formula is as follows: v(x)=Reshape(Conv 3×3 (x)); Among them, v(x) represents the output feature map, Reshape represents the shape adjustment function, Conv 3×3 Represents the convolution operation of a 3×3 convolution kernel; Use 3x3 convolution m, m function to reconstruct the fused context information, restore and enhance the features of the final context information, and output weighted features; S3, combines v and k through matrix multiplication to obtain the fused context feature h, and adjusts the importance of the feature through weighted mapping a. The formula is as follows: h(x)=Conv 3×3 (v(x)·k(x)); y(x)=m(h)·a(x); Among them, h(x) represents the fused context feature, x represents the input feature, Conv 3×3 represents the convolution operation of the 3×3 convolution kernel, v(x) and k(x) represent the feature map and the context features before fusion respectively, m function represents the function of reconstructing the context features, a(x) represents the weight mapping function, and y(x) represents the feature map of the final output; S4, use the ReLU activation function to perform nonlinear activation processing on the weighted features; S5, performs a residual connection between the original input and the weighted contextual features, and outputs the final aggregated features; the original input is: the feature image input by the contextual attention module.
7. A method for detecting small targets based on multi-scale feature extraction and multi-branch fusion according to any one of claims 1 to 6, characterized in that: The final aggregated features are further input into the detection head for small target detection, and the detection head is YOLOv5.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of a small target detection method based on multi-scale feature extraction and multi-branch fusion as described in any one of claims 1-7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a small target detection method based on multi-scale feature extraction and multi-branch fusion as described in any one of claims 1 to 7 are implemented.
10. A small target detection system based on multi-scale feature extraction and multi-branch fusion, characterized in that: include: The image acquisition and preprocessing module is configured to: acquire images and perform preprocessing; The small target detection module is configured as follows: the preprocessed image is input into the trained small target detection model for target detection to obtain the target detection result; wherein the small target detection model includes a multi-scale multi-branch feature extraction module, a contextual attention module and a detection head. Through the multi-scale multi-branch feature extraction module, from the perspective of increasing feature representation, a multi-branch convolution structure is used to extract semantic information of different scales; the global relationship between the small target and the background is constructed through the contextual attention module; and the small target detection is realized through the detection head.
Citation Information
Cited By
Magnetic mineral particle detection method and system fused with attention mechanism under microscope
CN120235882A