A small target detection method based on the super-resolution YOLO network
By combining the super-segment module and the yolo network, and introducing attention mechanism and recursive hollow convolution module, the problem of targets being easily lost in small object detection is solved, achieving higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202310255946.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-16
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-03-16
AI Technical Summary
The existing small-object detection technology is difficult to effectively identify and extract small-object objects, especially when the image pixels are low and there is a lot of interference information, resulting in the target being easily lost.
A small object detection method based on super-yolo network is adopted. By combining the yolo network with the super-segment module, a position attention mechanism, a self-attention mechanism and a recursive hollow convolution module are introduced to enhance the location information and detailed feature extraction of the small object.
Obtaining more object feature information through a larger receptive field significantly improves the accuracy and robustness of small object detection, and can effectively extract small objects.
Smart Images

Figure CN116245860B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of small target object detection, and in particular to a small target detection method based on a super-resolution-Yolo network. Background Art
[0002] Small object detection has always been a major difficulty in the field of object detection. Although object detection algorithms have emerged in an endless stream in the past few years, in the process of small object recognition, it is often encountered that the image pixels are too low and there is too much interference information, which makes the object very easy to lose. Small object detection has become one of the most challenging tasks in computer vision. In addition, large-scale benchmark datasets for small-size object detection are still not comprehensive enough.
[0003] With the in-depth study of deep learning, the target detection algorithm based on convolutional neural network has also made great progress, especially the detection algorithm for large and medium targets, which can basically meet the needs of various scenarios. Small targets also exist in large numbers in real life and have broad application prospects. For example, they play a huge role in many application fields such as remote sensing image processing, drone navigation, automatic driving, medical diagnosis, face recognition, etc. Due to the small scale of small targets themselves, the amount of information contained in the image is small, which easily causes the target to be blurred and the detailed features to be unclear, thus restricting the further development of small target detection performance.
[0004] The small target detection method based on deep learning is improved on the basis of the two-stage and single-stage algorithms. The two-stage method divides the detection problem into two steps. First, some possible candidate areas are screened out, and then the target features are extracted for each candidate area. This method is relatively inefficient and cannot meet the real-time requirements. The target detection process of the single-stage and two-stage algorithms is different. End-to-end detection can be performed without the need to screen candidate areas, and the running speed is faster. Common single-stage target detection algorithms include the YOLO series and the SSD series.
[0005] The Chinese patent application number is: CN202110618368.4, and the name is: A method for super-resolution and small target detection of infrared images. This method is a super-resolution reconstruction algorithm for infrared images assisted by visible light images. It improves the image resolution of the original input infrared image based on the super-resolution technology of visible light images; the infrared image with improved resolution is input into the designed generative adversarial network. In the designed generator, the original image can be directly input into the designed network, and the extracted high-level features and low-level features are combined to ensure the integrity of the image detail texture features as much as possible. Because this method refers to the generative adversarial network, the training process is long, it is easy to consume a lot of resources, and it starts from the infrared image unilaterally, and the features obtained are relatively simple.
[0006] The Chinese patent application number is: 202210736078.4, and the title is: A lightweight small target detection method. This method is a lightweight small target detection method. By enhancing, shrinking, and embedding the image to be detected into the enhanced image, and then using a small target detection network based on RetinaTarget for target detection, it can achieve the simultaneous detection of large targets at close range and small targets at long range. When this method embeds small pictures into large images and uses feature enhancement, some original details of the image will be lost during image scaling, and it is easy to miss detections when blurred images appear. Summary of the Invention
[0007] The present invention proposes a small target detection method based on a super-resolution-yolo network, which can obtain more object feature information through a larger receptive field. When applied to small target detection, it can effectively extract small objects.
[0008] The present invention adopts the following technical solutions.
[0009] A small target detection method based on a super-resolution-yolo network, the method uses the yolo network as the basic network architecture, combines with a super-resolution module and adds a position attention mechanism, a self-attention mechanism, and introduces a recursive dilated convolution module to highlight the small target position information and enhance the detail features; including the following steps;
[0010] Step 1: Collect images and preprocess the images to generate the original images;
[0011] Step 2: Connect the super-resolution fitting network with the improved yolo target recognition network to form a super-resolution-yolo network model, process the original images and output ultra-clear images with higher resolution, and then perform operations such as feature extraction, target box regression classification, and object detection and recognition on the ultra-clear images to obtain the recognized images and form a training dataset;
[0012] Step 3: Train the small target object detection model;
[0013] Step 4: Use the trained model for small target object detection.
[0014] The said Step 1 includes the following steps;
[0015] Step 1-1: Image acquisition: Place the camera at a specific position, and then obtain small object images through the camera, that is, images with small target objects;
[0016] Step 1-2: Image cutting: According to the captured images with small target objects, use bilinear interpolation to fill all images into small squares centered on each pixel point in the image, that is, perform expansion processing on each pixel of the picture;
[0017] Steps 1-3, Sample Set Creation: Manually add pixel-level labels to the small square images with small targets in the previous step to create a small target sample set.
[0018] In the second step, the super-resolution fitting network is composed of a feature extraction network and a super-resolution linear fitting network. The original image passes through the super-resolution fitting network to output a higher-resolution ultra-clear image. Specifically, it includes:
[0019] Step 2-1, The super-resolution fitting network is composed of a super-resolution feature extraction network Encoder and a super-resolution linear fitting network LIIF. The super-resolution feature extraction network Encoder uses recursive dilated convolution to obtain shallow features of the original image, then connects a residual group for deep feature extraction, then uses sub-pixel convolution for upsampling, and finally reconstructs the enlarged features through a convolutional layer to obtain a feature tensor.
[0020] The super-resolution linear fitting network first unfolds the feature tensor obtained by the super-resolution feature extraction network, aligns the local neighbors of the latent features and performs feature enrichment processing, and then decodes the transformed features to generate an ultra-clear image.
[0021] Step 2-2, Send the generated ultra-clear image to the improved yolo object recognition network for feature extraction, and then perform feature fusion processing on the extracted features through the network head module to obtain features of three sizes: large, medium, and small. The fused features are used to predict the object type and location, and finally the detection result is output.
[0022] The improved yolo object recognition network is improved in that
[0023] Improvement 1: Use three convolutional layers to input into Hornet through forward propagation, combine them into a new architecture ConvHB, and replace the ELAN module of the original yolo backbone network backbone with the ConvHB module, reducing the computational complexity and forming a highly effective spatial interaction modeling method.
[0024] Improvement 2: Connect a convolutional layer Conv and a Vision Transformer module ViT to form a new CViT module, and improve the detection of small target objects by introducing context information.
[0025] Improvement Point 3: Introduce a position attention mechanism to highlight the position information of small target objects. Input feature X into two convolutional layers with kernel sizes of 1×W and H×1 respectively to generate two vectors, A and B. Perform a matrix multiplication operation between A and B, and use a softmax layer to obtain the position attention map S. Multiply the original feature X by the position attention map S, and then perform element-wise summation with feature X to generate the final enhanced feature map Y. Finally, transplant the position attention mechanism WZ module into the last layer of the target recognition backbone network.
[0026] Step 3 includes the following steps;
[0027] Step 3-1: Parameter initialization settings: During training, the image size is adjusted to Hi×Wi; the number of training times, the initial learning rate, and the batch size are also set to a, b, and c; the cosine annealing learning rate adjustment method is used to adjust the learning rate b; Adam is selected as the optimizer to make the dataset obtain the optimal value faster; during the training phase, perform image enhancement operations such as random cropping, random brightness, and random rotation.
[0028] Step 3-2: Start training: Input the training samples into the SR-yolo detection model.
[0029] Step 3-3: Update parameters: The loss function used includes three parts: bounding box loss, confidence loss, and classification loss;
[0030] Step 3-4: Output the model: Repeat steps 3-2 and 3-3. When the number of training times reaches the set value, stop training and output the model with the minimum loss value.
[0031] In step 3-3, the loss function of yolov7 is used, that is: bounding box loss, confidence loss, and classification loss. The specific formula is as follows:
[0032] Loss = a×loss obj + b×loss rect + c×loss clc;
[0033] In the formula, CIOU loss is used to calculate the bounding box loss, and both the confidence loss loss rect and the classification loss loss clc are calculated using BCE loss, that is, calculated using the binary cross-entropy loss with log.
[0034] The improved yolo target recognition network has its main body adopting an improved yolov7 network architecture, including Input, Backbone, and Head. After passing the ultra-high-definition image through four CBS modules, it is sent into the ConvHB module with dilated recursive convolution.
[0035] The present invention uses the YOLO network as the basic network architecture, combines it with a super-resolution module, adds a position attention mechanism, a self-attention mechanism, and introduces a recursive dilated convolution module; it can better highlight the position information of small targets and enhance detailed features; since the residual block can fuse the receptive fields of multiple scales, this method can obtain more object feature information through a larger receptive field; when the small target detection method of the present invention is applied to small target detection, it can effectively realize the extraction of small objects.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] (1) The present invention proposes a small target detection method based on a super-resolution-YOLO network, realizing the intelligent detection of small target objects and improving the detection accuracy.
[0038] (2) A position attention mechanism is designed to highlight the position of small targets, which can tolerate the interference of other factors on the target and improve the robustness of detection.
[0039] (3) A super-resolution network architecture is designed, which can effectively combine with the target recognition network to improve the recognition rate of small target objects.
[0040] (4) Multiple residual blocks are used to fuse the receptive fields of multiple scales, combined with context information, to effectively determine the position information of the target object. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The present invention will be further described in detail below with reference to the drawings and specific embodiments:
[0042] Attached Figure 1 is a schematic flow chart of the present invention for completing image acquisition, super-resolution, and recognition;
[0043] Attached Figure 2 is a schematic network structure diagram of SR-YOLO of the present invention;
[0044] Attached Figure 3 is a schematic super-resolution network structure diagram of the present invention;
[0045] Attached Figure 4 is Figure 2 a supplementary schematic diagram of the target recognition network structure;
[0046] Attached Figure 5 is a schematic diagram of the position attention mechanism structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The technical solutions of the present invention will be specifically described below with reference to the drawings.
[0048] As shown in the figure, a small target detection method based on the super-resolution YOLO network. The method uses the YOLO network as the basic network architecture, combines it with the super-resolution module, adds a position attention mechanism, a self-attention mechanism, and introduces a recursive dilated convolution module to highlight the position information of small targets and enhance the detailed features. It includes the following steps;
[0049] Step 1: Collect images and preprocess the images to generate the original images;
[0050] Step 2: Connect the super-resolution fitting network with the improved YOLO target recognition network to form a super-resolution YOLO network model. Process the original images and output super-clear images with higher resolution. Then perform operations such as feature extraction, target box regression classification, and object detection and recognition on the super-clear images to obtain the recognized images and form a training dataset;
[0051] Step 3: Train the small target object detection model;
[0052] Step 4: Use the trained model for small target object detection.
[0053] The above Step 1 includes the following steps;
[0054] Step 1-1: Image acquisition: Place the camera at a specific position, and then obtain small object images through the camera, that is, images with small target objects;
[0055] Step 1-2: Image cutting: According to the images with small target objects captured, use bilinear interpolation to fill all images into small squares centered on each pixel point in the image, that is, perform expansion processing on each pixel of the picture;
[0056] Step 1-3: Sample set production: Manually add pixel-level labels to the small square images with small targets in the previous step to make a small target sample set.
[0057] In the above Step 2, the super-resolution fitting network is composed of a feature extraction network and a super-resolution linear fitting network. The original images pass through the super-resolution fitting network and output super-clear images with higher resolution. Specifically, it includes:
[0058] Step 2-1: The super-resolution fitting network is composed of a super-resolution feature extraction network Encoder and a super-resolution linear fitting network LIIF. The super-resolution feature extraction network Encoder uses recursive dilated convolution to obtain the shallow features of the original images, then connects a residual group for deep feature extraction, then performs upsampling in the form of sub-pixel convolution, and finally reconstructs the enlarged features through a convolution layer to obtain a feature tensor;
[0059] The super-resolution linear fitting network first unfolds the feature tensor obtained by the super-resolution feature extraction network, aligns the local neighbors of the hidden features and performs feature enrichment processing, and then decodes the features after the unfolding transformation to generate a super-clear image;
[0060] Step 2-2: Send the generated super-clear image to the improved yolo object recognition network for feature extraction, and then perform feature fusion processing on the extracted features through the network head module to obtain features of three sizes: large, medium, and small. The fused features are used to predict the object type and location, and finally the detection results are output.
[0061] The improved yolo object recognition network is improved in that,
[0062] Improvement 1: Use three convolutional layers to input into Hornet through forward propagation, combine them into a new architecture ConvHB, and replace the ELAN module of the original yolo backbone network backbone with the ConvHB module, reducing the computational complexity and forming a highly effective spatial interaction modeling method;
[0063] Improvement 2: Connect the convolutional layer Conv and the Vision Transformer module ViT to form a new CViT module, and improve the detection of small target objects by introducing context information;
[0064] Improvement 3: Introduce a position attention mechanism to highlight the position information of small target objects. Input the feature X into two convolutional layers with a convolutional kernel size of 1×W and H×1 respectively to generate two vectors A and B; perform a matrix multiplication operation between A and B, and use a softmax layer to obtain the position attention map S; multiply the original feature X by the position attention map S, and then perform an element-wise summation with the feature X to generate the final enhanced feature map Y; finally, transplant the position attention mechanism WZ module to the last layer of the object recognition backbone network.
[0065] The third step includes the following steps;
[0066] Step 3-1: Parameter initialization settings: During training, the image size is adjusted to Hi×Wi; the number of training times, the initial learning rate, and the batch size are also set to a, b, and c; the cosine annealing learning rate adjustment method is used to adjust the learning rate b; Adam is selected as the optimizer to make the dataset obtain the optimal value faster; during the training phase, image enhancement operations such as random cropping, random brightness, and random rotation are performed;
[0067] Step 3-2: Start training: Input the training samples into the SR-yolo detection model;
[0068] Step 3-3. Update parameters: The loss function adopted includes three parts: bounding box loss, confidence loss, and classification loss;
[0069] Step 3-4. Output the model: Repeat Step 3-2 and Step 3-3. When the number of training times reaches the set value, stop training and output the model with the minimum loss value.
[0070] In Step 3-3, the loss function of yolov7 is adopted, that is: bounding box loss, confidence loss, and classification loss, which are specifically expressed by the formula as follows:
[0071] Loss = a×loss_obj + b×loss_rect + c×loss_clc;
[0072] In the formula, CIOU loss is used to calculate the bounding box loss, and both the confidence loss loss_rect and the classification loss loss_clc are calculated using BCE loss, that is, calculated with the binary cross-entropy loss with log.
[0073] For the improved yolo object recognition network, its main body adopts the improved yolov7 network architecture, including Input, Backbone, and Head. After passing the ultra-high-definition image through four CBS modules, it is sent to the ConvHB module with dilated recursive convolution.
[0074] In this example, the super-resolution network structure is like the super-resolution fitting network part in Figure 2 and is composed of the feature extraction network Encoder and the super-resolution linear fitting network LIIF.
[0075] The feature extraction network Encoder uses dilated recursive convolution to obtain the shallow features of the original image, then connects a residual group to extract the deep features of the image, then uses sub-pixel convolution for upsampling, and finally reconstructs the enlarged features through a convolutional layer Conv. As shown in Figure 3 shown.
[0076] The super-resolution fitting network (LIIF) first unfolds the tensor after feature extraction, aiming to align the local neighbors of the latent features and enrich the features, and then performs a decoding operation on the unfolded and transformed features to output the ultra-high-definition image.
[0077] In this example, the main body of the object recognition network adopts the improved yolov7 network architecture, and the ultra-high-definition image obtained through the super-resolution fitting network is input into the improved yolov7 network. The yolov7 network mainly includes three parts: Input, Backbone, and Head, as shown in Figure 2 shown. The improvement steps for yolov7 are as follows:
[0078] (1) After passing the ultra-high-definition image through four CBS modules, it is sent to the ConvHB module with dilated recursive convolution, which replaces the original ELAN module, reduces the computational complexity, and effectively realizes spatial interaction modeling.
[0079] (2) After feature extraction by the Backbone, three sizes of feature vectors are output and sent to the position attention mechanism WZ module respectively to extract the position information of the target object. Then it enters the head part of the network for feature fusion. After passing through the CViT module with attention mechanism, three different sizes of feature maps are obtained again. Among them, the CViT module is connected by an ordinary convolutional layer and a Vision Tranformer module. Finally, the fused features are calculated by the REP module and the results are output.
[0080] (3) To highlight the position information of small target objects, a position attention mechanism is introduced, as Figure 5 shown. It inputs the feature X into two convolutional layers with kernel sizes of 1×W and H×1 respectively to generate two vectors A and B; a matrix multiplication operation is performed between A and B, and a softmax layer is used to obtain the position attention map S; the original feature X is multiplied by the position attention map S, and then element-wise summation is performed with the feature X to generate the final enhanced feature map Y. Finally, the position attention mechanism WZ module is transplanted into the last layer of the target recognition backbone network, such as the WZ module in Figure 2 .
Claims
1. A small target detection method based on the super-resolution YOLO network, characterized in that: The method uses the YOLO network as the basic network architecture, combines it with the super-resolution module, and adds a position attention mechanism, a self-attention mechanism, and introduces a recursive dilated convolution module to highlight the position information of small targets and enhance the detailed features; It includes the following steps; Step 1: Collect images and preprocess the images to generate the original images; Step 2: Connect the super-resolution fitting network with the improved YOLO object recognition network to form a super-resolution-YOLO network model. Process the original images and output ultra-clear images with higher resolution. Then, perform feature extraction, target box regression classification, and object detection and recognition operations on the ultra-clear images to obtain the recognized images, forming a training dataset; Step 3: Train the small target object detection model; Step 4: Use the trained model for small target object detection; In the second step, the super-resolution fitting network is composed of a feature extraction network and a super-resolution linear fitting network. The original images pass through the super-resolution fitting network and output ultra-clear images with higher resolution. Specifically, it includes: Step 2-1: The super-resolution fitting network is composed of a super-resolution feature extraction network Encoder and a super-resolution linear fitting network LIIF. The super-resolution feature extraction network Encoder uses recursive dilated convolution to obtain the shallow features of the original images, then connects a residual group for deep feature extraction, and then uses the sub-pixel convolution method for upsampling. Finally, through a convolutional layer, the enlarged features are reconstructed to obtain the feature tensor; The super-resolution linear fitting network first unfolds the feature tensor obtained by the super-resolution feature extraction network, aligns the local neighbors of the hidden features and performs feature enrichment processing, and then decodes the features after the unfolding transformation to generate ultra-clear images; Step 2-2: Send the generated ultra-clear images to the improved YOLO object recognition network for feature extraction, and then perform feature fusion processing on the extracted features through the network head module to obtain features of three sizes: large, medium, and small. The fused features are used to predict the target types and positions, and finally the detection results are output; For the improved YOLO object recognition network, the improvement points are as follows Improvement point 1: Use three convolutional layers to be input into Hornet through forward propagation, combine them into a new architecture ConvHB, and use the ConvHB module to replace the ELAN module of the original YOLO backbone network backbone, reducing the computational complexity and forming a highly effective spatial interaction modeling method; Improvement point 2: Connect the convolutional layer Conv and the VisionTransformer module ViT to form a new CViT module, and improve the detection of small target objects by introducing context information; Improvement Point 3: Introduce a position attention mechanism to highlight the position information of small target objects. Input feature X into two convolutional layers with kernel sizes of 1×W and H×1 respectively to generate two vectors A and B. Perform a matrix multiplication operation between A and B, and use a softmax layer to obtain the position attention map S. Multiply the original feature X by the position attention map S, and then perform element-wise summation with feature X to generate the final enhanced feature map Y. Finally, transplant the position attention mechanism WZ module into the last layer of the target recognition backbone network; Step 3 includes the following steps; Step 3-1: Parameter initialization settings: During training, the image size is adjusted to Hi×Wi; the number of training times, the initial learning rate, and the batch size are also set to a, b, and c; the cosine annealing learning rate adjustment method is used to adjust the learning rate b; Adam is selected as the optimizer to make the dataset obtain the optimal value faster; during the training phase, image enhancement operations such as random cropping, random brightness, and random rotation are performed; Step 3-2: Start training: Input the training samples into the SR-yolo detection model; Step 3-3: Update parameters: The loss function used includes three parts: bounding box loss, confidence loss, and classification loss; Step 3-4: Output the model: Repeat Step 3-2 and Step 3-3. When the number of training times reaches the set value, stop training and output the model with the minimum loss value; In Step 3-3, the loss function of yolov7 is used, that is: bounding box loss, confidence loss, and classification loss, which are specifically expressed by the formula as follows: Loss = a×loss obj + b×lossrect + c×loss clc; In the formula, CIOU loss is used to calculate the bounding box loss, and both the confidence loss loss rect and the classification loss loss clc are calculated using BCE loss, that is, calculated using the binary cross-entropy loss with log; For the improved yolo target recognition network, its main body adopts an improved yolov7 network architecture, including Input, Backbone, and Head. After passing the ultra-high-definition image through four CBS modules, it is sent to the ConvHB module with dilated recursive convolution.
2. The small target detection method based on the super-resolution YOLO network according to claim 1, characterized in that: Step 1 includes the following steps; Step 1-1: Image acquisition: Place the camera at a specific position, and then obtain small object images through the camera, that is, images with small target objects; Step 1-2: Image cutting: According to the captured images with small target objects, use bilinear interpolation to fill all images into small squares centered on each pixel point in the image, that is, perform expansion processing on each pixel of the picture; Step 1-3: Sample set production: Manually add pixel-level labels to the small square images with small targets in the previous step to produce a small target sample set.
Citation Information
Patent Citations
Infrared image super-resolution and small target detection method
CN113222824A
Lightweight small target detection method
CN115346064A
Lightweight small target detection method in combination with attention mechanism
CN113065558A
Small target detection method based on attention mechanism
CN114202672A
Cited By
Distribution network power transmission line insulator damage detection method based on YOLOv8 improvement
CN120599276A