YOLOv10-based small target detection model optimization method and detection method
By introducing dynamic upsampler and time-frequency domain feature extraction module into the neck network of the YOLOv10 model and adding a small object detection head, the problem of degradation of small target detection accuracy in drone perception scenarios is solved, and the balanced optimization of detection accuracy, processing speed and calculation cost is achieved.
Patent Information
- Application Number
- CN202510076953.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The existing YOLOv10 model has reduced accuracy for small target detection in drone-aware scenarios, and it is difficult to achieve balanced optimization between detection accuracy, processing speed and computing cost.
By introducing a multi-scale fusion structure of the dynamic upsampler DySample and the time-frequency domain feature extraction module into the neck network of the YOLOv10 model, and adding a small object detection head, the model is optimized to adapt to small object detection.
The accuracy and processing speed of small object detection are improved, while reducing the calculation cost, achieving the balanced optimization of the model between detection accuracy, processing speed and calculation cost.
Smart Images

Figure CN120014291A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer image processing, and in particular to a small target detection model optimization method and a detection method based on YOLOv10. Background Art
[0002] In the past few years, the YOLO series has become the main tool in the field of real-time object detection due to its effective balance between computational cost and detection performance. However, with the development of applications, the reliance on non-maximum suppression (NMS) in post-processing has hindered the end-to-end deployment of YOLO and has an adverse effect on reasoning. Therefore, some scholars have further improved the performance of YOLO from two aspects: post-processing and model architecture, and proposed a YOLO-free NMS training algorithm based on consistent dual assignment, namely YOLOv10. The algorithm reduces the cost of reasoning while ensuring training performance.
[0003] Although YOLOv10 solves the problems of inference cost and lightweight deployment, it still has some defects when applied to drone perception scenarios. In drone perception scenarios, the captured images are usually small and contain unevenly distributed objects, while drone images are usually high-resolution images, which makes target detection in drone images more specific and challenging than traditional detection tasks. At the same time, since drones collect natural images during the acquisition process, natural images will suffer from various image degradation phenomena caused by adverse atmospheric conditions or unique degradation mechanisms, which introduce noise during the detection process. This leads to a decrease in the detection accuracy of YOLOv10.
[0004] Other image algorithms, such as global modeling models like Transformer, have achieved state-of-the-art performance in various image restoration tasks, but require a lot of computational resources. Although many methods have been optimized for computational efficiency, their performance degrades and they fail to achieve a favorable balance between accuracy, speed, and computational cost. Summary of the invention
[0005] The present invention provides a small target detection model optimization method and a detection method based on YOLOv10, so as to solve the problem that the existing models or algorithms cannot achieve balanced optimization among detection accuracy, processing speed and computational cost when facing natural small target images collected by drones or the like.
[0006] The present invention is achieved through the following technical solutions:
[0007] A first aspect of the present invention provides a small target detection model optimization method based on YOLOv10, comprising:
[0008] A YOLOv10 model is used as a basic model, wherein the YOLOv10 model includes a backbone network, a neck network, and a detection head part;
[0009] A multi-scale fusion structure based on a dynamic upsampler and a time-frequency domain feature extraction module is used as the neck network of the YOLOv10 model; and a small object detection head is added to the detection head part of the YOLOv10 model to obtain an improved YOLOv10 model;
[0010] The improved YOLOv10 model is trained with small target samples to obtain an optimized small target detection model.
[0011] The solution proposed in the present invention is based on the YOLOv10 model, takes advantage of the lightweight and reasoning advantages of YOLOv10 target detection, and improves the neck network and detection head of the YOLOv10 model, so that it has excellent inspection performance, processing speed and computational cost when applied to small target image detection. On the one hand, the original neck network of the YOLOv10 model is improved by fusing the multi-scale fusion structure of the dynamic upsampler DySample and the time-frequency domain feature extraction module, and by extracting features in the time domain and frequency domain, capturing spatial and frequency information at different levels, and enhancing the robustness and representation ability of the model when processing image data. At the same time, the upsampling process of the neck is optimized through the dynamic upsampling technology, reducing the computational load. Combined with the added small object detection head, the detection performance of small and medium-sized target objects is improved, and the overall model does not increase the computational load, achieving a balance between detection accuracy, processing speed and computational cost in the overall model.
[0012] In one embodiment, the backbone network of the improved YOLOv10 model includes an initial convolution layer, a first convolution layer, a first residual convolution block, a second convolution layer, a second residual convolution block, a first downsampling module, a third residual convolution block, a second downsampling module, a fourth residual convolution block, an SPPF module and a PSA module connected in sequence.
[0013] In one embodiment, the neck network of the improved YOLOv10 model includes a first feature fusion network and a second feature fusion network;
[0014] The first feature fusion network includes a first connection layer, a first fusion structure, a second connection layer, a second fusion structure, a third connection layer and a first time-frequency domain feature extraction module connected in sequence; the first fusion structure and the second fusion structure are composed of the dynamic upsampler and the time-frequency domain feature extraction module connected;
[0015] The input of the first connection layer is connected to the third residual volume of the backbone network, i.e., the block, and the output of the PSA module, the input of the second connection layer is also connected to the output of the second residual volume of the backbone network, i.e., the block, the input of the third connection layer is also connected to the output of the first residual volume of the backbone network, i.e., the block, and the output of the first time-frequency domain feature extraction module is connected to the input of the small target detection head;
[0016] The second feature fusion network includes a first Ghost convolution block, a fourth connection layer, a second time-frequency domain feature extraction module, a second Ghost convolution block, a fifth connection layer, a third time-frequency domain feature extraction module, a third downsampling module, a sixth connection layer and a C2fCIB module connected in sequence;
[0017] The input of the fourth connection layer is also connected to the output of the second fusion structure, the input of the fifth connection layer is also connected to the output of the first fusion structure, and the input of the sixth connection layer is also connected to the output of the PSA module of the backbone network; the outputs of the C2fCIB module, the second time-frequency domain feature extraction module and the third time-frequency domain feature extraction module are respectively connected to the inputs of the three detection heads of the YOLOv10 model.
[0018] In one implementation, the resolutions of the three detection heads of the YOLOv10 model are 20×20, 40×40, and 80×80, respectively, and the resolution of the small target detection head is 160×160.
[0019] In one embodiment, the time-frequency domain feature extraction module includes an input convolution layer, multiple branch computing structures, a seventh connection layer and a first output convolution layer connected in sequence; wherein the branch computing structure includes a lightweight time-domain and frequency-domain convolution structure RFCSP, multiple Ghost convolution blocks and a convolution layer, and the input of the seventh connection layer is connected to the output of each of the branch computing structures.
[0020] In one implementation, the lightweight time-domain frequency-domain convolution structure RFCSP includes an input layer, a first calculation branch, a second calculation branch, and an output layer; the input layer includes two edge detection operators Gx, Gy and a first fast Fourier transform operator;
[0021] The input of the first calculation branch is connected to the outputs of the two edge detection operators Gx and Gy, and includes a first fusion layer, a first re-parameterization module, a second fusion layer and an eighth convolution layer connected in sequence; wherein the input of the first fusion layer is connected to the outputs of the two edge detection operators Gx and Gy, and the input of the second fusion layer is connected to the output of the first re-parameterization module and the input features of the lightweight time-domain and frequency-domain convolution structure RFCSP;
[0022] The input of the second calculation branch is connected to the output of the first fast Fourier transform operator, and includes a second parameterization module, a second Fourier transform operator and a third parameterization module connected in sequence;
[0023] The output layer includes a third fusion layer and a second output convolution layer, the input of the third fusion layer is connected to the outputs of the first calculation branch and the second calculation branch, and the input of the second output convolution layer is connected to the output of the third fusion layer.
[0024] In one embodiment, before training the improved YOLOv10 model using small target samples, the method further includes:
[0025] The second residual convolution block, the third residual convolution block and the fourth residual convolution block in the backbone network of the YOLOv10 model are replaced with the time-frequency domain feature extraction module.
[0026] In one embodiment, before training the improved YOLOv10 model using small target samples, the method further includes: replacing a PAS module in a backbone network of the YOLOv10 model with a BoTNet module.
[0027] In one embodiment, the small target samples include images taken by drones and remote sensing images.
[0028] A second aspect of the present invention provides a small target detection method based on an improved YOLOv10 model, comprising: performing small target detection on a target image through an improved YOLOv10 model, wherein the improved YOLOv10 model is obtained by the small target detection model optimization method based on YOLOv10 described in any one of the first aspects of the present invention.
[0029] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0030] A multi-scale feature fusion structure based on an upsampler and time-frequency domain feature extraction module is introduced into the neck network of the YOLOv10 model, and an additional detection head specifically for small object detection is added, which effectively improves the detection accuracy of small targets. Compared with traditional models, it can better capture detailed information in low-resolution feature maps and significantly enhance the recognition ability of small objects.
[0031] The multi-scale feature fusion structure optimizes the upsampling process of the neck through dynamic upsampling technology, reduces the computational load, and improves the detection performance of small and medium-sized targets;
[0032] In the feature extraction part, the time domain and frequency domain convolution module and the reparameterization module are introduced to improve the robustness and representation ability of the model when processing image data.
[0033] The fast Fourier transform mechanism is introduced into the time-domain frequency-domain convolution module to model image degradation from the frequency domain perspective, which improves the robustness to natural image noise and degradation phenomena, reduces the computational burden of the self-attention mechanism, and enhances the efficiency of feature extraction.
[0034] The lightweight design is adopted, and the time-domain and frequency-domain convolution modules and the re-parameterization modules are introduced in combination with BoTNet to replace the PSA module in YOLOv10, which effectively reduces the computational complexity and parameter quantity of the network while ensuring the inference speed and adapting to the UAV environment with limited computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative work. In the drawings:
[0036] Figure 1 It is a flowchart of a small target detection model optimization method based on YOLOv10 according to an embodiment of the present invention;
[0037] Figure 2 1 is a schematic diagram of a network structure of an improved YOLOv10 model according to an embodiment of the present invention;
[0038] Figure 3 This is a schematic diagram of the network structure of a time-frequency domain feature extraction module RFELAN according to an embodiment of the present invention;
[0039] Figure 4 It is a schematic diagram of a model structure of an improved time-domain and frequency-domain convolution structure RFCSP according to an embodiment of the present invention;
[0040] Figure 5 1 is a schematic diagram of a backbone network of an improved YOLOv10 model according to an embodiment of the present invention;
[0041] Figure 6 It is a schematic diagram of a BoTNet structure according to an embodiment of the present invention. DETAILED DESCRIPTION
[0042] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments and drawings. The exemplary embodiments of the present invention and their description are only used to explain the present invention and are not intended to limit the present invention.
[0043] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to or inherent to other steps or units of the device.
[0044] The terms used in various embodiments of the present invention are only used for the purpose of describing specific embodiments and are not intended to limit various embodiments of the present invention. As used herein, the singular form is intended to also include the plural form, unless the context clearly indicates otherwise. Unless otherwise limited, all terms used here (including technical terms and scientific terms) have the same meaning as the meaning generally understood by those of ordinary skill in the art to which the various embodiments of the present invention belong. The terms (such as the terms defined in the dictionary generally used) will be interpreted as having the same meaning as the contextual meaning in the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning, unless clearly defined in various embodiments of the present invention.
[0045] The embodiments of the present invention provide a small target detection model optimization method and a detection method based on YOLOv10, which are suitable for detecting small target objects in natural images such as drone images and remote sensing graphics, and are beneficial to improving detection accuracy and processing speed, and reducing computing costs.
[0046] In the drone scene, lightweight design is an important indicator of the drone detection model, and its goal is to reduce the model size while maintaining good accuracy under limited computing resources. This paper takes advantage of the lightweight and inference cost advantages of the YOLOv10 model, and improves its network architecture in order to improve the detection performance of the YOLOv10 model, aiming to minimize the model size and maintain high accuracy under the constraints of limited computing resources.
[0047] From the perspective of network structure optimization, the YOLOv10 model consists of three parts: the backbone network (Backbone), the neck network (Neck) and the detection head (Head). The backbone network is responsible for feature extraction, the neck network is used for feature fusion, and the head network is used for object classification and positioning. The present invention improves the backbone network, the neck network and the detection head of the YOLOv10 model respectively, effectively improving the detection accuracy of the YOLOv10 model for small targets without adding additional computational overhead.
[0048] like Figure 1 The figure is a flow chart of a small target detection model optimization method based on YOLOv10 according to an embodiment of the present invention, including the improvement of the YOLOv10 model and the training of the improved model.
[0049] S1, Improvements of YOLOv10 model.
[0050] The improvements include: taking the YOLOv10 model as the base model, replacing the neck network of the YOLOv10 model with a multi-scale fusion structure based on a dynamic upsampler and a time-frequency domain feature extraction module; and adding a small target detection head to the detection head part of the YOLOv10 model to obtain an improved YOLOv10 model.
[0051] S2, model training: Train the improved YOLOv10 model in S1 with small target samples to obtain an optimized small target detection model.
[0052] Add technical effect analysis to improve the model?
[0053] Among them, small target samples can be obtained from drone images or remote sensing images, or the public benchmark dataset NWPU VHR-10 dedicated to drone image target detection tasks. NWPU VHR-10 contains 800 remote sensing images of 10 types of ground objects. The Labelmg tool is used to manually annotate the bounding box of the dataset to obtain small target samples for model training. The improved YOLOv10 model is trained with small target samples, and the obtained model can achieve accurate detection of small target objects.
[0054] like Figure 2 The figure shows a schematic diagram of the network structure of the improved YOLOv10 model. The improved YOLOv10 model is based on the YOLOv10 model. YOLOv10 includes three parts: the backbone network (Backbone), the neck network (Neck) and the detection head (Head). In order to solve the challenge of detecting small objects, the most effective way is to add a small object detection head small Detect to the detection head part, but this inevitably increases the computational load. Therefore, this embodiment designs a multi-scale fusion structure DR-PAN in the neck structure that integrates the dynamic upsampler DySample and the time-frequency domain feature extraction module RFELAN, and uses DR-PAN to replace the original neck network of the YOLOv10 model.
[0055] The time-frequency domain feature extraction module is a time-frequency convolutional neural network, which replaces the convolution block in the YOLOv10 model. It captures spatial and frequency information at different levels by extracting features in the time and frequency domains, enhancing the robustness and representation ability of the model when processing image data. The time domain and frequency domain focus on local modeling and global modeling, respectively, to improve the model's ability to recognize natural images in drone scenarios. At the same time, the DR-PAN structure optimizes the upsampling process of the neck through the DySample technology, reducing the computational load. Combined with the added small object detection head Small Detect, it improves the detection performance of small and medium-sized target objects, and the overall model does not increase the computational load.
[0056] The detection head part of the YOLOv10 model includes three classification detection heads YOLOv10 Detect. YOLOv10Detect adopts a lightweight architecture and consists of two Dependency convolutions with a kernel size of 3×3, followed by a 1×1 convolution. The three detection heads have different resolutions, which significantly enhances the detection capabilities in various scenarios. However, its performance in detecting small objects is still not ideal. The reason is that small objects occupy fewer pixels and are easily ignored. However, as the network depth of the convolution layer increases, the resolution of the feature map will decrease, making the model more challenging in capturing the details of small objects. In order to solve this problem, the present invention additionally introduces a small object detection head smallDetect dedicated to small object detection in the Head part. Figure 2 As shown in the figure, the improved YOLOv10 model has three YOLOv10 Detects and a small object detection head small Detect in the Head part.
[0057] The structure of small Detect is the same as YOLOv10 Detect, that is, two Dependency convolutions with a kernel size of 3×3 are connected, followed by a 1×1 convolution, but its resolution is larger to improve the accuracy of identifying small objects.
[0058] Preferably, the resolutions of the three YOLOv10 Detects are 80×80, 40×40, and 20×20, respectively, and the resolution of smallDetect is set to 160×160.
[0059] Each detection head YOLOv10 Detect in the YOLOv10 model takes the fused features extracted from the backbone and neck as input, and finally outputs a vector consisting of the regressed bounding box, bounding box confidence, and object class prediction. Before generating the final bounding box, the k-means algorithm based on the dataset is used to generate anchor boxes, and three different scales are defined to accommodate the detection of small, medium, and large objects. Similarly, for small Detect, the k-means clustering algorithm is also applied to generate anchor boxes.
[0060] The backbone network of the improved YOLOv10 model in this embodiment is the same as the YOLOv10 model, including an initial convolution layer Conv-0, a first convolution layer Conv-1, a first volume residual product block C2f-1, a second convolution layer Conv-2, a second residual convolution block C2f-2, a first downsampling module SCDown-1, a third residual convolution block C2f-3, a second downsampling module SCDown-2, a fourth residual volume block C2f-4, an SPPF module and a PSA module connected in sequence.
[0061] The neck network DR-PAN of the improved YOLOv10 model is divided into the first feature fusion network and the second feature fusion network, wherein the first feature fusion network includes a fusion structure of two upsamplers DySample and a time-frequency domain feature extraction module RFELAN, represented by DR1 and DR2 respectively, and each fusion structure is connected by a connection layer. Specifically, see Figure 2 As shown, along the data processing direction, the first feature fusion network includes a first connection layer Concat-1, a first fusion structure DR-1, a second connection layer Concat-2, a second fusion structure DR-2, a third connection layer Concat-3 and a first time-frequency domain feature extraction module RFELAN-1, which are connected in sequence, wherein the input of the first connection layer Concat-1 is connected to the third residual volume of the backbone network, i.e., block C2f-3 and the output of the PSA module, the input of the second connection layer Concat-2 is connected to the second residual convolution block C2f-2 in the backbone network and the output of the first fusion structure DR-1, the input of the third connection layer Concat-3 is connected to the first residual volume in the backbone network, i.e., block C2f-1 and the output of the second fusion structure DR-2, and the output of the first time-frequency domain feature extraction module RFELAN-1 is connected to the input of the small target detection head smal1 Detect.
[0062] The second feature fusion network includes the first Ghost convolutional blocks GhostCon v -1, the fourth connection layer Concat-4, the second time-frequency domain feature extraction module RFELAN-2, the second Ghost convolution block GhostCon v-2, the fifth connection layer Concat-5, the third time-frequency domain feature extraction module RFELAN-3, the third downsampling module SCDown-3, the sixth connection layer Concat-6 and the volume block C2fCIB, wherein the input of the fourth connection layer Concat-4 is connected to the first Ghost convolution block GhostCon v -1 and the output of the second fusion structure DR-2, the input of the fifth connection layer Concat-5 is connected to the second Ghost convolution block GhostCon v -2 and the output of the first fusion structure DR-1, the input of the sixth connection layer Concat-6 is connected to the output of the PSA module of the backbone network; the outputs of the volume block C2fCIB, the second time-frequency domain feature extraction module RFELAN-2 and the third time-frequency domain feature extraction module RFELAN-3 are respectively connected to the inputs of the three detection heads YOLOv10 Detect.
[0063] In order to identify targets of different scales, the main branch feature of the neck network of the YOLOv10 model is to improve the resolution through upsampling, thereby forming a neck structure with sequence characteristics. The traditional upsampling method relies on bilinear interpolation, which is prone to checkerboard artifacts, resulting in the loss of small semantic information of drone targets. The present invention introduces a multi-scale fusion structure DR-PAN that integrates the upsampler DySample and the time-frequency domain feature extraction module as the neck network. Compared with the original neck network structure FPN-PAN of YOLOv10, the present invention upsamples the extracted feature map (80×80) through the upsampler DySample, and fuses the upsampled features with the large feature map (160×160) of the backbone part containing rich detail information through the time-frequency domain feature extraction module to obtain new features for small targets.
[0064] DySample is different from the traditional dynamic upsampling method based on convolution kernels. It is designed from the perspective of point sampling, splitting a point into multiple points to achieve clearer edges, and implementing the upsampling process through learning sampling. The use of point sampling and learning sampling can avoid the overhead caused by calculating dynamic convolution and sub-networks and improve the performance of the model, resulting in a model with minimal computational cost. The structure of DySample is as follows:
[0065] X = grid_sample(X, S)
[0066] S=G+O
[0067] O=0.5sigmoid(linear1(X).linear2(X))
[0068] Where S represents the sampling set, G is the original sampling network, and O is the offset. X represents the input feature map, and X′ is the sampled feature map. Specifically, given an input feature map X of size C×H×Q, a linear projection of X is performed, and a point-by-point dynamic range factor is generated using an S-shaped function with a static coefficient of 0.5. The offset is then reshaped to a size of 2×sH×sW using a pixel shuffling method, where 2 represents the x and y coordinates and s is the upsampling factor. The grid sampling function resamples the input features using a sampling point generator to generate the final upsampled feature map of size C×sH×sW.
[0069] In addition, since the number of large objects in drone images and remote sensing images is very small, the multi-scale fusion structure DR-PAN proposed in the present invention abandons the network structure for generating high-level features and retains a lightweight structure for generating mid-level features, so that the model reduces the number of parameters while enhancing the focus on identifying small and medium-sized objects.
[0070] In this embodiment, the time-frequency domain feature extraction module RFELAN adopts the design principle of GELAN in YOLOv9. In order to reduce the computational complexity, RGELAN omits the blocks in GELAN and uses the newly designed lightweight time-frequency domain convolution structure RFCSP as the calculation structure of the gradient branch.
[0071] like Figure 3 The figure shows a schematic diagram of the network structure of the time-frequency domain feature extraction module RFELAN, which includes an input convolution layer Conv-3, multiple branch computing structures, a seventh connection layer Concat-7 and an output convolution layer Conv-4 connected in sequence, wherein the branch computing structure includes a lightweight time-domain and frequency-domain convolution structure RFCSP, multiple Ghost convolution blocks GhostConv and a convolution layer Conv, and the input of the seventh connection layer Concat-7 is connected to the output of each branch computing structure (RFCSP, multiple GhostConv and convolution layer Conv).
[0072] Usually, a block consists of several convolutional structures. At the same time, since the images taken in drone scenes are usually natural images, natural images will inevitably introduce noise, which will reduce the main performance of the model.
[0073] Furthermore, the present invention proposes an improved time-domain and frequency-domain convolution structure RFCSP to solve the problem that the images taken in drone scenes are usually natural images, which inevitably introduce noise and reduce the main performance of the model.
[0074] like Figure 4The figure shows the model structure diagram of the improved time-domain frequency-domain convolution structure RFCSP. The improved RFCSP can be divided into an input layer, an intermediate layer, and an output layer. The input layer includes edge detection operators Gx and Gy and the first fast Fourier transform operator FFT2D-1, where Gx and Gy are 3x3 convolution kernels, which are used to calculate the grayscale weighted difference of the neighborhood of the central pixel of the image, Gx for the vertical direction, and Gy for the horizontal direction. The middle layer includes two calculation branches, the input of the first calculation branch is connected to the outputs of the two edge detection operators Gx and Gy, including the first fusion layer, the first re-parameterization module RepConv-1, the second fusion layer, and the eighth convolution layer Conv-8 connected in sequence, wherein the input of the first fusion layer is connected to the outputs of the edge detection operators Gx and Gy, and the input of the second fusion layer is connected to the output and input features of the first re-parameterization module RepConv-1 (i.e., the input of the RFCSP module); the input of the second calculation branch is connected to the output of the first fast Fourier transform operator FFT2D-1, including the second re-parameterization module RepConv-2, the second Fourier transform operator FFT2D-2 and the third re-parameterization module RepConv-3 connected in sequence; the output layer includes the third fusion layer and the second output convolution layer Conv-9, the input of the third fusion layer is connected to the outputs of the first calculation branch and the second calculation branch, and the input of the second output convolution layer Conv-9 is connected to the input of the second fusion layer.
[0075] In the above improved RFCSP, the reparameterization module RepConv uses three branches to capture features with different receptive fields during training. It reduces the parameters and computational requirements of the model during inference through reparameterization, ensuring efficient inference speed of lightweight models. During training, RepConv includes an identity branch, a 1×1 convolution, and a 3×3 convolution. This multi-branch structure allows the network to learn features from different receptive fields and enrich the extracted feature information. During inference, RepConv reparameterizes by converting the identity branch and 1×1 convolution into 3×3 convolution, which is then fused with the 3×3 convolution branch to produce a single branch, combining features from different input branches with the parameter count of the standard 3×3 convolution. The process between the convolution layer and the normalization layer can be seen as:
[0076]
[0077] The parameters of the merged convolution are obtained as:
[0078]
[0079] The fused convolution during inference is:
[0080]
[0081] This equation shows the combined convolution parameters and the parameters used in RepConv during inference, ω Conv and ω BN Represent the parameters of the convolution process and the BN layer respectively, and b Conv and b BN The bias and fused convolution parameters for convolution and BN operations are equivalent to those used in standard convolution operations.
[0082] In one embodiment, the improvement of the YOLOv10 model further includes replacing the second residual convolution block C2N2, the third residual convolution block C2N3, and the fourth residual convolution block C2N4 in the backbone network of the YOLOv10 model with a time-frequency domain feature extraction module. Figure 5 The figure shows the backbone network diagram of the improved YOLOv10 model. In the backbone network BackBone, a time-frequency domain feature extraction module RFELAN is designed as a new feature extraction structure for natural image degradation and drone small target detection. The C2f structure of the first layer is retained to enrich the detail information. The specific structure of the time-frequency domain feature extraction module RFELAN here can be found in the description of the above embodiment and Figure 3 , Figure 4 , I will not go into details here.
[0083] Furthermore, the PSA structure in the backbone network is replaced with a BoTNet (Bottleneck Transformer) structure. BoTNet is an innovative architecture that introduces a self-attention mechanism into ResNet. The improvement of this embodiment is that a multi-head attention mechanism (MHSA) is introduced into the ResNet architecture to form a new BoTNet structure. Figure 6 shown.
[0084] The core idea of the multi-head self-attention mechanism is to use multiple attention heads to simultaneously calculate the relationship between each position and all other positions in the input sequence, thereby capturing different feature subspaces. In contrast, the self-attention mechanism essentially captures global information and reduces the depth of the model. The memory requirements of the self-attention mechanism grow quadratically in the spatial dimension, resulting in very large memory and computational overheads when processing large-sized input images. Multi-head self-attention (MHSA) replaces the last three spatial (3×3) convolutions in ResNet. ResNet reduces memory requirements by first using convolutional layers to extract low-resolution features, and then inputs these features into the multi-head attention mechanism, thereby reducing the computational burden.
[0085] An embodiment of the present invention also provides a small target detection method based on an improved YOLOv10 model, wherein the improved YOLOv10 model in any of the above embodiments is used to perform small target detection on a target image, where the target image is an image collected by a drone or a remote sensing image.
[0086] An embodiment of the present invention further provides an electronic device, which includes a processor and a memory, and the number of processors may be one or more. The memory, as a computer-readable storage medium, can be used to store software programs, computer executable programs, and modules. The processor executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory, so as to implement the small target detection model optimization method based on YOLOv10 or the small target detection method based on the improved YOLOv10 model of any of the above embodiments of the present invention.
[0087] The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0088] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the small target detection model optimization method based on YOLOv10 or the small target detection method based on the improved YOLOv10 model of any of the above embodiments of the present invention is implemented.
[0089] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it.
[0090] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0091] An embodiment of the present invention also provides a computer program product. When the computer program product runs on a computer, it enables the computer to execute the small target detection model optimization method based on YOLOv10 or the small target detection method based on the improved YOLOv10 model of any of the above embodiments of the present invention.
[0092] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A small target detection model optimization method based on YOLOv10, characterized in that: include: A YOLOv10 model is used as a basic model, wherein the YOLOv10 model includes a backbone network, a neck network, and a detection head part; A multi-scale fusion structure based on a dynamic upsampler and a time-frequency domain feature extraction module is used as the neck network of the YOLOv10 model; and a small object detection head is added to the detection head part of the YOLOv10 model to obtain an improved YOLOv10 model; The improved YOLOv10 model is trained with small target samples to obtain an optimized small target detection model.
2. The small target detection model optimization method based on YOLOv10 according to claim 1, characterized in that: The backbone network of the improved YOLOv10 model includes an initial convolution layer, a first convolution layer, a first residual convolution block, a second convolution layer, a second residual convolution block, a first downsampling module, a third residual convolution block, a second downsampling module, a fourth residual convolution block, an SPPF module and a PSA module, which are connected in sequence.
3. The small target detection model optimization method based on YOLOv10 according to claim 2, characterized in that: The neck network of the improved YOLOv10 model includes a first feature fusion network and a second feature fusion network; The first feature fusion network includes a first connection layer, a first fusion structure, a second connection layer, a second fusion structure, a third connection layer and a first time-frequency domain feature extraction module connected in sequence; The first fusion structure and the second fusion structure are composed of the dynamic upsampler and the time-frequency domain feature extraction module connected; The input of the first connection layer is connected to the third residual volume of the backbone network, i.e., the block, and the output of the PSA module, the input of the second connection layer is also connected to the output of the second residual volume of the backbone network, i.e., the block, the input of the third connection layer is also connected to the output of the first residual volume of the backbone network, i.e., the block, and the output of the first time-frequency domain feature extraction module is connected to the input of the small target detection head; The second feature fusion network includes a first Ghost convolution block, a fourth connection layer, a second time-frequency domain feature extraction module, a second Ghost convolution block, a fifth connection layer, a third time-frequency domain feature extraction module, a third downsampling module, a sixth connection layer and a C2fCIB module connected in sequence; The input of the fourth connection layer is also connected to the output of the second fusion structure, the input of the fifth connection layer is also connected to the output of the first fusion structure, and the input of the sixth connection layer is also connected to the output of the PSA module of the backbone network; the outputs of the C2fCIB module, the second time-frequency domain feature extraction module and the third time-frequency domain feature extraction module are respectively connected to the inputs of the three detection heads of the YOLOv10 model.
4. The small target detection model optimization method based on YOLOv10 according to claim 3, characterized in that: The resolutions of the three detection heads of the YOLOv10 model are 20×20, 40×40, and 80×80, respectively, and the resolution of the small target detection head is 160×160.
5. The small target detection model optimization method based on YOLOv10 according to claim 3, characterized in that: The time-frequency domain feature extraction module includes an input convolution layer, multiple branch computing structures, a seventh connection layer and a first output convolution layer connected in sequence; wherein the branch computing structure includes a lightweight time-domain and frequency-domain convolution structure RFCSP, multiple Ghost convolution blocks and a convolution layer, and the input of the seventh connection layer is connected to the output of each of the branch computing structures.
6. The small target detection model optimization method based on YOLOv10 according to claim 5, characterized in that: The lightweight time-domain frequency-domain convolution structure RFCSP includes an input layer, a first calculation branch, a second calculation branch and an output layer; the input layer includes two edge detection operators Gx, Gy and a first fast Fourier transform operator; The input of the first calculation branch is connected to the outputs of the two edge detection operators Gx and Gy, and includes a first fusion layer, a first re-parameterization module, a second fusion layer and an eighth convolution layer connected in sequence; wherein the input of the first fusion layer is connected to the outputs of the two edge detection operators Gx and Gy, and the input of the second fusion layer is connected to the output of the first re-parameterization module and the input features of the lightweight time-domain and frequency-domain convolution structure RFCSP; The input of the second calculation branch is connected to the output of the first fast Fourier transform operator, and includes a second parameterization module, a second Fourier transform operator and a third parameterization module connected in sequence; The output layer includes a third fusion layer and a second output convolution layer, the input of the third fusion layer is connected to the outputs of the first calculation branch and the second calculation branch, and the input of the second output convolution layer is connected to the output of the third fusion layer.
7. The small target detection model optimization method based on YOLOv10 according to claim 5, characterized in that: Before training the improved YOLOv10 model with small target samples, the method further includes: replacing the second residual convolution block, the third residual convolution block and the fourth residual convolution block in the backbone network of the YOLOv10 model with the time-frequency domain feature extraction module.
8. The small target detection model optimization method based on YOLOv10 according to claim 7, characterized in that: Before training the improved YOLOv10 model using small target samples, the method further includes: replacing a PAS module in a backbone network of the YOLOv10 model with a BoTNet module.
9. The small target detection model optimization method based on YOLOv10 according to claim 1, characterized in that: The small target samples include images taken by drones and remote sensing images.
10. A small target detection method based on an improved YOLOv10 model, characterized in that: include: Small target detection is performed on a target image by using an improved YOLOv10 model, wherein the improved YOLOv10 model is obtained by the small target detection model optimization method based on YOLOv10 described in any one of claims 1 to 9.
Citation Information
Cited By
Current transformer secondary terminal small target detection method and system based on improved YOLOv10
CN121999201A