Small target detection method and system based on improved YOLOv8 structure

By improving the small object detection method of YOLOv8 structure, combining the scale sequence feature fusion module and soft non-maximum suppression algorithm, the accuracy and real-time problems of small object detection in the existing technology are solved, and the detection performance in complex backgrounds is improved, and it is suitable for multiple practical application scenarios.

CN120356060AActive Publication Date: 2025-07-22HUAZHONG NORMAL UNIV

Patent Information

Application Number
CN202510413963.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing target detection methods have shortcomings in small target detection accuracy, scale adaptability and real-time performance. Especially in dense or complex backgrounds, it is easy to miss detection and misdelete small targets, which are difficult to meet the needs of smart agriculture, medical diagnosis, security monitoring, autonomous driving and industrial testing.

Method used

The small object detection method with improved YOLOv8 structure is adopted, and the multi-scale feature fusion capability is enhanced by introducing a scale sequence feature fusion module and a three-feature encoder module, and the confidence of the candidate target box is adjusted using a soft non-maximum suppression algorithm to avoid accidentally deleting the real target.

Benefits of technology

It significantly improves the recognition ability and detection recall rate of small targets, improves the detection accuracy and robustness in dense scenarios, and is suitable for application scenarios such as drone aerial imagery, remote sensing monitoring, traffic management and agricultural plant protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356060A_ABST
    Figure CN120356060A_ABST
Patent Text Reader

Abstract

The invention provides a small target detection method and system based on an improved YOLOv8 structure, and the method comprises the following steps: carrying out the preprocessing of a to-be-detected image, so as to obtain an input standardized image; inputting the standardized image into a YOLOv8 backbone network, and extracting feature maps of different scales; the feature map is input to a YOLOv8 neck network for feature fusion, the neck network introduces an ASF mechanism, a scale sequence feature fusion module and a three-feature encoder module are integrated, cross-scale enhancement and integration are performed on features of different resolutions, and a multi-scale fusion feature map is generated; constructing P2, P3, P4 and P5 detection layers based on the fused feature map; the P2-P5 detection layer predicts candidate target frames and category confidence of the candidate target frames respectively; soft non-maximum suppression is adopted to process the candidate target frame, a linear or Gaussian attenuation function is utilized to adjust confidence, a redundant frame is suppressed, and a final detection result is output. According to the method, the recognition capability of small targets in dense and complex scenes is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image target recognition, and particularly relates to a small target detection method and system based on an improved YOLOv8 structure. Background Art

[0002] With the development of computer vision technology, target detection has been widely applied in multiple fields, including scenarios such as smart agriculture, security monitoring, autonomous driving, wildlife protection, medical image analysis, and industrial quality inspection. In many practical applications, the targets to be detected are often small in scale, such as pest detection in agricultural plant protection, lesion recognition in medical images, small target tracking in surveillance videos, and long-distance obstacle detection in autonomous driving. However, there are still many challenges in small target detection in the existing technology.

[0003] Target detection methods can generally be divided into two-stage detection and single-stage detection methods. Two-stage target detection methods, such as R-CNN and its variants, usually first generate a large number of candidate regions and then classify each candidate region one by one. Although such methods have high detection accuracy, due to the large amount of computation, their real-time performance is poor, and it is difficult to meet the requirements of scenarios with high real-time requirements, such as the needs of security monitoring and autonomous driving applications.

[0004] Single-stage target detection methods, such as YOLO (You Only Look Once), achieve higher real-time detection by directly predicting the bounding boxes and class probabilities in the entire image. However, the detection performance of such methods for small targets is often poor, mainly manifested in the following aspects:

[0005] First, the fixed grid division strategy adopted by the YOLO method is likely to cause the loss of feature information in small target regions, making it difficult to capture the fine-grained features of small targets, resulting in prominent problems of missed detection of small targets in dense or complex background environments.

[0006] Second, single-stage detection methods usually rely on predefined anchor boxes for target localization. These anchor boxes have fixed sizes and are difficult to adapt to small targets of different scales and densities. Especially in fields such as smart agriculture, medical images, and industrial quality inspection, the situations of diverse scales and dense distributions of small targets are relatively common, and the existing detection models cannot effectively adapt to such scenarios.

[0007] In addition, existing target detection methods generally adopt traditional non-maximum suppression (NMS) as a post-processing step to remove redundant candidate boxes. However, this method strictly suppresses overlapping detection boxes through a fixed threshold, and it is easy to misdelete real targets in small target dense scenarios, thus significantly reducing the recall rate and overall detection accuracy of small targets.

[0008] In summary, the current object detection methods still have significant deficiencies in the detection accuracy of small objects, scale adaptability, and real-time performance. These deficiencies limit their effective applications in fields such as smart agriculture, medical diagnosis, security monitoring, autonomous driving, and industrial inspection. There is an urgent need to propose a small object detection technical solution that combines real-time performance and high accuracy to address the above technical defects. Summary of the Invention

[0009] The purpose of the present invention is to solve the deficiencies existing in the above-mentioned background technology, and to provide a small object detection method and system based on an improved YOLOv8 structure, which significantly improves the recognition ability of small objects in dense and complex scenarios.

[0010] The technical solution adopted by the present invention is: a small object detection method based on an improved YOLOv8 structure, including the following steps:

[0011] Preprocess the image to be detected to obtain a standardized image;

[0012] Input the standardized image into the backbone network of YOLOv8 to extract multi-scale feature maps;

[0013] Input the feature maps into the neck network of YOLOv8 for fusion processing; the neck network introduces the ASF mechanism and integrates a scale sequence feature fusion module and a three-feature encoder module to perform multi-scale integration on features of different resolutions and output fused feature maps;

[0014] Generate P2, P3, P4, and P5 detection layers based on the fused feature maps for detecting targets of different sizes;

[0015] The P2–P5 detection layers respectively generate candidate target boxes and their class confidence scores;

[0016] Perform soft non-maximum suppression on the candidate target boxes: based on the intersection over union between the candidate target boxes, attenuate the confidence scores of the candidate target boxes and output the final detection results.

[0017] In the above technical solution, the middle-level features output by the backbone network are fused and then subjected to scale sequence integration operations, and then added and fused with the deep feature maps output by the backbone network processed by the scale sequence feature fusion module to obtain the feature maps corresponding to the resolution of the P3 detection layer.

[0018] In the above technical solution, the deep features output by the backbone network are processed by the scale sequence feature fusion module and the three-feature encoder module to obtain the feature maps corresponding to the resolution of the P4 detection layer.

[0019] In the above technical solution, after performing a scale sequence integration operation on the shallow features output by the backbone network, and the feature maps corresponding to the P3 detection layer and the P4 detection layer, the result is added and fused with the feature map corresponding to the P3 detection layer that has been successively upsampled, concatenated with the shallow features output by the backbone network, and processed by the residual module, to obtain the feature map corresponding to the P2 detection layer.

[0020] In the above technical solution, the deep features output by the backbone network are processed by the scale sequence feature fusion module and the three-feature encoder module to obtain the feature map corresponding to the P5 detection layer.

[0021] In the above technical solution, the soft non-maximum suppression processing includes: calculating the intersection over union (IoU) for the overlapping regions between multiple candidate target boxes, and linearly or Gaussianly attenuating the corresponding confidence levels according to the degree of overlap, so as to retain some neighboring small target candidate boxes and improve the detection recall rate in the small target dense scene.

[0022] In the above technical solution, the method is applicable to aerial images collected by a camera device carried by a drone, or used in remote sensing monitoring, traffic management, and small target dense detection application scenarios for agricultural plant protection.

[0023] The present invention also provides a small target detection system based on an improved structure of YOLOv8, including:

[0024] An image input module, configured to receive the image to be detected;

[0025] A backbone network module, configured to extract multi-scale features of the image;

[0026] A neck network module, introducing the ASF mechanism, integrating the scale sequence feature fusion module and the three-feature encoder module, configured to perform deep fusion on the shallow, middle, and deep features extracted by the backbone network, and generate high-quality multi-scale fusion feature maps;

[0027] A feature generation module, constructing P2, P3, P4, and P5 detection layers based on the multi-scale fusion feature maps to ensure that targets of different scales can be fully detected;

[0028] A detection head module, configured to generate candidate target boxes corresponding to the P2–P5 detection layers and their class confidence levels;

[0029] A soft non-maximum suppression module, configured to perform attenuation processing on the confidence levels of the candidate target boxes based on the intersection over union between the candidate target boxes, and output the final detection result.

[0030] In the above technical solution, the generation process of the feature map corresponding to the resolution of the P2 detection layer includes: after performing a scale sequence integration operation on the shallow features output by the backbone network, and the feature maps corresponding to the resolutions of the P3 detection layer and the P4 detection layer, adding and fusing them with the feature map corresponding to the resolution of the P3 detection layer that has been successively upsampled, concatenated with the shallow features output by the backbone network, and processed by a residual module, to obtain the feature map corresponding to the resolution of the P2 detection layer.

[0031] In the above technical solution, the soft non-maximum suppression module calculates the intersection over union between candidate target boxes, and adjusts the confidence of the candidate boxes using a linear or Gaussian decay function according to the degree of overlap, so as to improve the detection recall rate in the small target dense scenario.

[0032] The beneficial effects of the present invention are as follows: By improving YOLOv8 and introducing a scale sequence feature fusion module and a three-feature encoder module, the method of the present invention significantly enhances the network's ability to fuse multi-scale features, and further improves the accuracy of capturing and recognizing small target features. In addition, by adding a new P2 detection layer, the network has a higher-resolution feature representation, which is specifically used for small-scale target detection. The soft non-maximum suppression (Soft-NMS) processing avoids the problem of missed detection of small targets that may be caused by the hard suppression of traditional NMS, thus improving the detection recall rate and accuracy in the dense scenario.

[0033] Furthermore, by performing a scale sequence integration operation on the middle-layer features output by the backbone network and then adding and fusing them with the deep-layer feature map, the present invention can more effectively fuse the middle-layer and deep-layer semantic information, strengthen the representation ability of the middle-scale target features, further optimize the accuracy of small target detection, and reduce the missed detection rate.

[0034] Furthermore, the present invention processes the deep-layer features output by the backbone network through a scale sequence feature fusion module and a three-feature encoder module to obtain a more representative P4 detection layer, so that the network can enhance the ability to capture deep-layer semantic information and improve the robustness and accuracy of target detection in the scenarios with complex backgrounds and dense small targets.

[0035] Furthermore, by fusing the shallow features with the feature maps corresponding to the resolutions of the P3 and P4 detection layers that have been processed by scale integration, and then further adding and fusing them with the feature map corresponding to the resolution of the P3 detection layer that has been upsampled, concatenated, and processed by a residual module, the present invention obtains a feature map corresponding to the resolution of the P2 detection layer with a higher resolution, significantly enhancing the network's ability to express fine-grained spatial features. When the network detects extremely small-sized targets, it can more effectively extract and utilize fine spatial information, improving the accuracy and recall ability of small target detection.

[0036] Furthermore, the present invention defines that the deep features of the backbone network are processed by the scale sequence feature fusion module and the three-feature encoder module to obtain the feature map corresponding to the resolution of the P5 detection layer, ensuring that the network can make full use of the feature representations at higher abstraction levels, better improving the detection performance of small targets in large-range scenarios, and further enhancing the overall detection performance of the model in multi-scale target mixed scenarios.

[0037] Furthermore, the Soft-NMS of the present invention attenuates the confidence of candidate target boxes in a linear or Gaussian manner. Compared with the traditional Non-Maximum Suppression (NMS) method, it avoids over-suppression of highly overlapping target boxes, enables some real small target boxes to be retained, effectively solves the problem of missed detection caused by misdeleting real targets in dense small target scenarios, and greatly improves the detection recall rate in small target dense scenarios.

[0038] Furthermore, the present invention effectively broadens the practical application scope of the detection method by explicitly applying the method of the present invention to small target dense detection scenarios such as drone aerial images, remote sensing monitoring, traffic management, and agricultural plant protection, enabling the technology of the present invention to be more widely applied to multiple important fields requiring real-time and high-precision small target detection, and having good practical value and application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a schematic flow chart of the method of the present invention;

[0040] Figure 2 is a schematic diagram of the improved YOLOv8 network structure of the present invention;

[0041] Figure 3 is a schematic diagram of the experimental effect using the prior art;

[0042] Figure 4 is a schematic diagram of the experimental effect using the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, which are convenient for clearly understanding the present invention, but they do not constitute a limitation to the present invention.

[0044] Embodiment 1

[0045] As Figure 1 shown, a small target detection method based on an improved YOLOv8 structure of the present invention includes the following steps:

[0046] A. Preprocess the image to be detected to obtain a standardized image;

[0047] B. Input the standardized image into the backbone network of YOLOv8 to extract multi-scale feature maps including shallow, middle, and deep features;

[0048] C. Input the multi-scale feature maps into the neck network for feature fusion. The neck network introduces the ASF mechanism and integrates a Scale Sequence Feature Fusion (SSF) module and a Triple Feature Encoding (TFE) module. The SSF module is used to normalize, upsample, and splice and fuse features from different scales, and the TFE module is used to encode the fused features to form a multi-scale fusion feature map;

[0049] D. Generate multi-resolution detection layers including P2, P3, P4, and P5 based on the multi-scale fusion feature map, where:

[0050] The feature map corresponding to the P3 detection layer resolution is obtained by adding the middle layer features output by the backbone network after scale sequence integration operation to the deep layer feature map processed by the SSF module;

[0051] The feature map corresponding to the P4 detection layer resolution is obtained by processing the deep layer features output by the backbone network through the SSF module and the TFE module;

[0052] The feature map corresponding to the P2 detection layer resolution is obtained by first performing a scale sequence integration operation on the shallow layer features output by the backbone network, the feature maps corresponding to the P3 and P4 detection layer resolutions, and then adding and fusing them with the feature map corresponding to the P3 detection layer resolution that has been upsampled, spliced with the shallow layer features output by the backbone network, and processed by the residual module;

[0053] The feature map corresponding to the P5 detection layer resolution is obtained by processing the deep layer features output by the backbone network through the SSF module and the TFE module;

[0054] E. Input the feature maps corresponding to the P2 to P5 detection layer resolutions into the corresponding detection heads respectively to generate candidate target boxes and their class confidence scores;

[0055] F. Perform Soft-Non-Maximum Suppression (Soft-NMS) processing on the candidate target boxes. According to the Intersection over Union (IoU) between the candidate boxes, use a linear or Gaussian decay function to adjust the confidence scores, suppress redundant boxes, and at the same time improve the detection recall rate in the small target dense scene, and output the final detection result. The specific process includes:

[0056] Calculate the Intersection over Union (IoU) between the candidate target boxes;

[0057] When the IoU between the candidate target boxes is less than a predetermined threshold, keep their confidence scores unchanged;

[0058] When the IoU between candidate target boxes is greater than or equal to a predetermined threshold, the confidence of the candidate boxes is attenuated according to a linear or Gaussian attenuation function, and the final detection result is output.

[0059] As Figure 2 shown, the backbone part of YOLOv8 successively includes multiple layers of Conv, several C2f modules, and an SPPF module from top to bottom. The output of the upper layer module is used as the input of the lower layer module, specifically as follows:

[0060] The first layer of Conv performs preliminary convolutional feature extraction on an image with an input size of 640×640×3, including basic spatial receptive field expansion and channel transformation. It outputs a low-level feature map, which usually retains preliminary visual information such as edges and corners, preparing for subsequent layers.

[0061] The second layer of Conv further stacks convolutional operations to enhance local feature representation or appropriately adjust channels / resolution, extracting more fine-grained texture information. It outputs a more abundant local feature map compared to the first layer, but still at a relatively high resolution and not significantly downsampled.

[0062] The third layer of C2f module combines the residual idea with partial feature reuse (Cross Stage). It can divide the input features into several branches, perform convolutional processing on each branch separately, and then splice or sum the branch results, thereby enhancing the feature expression ability and reducing redundant calculations. It outputs a shallow feature map with a relatively high resolution and more details retained, which is crucial for detecting small targets.

[0063] The fourth layer of Conv performs convolutional operations again to extract higher-level patterns on the basis of shallow features or further scale the resolution. It outputs a feature map after deepening the shallow features, serving as the input for the subsequent middle layer C2f module.

[0064] The fifth layer of C2f, based on the existing convolutional output, extracts more semantic features through multi-branch parallel convolution and residual fusion. It outputs the first middle layer feature map, with a slightly reduced resolution compared to the shallow layer and containing more abstract semantic information.

[0065] The sixth layer of Conv continues to perform convolutional transformation on the middle layer features, which may include downsampling or channel adjustment, preparing for subsequent middle layer or deeper C2f. It outputs a more compact but semantically more concentrated feature map.

[0066] The seventh layer of C2f uses the multi-branch convolution and fusion strategy again to further refine the details and semantic information in the middle layer features; it outputs the second middle layer feature map, which is more in-depth semantically but still retains certain spatial details.

[0067] After continuous downsampling or convolution processing in the eighth - layer Conv, the features are taken to deeper levels, usually significantly reducing the resolution while retaining more concentrated semantic information, serving as the input for the final deep - layer C2f.

[0068] The ninth - layer C2f performs multi - branch convolution and residual fusion based on the existing features, further refining the high - order features. The outputs all take the middle - layer features, which can continue to perform multi - scale fusion with the shallow - layer / deep - layer features in the Neck.

[0069] The tenth - layer SPPF (Spatial Pyramid Pooling Fast) expands the network receptive field and captures global context by using multi - scale pooling operations (such as max - pooling with k = 5, 9, 13) on the feature map and concatenating the results. It outputs a deep - layer feature map with the lowest resolution but the strongest semantics, which can be fused with other feature layers in the subsequent Neck structure, especially helpful for detecting large objects or objects in complex backgrounds.

[0070] The neck part of YOLOv8 aims to improve its ability to detect small objects in complex backgrounds (such as long - distance and cluttered backgrounds) by integrating the ASF module into YOLOv8. The ASF module includes the Scale - Sequence Feature Fusion (SSFF) module and the Triple Feature Encoder (TFE) module.

[0071] The Scale - Sequence Feature Fusion (SSFF) module consists of:

[0072] Input: Deep features from the final downsampling of the backbone network (such as the SPPF output), or the output from the previous SSFF module. If multiple SSFF modules are cascaded, the input of the latter SSFF is the output of the former SSFF, achieving segmented multiple - time deep multi - scale fusion;

[0073] Conv: Inside the SSFF, first perform 2 convolution (Conv) operations on the input features for channel adjustment, feature refinement, or activation processing. This convolution can reduce the computational burden of directly scaling and concatenating large - channel, multi - scale features and preliminarily screen target information;

[0074] ZoomCat operation: Scale the feature map after convolution processing to different degrees (Zoom), such as up / downsampling or dilated convolution, and then concatenate (Concat) in the channel / scale dimension to form three - dimensional or multi - branch features. This step realizes the "adaptive multi - scale reconstruction" of the deep features in a single path, taking into account both high - level semantics and local details;

[0075] C2f or lightweight residual fusion: The multi-scale features after ZoomCat splicing usually go through a C2f (or similar residual / convolution) module for fusion to generate the final output SSFF features. If there is a next SSFF, this output needs to be passed to the next SSFF; otherwise, it can be passed to TFE or subsequent branches.

[0076] Since the input of SSFF can come from the deep features of the backbone network or the output of the previous SSFF in sequence, the multi-scale representation ability of deep features can be enhanced through multiple iterations. By first performing convolutional processing on the features, it is possible to better filter noise, reduce channels, and reduce the dimensionality burden during ZoomCat splicing, while highlighting the key features related to small targets. The features output by SSFF can be used by the downstream TFE module or detection head (Head), providing rich global context and multi-scale information for the recognition of large targets, complex backgrounds, and small targets. SSFF normalizes, upsamples, and splices features at multiple scales, and fuses them through 3D convolutional operations. This operation cross-scale fuses global semantic information, enabling YOLOv8 to better recognize targets of different sizes, orientations, and aspect ratios, which is particularly important in aerial scenarios where targets vary greatly due to height and perspective changes.

[0077] The triple feature encoder (TFE) module consists of:

[0078] Input: The output from an SSFF module, or the output from the previous TFE. TFE modules can be stacked in multiple levels or connected in series with SSFF to form a multiple fusion pipeline; if this TFE is the first TFE, the input directly comes from a certain SSFF. This input feature may already contain some multi-scale or deep information, but still needs to be further refined and fused in TFE.

[0079] Conv (preprocessing convolution): Before the multi-branch convolution and concatenation (Concat), a convolution (Conv) operation is first performed: to initially adapt the input feature channels, resolution, or feature distribution, reduce the channel dimension conflict in subsequent parallel branch processing; extract higher-level or more aggregated feature information to facilitate more focus on key targets during three-branch parallel encoding; provide a relatively unified input basis so that subsequent branches can perform convolution in a "similar" feature space.

[0080] Three-branch parallel convolution and concatenation (Concat): TFE splits or duplicates the input feature into multiple branches, and different convolution kernel sizes, dilation rates, or lightweight attention operations are respectively used to obtain feature descriptions under large, medium, and small receptive fields, taking into account both fine-grained and global context; the branch outputs are fused in a C2f (or residual structure) after channel dimension concatenation (Concat) to generate the final output.

[0081] C2f / Residual Fusion: The concatenated features can then enter a lightweight C2f (or other residual structure) for convolutional fusion; the gradient vanishing problem is alleviated and the network's expressive ability is enhanced through the residual / parallel channel strategy to form the final TFE output. The fused features after TFE multi-branch encoding can be passed to the next TFE (if any) or sent to other branches of the Neck / detection head. This can gradually enhance the network's ability to recognize multi-scale targets, retain more details and combine deep semantics.

[0082] In this embodiment, the TFE module is introduced at the standard concatenation operation between the YOLOv8 neck networks. The TFE module selectively integrates feature maps of small, medium, and large scales, improving the limited aggregation of YOLOv8 through simple summation and concatenation. By explicitly encoding the features of each scale, the TFE module enhances the retention of fine-grained spatial details, which is crucial for recognizing tiny targets in cluttered backgrounds.

[0083] Compared with single-path convolution, multi-branch convolution can characterize the input features with multiple receptive fields at the same level, which is beneficial for simultaneously considering the semantics of large targets and the details of small targets.

[0084] By having the TFE input from the previous TFE or SSFF, the network can alternately perform "deep multi-scale construction" (SSFF) and "multi-branch encoding fusion" (TFE) to gradually refine or magnify the target recognition ability.

[0085] If multiple TFE are configured in the network design, the subsequent TFE will further refine the feature representation on the basis of the results of the previous TFE, enabling the detection head to obtain higher-quality input features, which is particularly suitable for fine positioning and recognition in scenarios with dense small targets.

[0086] SSFF is only inserted after the final downsampled features of the backbone network (or the previous SSFF), and realizes the multi-scale adaptive reconstruction of a single deep feature through "Conv + ZoomCat + residual fusion". The TFE input comes from SSFF or the previous TFE and can be stacked multiple times. Through triple parallel encoding and concatenation, the feature's ability to express targets of different sizes is enhanced. The two are usually used in conjunction in the Neck: SSFF is responsible for the multi-scale scaling and concatenation of deep features, and TFE is responsible for triple parallel encoding and fusion, forming a progressive feature enhancement to finally output high-precision target localization and classification results for the detection head.

[0087] The generation path of the P3 feature layer mainly includes two branches:

[0088] Deep branch: Processed by cascading two SSFF modules from the final deep features of the backbone network (SPPF output);

[0089] Middle-level branch: The middle-level features generated by three C2f modules output from the backbone network are processed by the ScalSeq module.

[0090] Finally, the deep branch and the middle-level branch perform an Add operation to achieve feature fusion, and output the feature map corresponding to the resolution of the P3 detection layer.

[0091] Specifically, the deep branch is: SPPF → SSFF → SSFF:

[0092] The spatial pyramid pooling (SPPF) module at the very end of the Backbone outputs a deep feature map, which contains a global receptive field and the strongest semantic information, but has a low resolution. It provides deep semantic support for subsequent modules to help detect targets in complex or large-scale background scenes.

[0093] In the first SSFF module of the Neck part, in the SSFF, first perform convolution on the deep features to adjust the number of channels and extract key semantics, reducing the redundancy of subsequent scaling and concatenation. Scale (zoom) or interpolate the convolved features by different multiples, and concatenate (concat) them in the newly added scale / channel dimension to form a feature tensor with more multi-scale semantics. The output of ZoomCat is further fused through C2f (or residual convolution) to obtain the output features of the first SSFF.

[0094] The input of the second SSFF module comes from the output features of the previous SSFF; the process is similar to the previous SSFF: first convolution (Conv), then ZoomCat concatenation; further fusion through C2f / residual; output the final feature map after multiple deep multi-scale fusions, which is the result of the deep branch. This result has richer global context and the ability to recover some details.

[0095] Specifically, the middle-level branch is: Middle-level C2f output of the backbone network → ScalSeq

[0096] After the shallow C2f in the backbone network, it will perform downsampling and several C2f operations in sequence (there are three middle-level C2fs in total) to output three middle-level feature maps, which are roughly similar in resolution and semantic abstraction degree, but have their own differences in different local regions and textures.

[0097] The ScalSeq module receives the three middle-level C2f outputs at the same time; slightly scales or aligns the middle-level features (if there are slight differences in resolution), and then stacks or concatenates them in the scale dimension or channel dimension; fuses the three-way middle-level features through convolution or lightweight attention operations, so that the output features have both more details and medium-level semantics. Output the middle-level feature map after ScalSeq fusion, which is the result of the middle-level branch.

[0098] The way of the Add operation is: Deep branch + Middle-level branch → P3

[0099] The output of the deep branch is the feature map obtained after the cascade processing of SPPF + SSFF + SSFF. It has strong semantics and low resolution, but after multiple ZoomCat and residual operations, it has certain multi-scale information.

[0100] The output of the middle branch is the feature map obtained by integrating the outputs of three middle-level C2f through ScalSeq. Its resolution is relatively higher than that of the deep features, and it has both considerable semantic and texture expressions.

[0101] After aligning the two feature maps in the same spatial dimension, perform an element-wise addition operation: retain and stack the global semantic information of the deep branch; at the same time, supplement the details and medium-scale information of the middle branch.

[0102] The feature map corresponding to the resolution of the output P3 detection layer integrates the information of the deep branch and the middle branch. It has both context awareness and retains a certain spatial resolution, and has good adaptability for detecting small and medium-scale targets.

[0103] The generation path of the P4 feature layer is: SPPF → SSFF → SSFF → TFE:

[0104] The SPPF inputs the deep feature map from the end of the backbone network. It has low resolution but contains strong semantics and global context information. Integrate the context through multi-scale pooling to improve the understanding of the large-scale background; provide deep base features for subsequent modules to construct higher-level detection feature maps.

[0105] The first SSFF module first performs a convolution (Conv) operation on the deep features output by the SPPF to adjust the number of channels or perform preliminary refinement, reduce noise and focus on target-related features. Scale (Zoom) or dilated convolution processing is performed on the convolved features in the scale dimension, and features of different resolutions are then concatenated (Concat) in the channel / new dimension to form a multi-scale fusion representation. After concatenation, perform convolution and residual integration (C2f) on the multi-scale features to output the deep fusion features of the first SSFF.

[0106] The output features of the first SSFF module continue to be input into the second SSFF to form a cascade. The second SSFF is the same as the previous SSFF: first perform Conv, then execute ZoomCat, and after multi-scale concatenation, fuse with C2f or a residual structure. By stacking multiple SSFFs, further enhance the perception ability of the deep features for small targets and compensate for the detail loss caused by low resolution.

[0107] The deep multi-scale fusion features obtained after stacking multiple SSFFs are the input for the subsequent TFE module.

[0108] The deep multi-scale fusion features from the second SSFF module already have quite a lot of global semantics and certain detailed information, but still need to be encoded in parallel with multiple branches within the TFE. Before the parallel branches, a convolution is first performed on the input features to unify the feature channels or filter out key textures, reducing redundancy between branches. The convolved features are split or copied into three parallel branches, and each branch may have different convolution kernel sizes or dilation rates:

[0109] Large receptive field branch: More suitable for capturing the semantics of larger or scattered targets;

[0110] Medium receptive field branch: Conventional 3×3 convolution, taking into account both local details and context;

[0111] Small receptive field branch: May be 1×1 convolution or a lighter convolution, highlighting the edges and textures of small targets.

[0112] The outputs of the parallel branches are concatenated in the channel dimension to obtain a feature map that combines the information of large, medium, and small receptive fields.

[0113] The concatenated multi-branch features are then convolved and integrated using C2f or a lightweight residual structure to output the final triple feature encoding result.

[0114] The feature map corresponding to the output P4 detection layer has the advantages of both deep multi-scale fusion (from SSFF) and triple parallel encoding (TFE). The generation path of the P4 detection layer undergoes multiple deep feature enhancements (double SSFF) and triple encoding (TFE), which not only strengthens the context understanding but also takes into account details in the multi-branch parallel convolution, making it suitable for detecting medium and large targets or subtle targets in a more complex background.

[0115] The generation path of the P5 feature layer is: SPPF→SSFF→SSFF→TFE→TFE:

[0116] The SPPF module at the end of the backbone network outputs a deep feature map with the lowest resolution but the strongest global receptive field and abstract semantics. Compressing the context information into the deep features through multi-scale pooling helps to identify large targets or make global judgments in a complex background, providing a highly condensed basic representation for subsequent SSFF and TFE.

[0117] The first SSFF module performs convolution on the deep features output by SPPF to filter out target key information, reduce noise, and adjust the channels or resolution. Then, the Zoom-Cat operation is performed: the convolved feature maps are scaled by different multiples (zoom) and concatenated (Concat) in the scale or channel dimension to form a multi-scale fusion representation. This enables multi-scale adaptive processing within the same deep feature. The concatenated features are then comprehensively refined through C2f or residual convolution to output the multi-scale deep features of the first SSFF.

[0118] The deep fusion features output by the first SSFF contain a certain amount of multi-scale information but can still be further enhanced. The second SSFF module is similar to the previous SSFF: first convolution (Conv), then Zoom-Cat, and finally C2f residual fusion. After stacking multiple SSFFs, the deep features gain stronger semantics and a certain ability to restore details in the multi-scale dimension.

[0119] The input to the first TFE module is the fusion feature map output by the second SSFF module. Before the parallel branches, the input features are convolved once to integrate the channel dimension and aggregate important features. The convolved features are copied / divided into three parallel paths for convolutional encoding with small, medium, and large receptive fields respectively:

[0120] Large receptive field branch: More suitable for large targets or global backgrounds;

[0121] Medium receptive field branch: Conventional 3×3 convolution balances local and global;

[0122] Small receptive field branch: 1×1 or lightweight convolution retains detailed information.

[0123] The outputs of the parallel branches are concatenated in the channel dimension to fuse multi-receptive field features. The concatenated three-way features are then convolved and residually integrated to output the result features of the first TFE.

[0124] The input to the second TFE module comes from the output features of the first TFE. Repeat the steps similar to the previous TFE: first convolution, then three-branch parallel convolution and concatenation, and finally C2f or residual fusion. By stacking multiple levels of TFE, the expressive ability for multi-size targets can be further enriched each time.

[0125] Finally, a feature map that has been deeply encoded by two TFEs is formed as the feature map corresponding to the resolution of the P5 detection layer.

[0126] After the P5 detection layer undergoes multi-scale enhancement of multiple SSFFs and triple parallel encoding of two TFE based on the deepest layer features (SPPF), it has the fusion of global semantics and multi-scale information. In the detection head (Head), it has high stability for large targets or scenes with complex backgrounds; it can also provide certain support for small targets because multiple residuals and small receptive field branches may still retain details. As the deepest detection layer in the network, P5, together with P2, P3, P4, etc., covers the detection requirements of different-sized targets, and can take into account the global context and detail description in small target detection tasks such as UAV images, thus achieving better results.

[0127] The generation path of the P2 detection layer is based on two main fusion branches:

[0128] The first branch: [P3 detection layer, P4 detection layer, shallow feature map] → ScalSeq

[0129] The second branch: P3 detection layer → Upsample → Concat → C2f

[0130] Then, the outputs of these two branches are subjected to the Add operation to finally generate the P2 detection layer.

[0131] The input features of the first branch include:

[0132] The feature map corresponding to the resolution of the P3 detection layer: the feature generated after medium-depth fusion, with both certain semantics and resolution;

[0133] The feature map corresponding to the resolution of the P4 detection layer: generated after deeper processing, with stronger semantics but lower resolution;

[0134] The first C2f output (shallow feature map) of the Backbone: the C2f in the earliest segment of the network, with the highest resolution and the most retained details.

[0135] The above three features input to the ScalSeq module each have different numbers of channels or spatial dimensions. Inside ScalSeq, usually, a minimum degree of scaling / channel adaptation is first performed to ensure dimension matching during subsequent splicing / fusion. The features at three different levels are spliced (Concat) in the newly added "scale" or channel dimension to form a "multi-scale sequence" feature. Subsequently, it can be uniformly refined through convolution or attention processing to condense shallow details, middle-level features, and deeper semantics on the same feature map. Output a multi-scale feature map that fuses the shallow feature (the first C2f output of the Backbone), P3, and P4 information as the result of the first branch.

[0136] The second branch uses the feature map corresponding to the resolution of the P3 detection layer as the input. The resolution of the feature map corresponding to the P3 detection layer is increased (such as 2x interpolation or transposed convolution) through upsampling (Upsample), making it closer to the shallow features in terms of spatial dimensions, retaining / restoring more spatial details for convenient subsequent fusion. It is possible to concatenate the upsampled feature map corresponding to the resolution of the P3 detection layer with some local features (such as shallow features or the output of the intermediate branch) to form a richer channel combination; the specific concatenation source is determined according to the network design, such as Figure 2 marked with "concat" in the figure, which may indicate channel merging with a certain shallow or intermediate layer feature. After concatenation, the C2f layer is used to further fuse the channel features and extract key information; this layer retains the key textures of small targets without significantly increasing the computational load and enhances the ability to integrate multi-scale features. The final output of the second branch is a feature map processed by Upsample + Concat + C2f, retaining the original medium-level semantics of the feature map corresponding to the resolution of the P3 detection layer and further supplementing detailed information due to the upsampling and concatenation operations.

[0137] The feature map corresponding to the resolution of the P2 detection layer is obtained by performing the Add operation on the first and second branches. After aligning the spatial resolution and the number of channels, the two-way feature maps are subjected to element-wise addition; combining the multi-scale information integrated by ScalSeq across layers with the P3 features processed by Upsample + Concat + C2f; it has both the high-resolution shallow / middle-level details and more deep-level semantics fused from P4.

[0138] The finally generated P2 detection layer has the highest resolution, which is particularly important in cluttered scenes or small target detection; due to the effective accumulation of multi-way information by the Add operation, P2 has the fine-grained features and cross-layer semantic information fusion required for localizing and identifying tiny targets. This design of multi-branch fusion + final addition demonstrates the deep reinforcement idea of SOD-YOLO for small target detection at the P2 layer, maximizing the integration of information in each layer and each operation path.

[0139] In this embodiment, a new detection layer is introduced at the P2 level of the YOLOv8 architecture, leveraging higher-resolution feature maps to capture details crucial for accurately identifying small targets.

[0140] First, the feature maps of the backbone network are upsampled to preserve high-resolution details. These upsampled features are then concatenated with the early backbone features to enrich the representation by combining spatial details and deeper abstract features. The concatenated features are processed through the C2f layer (enhanced residual block) to refine and enhance the combined feature map. This step is crucial for maintaining the integrity of spatial details while integrating information from different feature scales. Through processing by the C2f layer, the model can effectively retain fine-grained spatial information, which is essential for accurately detecting small objects in cluttered and complex backgrounds. After feature processing, the Scale Sequence (ScalSeq) module integrates features across multiple scales to ensure that small object details are retained and emphasized. Finally, the processed P2 features are combined with the existing detection heads (P3, P4, P5), enabling the model to detect objects at various scales and improving the detection accuracy of small objects.

[0141] The above enhancement measures significantly improve the model's performance in detecting small objects, making it more robust and effective in drone-based applications. By introducing the P2 detection head, the improved YOLOv8 architecture in this embodiment can better address the unique challenges of small object detection in aerial images, capture finer details, and improve the overall detection accuracy in complex backgrounds.

[0142] In this embodiment, the feature maps corresponding to the resolutions of the multi-scale detection layers (P2–P5) are respectively input into the conventional detection heads to form their corresponding detection layers. Bounding box regression and classification prediction are performed for each scale to obtain candidate target boxes and confidence scores. This step can be implemented using conventional YOLO detection heads or head structures such as FCOS / RetinaNet in the art.

[0143] To address the problem of potentially discarding detection boxes containing objects, this embodiment employs the soft non-maximum suppression (soft-NMS) algorithm. Different from traditional NMS, the soft-NMS algorithm adjusts the score of the current detection box by applying a weighting function that reduces the scores of adjacent boxes according to the degree of overlap with the highest-scoring box. The greater the degree of overlap, the faster the score decays. This method prevents boxes containing objects from being removed and avoids the situation where two similar boxes detect the same object simultaneously. The soft-NMS algorithm used in this embodiment takes into account the overlap between detection boxes, reducing the possibility of false negatives compared to traditional NMS algorithms, thereby improving the accuracy and reliability of small object detection.

[0144] The implementation process of the soft non-maximum suppression (Soft-NMS) algorithm specifically includes:

[0145] 1. Candidate box initialization

[0146] After network inference is completed, multiple candidate target boxes {B1, B2, …, Bn} and its corresponding scores {S1, S2, …, S n}. Among them, the scores usually represent the detection confidence or the classification confidence.

[0147] Store these candidate bounding boxes in a list or a priority queue and perform a preliminary sorting in descending order of scores.

[0148] 2. Select the current highest-score detection bounding box

[0149] Select the detection bounding box with the highest score from the candidate bounding box list and denote it as A, regarding it as the current "reference detection bounding box"; output the detection bounding box A to the final detection result list as a valid detection.

[0150] 3. Calculate IoU and attenuate the scores of adjacent candidate bounding boxes

[0151] For the remaining candidate bounding boxes B i (i ≠ A), calculate their intersection over union (IoU(A, B i )) one by one;

[0152] Update the scores of adjacent candidate bounding boxes according to the following linear attenuation formula:

[0153]

[0154] S i represents the score of the i-th detection bounding box; A represents the detection bounding box with the highest confidence in the region of interest; B i represents the i-th detection bounding box; IoU represents the overlapping degree of the i-th detection bounding box and A; N t represents the calibrated overlapping threshold, which can usually be selected as 0.5 - 0.6 (can also be adjusted to 0.3 - 0.4 for small target scenarios in UAV images), used to control what degree of IoU requires score attenuation; when IoU is less than this threshold, keep the original score, and when it is greater than or equal to this threshold, linearly attenuate the confidence of this bounding box.

[0155] 4. Update the candidate bounding box list

[0156] After the attenuation is completed, keep the candidate bounding boxes with updated scores in the list, but their confidences have been reduced;

[0157] If the score of a candidate bounding box is lower than a certain minimum detection threshold (which can be set according to the application scenario, such as 0.01 or 0.1) after attenuation, it can be removed from the list to reduce the computational amount of subsequent iterations.

[0158] 5. Repeat the iteration

[0159] After removing the output detection box A, the remaining candidate boxes are sorted again from high to low, and the new highest-score detection box is selected as the current reference box, repeating steps 3 and 4; the iteration continues until the candidate box list is empty or all scores are lower than the minimum detection threshold.

[0160] 6. Final Output

[0161] All candidate boxes output to the final detection result list through the above process are the target boxes after Soft-NMS screening; the difference between this process and the "removing overlapping boxes" of traditional NMS is that: Soft-NMS does not directly discard the candidate boxes with a large overlap degree with the highest-score box, but retains their potential contributions through a linear attenuation method, thus reducing the false negative situation.

[0162] In specific scenarios (such as extremely dense targets), the linear attenuation can be replaced by Gaussian attenuation:

[0163] S i ←S i ×exp(-αIoU(A, B i ) 2 )

[0164] where α is an adjustable hyperparameter. Different scenarios can adjust N t , the minimum detection threshold or the attenuation function respectively to obtain better detection effects.

[0165] To comprehensively evaluate the small target detection performance of the present invention (SOD-YOLO) for drones, this embodiment selects the VisDrone2019-DET dataset widely recognized in the academic community. This dataset aims to evaluate the target detection and visual perception capabilities under the drone platform and consists of 288 video clips (a total of 261,908 frames) taken at multiple locations and an additional 10,209 static images. The shooting equipment is cameras mounted on different types of drones, covering various environments such as cities and villages, and multi-target scenarios such as pedestrians, vehicles (cars, trucks, etc.) and bicycles. The targets are densely distributed in the images and have a huge size difference, which is extremely challenging.

[0166] The division of VisDrone2019-DET is as follows:

[0167] Training set: 6,471 images;

[0168] Validation set: 548 images;

[0169] Test set: 1,610 images.

[0170] The dataset defines a total of 10 target categories: pedestrians, crowds, bicycles, cars, vans, trucks, tricycles, covered tricycles, buses, and motorcycles. Approximately 75% of the targets are smaller than 0.001 times the image size, which further highlights the difficulty of small target detection. The distribution of target annotations also shows an obvious aggregation in the center of the image. Therefore, adopting a central cropping strategy in data augmentation can effectively improve the detection ability of the model.

[0171] During the experiment, in this embodiment, a training environment was deployed on a workstation with 1 RTX4090 GPU, and CSPDarknet53 was selected as the backbone network. The Stochastic Gradient Descent (SGD) optimizer was used during training, and the batch size was 8. In addition, this embodiment first performed 3 warm-up epochs, gradually increasing the learning rate from 0 to 0.005. The main training hyperparameters are as follows:

[0172] Weight Decay: 0.0005

[0173] Momentum: 0.937

[0174] Total Epochs: 200

[0175] Image Size: 640×640

[0176] Batch Size: 8

[0177] During the training and validation phases, this embodiment was evaluated using widely adopted object detection metrics, specifically including Intersection over Union (IoU), Precision, Recall, and mean Average Precision (mAP):

[0178] Intersection over Union (IoU): Measures the degree of overlap of the target bounding box by the ratio of the intersection to the union of the predicted region and the actual annotation region. The value range is [0,1], and the larger the value, the more accurate the localization.

[0179] Precision represents the proportion of true positive samples among the predicted positive samples.

[0180] Recall represents the proportion of actual positive samples that are correctly detected.

[0181] Mean Average Precision (mAP) calculates the area under the Precision-Recall (PR) curve at different IoU thresholds, and the mAP is obtained by taking the average of the AP values of each category:

[0182] mAP50: The average precision when IoU = 0.5;

[0183] mAP50:95: The average value when the IoU threshold increases from 0.5 to 0.95, which is more stringent and can reflect the performance of the model at higher IoUs.

[0184] This embodiment is based on YOLOv8-m as the baseline, compares multiple advanced object detectors, and conducts experiments on the VisDrone2019-DET dataset mainly composed of small objects.

[0185] Table 1 Detection results of different models on the VisDrone2019-DET validation set (val)

[0186]

[0187]

[0188] As shown in Table 1, compared with YOLOv8-m, the SOD-YOLO proposed in this embodiment shows obvious advantages in detecting small objects, specifically reflected in:

[0189] mAP50:95: SOD-YOLO reaches 0.351, an increase of 0.093 compared to 0.258 of YOLOv8-m;

[0190] mAP50: SOD-YOLO reaches 0.526, an increase of 0.090 compared to 0.436 of YOLOv8-m.

[0191] These results indicate that at more stringent IoU thresholds, SOD-YOLO can maintain higher detection accuracy and recall rate, significantly enhancing the localization ability of small objects. Even compared with other models such as YOLOv9-gelan-c and YOLOv10-l, SOD-YOLO also shows strong competitiveness. It exceeds YOLOv9-gelan-c in both mAP50:95 (0.351 vs. 0.305) and mAP50 (0.526 vs. 0.489). In addition, the number of parameters of SOD-YOLO is about 22.6M, and the FLOPs is about 94.9G, which is more lightweight and efficient compared to YOLOv7-m (36.9M parameters, 103.5G FLOPs), Edge-YOLO, etc.

[0192] Visualization comparison such as Figure 3 and Figure 4 shown, SOD-YOLO can detect more small objects and has higher bounding box accuracy than the baseline YOLOv8-m, demonstrating robustness in more complex backgrounds.

[0193] This embodiment further conducts ablation experiments (see Table 2) to investigate the contribution of each component in SOD-YOLO to performance improvement, with the baseline being YOLOv8:

[0194] Table 2 Ablation Experiment Results Diagram

[0195]

[0196]

[0197] The evaluation metrics of the baseline YOLOv8 are as follows:

[0198] FLOPs: 78.7G; mAP50:95: 0.258; mAP50: 0.436.

[0199] The evaluation metrics of introducing ASF (neck attention fusion strategy) are as follows:

[0200] FLOPs: 82.7G; mAP50:95: 0.265 (+0.007); mAP50: 0.440 (+0.004).

[0201] Although the FLOPs have increased, there is an improvement in detecting small targets.

[0202] The evaluation metrics of adding the P2 layer (high-resolution detection head) are as follows:

[0203] FLOPs: 94.9G; mAP50:95: 0.294 (+0.036); mAP50: 0.476 (+0.040).

[0204] Using the high-resolution feature map for specialized small target detection significantly improves the detection performance.

[0205] The evaluation metrics of integrating Soft-NMS are as follows: FLOPs: 94.9G (not increased further); mAP50:95: 0.352 (+0.094); mAP50: 0.526 (+0.090).

[0206] By gently attenuating the confidence of overlapping bounding boxes, the positive sample retention rate is improved, and at the same time, the false detection is reduced, resulting in a significant performance improvement.

[0207] In summary, the combination of the ASF mechanism, the P2 detection layer, and the Soft-NMS algorithm brings the most significant performance gain. In particular, the Soft-NMS algorithm achieves a large improvement from mAP50:95 = 0.294 to 0.352 with almost no increase in computational overhead.

[0208] When the SOD-YOLO model proposed in this embodiment is used to detect small targets in drone images, it shows excellent improvement in accuracy and recall. Its key improvements include:

[0209] 1. ASF mechanism: Introduce the attention scale sequence fusion strategy to focus on key regions in the neck network;

[0210] 2. P2 detection layer: Utilize high-resolution feature maps to specifically process small targets, significantly enhancing the ability to capture extremely small targets;

[0211] 3. Soft-NMS algorithm: Retain more positive samples by implementing soft confidence decay for overlapping candidate boxes, further improving the recall rate.

[0212] Experimental results on the VisDrone2019-DET dataset show that SOD-YOLO improves by 36% in mAP50:95 (from 0.258 to 0.352) and 20% in mAP50 (from 0.436 to 0.526) compared to YOLOv8-m, and the number of parameters and FLOPs remain at a reasonable level, demonstrating the effectiveness and efficiency of this method in the scenario of detecting dense small targets in drones.

[0213] Example 2

[0214] The present invention also provides a small target detection system based on the improved structure of YOLOv8, including:

[0215] An image input module for receiving the image to be detected;

[0216] A backbone network module for extracting multi-scale features of the image;

[0217] A neck network module that introduces the ASF mechanism and integrates a scale sequence feature fusion module and a three-feature encoder module for fusing shallow, middle, and deep features from the backbone network to generate a multi-scale fusion feature map;

[0218] A feature generation module for generating multi-resolution detection feature maps including P2, P3, P4, and P5 based on the fusion feature map;

[0219] A detection head module for inputting the P2–P5 detection layers into the corresponding detection heads respectively to generate candidate target boxes and their class confidence scores;

[0220] A soft non-maximum suppression module for performing attenuation processing on the confidence scores of the candidate target boxes and outputting the final detection results.

[0221] Example 3

[0222] The present invention provides a computer-readable storage medium with a computer program stored thereon, and when the computer program is executed by a processor, it implements the small target detection method based on the improved structure of YOLOv8 described in the above technical solution.

[0223] Example 4

[0224] The present invention provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the small target detection method based on the improved structure of YOLOv8 described in the above technical solution.

[0225] The SOD-YOLO detection method and system provided by the present invention have significant advantages in identifying small targets, and are particularly suitable for the following typical task scenarios executed by unmanned aerial vehicle (UAV) platforms:

[0226] Urban traffic monitoring: In the urban central area or highway intersections, UAVs can monitor vehicle and pedestrian flows through real-time aerial images. When the target sizes of vehicles or pedestrians are small and densely distributed, conventional detection methods are prone to target missed detection and false detection. The integrated ASF mechanism and P2 detection layer of the present invention can effectively meet the small target detection requirements in complex backgrounds and high-density environments, and improve the recognition and response efficiency of the traffic monitoring system for congestion conditions and accident scenes.

[0227] Public safety and emergency rescue: UAVs can be used to search for people or trapped targets in large-scale events or disaster sites (such as earthquakes, floods, etc.). When the target has extremely low resolution in long-distance aerial images and the scene has a high degree of clutter, Soft-NMS can retain more valid candidates when identifying highly overlapping bounding boxes, reducing the false deletion of potential targets, thereby improving the search efficiency and reliability for key areas.

[0228] Agriculture and forestry monitoring: When conducting UAV cruises in areas such as farmlands, orchards, and forest farms, it is often necessary to monitor small pests, animals, or disease spots. When the present invention detects extremely small and irregularly shaped targets, the P2 detection layer can capture rich details, which greatly promotes the early detection and precise positioning of pests and diseases, and assists in the precise management of modern agriculture.

[0229] Power line inspection and infrastructure monitoring: For structures such as high-altitude transmission lines, wind turbine blades, and bridges, UAVs are commonly used for inspection and shooting. The detection objects may be tiny cracks, loose components, or metal corrosion points. The present invention can retain details in high-resolution feature maps, has advantages in detecting defects that are far away and small in size, and focuses on key areas through ASF to reduce background clutter interference.

[0230] Environment and wildlife protection: When UAVs take pictures in nature reserves or wild environments, the scenes are complex and the vegetation is dense, and wild animals are small in size and easy to hide. SOD-YOLO can improve the detection rate of small animals through multi-layer feature fusion, and combine with Soft-NMS to avoid missed detection when adjacent targets overlap, which is of great significance for endangered species monitoring and ecological patrols.

[0231] Logistics and warehousing management: When using drones to detect goods, small packages or labels in indoor or outdoor warehouses or during transportation, it is necessary to capture extremely small and insignificantly marked objects. The detection ability of the present invention for small-sized and low-contrast targets can effectively reduce missed detections and improve the efficiency and accuracy of logistics turnover.

[0232] With the continuous improvement of the imaging resolution and flight control accuracy of drones, the application of small target detection in fields such as security and defense, unmanned patrol, and smart cities will be more in-depth. The multi-scale feature fusion method and Soft-NMS technology of the present invention can also be combined with other vision algorithms, such as behavior recognition and target tracking, to achieve diverse intelligent vision applications for drones.

[0233] In summary, the present invention is not only applicable to the rapid detection of small ground targets on drone platforms, but also widely applicable to multi-scenarios such as satellite remote sensing, traffic monitoring, industrial inspection, and public safety. The combination of the ASF mechanism, P2 detection layer, and Soft-NMS in it takes into account both algorithm efficiency while improving detection accuracy, enabling it to meet the detection requirements for various harsh environments with high resolution requirements, complex background noise, a large number of small targets, or high target overlap, and has significant practical value and promotion potential.

[0234] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0235] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0236] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the function specified in one process Figure 1 one process or more processes and / or blocks Figure 1 in one block or more blocks.

[0237] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one process Figure 1 one process or more processes and / or blocks Figure 1 in one block or more blocks.

[0238] Embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.

[0239] What is not described in detail in this specification belongs to the known prior art of those skilled in the art.

Claims

1. A small target detection method based on an improved YOLOv8 structure, characterized in that, It includes the following steps: Preprocess the image to be detected to obtain a standardized image; Input the standardized image into the backbone network of YOLOv8 to extract multi-scale feature maps; Input the feature maps into the neck network of YOLOv8 for fusion processing; the neck network introduces the ASF mechanism and integrates a scale sequence feature fusion module and a three-feature encoder module to perform multi-scale integration on features of different resolutions and output fused feature maps; Generate P2, P3, P4, and P5 multi-resolution detection layers based on the fused feature maps to enhance the detection ability for targets of different scales; The P2–P5 detection layers respectively generate candidate target boxes and their class confidence scores; Perform soft non-maximum suppression on the candidate target boxes: based on the intersection over union between candidate target boxes, attenuate the confidence scores of the candidate target boxes and output the final detection results.

2. The method according to claim 1, wherein The intermediate-level features output by the backbone network are fused and then perform a scale sequence integration operation, and then are added and fused with the deep feature maps output by the backbone network processed by the scale sequence feature fusion module to obtain the feature maps corresponding to the resolution of the P3 detection layer.

3. The method according to claim 1, characterized in that, The deep features output by the backbone network are processed by the scale sequence feature fusion module and the three-feature encoder module to obtain the feature maps corresponding to the resolution of the P4 detection layer.

4. The method according to claim 1, wherein After performing a scale sequence integration operation on the shallow features output by the backbone network, the feature maps corresponding to the resolutions of the P3 and P4 detection layers, and then adding and fusing them with the feature maps corresponding to the resolution of the P3 detection layer that are successively upsampled, concatenated with the shallow features output by the backbone network, and processed by the residual module, to obtain the feature maps corresponding to the resolution of the P2 detection layer.

5. The method according to claim 1, characterized in that The deep features output by the backbone network are processed by the scale sequence feature fusion module and the three-feature encoder module to obtain the feature maps corresponding to the resolution of the P5 detection layer.

6. The method according to claim 1, characterized in that, The soft non-maximum suppression processing includes: calculating the intersection over union for the overlapping regions between multiple candidate target boxes, and linearly or Gaussianly attenuating the corresponding confidence scores according to the degree of overlap, so as to retain some neighboring small target candidate boxes and improve the detection recall rate in the small target dense scene.

7. The method according to any one of claims 1 to 6, characterized in that The method is applicable to aerial images collected by a camera device carried by a drone, or used in remote sensing monitoring, traffic management, and agricultural plant protection small target dense detection application scenarios.

8. A small target detection system based on an improved structure of YOLOv8, characterized in that, It includes: An image input module for receiving the image to be detected; A backbone network module for extracting multi-scale features of the image; A neck network module, the neck network introduces the ASF mechanism and integrates a scale sequence feature fusion module and a three-feature encoder module, for deeply fusing the shallow, intermediate-level, and deep features from the backbone network to generate multi-scale fused feature maps; A feature generation module for generating P2, P3, P4, and P5 detection layers based on the multi-scale fused feature maps to ensure that targets of different scales can be fully detected; A detection head module for generating candidate target boxes and their class confidence scores corresponding to the P2–P5 detection layers; A soft non-maximum suppression module for attenuating the confidence scores of the candidate target boxes based on the intersection over union between candidate target boxes and outputting the final detection results.

9. The system according to claim 8, wherein The generation process of the feature map corresponding to the resolution of the P2 detection layer includes: after performing a scale sequence integration operation on the shallow features output by the backbone network, the feature maps corresponding to the resolutions of the P3 detection layer and the P4 detection layer, and adding them to the feature map corresponding to the resolution of the P3 detection layer that has been upsampled, concatenated with the shallow features output by the backbone network, and processed by a residual module, a summation fusion operation is performed to obtain the feature map corresponding to the resolution of the P2 detection layer.

10. The system according to claim 8, characterized in that The soft non-maximum suppression module calculates the intersection over union between candidate target boxes and adjusts the confidence of candidate boxes using a linear or Gaussian decay function based on the degree of overlap, thereby improving the detection recall rate in scenarios with dense small targets.

Citation Information

Patent Citations

  • Unmanned aerial vehicle image detection method based on multi-scale feature fusion and context enhancement

    CN117037004A

  • Metal smelting personnel safety equipment identification method and device based on double machine positions

    CN118823637A

  • Traffic tracking detection system for view angle of unmanned aerial vehicle

    CN118918148A

  • Multi-scale pest type detection method and device and electronic equipment

    CN119274026A

Cited By

  • Enhanced high-resolution small target detection method and system based on full-scale semantics

    CN121708488A

  • Fishbone removing method, fishbone removing device, fishbone removing equipment and fishbone removing control system

    CN121904350A

  • A fishbone removing method, device, and apparatus, and a control system

    CN121904350B

  • Industrial grade ore belt boundary identification method based on multi-frame time sequence fusion

    CN122115444A

  • Method for detecting real-time moving target on unmanned aerial vehicle in open area and medium

    CN122176286A