Small target detection method in complex scene and related equipment
By introducing MSGEC and CSP-OmniKernel modules into the small object detection model, combined with the Yolov8 network, the balance between small object detection accuracy and model lightweight in complex scenarios is solved, and efficient small object detection performance and lightweight model are achieved.
Patent Information
- Application Number
- CN202510090496.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
AI Technical Summary
In complex scenarios, small-object detection tasks face the problem of balancing detection accuracy and model lightweighting. The existing technology is difficult to effectively control the amount of model parameters while improving detection accuracy.
Multi-scale packet efficient convolution module (MSGEC) and CSP-OmniKernel module are used for feature extraction and fusion. Combined with the Yolov8 network as the basic framework, multi-scale features are efficiently extracted through the MSGEC module, and P2 feature layer and P3 detection layer are fused in the neck network, and multi-scale feature fusion is used for multi-scale feature fusion.
A good balance between detection accuracy and model lightweighting is achieved, and the detection performance of small targets is improved. The accuracy rate, recall rate, mAP50, mAP50:95 is increased to 1.2%, 2.1%, 2.6% and 1.3%, respectively. The model size is only 6.1MB and the parameter volume is only 3.02M.
Smart Images

Figure CN119942079A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a small target detection method in complex scenes and related equipment. Background Art
[0002] In recent years, with the rapid development of deep learning technology, the research on target detection has gradually deepened, especially the detection of small targets has attracted widespread attention. In many practical applications such as security monitoring, intelligent transportation, and drone military reconnaissance, the accurate detection of small targets is crucial. However, due to the small area occupied by small targets in the image, they are easily disturbed by various factors such as background noise, image blur, and lighting, which makes the small target detection task face huge challenges.
[0003] At present, target detection algorithms based on deep learning are mainly divided into two-stage and single-stage detection algorithms. Two-stage algorithms, such as Fast R-CNN, Faster R-CNN, and Mask R-CNN, have large computational complexity and slow detection rates, making it difficult to meet real-time detection requirements. In contrast, single-stage algorithms such as SSD and the Yolo series only require a single stage to complete detection, with faster detection speeds and suitable for practical applications. Many scholars have conducted research on small target detection algorithms and have achieved a series of remarkable results. Zhang Runmei (“Zhang Runmei, Xiao Yufei, Jia Zhennan, et al. Improved YOLOv7 algorithm for target detection in complex environments from the perspective of drones [J]. Optoelectronic Engineering, 2024, 51(05):92-103”) et al. improved the Yolov7 algorithm by using a parameter-free mechanism and a reconstructed feature extraction module, which improved the detection accuracy, but did not effectively control the model parameter quantity; GONG (“Gong, Yuqi et al. “Effective Fusion Factor in FPN for Tiny Object Detection.” 2021IEEE WinterConference on Applications of Computer Vision (WACV) (2020):1159-1167”) et al. proposed a multi-scale convolution module feature fusion technology for tiny target features, and adjusted the degree of fusion of shallow and deep layers by designing the fusion factor to determine the optimal parameter value. Yin (“Zhang, Yin et al. “FFCA-YOLO for Small Object Detection in Remote Sensing Images.” IEEE Transactions on Geoscience and Remote Sensing. 62(2024): 1-15.”) et al. improved the network’s local perception ability, multi-scale feature fusion ability, and cross-channel and cross-space global correlation ability by improving the feature enhancement and fusion module and optimizing the spatial context perception module, but still failed to effectively solve the problem of increased parameters.
[0004] Although the above methods have optimized the small target detection performance in many aspects and improved the small target detection ability to a certain extent, it is still difficult to achieve a good balance between detection accuracy and model lightweight. Summary of the invention
[0005] For small target detection in complex scenes, in order to achieve a good balance between detection accuracy and model lightweight as much as possible, the present invention provides a small target detection method and related equipment in complex scenes.
[0006] In a first aspect, the present invention provides a small target detection method in a complex scene, comprising:
[0007] Acquire the image to be detected;
[0008] The remote sensing image to be tested is input into a preset target detection model, and the detection result is output; the detection result includes the target type and the corresponding confidence; the target detection model includes a backbone network, a neck network and a detection head; wherein, the MSGEC module is used in the backbone network and the neck network to extract multi-scale features; the process of feature extraction by the MSGEC module includes: splitting the channel of the input feature map into a first part, a second part and a third part, performing identity mapping on the first part, convolving the second part and the third part through convolution layers of different sizes respectively, using point-by-point convolution to restore the output after identity mapping and the output of the two convolution layers to the original number of channels, splicing and reorganizing the output of the point-by-point convolution layer to obtain an output feature map.
[0009] Furthermore, in the neck network, the features of the P2 feature layer after being processed by SPDConv are fused with the P3 detection layer.
[0010] Furthermore, in the neck network, a CSP-OmniKernel module is added before the P3 detection layer, and the CSP-OmniKernel module is used to learn feature representation from global to local and perform multi-scale feature fusion; the process of the CSP-OmniKernel module learning feature representation from global to local and performing multi-scale feature fusion includes:
[0011] The input feature map is processed by a 1×1 convolution layer to generate a first intermediate feature map; the channel of the first intermediate feature map is split into the fourth part and the fifth part, the fourth part is sent to the OmniKernel module for processing, and the second intermediate feature map is output; the second intermediate feature map is spliced with the fifth part to obtain the output feature map.
[0012] Furthermore, the number of convolution kernels of the depthwise convolution in the large branch of the OmniKernel module is set to 31.
[0013] Furthermore, the target detection model uses the Yolov8 network as a basic framework.
[0014] In a second aspect, the present invention provides a small target detection device in a complex scene, comprising:
[0015] An acquisition module, used for acquiring an image to be detected;
[0016] A detection module is used to input the remote sensing image to be tested into a preset target detection model and output a detection result; the detection result includes a target type and a corresponding confidence level; the target detection model includes a backbone network, a neck network and a detection head; wherein, the MSGEC module is used in the backbone network and the neck network to extract multi-scale features; the process of feature extraction by the MSGEC module includes: splitting the channel of the input feature map into a first part, a second part and a third part, performing identity mapping on the first part, convolving the second part and the third part through convolution layers of different sizes respectively, using point-by-point convolution to restore the output after identity mapping and the output of the two convolution layers to the original number of channels, splicing and reorganizing the output of the point-by-point convolution layer to obtain an output feature map.
[0017] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.
[0018] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the first aspect.
[0019] The beneficial effects of the present invention are:
[0020] (1) The present invention introduces a multi-scale grouping efficient convolution module MSGEC into the target detection model to efficiently extract multi-scale features while maintaining a controllable computational cost, thereby reducing the number of parameters.
[0021] (2) In order to make full use of the feature information of small targets, based on the original PANFPN, the present invention fuses the features rich in small target information in the P2 feature layer with the P3 detection layer through the SPDConv operation, and constructs a feature pyramid for small targets.
[0022] (3) The present invention utilizes the CSP concept and OmniKernel to obtain CSP-OmniKernel for feature fusion, so as to effectively learn feature representation from global to local, and further improve the detection performance of small targets.
[0023] (4) By applying the proposed method on Yolov8n and applying the obtained improved model to the VisDrone2019 dataset, the proposed method, compared with the baseline model, has improved precision, recall, mAP50, and mAP50:95 by 1.2%, 2.1%, 2.6%, and 1.3%, respectively, which is significantly better than other mainstream algorithms, indicating that the proposed method has good detection performance and generalization ability. In addition, the model size of the proposed algorithm is only 6.1MB, and the number of parameters is only 3.02M, which has the advantage of being lightweight while improving accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 A schematic diagram of a flow chart of a small target detection method in a complex scene provided by an embodiment of the present invention;
[0025] Figure 2 A structural diagram of the MSGEC module provided in an embodiment of the present invention;
[0026] Figure 3 A structural diagram of a CSP-OmniKernel module provided in an embodiment of the present invention;
[0027] Figure 4 A target detection model framework diagram formed by applying the method of the present invention in the Yolov8n network provided by the embodiment of the invention;
[0028] Figure 5 A schematic diagram of the structure of a small target detection device in a complex scene provided by an embodiment of the present invention;
[0029] Figure 6 A structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0031] Small target detection is of great significance in both military and civilian fields. Therefore, in order to solve the problems of small target aggregation and scale diversification in complex scenes, the present invention proposes a small target detection method and related equipment in complex scenes. The present invention designs a multi-scale group efficient convolution module MSGEC (Multi-Scale Group Efficient Conv) to efficiently extract multi-scale features; constructs a feature pyramid network for small targets, and constructs a CSP-OmniKernel feature fusion module based on the CSP and OmniKernel ideas to retain the information of small targets to a greater extent and improve the target detection capability.
[0032] In one embodiment, Figure 1 As shown, an embodiment of the present invention provides a small target detection method in a complex scene, comprising the following steps:
[0033] S101: Acquire an image to be detected;
[0034] S102: Input the remote sensing image to be tested into a preset target detection model and output the detection result; the detection result includes the target type and the corresponding confidence; the target detection model includes a backbone network, a neck network and a detection head; wherein, the MSGEC module is used in the backbone network and the neck network to extract multi-scale features; the process of feature extraction by the MSGEC module includes: splitting the channel of the input feature map into a first part, a second part and a third part, performing identity mapping on the first part, convolving the second part and the third part through convolution layers of different sizes respectively, using point-by-point convolution to restore the output after identity mapping and the output of the two convolution layers to the original number of channels, splicing and reorganizing the output of the point-by-point convolution layer to obtain an output feature map.
[0035] Specifically, the main task of the backbone network is to extract rich features from the input image to be detected. Generally, it can output multiple feature maps of different scales, provide basic features for the neck network and detection head after the backbone network, and serve as the basis for subsequent target positioning and classification. The main task of the neck network is to fuse the feature maps of different scales output by the backbone network. Through feature fusion, feature information at different levels interacts to enhance the expression ability of features. The main task of the detection head is to output a bounding box based on the fused feature map, give the position of the predicted target through the bounding box, identify the category of the target and give the confidence of the bounding box. The Yolo series network generally includes a backbone network, a neck network and a detection head. Therefore, the target detection model in the small target detection method of hybrid feature and multi-scale fusion provided in the embodiment of the present invention can use the Yolo series network as the basic framework.
[0036] In GhostNet, it has been proved that there is similarity in information between different channels of the intermediate feature map of the network, namely "Ghost pairs". It is speculated that the powerful feature extraction ability of CNN is positively correlated with these similar feature maps (Ghost pairs). Therefore, the present invention does not deliberately avoid the generation of Ghost pairs, but utilizes them through simple linear operations. In the present invention, redundant information is regarded as an important part of the successful model, which helps to fully understand the input data. Therefore, when designing the model, the present invention chooses not to remove these redundant feature maps, but to obtain them at a lower computational cost. To this end, the present invention designs a multi-scale grouped efficient convolution module (MSGEC), which only performs convolution operations of different scales on some channels, aiming to efficiently extract multi-scale features while controlling computational costs. The module adopts the idea of grouped convolution and uses convolution kernels of different sizes to capture multi-scale contextual information, thereby enhancing the feature expression capability.
[0037] The structure of the MSGEC module is as follows Figure 2 As shown. In an exemplary embodiment, first, the number of channels of the input feature map C is grouped based on the idea of grouped convolution, and the channels of the input feature map C are divided into three parts: C / 2, C / 4, and C / 4. The C / 2 part is used for identity mapping, and the other two C / 4 parts are grouped with 3×3 and 5×5 convolutions respectively; after multiple convolution operations, these features are completed on a single feature channel, and the information between each channel feature is independent. In order to achieve the fusion of each channel feature, the present invention uses pointwise convolution (PWConv) to adjust the number of channels of the feature map to the original number, which not only reduces the computational burden, but also improves the efficiency of feature fusion. Finally, the convolution results of the same scale are spliced and reorganized.
[0038] In the above embodiment, the parameters of the MSGEC module are:
[0039]
[0040] The parameters of the C2f module in the original yolov8n are:
[0041]
[0042] Where c is the number of convolution channels and n is the number of iterations of the module. In yolov8n, since the depth factor is 0.33, the value of n in the network is 1 or 2. It can be seen that the MSGEC method can significantly reduce the number of network parameters while ensuring accuracy. Therefore, in some exemplary embodiments, the MSGEC module can be used to replace some traditional C3 modules and C2f modules in the backbone network and the neck network.
[0043] Through the above design, the MSGEC module has achieved significant improvements in performance and efficiency. The experimental results below show that MSGEC outperforms traditional convolutional modules in feature extraction while maintaining a low cost in computing resource overhead. This module provides an effective solution for multi-scale feature extraction and has good application potential.
[0044] In one embodiment, in order to effectively extract small targets in the image to be detected, in the neck network, the features of the P2 feature layer after SPDConv processing are fused with the P3 detection layer.
[0045] Specifically, due to the downsampling operation in the target detection network, the features of small targets are lost, and small targets are slightly struggling on the normal P3, P4, and P5 detection layers. The P2 detection layer is usually located in the shallower layer of the network, and the feature map has a higher resolution, which means that small targets occupy more pixels on the feature map and have richer detail information. The traditional detection method is to add an additional P2 detection layer on the basis of the original P3, P4, and P5 detection layers to improve the detection ability of small targets. However, this approach will also bring a series of problems, such as excessive calculation after adding the P2 detection layer, and more time-consuming post-processing. To address this problem, this embodiment is based on the original PANFPN. By fusing the features of the P2 feature layer after SPDConv processing with the P3 detection layer, the detailed information of the small target on the feature map is retained without increasing the amount of calculation, forming an effective feature pyramid for small targets.
[0046] In one embodiment, in order to further improve the detection performance of small targets, a CSP-OmniKernel module is added before the P3 detection layer, and the CSP-OmniKernel module is used to learn feature representation from global to local and perform multi-scale feature fusion; the process of the CSP-OmniKernel module learning feature representation from global to local and performing multi-scale feature fusion includes: processing the input feature map through a 1×1 convolution layer to generate a first intermediate feature map; splitting the channel of the first intermediate feature map into a fourth part and a fifth part, sending the fourth part to the OmniKernel module for processing, and outputting a second intermediate feature map; the second intermediate feature map is spliced with the fifth part to obtain an output feature map.
[0047] Specifically, this embodiment is based on the CSP idea and OmniKernel to obtain a CSP-OmniKernel (CSPOK) module for feature fusion, so as to effectively learn feature representation from global to local, and ultimately improve the detection performance of small targets.
[0048] In an exemplary embodiment, the CSP-OmniKernel structure is as follows Figure 3 As shown in the figure, the processing of this module includes: first, the input features are adjusted through 1×1 convolution, and the channels of the output features are split, one quarter of which are sent to the OmniKernel module, which inputs the features into three branches, namely the local branch, the large branch, and the global branch, to enhance the multi-scale representation. The results of the three branches are added and fused, and then modulated by another 1×1 convolution. Finally, the output of the OmniKernel module is spliced and output with the other three quarters of the split channels.
[0049] In the above exemplary embodiment, the convolution kernel of the depthwise convolution in the large branch of the OmniKernel module is further adjusted from 63 to 31 to reduce the amount of calculation.
[0050] Yolov8 is an end-to-end object detection network, which is built on the basis of the historical version of the Yolo series, uses a new backbone network and anchor-free detection, and has the advantages of fast detection speed and good real-time performance. Therefore, in one embodiment, Yolov8n is used as the basic framework, and the method of the present invention is used to improve it. The overall framework is as follows Figure 4 shown.
[0051] Specifically, the embodiments of the present invention make improvements in the following three aspects: (1) Backbone network feature extraction. The multi-scale group efficient convolution module (MSGroupEfficientConv, MSGEC) of this embodiment replaces part of the C2f module in the original network, which can extract features more effectively while reducing the amount of calculation. (2) Small target feature enhancement. The small target features are lost due to downsampling in the network. In order to retain more small target information, the P2 feature layer and the P3 feature layer are fused to retain the information of the small target to a greater extent. (3) Multi-scale feature fusion. Construct a CSP-OmniKernel feature fusion module to effectively learn feature representations from global to local, and perform multi-scale feature fusion, thereby improving the detection performance of small targets.
[0052] In order to verify the effectiveness of the solution of the present invention, the present invention also provides the following comparative experiments.
[0053] 1.1 Dataset and Experimental Settings
[0054] This experiment uses the VisDrone2019 dataset for training and performance evaluation. This dataset was taken by drones of different models in 14 cities in China, including different scenes, altitudes, weather and lighting conditions, covering 10 categories such as pedestrians and cars, with a total of 10,209 static images and about 2.6 million target instances. Compared with other datasets, VisDrone2019 has more obvious differences in the number and type scale of small targets, which is of great research significance.
[0055] The experiment was conducted on Ubuntu 18.04 operating system, using NVIDIA GeForce RTX 4090 GPU, and configured with Python 3.8.10, PyTorch 2.0.0 and CUDA 11.8. Yolov8n was used as the basic framework, and the method of the present invention was used to improve it, and the improved target detection model was obtained as follows: Figure 4 The model training parameters are set as follows: batch size 8, number of training rounds 27, learning rate 0.01, weight decay coefficient 0.0005, input image size 1024 × 1024. The pre-trained weights of yolov8n are used, and the rest of the parameters remain default.
[0056] Performance evaluation indicators include precision P (Precision), recall R (Recall), and mean average precision mAP (Mean Average Precision). In target detection, TP (True Positive) is the number of correctly detected targets, FP (False Positive) is the number of incorrectly detected targets, and FN (False Negative) is the number of missed targets. The calculation formula for precision P is:
[0057]
[0058] The calculation formula of recall rate R is:
[0059]
[0060] The area enclosed by the precision-recall (PR) curve represents AP, which is calculated as follows:
[0061]
[0062] mAP is the average of all APs of different categories. mAP@0.5 refers to the average detection accuracy of all target categories when the IoU threshold is 0.5. mAP@0.5:0.95 means the average of the detection accuracy calculated from all 10 IoU thresholds (from 0.5 to 0.95) with a step size of 0.05. The formula is as follows:
[0063]
[0064] Among them, N is the number of categories in the dataset, and a higher mAP value indicates that the model performs better in the object detection task. In addition, model size and number of parameters are also used as evaluation indicators of the model.
[0065] (II) Ablation Experiment Results
[0066] In order to verify the effectiveness of the improved modules of the method of the present invention, each module was evaluated on the VisDrone2019 dataset. The evaluation results are shown in Table 1. It can be seen that after replacing the C2f module with the MSGEC module, although the precision rate is slightly reduced, the recall rate, mAP50, and mAP50:95 accuracy are significantly improved, and the model size and parameter amount are also greatly reduced. Subsequently, after adding the CSP-OK module, the accuracy of P, R, mAP50, and mAP50:95 is further improved, while the model size and parameter amount are only 6.1MB and 3.02M. This proves the effectiveness of the method of the present invention.
[0067] Table 1 Evaluation results of detection accuracy of each module of the algorithm
[0068]
[0069] (III) Comparative experimental results
[0070] Table 2 Performance comparison results of the proposed algorithm and other algorithms on the VisDrone2019 dataset
[0071]
[0072] Note: Bold font indicates the best result, and underline “_” indicates the second best result.
[0073] In order to further explore the detection performance of the algorithm, this experiment compares the detection accuracy, model size and parameter amount with other classic target detection algorithms Yolov3-tiny, Yolv5s, Yolov6n, Yolov7-tiny and FFCA-YOLO on the VisDrone2019 dataset. The results are shown in Table 2. It can be clearly seen that the algorithm of the present invention has the best detection accuracy on the dataset, and the model size and network parameters are relatively small, which is still competitive compared to the newer Yolov10n. This shows that the algorithm of the present invention can maintain good detection accuracy and speed under lightweight conditions, meeting the needs of real-time detection.
[0074] Based on the same inventive concept, Figure 5 As shown, an embodiment of the present invention further provides a small target detection device in a complex scene, including an acquisition module and a detection module.
[0075] The acquisition module is used to acquire the image to be detected; the detection module is used to input the remote sensing image to be detected into a preset target detection model and output the detection result; the detection result includes the target type and the corresponding confidence; the target detection model includes a backbone network, a neck network and a detection head; wherein, the MSGEC module is used in the backbone network and the neck network to extract multi-scale features; the process of feature extraction by the MSGEC module includes: splitting the channel of the input feature map into a first part, a second part and a third part, performing identity mapping on the first part, convolving the second part and the third part through convolution layers of different sizes respectively, using point-by-point convolution to restore the output after identity mapping and the output of the two convolution layers to the original number of channels, splicing and reorganizing the output of the point-by-point convolution layer to obtain an output feature map.
[0076] It should be noted that the small target detection device provided in the embodiment of the present invention is for implementing the above method. The specific functions thereof can be referred to the above method embodiments, which will not be described in detail here.
[0077] Figure 6 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 6 As shown, the electronic device may include: a processor (processor) 601, a communication interface (Communications Interface) 602, a memory (memory) 603 and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604. The processor 601 can call the logic instructions in the memory 603 to execute the small target detection method, which includes: obtaining an image to be detected; inputting the remote sensing image to be detected into a preset target detection model, and outputting the detection result; the detection result includes the target type and the corresponding confidence; the target detection model includes a backbone network, a neck network and a detection head; wherein, the MSGEC module is used in the backbone network and the neck network to extract multi-scale features; the process of feature extraction by the MSGEC module includes: splitting the channel of the input feature map into a first part, a second part and a third part, performing identity mapping on the first part, convolving the second part and the third part through convolution layers of different sizes respectively, using point-by-point convolution to restore the output after identity mapping and the output of the two convolution layers to the original number of channels, splicing and reorganizing the output of the point-by-point convolution layer to obtain an output feature map.
[0078] In addition, when the logic instructions in the above-mentioned memory 603 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0079] An embodiment of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the small target detection method provided by the above-mentioned method embodiments.
[0080] An embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the small target detection method provided by the above-mentioned method embodiments is implemented.
[0081] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A small target detection method in complex scenes, characterized in that: include: Acquire the image to be detected; Inputting the remote sensing image to be tested into a preset target detection model and outputting the detection result; The detection result includes the target type and the corresponding confidence; the target detection model includes a backbone network, a neck network and a detection head; wherein, the MSGEC module is used in the backbone network and the neck network to extract multi-scale features; the process of feature extraction by the MSGEC module includes: splitting the channel of the input feature map into a first part, a second part and a third part, performing identity mapping on the first part, convolving the second part and the third part through convolution layers of different sizes respectively, using point-by-point convolution to restore the output after identity mapping and the output of the two convolution layers to the original number of channels, splicing and reorganizing the output of the point-by-point convolution layer to obtain an output feature map.
2. The small target detection method in a complex scene according to claim 1, characterized in that: In the neck network, the features of the P2 feature layer after SPDConv processing are fused with the P3 detection layer.
3. The small target detection method in a complex scene according to claim 2, characterized in that: In the neck network, a CSP-OmniKernel module is added before the P3 detection layer, and the CSP-OmniKernel module is used to learn feature representation from global to local and perform multi-scale feature fusion; The process of the CSP-OmniKernel module learning feature representation from global to local and performing multi-scale feature fusion includes: The input feature map is processed by a 1×1 convolution layer to generate a first intermediate feature map; the channel of the first intermediate feature map is split into the fourth part and the fifth part, the fourth part is sent to the OmniKernel module for processing, and the second intermediate feature map is output; the second intermediate feature map is spliced with the fifth part to obtain the output feature map.
4. The small target detection method in a complex scene according to claim 3, characterized in that: The number of convolution kernels for the depthwise convolution in the large branch of the OmniKernel module is set to 31.
5. A small target detection method in a complex scene according to any one of claims 1 to 4, characterized in that: The target detection model uses the Yolov8 network as the basic framework.
6. A small target detection device in complex scenes, characterized in that: include: An acquisition module, used for acquiring an image to be detected; A detection module, used to input the remote sensing image to be detected into a preset target detection model and output a detection result; The detection result includes the target type and the corresponding confidence; the target detection model includes a backbone network, a neck network and a detection head; wherein, the MSGEC module is used in the backbone network and the neck network to extract multi-scale features; the process of feature extraction by the MSGEC module includes: splitting the channel of the input feature map into a first part, a second part and a third part, performing identity mapping on the first part, convolving the second part and the third part through convolution layers of different sizes respectively, using point-by-point convolution to restore the output after identity mapping and the output of the two convolution layers to the original number of channels, splicing and reorganizing the output of the point-by-point convolution layer to obtain an output feature map.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Cited By
Forge piece surface defect detection method based on YOLO model
CN121481935A