Lightweight target detection method and unmanned aerial vehicle image target detection method
By building a lightweight object detection model including backbone network, feature fusion network and detection network, the existing lightweight object detection solutions are solved in terms of reliability and accuracy. Especially when small object detection, more efficient feature extraction and fusion are achieved, improving detection accuracy and reliability.
Patent Information
- Application Number
- CN202510605413.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
While reducing the complexity of the model, the existing lightweight object detection schemes lead to a reduction in reliability and accuracy of the detection scheme, especially when small object detection is weak.
A lightweight object detection initial model including backbone network, feature fusion network and detection network is adopted to extract and fusion features through 2D convolutional network, feature extraction and fusion of spatial pyramid pooled fusion network, and combined with convolution and deformable convolution modules, the capture and weighted fusion of multi-scale spatial features are achieved.
It improves the reliability and accuracy of lightweight object detection, especially when small object detection, it can better capture subtle features, and enhances detection capabilities in complex scenarios.
Smart Images

Figure CN120107572A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and in particular relates to a lightweight target detection method and a drone image target detection method. Background Art
[0002] With the development of economy and technology, target detection technology has been widely used in people's production and life. Target detection technology can identify and locate target objects in images in real time, thereby greatly improving the perception ability and intelligence level of smart devices.
[0003] In recent years, with the rapid development of edge devices and related technologies, lightweight object detection solutions have become the focus of researchers. At present, although researchers have proposed many lightweight object detection solutions, these solutions all have the following defects: Existing lightweight target detection solutions often simplify traditional target detection solutions. Although such simplification reduces the complexity of the model, it also reduces the reliability and accuracy of the detection solution. In addition, the existing lightweight target detection solutions have relatively weak feature extraction capabilities during small target detection and are unable to capture the subtle features of small targets.
[0004] Therefore, existing lightweight object detection solutions all have the defects of poor reliability and accuracy. Summary of the invention
[0005] One of the purposes of the present invention is to provide a lightweight target detection method with high reliability and good accuracy.
[0006] A second object of the present invention is to provide a drone image target detection method that includes the lightweight target detection method.
[0007] The lightweight target detection method provided by the present invention comprises the following steps: S1. Obtaining existing target detection image data information; S2. Preprocessing the data information obtained in step S1 to construct a training data set; S3. Construct a lightweight target detection initial model including a backbone network, a feature fusion network, and a detection network; Among them, a backbone network is constructed based on a 2D convolutional network, a feature extraction network and a spatial pyramid pooling fusion network to extract feature map information of input image data at different scales; a feature fusion network is constructed based on the feature extraction network to fuse the feature map information of different scales and obtain fusion features; a detection network is constructed based on the convolutional network to extract and predict the location information and category information of the target according to the obtained fusion features; The feature extraction network uses convolution and diversified branch modules for hierarchical parallel processing to extract multi-branch features of the input image from the perspective of the spatial domain, while ensuring the diversity and efficiency of the extracted features; the feature extraction network also uses convolution and deformable convolution to extract features from the perspective of the geometric domain to ensure the diversity and pertinence of feature fusion; The detection network captures multi-scale spatial features through several convolutions of different scales in parallel, and performs weighted fusion of the captured multi-scale spatial features to achieve target detection; S4. Using the training data set constructed in step S2, the lightweight target detection initial model constructed in step S3 is trained to obtain a lightweight target detection model; S5. Using the lightweight target detection model obtained in step S4, actual lightweight target detection is performed.
[0008] The backbone network constructed includes the following contents: The backbone network includes an input 2D convolutional network, a first 2D convolutional network, a first feature extraction network, a second 2D convolutional network, a second feature extraction network, a third 2D convolutional network, a third feature extraction network, a fourth 2D convolutional network, a fourth feature extraction network and a spatial pyramid pooling fusion network, which are connected in series in sequence; Among them, the structures of the first feature extraction network to the fourth feature extraction network are the same; the feature extraction networks are all based on diversified branch modules, convolution modules and deformable convolution modules to achieve efficient feature extraction; The backbone network is used to extract feature map information of the input image data at different scales, and finally output feature maps of three different sizes; the feature map includes the third feature map , the fourth characteristic graph And the fifth characteristic diagram .
[0009] The constructed feature fusion network specifically includes the following contents: The input of the feature fusion network is the third feature map output by the backbone network , the fourth characteristic graph And the fifth characteristic diagram ; Fifth characteristic diagram And the fourth characteristic diagram After the connection is performed, feature extraction is performed through a fifth feature extraction network to obtain a fifth extracted feature; Fourth characteristic diagram , the third characteristic map After being connected with the fifth extracted feature, feature extraction is performed through a sixth feature extraction network to obtain a sixth extracted feature; The third characteristic diagram After being connected with the sixth extracted feature, feature extraction is performed through a ninth feature extraction network to obtain a ninth extracted feature; After the sixth extracted feature and the ninth extracted feature are connected, feature extraction is performed through an eighth feature extraction network to obtain an eighth extracted feature; After the fifth extracted feature, the sixth extracted feature and the eighth extracted feature are connected, feature extraction is performed through a seventh feature extraction network to obtain a seventh extracted feature; The feature fusion network is used to fuse the feature map information of the three different scales and output the fusion features of three different scales; the fusion features include the third fusion feature , the fourth fusion feature and the fifth fusion feature , and the ninth extracted feature is used as the third fusion feature , the eighth extracted feature is used as the fourth fusion feature , the seventh extracted feature is used as the fifth fusion feature ; The third fusion feature With the third characteristic diagram Corresponding to the fourth fusion feature With the fourth characteristic diagram Correspondingly, the fifth fusion feature With the fifth characteristic diagram correspond; The structures of the fifth feature extraction network to the ninth feature extraction network are the same, and are the same as the structures of the first feature extraction network to the fourth feature extraction network.
[0010] The processing process of the first feature extraction network to the ninth feature extraction network specifically includes the following contents: The input feature map is divided into two identical sub-feature maps along the channel dimension, and the number of channels of each sub-feature map is half of the input feature map; The first sub-feature map obtained is passed through the first The convolutional layer is processed, activated by the activation function, and then by the second The convolutional layer is processed to generate an enhanced first local feature representation; Processing the obtained second sub-feature map through a diversified branch module to generate a diversified feature representation; After the first sub-feature map, the first local feature representation, the second sub-feature map and the diversified feature representation are concatenated along the channel dimension, they are sequentially passed through The convolutional layer and deformable convolution module are processed to obtain the output of the feature extraction network.
[0011] The constructed detection network includes the following contents: The input of the detection network is the third fusion feature , the fourth fusion feature and the fifth fusion feature ; The detection network includes a first detection head, a second detection head, a third detection head and a screening layer; the first detection head is used to input the third fusion feature The second detection head is used to detect and output the location information and category information of the corresponding target; the fourth fusion feature of the input The third detection head is used to detect the fifth fusion feature of the input Perform detection and output the location information and category information of the corresponding target; the screening layer is used to screen the location information and category information of the target output by the first detection head, the second detection head and the third detection head to output the final target detection result; The structures of the first detection head, the second detection head and the third detection head are the same. The detection head is constructed based on the convolutional network. The screening layer uses the non-maximum suppression method to screen the honor prediction box and retain the detection results with confidence higher than the set value. Finally, the detection box and the category label are superimposed on the original image to complete the result display.
[0012] The processing of the detection head includes the following steps: The input feature map is divided into a first sub-feature map, a second sub-feature map, a third sub-feature map and a fourth sub-feature map along the channel dimension, and the first sub-feature map, the second sub-feature map, the third sub-feature map and the fourth sub-feature map are ensured to have the same number of channels; The first sub-feature map is passed through The convolution kernel is processed; the second sub-feature map is processed by The convolution kernel is processed; the third sub-feature map is processed by The convolution kernel is processed; the fourth sub-feature map is processed by The outputs of the four convolution kernels are concatenated along the channel dimension to obtain a new feature map, and then passed through 2 The convolutional layer is used to process the input feature map to obtain the location information and category information of the target corresponding to the input feature map.
[0013] The present invention also provides a method for detecting targets in drone images including the lightweight target detection method, comprising the following steps: A. Obtain the image data information taken by the drone during flight in real time; B. Using the lightweight target detection method, perform target detection on the image data obtained in step A; C. Based on the target detection results obtained in step B, complete the target detection of the drone's shooting picture.
[0014] The lightweight target detection method and the drone image target detection method provided by the present invention construct a lightweight target detection model based on a convolutional network, a spatial pyramid pooling fusion network and a feature extraction network. The model can not only ensure the efficiency of feature extraction and feature fusion, but also has good detection capabilities in complex scenes. Therefore, the present invention can not only realize lightweight target detection, but also has higher reliability, better accuracy and better applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 Schematic diagram of the method flow of the target detection method of the present invention.
[0016] Figure 2 It is a structural schematic diagram of the feature extraction network of the target detection method of the present invention.
[0017] Figure 3 It is a schematic diagram of the structure of the detection head of the target detection method of the present invention.
[0018] Figure 4 Schematic diagram comparing the detection results of the target detection method of the present invention and the existing target detection method; wherein, Figure 4 (a) is the original image, Figure 4 (b) is a schematic diagram of the results of the YOLOV5 model. Figure 4 (c) is a schematic diagram of the results of the YOLOV7 model. Figure 4 (d) is a schematic diagram of the results of the target detection method of the present invention.
[0019] Figure 5 Schematic diagram of heat map comparison of ablation experiment of target detection method of the present invention, wherein: Figure 5 (a) is the original image, Figure 5 (b) is the heat map when the feature extraction network is not used. Figure 5 (c) is the heat map after using the feature extraction network.
[0020] Figure 6 The figure is a schematic diagram of the method flow of the drone image target detection method of the present invention. DETAILED DESCRIPTION
[0021] like Figure 1 The method flow diagram of the target detection method of the present invention is shown as follows: The lightweight target detection method disclosed in the present invention comprises the following steps: S1. Obtain existing target detection image data information.
[0022] S2. Preprocess the data information obtained in step S1 to construct a training data set.
[0023] S3. Construct a lightweight target detection initial model including a backbone network, a feature fusion network, and a detection network; Among them, a backbone network is constructed based on a 2D convolutional network, a feature extraction network and a spatial pyramid pooling fusion network to extract feature map information of input image data at different scales; a feature fusion network is constructed based on the feature extraction network to fuse the feature map information of different scales and obtain fusion features; a detection network is constructed based on the convolutional network to extract and predict the location information and category information of the target according to the obtained fusion features; The feature extraction network uses convolution and diversified branch modules for hierarchical parallel processing to extract multi-branch features of the input image from the perspective of the spatial domain, while ensuring the diversity and efficiency of the extracted features; the feature extraction network also uses convolution and deformable convolution to extract features from the perspective of the geometric domain to ensure the diversity and pertinence of feature fusion; The detection network captures multi-scale spatial features through several convolutions of different scales in parallel, and weightedly fuses the captured multi-scale spatial features to achieve target detection.
[0024] The backbone network constructed specifically includes the following contents: The backbone network includes an input 2D convolutional network, a first 2D convolutional network, a first feature extraction network, a second 2D convolutional network, a second feature extraction network, a third 2D convolutional network, a third feature extraction network, a fourth 2D convolutional network, a fourth feature extraction network and a spatial pyramid pooling fusion network, which are connected in series in sequence; Among them, the structures of the first feature extraction network to the fourth feature extraction network are the same; the feature extraction networks are all based on diversified branch modules, convolution modules and deformable convolution modules to achieve efficient feature extraction; The backbone network is used to extract feature map information of the input image data at different scales, and finally output feature maps of three different sizes; the feature map includes the third feature map , the fourth characteristic graph And the fifth characteristic diagram ; Among them, the first feature extraction network to the fourth feature extraction network can efficiently extract and learn features from multi-scale feature maps, while enhancing the fine-grained semantic information in high-resolution feature maps and the high-level semantic information in low-resolution feature maps; at the same time, the spatial pyramid pooling fusion network applies pooling operations of different sizes to the input feature maps, captures multi-scale spatial information, improves the spatial information expression ability of the feature maps, and lays the foundation for subsequent feature fusion.
[0025] The constructed feature fusion network specifically includes the following contents: The input of the feature fusion network is the third feature map output by the backbone network , the fourth characteristic graph And the fifth characteristic diagram ; Fifth characteristic diagram And the fourth characteristic diagram After the connection is performed, feature extraction is performed through a fifth feature extraction network to obtain a fifth extracted feature; Fourth characteristic diagram , the third characteristic map After being connected with the fifth extracted feature, feature extraction is performed through a sixth feature extraction network to obtain a sixth extracted feature; The third characteristic diagram After being connected with the sixth extracted feature, feature extraction is performed through a ninth feature extraction network to obtain a ninth extracted feature; After the sixth extracted feature and the ninth extracted feature are connected, feature extraction is performed through an eighth feature extraction network to obtain an eighth extracted feature; After the fifth extracted feature, the sixth extracted feature and the eighth extracted feature are connected, feature extraction is performed through a seventh feature extraction network to obtain a seventh extracted feature; The feature fusion network is used to fuse the feature map information of the three different scales and output the fusion features of three different scales; the fusion features include the third fusion feature , the fourth fusion feature and the fifth fusion feature , and the ninth extracted feature is used as the third fusion feature , the eighth extracted feature is used as the fourth fusion feature , the seventh extracted feature is used as the fifth fusion feature ; The third fusion feature With the third characteristic diagram Corresponding to the fourth fusion feature With the fourth characteristic diagram Correspondingly, the fifth fusion feature With the fifth characteristic diagram correspond; The structures of the fifth feature extraction network to the ninth feature extraction network are the same, and are the same as the structures of the first feature extraction network to the fourth feature extraction network; Among them, the efficient feature extraction capabilities of the fifth to ninth feature extraction networks can reduce the computational complexity while maintaining the accuracy of feature expression, while also being able to mine the potential information in the feature graph and improve the feature discrimination capability.
[0026] In the backbone network and feature fusion network, the processing of the first feature extraction network to the ninth feature extraction network (such as Figure 2 ), specifically including the following: The input feature map is divided into two identical sub-feature maps along the channel dimension, and the number of channels of each sub-feature map is half of the input feature map; The first sub-feature map obtained is passed through the first The convolutional layer is processed, activated by the activation function, and then by the second The convolutional layer is processed to generate an enhanced first local feature representation; this processing step can enhance the expression ability of local features; Processing the obtained second sub-feature map through a diversified branch module to generate a diversified feature representation; After the first sub-feature map, the first local feature representation, the second sub-feature map and the diversified feature representation are concatenated along the channel dimension, they are sequentially passed through The convolutional layer and deformable convolution module are processed to obtain the output of the feature extraction network; The processing of the convolution layer not only adjusts the channel dimension of the feature map, but also realizes the dynamic weighting between different features through the parameters learned through training, thereby optimizing the fusion effect of multi-scale features and generating a more compact and information-rich feature representation; at the same time, the deformable convolution dynamically adjusts the sampling point position of the convolution kernel through adaptive offset, so that it can adaptively capture features according to the shape and posture changes of the target, thereby improving the feature extraction capability of irregular targets; The feature extraction network combines hierarchical processing with the original feature retention strategy, and introduces deformable convolution to enhance the geometric feature modeling capability. The focus is on multi-branch feature extraction and channel dimension splicing to balance feature diversity and computational efficiency. The feature extraction network starts from the spatial domain and the geometric domain respectively, and optimizes the feature fusion process according to different scene requirements. The former focuses more on the enhancement of semantic information, while the latter focuses on the modeling of geometric features, which reflects the diversity and pertinence of feature extraction and feature fusion strategies. The feature extraction network combines hierarchical processing with the original feature retention strategy, and introduces deformable convolution to enhance the geometric feature modeling capability. The focus is on multi-branch feature extraction and channel dimension splicing to balance feature diversity and computational efficiency. The feature extraction network achieves a significant improvement in feature extraction efficiency through three key designs: its hierarchical feature reuse architecture balances the extraction of shallow detail information and deep semantic features through parallel paths, avoiding information loss; the adaptive feature fusion mechanism optimizes multi-scale feature integration through channel splicing and dynamic weighting; the introduction of deformable convolution enhances the module's ability to model target geometric features in complex scenes; the network significantly improves feature expression capabilities through the synergy of feature extraction, fusion and enhancement, providing a more robust and effective feature representation for target detection tasks.
[0027] The constructed detection network includes the following contents: The input of the detection network is the third fusion feature , the fourth fusion feature and the fifth fusion feature ; The detection network includes a first detection head, a second detection head, a third detection head and a screening layer; the first detection head is used to input the third fusion feature The second detection head is used to detect the fourth fusion feature of the input The third detection head is used to detect the fifth fusion feature of the input Perform detection and output the location information and category information of the corresponding target; the screening layer is used to screen the location information and category information of the target output by the first detection head, the second detection head and the third detection head to output the final target detection result; The structures of the first detection head, the second detection head and the third detection head are the same. The detection head is constructed based on the convolutional network. The screening layer uses the non-maximum suppression method to screen the honor prediction box and retain the detection results with confidence higher than the set value. Finally, the detection box and the category label are superimposed on the original image to complete the detection, which can also be used for result display.
[0028] Among them, the processing process of the detection head (such as Figure 3 ), including the following steps: The input feature map is divided into a first sub-feature map, a second sub-feature map, a third sub-feature map and a fourth sub-feature map along the channel dimension, and the first sub-feature map, the second sub-feature map, the third sub-feature map and the fourth sub-feature map are ensured to have the same number of channels; The first sub-feature map is passed through The convolution kernel is processed; the second sub-feature map is processed by The convolution kernel is processed; the third sub-feature map is processed by The convolution kernel is processed; the fourth sub-feature map is processed by The outputs of the four convolution kernels are concatenated along the channel dimension to obtain a new feature map, and then passed through 2 The convolutional layer is processed to obtain the location information and category information of the target corresponding to the input feature map; Among them, by adopting convolution kernels of different sizes, the module can capture multi-scale features from local to global, among which the larger convolution kernels (such as The convolution kernel and The convolution kernel of the proposed method effectively expands the receptive field, thereby enhancing the perception of the overall structure and contextual information of the image. At the same time, the detection head can fully integrate multi-scale feature information, making the output feature map more comprehensive and accurate in expressing the input image features. The detection head improves the discriminability of spatial features through the diversity of convolution kernels, and through the hierarchical fusion of feature maps and the multi-scale combination of convolution kernels, they together constitute a complementary solution for lightweight target detection; the detection head uses multi-scale convolution to capture multi-scale spatial features and performs weighted fusion in the channel dimension, aiming to improve the discriminability of spatial features through the diversity of convolution kernels; by using a variety of convolution kernels of different sizes to process feature maps, it has many outstanding advantages: on the one hand, reasonable segmentation of feature maps and then combining them with different convolution kernels can greatly reduce the number of parameters and calculations, and improve the efficiency of model operation, which plays a vital role in resource-limited environments; on the other hand, the use of large-size convolution kernels can effectively expand the receptive field, and then extract global features from the input feature map, which helps to deeply understand the overall structure and contextual information of the image.
[0029] S4. Using the training data set constructed in step S2, the lightweight object detection initial model constructed in step S3 is trained to obtain a lightweight object detection model.
[0030] S5. Using the lightweight target detection model obtained in step S4, actual lightweight target detection is performed.
[0031] The target detection method of the present invention is further described below in conjunction with an embodiment: A series of comparative experiments were conducted on the MS COCO public dataset to verify the performance on the MS COCO target dataset, and the performance of the algorithm was evaluated based on two indicators: recall and precision.
[0032] The MS COCO dataset used is a large dataset widely used in the field of computer vision, especially for tasks such as object detection, segmentation, and image description generation. The images come from different scenes and environments, covering various objects in daily life, ensuring the diversity and richness of the dataset; the MS COCO dataset contains 118,000 training images and 5,000 verification images, and is widely used in various computer vision tasks.
[0033] The method of the present invention is compared with several state-of-the-art existing methods, including: YOLOV5 scheme, YOLOV7 scheme, YOLOV8 scheme and Gold-YOLO scheme. Among them, the YOLOV5 scheme is the scheme used by Bharat Mahaur et al. in the paper "Small-object detection based on YOLOv5 in autonomous driving systems" in 2023; the YOLOV7 scheme is the scheme proposed by Chien-Yao Wang et al. in the paper "YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors" in 2022; the YOLOV8 scheme is the scheme used by Dillon Reis et al. in the paper "Real-time flying object detection with YOLOv8" in 2023; the Gold-YOLO scheme is the scheme proposed by ChengCheng Wang et al. in the paper "Gold-YOLO: Efficient Object Detector via Gather-and-Distribute Mechanism".
[0034] Figure 4 The following is a schematic diagram of the comparison effect. The corresponding comparative experimental data are shown in Table 1: Table 1 Schematic diagram of comparative experimental data
[0035] It can be seen from experiments that in terms of the same level of parameter quantity and calculation amount, the method of the present invention has higher detection accuracy.
[0036] In addition, the effect of the feature extraction network and the detection head in the method of the present invention is proved by ablation experiments; Figure 5 The effect diagram of the ablation experiment of the present invention is shown in FIG. 2, and the specific experimental data are shown in Table 2: Table 2 Ablation experiment data diagram
[0037] Through ablation experiments, it can be seen that the feature extraction network and detection head in the method of the present invention can obtain better detection accuracy without significantly increasing the number of parameters and the amount of calculation.
[0038] like Figure 6 The method flow diagram of the drone screen target detection method of the present invention is shown as follows: The drone screen target detection method including the lightweight target detection method disclosed in the present invention comprises the following steps: A. Obtain the image data information taken by the drone during flight in real time; B. Using the lightweight target detection method, perform target detection on the image data obtained in step A; C. Based on the target detection results obtained in step B, complete the target detection of the drone's shooting picture.
[0039] Through the drone image target detection method provided by the present invention, the drone can perform target detection on the captured image in real time, thereby realizing various functions such as drone navigation, path planning, drone monitoring, target tracking, etc.
[0040] In addition, the lightweight target detection method provided by the present invention is also applicable to various target detection fields such as target detection in autonomous driving processes, target detection by cameras in the monitoring field, target detection in medical impact analysis, target detection in industrial automation processes, target detection in the field of motion analysis, target detection in augmented reality (AR) devices, and target detection in real-time human-computer interaction.
Claims
1. A lightweight target detection method, characterized in that The steps include: S1. Obtaining existing target detection image data information; S2. Preprocessing the data information obtained in step S1 to construct a training data set; S3. Construct a lightweight target detection initial model including a backbone network, a feature fusion network, and a detection network; Among them, a backbone network is constructed based on a 2D convolutional network, a feature extraction network and a spatial pyramid pooling fusion network to extract feature map information of input image data at different scales; a feature fusion network is constructed based on the feature extraction network to fuse the feature map information of different scales and obtain fusion features; a detection network is constructed based on the convolutional network to extract and predict the location information and category information of the target according to the obtained fusion features; The feature extraction network uses convolution and diversified branch modules for hierarchical parallel processing to extract multi-branch features of the input image from the perspective of the spatial domain, while ensuring the diversity and efficiency of the extracted features; the feature extraction network also uses convolution and deformable convolution to extract features from the perspective of the geometric domain to ensure the diversity and pertinence of feature fusion; The detection network captures multi-scale spatial features through several convolutions of different scales in parallel, and performs weighted fusion of the captured multi-scale spatial features to achieve target detection; S4. Using the training data set constructed in step S2, the lightweight target detection initial model constructed in step S3 is trained to obtain a lightweight target detection model; S5. Using the lightweight target detection model obtained in step S4, actual lightweight target detection is performed.
2. The lightweight target detection method according to claim 1, characterized in that The backbone network constructed includes the following contents: The backbone network includes an input 2D convolutional network, a first 2D convolutional network, a first feature extraction network, a second 2D convolutional network, a second feature extraction network, a third 2D convolutional network, a third feature extraction network, a fourth 2D convolutional network, a fourth feature extraction network and a spatial pyramid pooling fusion network, which are connected in series in sequence; Among them, the structures of the first feature extraction network to the fourth feature extraction network are the same; the feature extraction networks are all based on diversified branch modules, convolution modules and deformable convolution modules to achieve efficient feature extraction; The backbone network is used to extract feature map information of the input image data at different scales, and finally output feature maps of three different sizes; the feature map includes the third feature map , the fourth characteristic graph And the fifth characteristic diagram .
3. The lightweight target detection method according to claim 2, characterized in that The constructed feature fusion network specifically includes the following contents: The input of the feature fusion network is the third feature map output by the backbone network , the fourth characteristic graph And the fifth characteristic diagram ; Fifth characteristic diagram And the fourth characteristic diagram After the connection is performed, feature extraction is performed through a fifth feature extraction network to obtain a fifth extracted feature; Fourth characteristic diagram , the third characteristic map After being connected with the fifth extracted feature, feature extraction is performed through a sixth feature extraction network to obtain a sixth extracted feature; The third characteristic diagram After being connected with the sixth extracted feature, feature extraction is performed through a ninth feature extraction network to obtain a ninth extracted feature; After the sixth extracted feature and the ninth extracted feature are connected, feature extraction is performed through an eighth feature extraction network to obtain an eighth extracted feature; After the fifth extracted feature, the sixth extracted feature and the eighth extracted feature are connected, feature extraction is performed through a seventh feature extraction network to obtain a seventh extracted feature; The feature fusion network is used to fuse the feature map information of three different scales and output fusion features of three different scales; The fusion features include the third fusion feature , the fourth fusion feature and the fifth fusion feature , and the ninth extracted feature is used as the third fusion feature , the eighth extracted feature is used as the fourth fusion feature , the seventh extracted feature is used as the fifth fusion feature ; The third fusion feature With the third characteristic diagram Corresponding to the fourth fusion feature With the fourth characteristic diagram Correspondingly, the fifth fusion feature With the fifth characteristic diagram correspond; The structures of the fifth feature extraction network to the ninth feature extraction network are the same, and are the same as the structures of the first feature extraction network to the fourth feature extraction network.
4. The lightweight target detection method according to claim 3, characterized in that The processing process of the first feature extraction network to the ninth feature extraction network specifically includes the following contents: The input feature map is divided into two identical sub-feature maps along the channel dimension, and the number of channels of each sub-feature map is half of the input feature map; The first sub-feature map obtained is passed through the first The convolutional layer is processed, activated by the activation function, and then by the second The convolutional layer is processed to generate an enhanced first local feature representation; Processing the obtained second sub-feature map through a diversified branch module to generate a diversified feature representation; After the first sub-feature map, the first local feature representation, the second sub-feature map and the diversified feature representation are concatenated along the channel dimension, they are sequentially passed through The convolutional layer and deformable convolution module are processed to obtain the output of the feature extraction network.
5. The lightweight target detection method according to claim 4, characterized in that The constructed detection network includes the following contents: The input of the detection network is the third fusion feature , the fourth fusion feature and the fifth fusion feature ; The detection network includes a first detection head, a second detection head, a third detection head and a screening layer; the first detection head is used to input the third fusion feature The second detection head is used to detect the fourth fusion feature of the input The third detection head is used to detect the fifth fusion feature of the input Perform detection and output the location information and category information of the corresponding target; the screening layer is used to screen the location information and category information of the target output by the first detection head, the second detection head and the third detection head to output the final target detection result; The structures of the first detection head, the second detection head and the third detection head are the same. The detection head is constructed based on the convolutional network. The screening layer uses the non-maximum suppression method to screen the honor prediction box, retain the detection results with confidence higher than the set value, and finally superimpose the detection box and the category label on the original image to complete the detection.
6. The lightweight target detection method according to claim 5, characterized in that The processing of the detection head includes the following steps: The input feature map is divided into a first sub-feature map, a second sub-feature map, a third sub-feature map and a fourth sub-feature map along the channel dimension, and the first sub-feature map, the second sub-feature map, the third sub-feature map and the fourth sub-feature map are ensured to have the same number of channels; The first sub-feature map is passed through The convolution kernel is processed; the second sub-feature map is processed by The convolution kernel is processed; the third sub-feature map is processed by The convolution kernel is processed; the fourth sub-feature map is processed by The outputs of the four convolution kernels are concatenated along the channel dimension to obtain a new feature map, and then passed through 2 The convolutional layer is used to process the input feature map to obtain the location information and category information of the target corresponding to the input feature map.
7. A method for detecting targets in drone images comprising the lightweight target detection method according to any one of claims 1 to 6, characterized in that The steps include: A. Obtain the image data information taken by the drone during flight in real time; B. using the lightweight target detection method according to any one of claims 1 to 6 to perform target detection on the image data obtained in step A; C. Based on the target detection results obtained in step B, complete the target detection of the drone's shooting picture.
Citation Information
Patent Citations
Improved YOLOv8 unmanned aerial vehicle aerial target detection method
CN117557922A
Colonoscope polyp image detection method based on Mamba and YOLOv8
CN118762009A
Remote sensing target detection method based on convolutional neural network
CN119152367A
Low-illumination target detection method and device based on YOLOv7-tiny improvement
CN119169267A
Remote sensing SAR image rotating vehicle target detection method based on YOLOX
CN119206343A
Cited By
Lightweight AI-based distribution line unmanned aerial vehicle edge end real-time visual identification and target detection method and system
CN121459227A