Object Detection Method, Device, Electronic Device, and Storage Medium

Through labeling sample expansion and data augmentation, combined with lightweight networks and multi-scale attention aggregation module, the network structure is optimized, and the problems of insufficient data and poor feature perception capabilities in "low small and slow" object detection under the air perspective are solved, achieving high-precision and fast object detection.

CN116797779BActive Publication Date: 2025-07-08NORTH CHINA UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310584488.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-23
Publication Date
2025-07-08
Estimated Expiration
2043-05-23

AI Technical Summary

Technical Problem

The prior art has problems such as small samples, small targets, and complex backgrounds in the "low and slow" object detection from the empty perspective, resulting in insufficient data of the detection model, overfitting, poor feature perception ability, resulting in missed detection, missed detection and slow detection speed.

Method used

Through labeling sample expansion and data augmentation, a lightweight single-stage small-object detection network with multi-scale aggregation is used, combining a lightweight network and a multi-scale attention aggregation module to optimize the network structure, enhance the extraction ability of small-object features, and improve the anti-interference and accuracy of the detection model.

Benefits of technology

It provides sufficient data support, improves the anti-interference and detection accuracy of the detection model, solves the detection problem of small targets in complex backgrounds, and improves the generalization ability and detection speed of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116797779B_ABST
    Figure CN116797779B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing technology, and provides an object detection method, apparatus, electronic device, and storage medium. The method includes: performing annotation sample augmentation on the images in the dataset of the target object to obtain augmented images; performing data augmentation on the augmented images, and randomly reconstructing the first augmented images obtained by the data augmentation with background images to obtain reconstructed images; inputting the reconstructed images into a detection model to obtain the target object output by the detection model; wherein the detection model is trained by using sample images for a lightweight single-stage small object detection network with multi-scale aggregation. The present application provides sufficient data support for the training of the detection model through annotation sample augmentation and data augmentation, improves the anti-interference ability of the detection model, and at the same time improves the detection accuracy of the detection model by optimizing the object detection network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and particularly to an object detection method, apparatus, electronic device, and storage medium. Background Art

[0002] Due to the characteristics of small aircraft at low altitude and low speed (abbreviated as "low, small, and slow"), such as being lightweight, easy to operate, having low flight requirements, and good suddenness of takeoff, it is convenient for illegal photography and mapping, resulting in frequent occurrence of illegal and irregular use of small unmanned aerial vehicles, which increases the difficulty of important airspace protection and social security management.

[0003] Currently, in the detection task of "low, small, and slow" targets, the characteristics of small and weak targets in the air are not obvious and are not easily captured, which may lead to missed detections, false detections, and other situations resulting in detection failures. Therefore, in different scenarios, it is necessary to solve problems such as too small target size and complex background interference in the detection task. Based on this, various detection algorithms for "low, small, and slow" targets have been proposed. However, for special scenarios, such as the detection task of small sample targets from an aerial perspective, the effect is not ideal. Since there is no publicly available dataset for "low, small, and slow" targets from an aerial perspective, there are problems such as small samples, small targets, and complex backgrounds in the detection task, resulting in severely insufficient data input into the detection network, causing the problem of overfitting of the detection model; the detection model has poor ability to perceive target features and cannot fully extract small target features, resulting in situations such as missed detections and false detections of targets, further affecting the speed of the detection model. Based on this, there is an urgent need for a detection method for "low, small, and slow" targets from an aerial perspective. Summary of the Invention

[0004] The present application provides an object detection method, apparatus, electronic device, and storage medium to solve the problem of low accuracy in detecting "low, small, and slow" targets. By expanding labeled samples and augmenting data, sufficient data support is provided for the training of the detection model, improving the anti-interference ability of the detection model. At the same time, by optimizing the object detection network, the detection accuracy of the detection model is improved.

[0005] The present application provides an object detection method, including:

[0006] Expanding labeled samples for images in the dataset of the target object to obtain augmented images;

[0007] Augmenting the augmented images, and randomly reconstructing the first augmented images obtained by data augmentation with background images to obtain reconstructed images;

[0008] Inputting the reconstructed images into a detection model to obtain the target object output by the detection model;

[0009] Among them, the detection model is obtained by training a lightweight single-stage small target detection network with multi-scale aggregation using sample images; the lightweight single-stage small target detection network with multi-scale aggregation includes a lightweight network and a multi-scale attention aggregation module; the output layer that cannot recognize the feature information of the target object has been deleted in the lightweight network; the lightweight network is connected to the multi-scale attention aggregation module to aggregate the feature information of the reconstructed image.

[0010] In one embodiment, the lightweight network includes a first path and a second path. The first path is spliced with the multi-scale attention aggregation module to increase the semantic feature information of the reconstructed image; the second path is convolutionally connected to the multi-scale attention aggregation module to aggregate the feature information of the reconstructed image.

[0011] In one embodiment, the data augmentation of the augmented image includes:

[0012] Slice the augmented image to obtain the background image and the target image carrying the target object;

[0013] Perform data augmentation on the target image.

[0014] In one embodiment, the data augmentation includes forward augmentation and reverse augmentation;

[0015] The data augmentation of the target image includes:

[0016] Determine the first noise signal for forward augmentation and the second noise signal for reverse augmentation;

[0017] Perform forward augmentation on the target image based on the first noise signal to obtain a second augmented image;

[0018] Perform reverse augmentation on the second augmented image based on the second noise signal to obtain the first augmented image.

[0019] In one embodiment, the expression of the first noise signal is:

[0020]

[0021] Where q(X t |X0) represents the first noise signal, X0 represents the input of the diffusion model, X t represents the output of the diffusion model, I represents the identity matrix, N represents a constant, represents the product result of Markov;

[0022] The expression of the second noise signal is:

[0023]

[0024] Among them, q(X t-1 |X t X0) represents the second noise signal, X0 represents the input of the diffusion model, X t represents the output of the diffusion model, X t-1 represents adding noise once based on X t α t represents the distance between 1 and the step size variable, β t represents the variable affecting the diffusion step size, t represents the variable affecting the diffusion step size, represents the multiplicative result of Markov, represents the multiplicative result of Markov, Z ∼ N(0, I) represents the normal distribution, I represents the identity matrix, and Z and N represent constants.

[0025] In one embodiment, the step of performing annotation sample augmentation on the images in the dataset of the target object to obtain augmented images includes:

[0026] Obtain the annotated images and unannotated images in the dataset;

[0027] Use the annotated images to train a classifier, and based on the classifier, annotate the unannotated images to obtain the augmented images.

[0028] In one embodiment, the step of randomly reconstructing the first augmented image obtained by data augmentation and the background image to obtain a reconstructed image includes:

[0029] Randomly combine the first augmented image and the background image to obtain a combined image, and use the combined image as the reconstructed image.

[0030] This application also provides a target detection device, including:

[0031] A sample augmentation module, configured to perform annotation sample augmentation on the images in the dataset of the target object to obtain augmented images;

[0032] A data augmentation module, configured to perform data augmentation on the augmented images, and randomly reconstruct the first augmented image obtained by data augmentation and the background image to obtain a reconstructed image;

[0033] A target detection module, configured to input the reconstructed image into a detection model and obtain the target object output by the detection model;

[0034] Among them, the detection model is obtained by training a lightweight single-stage small object detection network with multi-scale aggregation using sample images; the lightweight single-stage small object detection network with multi-scale aggregation includes a lightweight network and a multi-scale attention aggregation module; the output layer that cannot recognize the feature information of the target object has been deleted in the lightweight network; the lightweight network is connected to the multi-scale attention aggregation module to aggregate the feature information of the reconstructed image.

[0035] The present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the target detection method as described in any one of the above.

[0036] The present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the target detection method as described in any one of the above.

[0037] The target detection method, device, electronic device, and storage medium provided by the present application obtain augmented images by performing annotation sample augmentation on the images in the dataset of the target object; perform data augmentation on the augmented images, and randomly reconstruct the first augmented images obtained by data augmentation with background images to obtain reconstructed images; input the reconstructed images into the detection model to obtain the target object output by the detection model; among them, the detection model is obtained by training a lightweight single-stage small object detection network with multi-scale aggregation using sample images; the lightweight single-stage small object detection network with multi-scale aggregation includes a lightweight network and a multi-scale attention aggregation module; the output layer that cannot recognize the feature information of the target object has been deleted in the lightweight network; the lightweight network is connected to the multi-scale attention aggregation module to aggregate the feature information of the reconstructed image. The present application provides sufficient data support for the training of the detection model through annotation sample augmentation and data augmentation, improves the anti-interference ability of the detection model, and at the same time, improves the detection accuracy of the detection model by optimizing the target detection network. Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 is one of the flow diagrams of the target detection method provided by the present application;

[0040] Figure 2It is a schematic structural diagram of the lightweight network provided by the present application;

[0041] Figure 3 It is a schematic structural diagram of the multi-scale attention aggregation module provided by the present application;

[0042] Figure 4 It is a schematic structural diagram of the lightweight single-stage small object detection network with multi-scale aggregation provided by the present application;

[0043] Figure 5 It is a schematic flow diagram of generating an image by a diffusion process provided by the present application;

[0044] Figure 6 It is a schematic flow diagram of self-training sample augmentation provided by the present application;

[0045] Figure 7 It is the second schematic flow diagram of the object detection method provided by the present application;

[0046] Figure 8 It is a schematic structural diagram of the object detection device provided by the present application;

[0047] Figure 9 It is a schematic structural diagram of the electronic device provided by the present application. Detailed implementation manners

[0048] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0049] The following combines Figures 1-9 to describe the object detection method, device, electronic device and storage medium of the present application.

[0050] Specifically, the present application provides an object detection method. Referring to Figure 1 , Figure 1 It is the first schematic flow diagram of the object detection method provided by the present application.

[0051] The object detection method provided by the embodiments of the present application includes:

[0052] S100, performing annotation sample augmentation on the images in the dataset of the target object to obtain augmented images;

[0053] It should be noted that this application mainly detects "low, small, and slow" targets in an empty view (or empty video), that is, it performs object detection on images of "low, small, and slow" targets taken towards the sky. In the embodiments of this application, the object detection method is analyzed and explained with "low, small, and slow" as the target object. "Low, small, and slow" refers to small aircraft at low altitude and with low speed, including light and ultra-light aircraft (including light and ultra-light helicopters), gliders, delta wings, powered delta wings, manned balloons (hot air balloons), airships, paragliders, powered paragliders, drones, aircraft models, unmanned free balloons, and tethered balloons, etc.

[0054] Since there is no publicly available dataset for "low, small, and slow" targets in an empty view, there are problems such as small samples, small targets, and complex backgrounds in the detection task, resulting in severely insufficient data input into the detection network, causing problems such as overfitting of the detection model, and resulting in poor target feature perception ability of the detection model, being unable to fully extract small target features, thus causing situations such as target missed detection and target misdetection, and then affecting the detection speed of the detection model. Based on this, to address the problem of the lack of a "low, small, and slow" dataset, this application expands the sample size and augments the data to provide sufficient data support for the training of the detection model.

[0055] In the sample processing stage, a semi-supervised learning (SSL) sample augmentation technique is introduced to complement the missing annotation information in the self-collected dataset, thus avoiding the problem of poor generalization ability of the detection model due to insufficient "low, small, and slow" samples.

[0056] Specifically, images of the target object in an empty view are collected, and a small number of images are labeled. A dataset is constructed through a small number of images with sample labels (i.e., labeled images) and a large number of images without sample labels (i.e., unlabeled images). Then, based on the semi-supervised learning sample augmentation technique, the images in the dataset are augmented with labeled samples to obtain augmented images. For example, based on semi-supervised learning, the classifier can automatically use unlabeled images to improve learning performance without relying on external interactions. That is, under the guidance of a small number of sample labels, a large number of unlabeled samples can be fully utilized to improve learning performance, thereby realizing the annotation of unlabeled images.

[0057] S200, perform data augmentation on the augmented images, and randomly reconstruct the first augmented images obtained by data augmentation and background images to obtain reconstructed images;

[0058] It should be noted that in deep learning, data augmentation can expand the dataset by performing a series of random transformations on the original data, thereby improving the robustness and generalization ability of the model.

[0059] Before data augmentation, it is necessary to slice the augmented image to separate the target object from the background, so as to obtain the background image of the target object and the target image carrying the target object. For example, based on the size of the target object, perform a slicing operation on the augmented image, and set the slicing resolution range between 15*15 and 30*30 to ensure that the slice contains the complete target object and its important features.

[0060] Furthermore, perform data augmentation on the target image. For example, use noise data augmentation to augment the target image, that is, increase the diversity of the dataset by adding random noise signals, thereby obtaining the first augmented image. Then randomly reconstruct the first augmented image and the background image to obtain a reconstructed image. Specifically, randomly combine the first augmented image and the background image to obtain a combined image, and use the combined image as the reconstructed image. For example, assume that the first augmented image is image A, and the background image includes image B and image C. Then randomly combine the first augmented image with the complex background image to obtain the combined images (i.e., the reconstructed images) as image AB and image AC. At the same time, scale the reconstructed images by 10%-30% to simulate target objects of different scales in the near and far views from an empty perspective, so as to restore the interference factors of background objects in the image regions irrelevant to the detection, and improve the anti-interference ability of the detection model.

[0061] Optionally, the first augmented image and the background image can also be randomly spliced, and the spliced image is used as the reconstructed image.

[0062] S300, input the reconstructed image into the detection model to obtain the target object output by the detection model; wherein, the detection model is trained by using sample images on a lightweight single-stage small target detection network with multi-scale aggregation; the lightweight single-stage small target detection network with multi-scale aggregation includes a lightweight network and a multi-scale attention aggregation module; the output layer that cannot recognize the feature information of the target object has been deleted in the lightweight network; the lightweight network is connected to the multi-scale attention aggregation module to aggregate the feature information of the reconstructed image.

[0063] It should be noted that according to the YoloV5 network, at the output end of the feature layer, the P5 - P3 layers are arranged from top to bottom. After three downsamplings of a 640*640 input image, three feature map layers with sizes of 19*19, 38*38, and 76*76 are finally output. The size of the feature map is closely related to the corresponding area size of each grid unit in the input image, and the two are inversely proportional. That is, the 76*76 output feature layer can retain more small - target feature information. The size of the "low - slow - small" targets detected in this application is between 15*15 and 30*30, and the size of the feature map output by P5 is 19*19. For some "low - slow - small" targets with relatively large sizes, the feature map output by P5 cannot contain complete target information, which is likely to cause certain interference to the judgment of the detection model. Therefore, in this application, the output layer of P5, which is not sensitive to capturing features of small targets, is trimmed, that is, the output layer that cannot recognize the feature information of the target object is deleted, so as to optimize the network structure and maximize the improvement of the network model training and target detection rate. Among them, the lightweight network structure diagram is as Figure 2 shown.

[0064] Furthermore, for the detection task of "low - slow - small" targets, since the target size itself is too small and accounts for a very small proportion in the image of the air - to - air perspective, it is difficult to capture and learn its features. After introducing a lightweight network for lightweight operations to remove irrelevant parameters, the number of output feature maps is reduced from three to two, which is not sufficient to support the detection model to fully perceive the feature information of small targets, and it is easy to make the robustness of the detection model insufficient, and it is difficult to accurately and efficiently identify small targets. Therefore, this application proposes a multi - scale attention aggregation module to output a feature map that better matches the feature information of small targets, providing more learnability for the detection model. The multi - scale attention aggregation module, among which, the structure schematic diagram of the multi - scale attention aggregation module is as Figure 3 shown.

[0065] The lightweight network includes a first path and a second path. The first path is spliced with the multi - scale attention aggregation module to increase the semantic feature information of the reconstructed image; the second path is convolutionally connected to the multi - scale attention aggregation module to aggregate the feature information of the reconstructed image. Among them, the first path refers to the top - down path, and the second path refers to the bottom - up path.

[0066] For example, refer to Figure 4, first, perform upsampling and connect it to the top-down path in the lightweight network to enrich the high-level strong semantic feature information in the network; secondly, achieve information sharing through horizontal connections with the bottom-up path in the lightweight network and perform information aggregation to complement the target location information ignored in the top-down path. Through the above two information aggregations, the multi-scale attention aggregation module finally outputs a feature map with a size of 152*152 and 255 channels. The feature map of this size matches the "low, small, and slow" targets and can fully express the feature information of small targets, providing a reliable training basis for the detection network and solving the problem of insufficient robustness of the network for small target detection.

[0067] The object detection method provided by the embodiments of the present application obtains augmented images by performing labeled sample augmentation on the images in the dataset of the object; performs data augmentation on the augmented images, and randomly reconstructs the first augmented image obtained by data augmentation and the background image to obtain reconstructed images; inputs the reconstructed images into the detection model to obtain the objects output by the detection model; wherein, the detection model is trained by using sample images on the lightweight single-stage small object detection network with multi-scale aggregation; the lightweight single-stage small object detection network with multi-scale aggregation includes a lightweight network and a multi-scale attention aggregation module; the output layer that cannot identify the feature information of the object has been deleted in the lightweight network; the lightweight network is connected to the multi-scale attention aggregation module to aggregate the feature information of the reconstructed images. The present application provides sufficient data support for the training of the detection model through labeled sample augmentation and data augmentation, improves the anti-interference ability of the detection model, and at the same time improves the detection accuracy of the detection model by optimizing the object detection network.

[0068] Based on the above embodiments, the data augmentation includes forward augmentation and reverse augmentation; the performing data augmentation on the target image includes:

[0069] S210, determine the first noise signal for forward augmentation and the second noise signal for reverse augmentation;

[0070] S220, perform forward augmentation on the target image based on the first noise signal to obtain a second augmented image;

[0071] S230, perform reverse augmentation on the second augmented image based on the second noise signal to obtain the first augmented image.

[0072] Data augmentation includes forward augmentation and reverse augmentation, and data augmentation can also be understood as data diffusion, that is, data augmentation includes two processes: a diffusion process and a reverse diffusion process. Among them, data augmentation methods include cropping, translation, changing brightness, adding noise, rotating angles, and mirroring, etc. In the embodiments of the present application, data augmentation is analyzed and explained by taking adding noise as an example. Noise augmentation refers to introducing different degrees of noise into the data to increase the robustness of the model, that is, adding noise to the training samples so that the model can learn more anti-interference capabilities during the learning process, thereby improving the generalization ability of the model.

[0073] Specifically, determine the first noise signal for forward augmentation and the second noise signal for reverse augmentation, and then perform forward augmentation on the target image based on the first noise signal to obtain a second augmented image, and further perform reverse augmentation on the second augmented image based on the second noise signal to obtain a first augmented image.

[0074] For example, referring to Figure 5 , Figure 5 in which X0 on the right side is the target image input to the diffusion model, that is, the slice containing the target object. The diffusion process is (X0 → X T ), that is, from right to left, the process of adding T times of noise to the input target image step by step to make it blurred. Among them, the X T+1 slice is obtained by adding noise to the X T slice, and X T+1 is only related to X T . Therefore, during the diffusion process from the input target image to the output second augmented image, each diffusion step size is affected by the variable {β t ∈0,1)}, and X T obeys a normal distribution with a mean of and a variance of β t I. Therefore, the noise q(X t |X t-1 ) added each time is known and can be represented by formula (1):

[0075]

[0076] Among them, q(X t |X t-1 ) represents the noise signal, represents the mean, β t I represents the variance, X t represents the output of the diffusion model, X t-1 represents removing the noise once on the basis of X t , and N represents a constant.

[0077] The form of expressing X t -X1 using the reparameterization technique is shown in the following formulas (2)-(5):

[0078]

[0079]

[0080]

[0081]

[0082] Among them, X0 represents the input of the diffusion model, X t represents the output of the diffusion model, X t-1 represents adding noise once on the basis of X t X t-2 represents adding noise once on the basis of X t-1 X t-3 represents adding noise once on the basis of X t-2 X1 represents adding noise once on the basis of X0, α t = 1 - β t and Z t ~N(0, I), t ≥ 0, α t represents the distance between 1 and the step variable of X t α t-1 represents the distance between 1 and the step variable of X t-1 α t-2 represents the distance between 1 and the step variable of X t-2 α1 represents the distance between 1 and the step variable of X1, Z t-1 represents the normal distribution of X t-1 from 0 to the identity matrix I, Z t-2 represents the normal distribution of X t-2 from 0 to the identity matrix I, Z t-3 represents the normal distribution of X t-3 from 0 to the identity matrix I, Z0 represents the normal distribution of X0 from 0 to the identity matrix I.

[0083] Based on the above formulas (2)-(5), X t is in the form shown in the following formula (6):

[0084]

[0085] Among them, represents the product result of Markov.

[0086] Therefore, X t and q(X t |X0) are in the forms shown in the following formulas (7) and (8):

[0087]

[0088]

[0089] Among them, q(X t |X0) represents the first noise signal, X0 represents the input of the diffusion model, and X t represents the output of the diffusion model. I represents the identity matrix, and N represents a constant. represents the product result of Markov. represents the random variable related to Z t-1 Z~N(0, I) represents the normal distribution, and Z and N represent constants.

[0090] Through the above derivation of the first noise signal q(X t |X0), any parameter in the diffusion model can be obtained.

[0091] As Figure 5 shown, the reverse path is the reverse diffusion process. Its essence is a process of generating a realistic image by inputting any noise image (i.e., the second augmented image), and the size of the generated image is the same as that of the input target image. In the reverse diffusion process, p θ (X t-1 |X t ) is used to approximately represent q(X t-1 |X t ). Therefore, the process of generating an image is actually a step-by-step derivation of q(X t-1 |X t X0), and then the derivation result is used to guide how to train the reverse diffusion network. Finally, a large number of realistic images of the target object, that is, the first augmented image, can be obtained through any noise map.

[0092] The second noise signal q(X t-1 |X t X0) can be obtained by using q(X t |X0) deduced in the diffusion process and q(X t-1 |X t ), and its form is shown in the following formula (9):

[0093]

[0094] Substituting the above formulas (6) and (8) can deduce the final expression form of the second noise signal q(X t-1 |X t X0) as shown in the following formula (10):

[0095]

[0096] Among them, q(X t-1 |X tX0) represents the second noise signal, X0 represents the input of the diffusion model, X t represents the output of the diffusion model, X t-1 represents adding noise once on the basis of X t α represents the distance between 1 and the step variable of X t β represents the variable affecting the diffusion step size t α represents the distance between 1 and the step variable of X t β represents the variable affecting the diffusion step size represents the product result of Markov represents the product result of Markov, Z~N(0,I) represents the normal distribution, I represents the identity matrix, and Z and N represent constants

[0097] Denoise the random noise image through the above formula to generate a large number of high-quality first augmented images to solve the problem of insufficient data volume, and effectively avoid the problem that the detection model has large differences in data detection results for different target objects due to poor generalization

[0098] In the embodiments of the present application, by introducing the target background reconstruction method, the interference of complex backgrounds in the data augmentation stage on the generative model is excluded. At the same time, to solve the problem of poor generalization ability of the detection model caused by insufficient data, the data is augmented through the diffusion model, the problem of missing samples in the dataset is solved, and finally the background complexity is restored through random reconstruction of the target background to improve the anti-interference ability of the subsequent detection model, thereby improving the generalization ability of the model

[0099] Based on the above embodiments, the method for augmenting the labeled samples of the images in the dataset of the target object to obtain the augmented images includes:

[0100] S110, obtaining the labeled images and unlabeled images in the dataset

[0101] S120, training a classifier using the labeled images, and labeling the unlabeled images based on the classifier to obtain the augmented images

[0102] It should be noted that the dataset includes a small number of labeled images and a large number of unlabeled images. When performing labeled sample augmentation, a small number of labeled images are used to train the classifier, and then the trained classifier is used to label the unlabeled images to obtain the augmented images

[0103] For example, refer to Figure 6, in the self-training sample augmentation technique of the embodiment of the present application, YoloV5 is selected as the classifier. Some images in the dataset are manually labeled as the input for training the classifier model, and the remaining images lacking annotation information are captured. At the same time, the output threshold of the classifier is set to 0.8 to obtain accurate target expression information. The samples that meet the threshold are used to train a new classifier, and the above operations are repeated until the proportion of correct samples obtained according to the classification criterion reaches the set value in the total output samples. At this time, the classifier has tended to be stable and can accurately predict the position information of the target object, thereby increasing the proportion of correct label outputs to obtain a large number of effective annotated samples.

[0104] In the embodiment of the present application, a small number of annotated images are used to train the classifier to annotate the unannotated images based on the classifier, and the augmented images are obtained. Based on this, the problem that the detection model has poor generalization ability due to insufficient "low, small, and slow" samples is avoided, and the generalization ability of the detection model is improved.

[0105] To further illustrate the object detection method provided by the present application, refer to Figure 7 , Figure 7 is the second flow schematic diagram of the object detection method provided by the present application.

[0106] The purpose of the embodiment of the present application is to improve based on the YoloV5 network, so that the improved network is more suitable for the "low, small, and slow" target detection task in the air perspective, especially for problems such as small samples, small targets, and complex backgrounds. Aiming at the problem of lack of "low, small, and slow" datasets, the embodiment of the present application provides sufficient data support for the training of the detection model by augmenting the sample size and augmenting the data; aiming at the problem that the "low, small, and slow" target size is too small, the embodiment of the present application strengthens the feature extraction ability of the model for small targets in the air.

[0107] In the embodiment of the present application, based on the YoloV5 network model, in the sample processing stage, a semi-supervised sample augmentation technique is introduced to complement the missing annotation information in the self-collected dataset, avoiding the problem of poor generalization ability of the model caused by insufficient "low, small, and slow" samples. In the data augmentation stage, a diffusion model is introduced to provide sufficient data for the detection network. First, through the target-background separation technique, the diffusion model augments the target itself, excluding the interference of complex backgrounds on generating realistic images. Then, the augmented "low, small, and slow" images are randomly reconstructed with the original background information to restore the complexity of the empty images and improve the anti-interference ability of the subsequent detection network. In the target detection stage, according to the specific size of the "low, small, and slow" targets, a small target detection network based on YoloV5 is designed, and the P5 output layer that does not match the size of the "low, small, and slow" targets is trimmed to optimize the network structure, maximizing the training of the network model and the target detection rate. That is, by optimizing the network structure to design a lightweight framework, reducing parameters to reduce the computational amount. At the same time, a multi-scale attention aggregation module with strong feature perception ability for small-size targets is designed to enhance the functional and expressive capabilities of network feature extraction, effectively solving the problems brought by small and weak targets in the detection stage.

[0108] As Figure 7 shown, the steps of target detection by the YoloV5 small aerial target detection network based on the diffusion model provided by the embodiment of the present application include:

[0109] (1) Semi-supervised annotation sample augmentation.

[0110] The self-training sample augmentation technique in the embodiment of the present application selects YoloV5 as the base classifier, manually annotates some images in the self-collected dataset as the input for the training of the classifier model, captures the remaining "low, small, and slow" targets lacking annotation information, sets the output threshold of the classifier to 0.8 to obtain accurate target expression information, and uses the samples that meet the threshold to train a new classifier. Repeat the above operations until the proportion of correct samples obtained according to the classification criterion in the total output samples reaches the set value. At this time, the classifier has tended to be stable and can accurately predict the position information of the target object, thereby increasing the proportion of correct label output to obtain a large number of effective annotation samples.

[0111] (2) Data augmentation based on the diffusion model.

[0112] First, slice the images augmented by the semi-supervised learning model. According to the size of the "low, small, and slow" target, set the slice resolution range from 15*15 to 30*30 to ensure that the slices contain the complete "low, small, and slow" target and its important features. Then, use the "low, small, and slow" slices as the input of the diffusion model, and perform data augmentation on the "low, small, and slow" slices through the diffusion model. Finally, randomly reconstruct the "low, small, and slow" target slices generated by the diffusion model with the original images, such as randomly combining the slices with complex backgrounds, and at the same time reducing and enlarging the slices by 10%-30% to simulate the "low, small, and slow" targets of different scales in the near and far scenes from the air perspective. Its essence is to restore the interference factors of background objects in the image area irrelevant to the detection, and improve the anti-interference ability of the detection model.

[0113] (3) Single-stage small target detection network based on multi-scale aggregation.

[0114] Crop the P5 output layer that cannot sensitively capture features for small targets, and at the same time introduce a multi-scale attention aggregation module to output a feature map that better matches the feature information of small targets, providing more learnability for the detection model. As Figure 7 shown, first perform upsampling and connect it to the top-down path in the lightweight network to enrich the strong semantic feature information in the high-level of the network. Secondly, achieve information sharing through horizontal connection and the bottom-up path in the lightweight network, and perform information aggregation to complement the target location information ignored in the top-down path. Through the above two information aggregations, the multi-scale attention aggregation module finally outputs a feature map with a size of 152*152 and 255 channels. The feature map of this size matches the "low, small, and slow" target and can fully express the feature information of small targets, providing a reliable training basis for the detection network and solving the problem that the network is not robust enough for small target detection.

[0115] The deep learning detection method for "low, small, and slow" targets based on small samples proposed in the embodiments of this application optimizes the network structure without significantly increasing the number of parameters, improves the detection accuracy of small targets in the air, and is superior to other network models, such as YoloV5, Resnet50, Faster R-CNN, etc. The embodiments of this application not only solve the problems existing in the detection of "low, small, and slow" targets from the air perspective, but also enhance the generalization ability of the network and improve the detection accuracy of the model to a certain extent.

[0116] Figure 8 is a schematic structural diagram of the target detection device provided by this application. Referring to Figure 8 , the embodiments of this application provide a target detection device, including a sample augmentation module 801, a data augmentation module 802, and a target detection module 803.

[0117] The sample augmentation module 801 is used to perform labeled sample augmentation on the images in the dataset of the target object to obtain augmented images;

[0118] The data augmentation module 802 is used to perform data augmentation on the augmented images, and randomly reconstruct the first augmented images obtained by data augmentation and the background images to obtain reconstructed images;

[0119] The target detection module 803 is used to input the reconstructed images into the detection model to obtain the target object output by the detection model;

[0120] Among them, the detection model is obtained by training a lightweight single-stage small target detection network with multi-scale aggregation using sample images; the lightweight single-stage small target detection network with multi-scale aggregation includes a lightweight network and a multi-scale attention aggregation module; the output layer that cannot recognize the feature information of the target object has been deleted in the lightweight network; the lightweight network is connected to the multi-scale attention aggregation module to aggregate the feature information of the reconstructed images.

[0121] In one embodiment, the lightweight network includes a first path and a second path. The first path is spliced with the multi-scale attention aggregation module to increase the semantic feature information of the reconstructed images; the second path is convolutionally connected to the multi-scale attention aggregation module to aggregate the feature information of the reconstructed images.

[0122] In one embodiment, the data augmentation module 802 specifically includes:

[0123] Slice the augmented images to obtain the background images and target images carrying the target object;

[0124] Perform data augmentation on the target images.

[0125] In one embodiment, the data augmentation includes forward augmentation and reverse augmentation;

[0126] The data augmentation module 802 specifically includes:

[0127] Determine the first noise signal for forward augmentation and the second noise signal for reverse augmentation;

[0128] Perform forward augmentation on the target images based on the first noise signal to obtain second augmented images;

[0129] Perform reverse augmentation on the second augmented images based on the second noise signal to obtain the first augmented images.

[0130] In one embodiment, the expression of the first noise signal is:

[0131]

[0132] Among them, q(X t |X0) represents the first noise signal, X0 represents the input of the diffusion model, and X t represents the output of the diffusion model, I represents the identity matrix, N represents a constant, represents the product result of Markov;

[0133] The expression of the second noise signal is:

[0134]

[0135] Among them, q(X t-1 |X t X0) represents the second noise signal, X0 represents the input of the diffusion model, and X t represents the output of the diffusion model, and X t-1 represents adding noise once based on X t , α t represents the distance between 1 and the step variable of X t , β t represents the variable affecting the diffusion step size, represents the product result of Markov, represents the product result of Markov, Z~N(0, I) represents the normal distribution, I represents the identity matrix, and Z and N represent constants.

[0136] In one embodiment, the sample augmentation module 801 specifically includes:

[0137] Obtain the labeled images and unlabeled images in the dataset;

[0138] Use the labeled images to train a classifier, and label the unlabeled images based on the classifier to obtain the augmented images.

[0139] In one embodiment, the data augmentation module 802 specifically includes:

[0140] Randomly combine the first augmented image with the background image to obtain a combined image, and use the combined image as the reconstructed image.

[0141] Figure 9 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 9As shown in the figure, the electronic device may include: a processor 910, a communications interface 920, a memory 930, and a communication bus 940. Among them, the processor 910, the communications interface 920, and the memory 930 complete communication with each other through the communication bus 940. The processor 910 may call the logical instructions in the memory 930 to execute a target detection method, which includes:

[0142] Performing annotation sample augmentation on the images in the dataset of the target object to obtain augmented images;

[0143] Performing data augmentation on the augmented images, and randomly reconstructing the first augmented images obtained by data augmentation and background images to obtain reconstructed images;

[0144] Inputting the reconstructed images into a detection model to obtain the target object output by the detection model;

[0145] Among them, the detection model is obtained by training a lightweight single-stage small target detection network with multi-scale aggregation using sample images; the lightweight single-stage small target detection network with multi-scale aggregation includes a lightweight network and a multi-scale attention aggregation module; the output layer that cannot recognize the feature information of the target object has been deleted in the lightweight network; the lightweight network is connected to the multi-scale attention aggregation module to aggregate the feature information of the reconstructed images.

[0146] In addition, when the logical instructions in the above-mentioned memory 930 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application essentially or the part that contributes to the prior art or a part of this technical solution may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0147] On the other hand, this application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the target detection method provided by the above-mentioned various methods. The method includes:

[0148] Perform annotation sample augmentation on the images in the dataset of the target object to obtain augmented images;

[0149] Perform data augmentation on the augmented images, and randomly reconstruct the first augmented images obtained by data augmentation and the background images to obtain reconstructed images;

[0150] Input the reconstructed images into the detection model to obtain the target object output by the detection model;

[0151] Among them, the detection model is trained by using sample images on a lightweight single-stage small target detection network with multi-scale aggregation; the lightweight single-stage small target detection network with multi-scale aggregation includes a lightweight network and a multi-scale attention aggregation module; the output layer that cannot recognize the feature information of the target object has been deleted in the lightweight network; the lightweight network is connected to the multi-scale attention aggregation module to aggregate the feature information of the reconstructed images.

[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0153] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the essence of the above technical solutions or the part that contributes to the prior art can be embodied in the form of a software product, and this computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0154] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A target detection method, characterized in that, Including: Performing annotation sample augmentation on the images in the dataset of the target object to obtain augmented images; Performing data augmentation on the augmented images, and randomly reconstructing the first augmented images obtained by data augmentation and background images to obtain reconstructed images; Inputting the reconstructed images into a detection model to obtain the target object output by the detection model; Among them, the detection model is trained by using sample images on a lightweight single-stage small target detection network with multi-scale aggregation; the lightweight single-stage small target detection network with multi-scale aggregation includes a lightweight network and a multi-scale attention aggregation module; the output layer that cannot recognize the feature information of the target object has been deleted in the lightweight network; the lightweight network is connected to the multi-scale attention aggregation module to aggregate the feature information of the reconstructed images; The performing data augmentation on the augmented images includes: Performing image slicing on the augmented images to obtain the background images and target images carrying the target object; Performing data augmentation on the target images; the data augmentation includes forward augmentation and reverse augmentation; The performing data augmentation on the target images includes: Determining the first noise signal for forward augmentation and the second noise signal for reverse augmentation; Performing forward augmentation on the target images based on the first noise signal to obtain second augmented images; Performing reverse augmentation on the second augmented images based on the second noise signal to obtain the first augmented images; The expression of the first noise signal is: ; Among them, represents the first noise signal, represents the input of the diffusion model, represents the output of the diffusion model, represents the identity matrix, represents a constant, represents the product result of Markov; The expression of the second noise signal is: ; Among them, represents the second noise signal, represents the input of the diffusion model, represents the output of the diffusion model, represents at the basis of removing noise once, represents the distance between 1 and the step size variable, represents the variable affecting the diffusion step size, represents the product result of Markov, represents the product result of Markov, represents the normal distribution, represents the identity matrix, and represents a constant; The randomly reconstructing the first augmented images obtained by data augmentation and background images to obtain reconstructed images includes: Randomly combining the first augmented images and the background images to obtain combined images, and using the combined images as the reconstructed images.

2. The object detection method according to claim 1, wherein The lightweight network includes a first path and a second path. The first path is spliced with the multi-scale attention aggregation module to increase the semantic feature information of the reconstructed images; the second path is convolutionally connected to the multi-scale attention aggregation module to aggregate the feature information of the reconstructed images.

3. The object detection method according to claim 1, wherein The performing annotation sample augmentation on the images in the dataset of the target object to obtain augmented images includes: Obtaining the annotated images and unannotated images in the dataset; Training a classifier with the annotated images to annotate the unannotated images based on the classifier to obtain the augmented images.

4. A target detection device, characterized in that, Including: A sample augmentation module for performing annotation sample augmentation on the images in the dataset of the target object to obtain augmented images; A data augmentation module for performing data augmentation on the augmented images, and randomly reconstructing the first augmented images obtained by data augmentation and background images to obtain reconstructed images; A target detection module for inputting the reconstructed images into a detection model to obtain the target object output by the detection model; Among them, the detection model is obtained by training a lightweight single-stage small object detection network with multi-scale aggregation using sample images; the lightweight single-stage small object detection network with multi-scale aggregation includes a lightweight network and a multi-scale attention aggregation module; the output layer that cannot recognize the feature information of the target object has been deleted in the lightweight network; the lightweight network is connected to the multi-scale attention aggregation module to aggregate the feature information of the reconstructed image; The data augmentation module is further configured to perform image slicing on the augmented image to obtain the background image and the target image carrying the target object; perform data augmentation on the target image; the data augmentation includes forward augmentation and reverse augmentation; The data augmentation module is further configured to determine a first noise signal for the forward augmentation and a second noise signal for the reverse augmentation; perform forward augmentation on the target image based on the first noise signal to obtain a second augmented image; perform reverse augmentation on the second augmented image based on the second noise signal to obtain the first augmented image; The expression of the first noise signal is: ; Among them, represents the first noise signal, represents the input of the diffusion model, represents the output of the diffusion model, represents the identity matrix, represents a constant, represents the product result of Markov; The expression of the second noise signal is: ; Among them, represents the second noise signal, represents the input of the diffusion model, represents the output of the diffusion model, represents at the basis of removing noise once, represents the distance between 1 and the step size variable, represents the variable affecting the diffusion step size, represents the product result of Markov, represents the product result of Markov, represents the normal distribution, represents the identity matrix, and represents a constant; The data augmentation module is further configured to randomly combine the first augmented image and the background image to obtain a combined image, and use the combined image as the reconstructed image.

5. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the object detection method according to any one of claims 1 to 3.

6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the object detection method according to any one of claims 1 to 3.