Unmanned aerial vehicle target detection method, device and equipment based on improved YOLOv9 model, medium and product

By improving the YOLOv9 model to the RepViT network and introducing a two-layer routing attention mechanism and Slide Loss loss function, the accuracy and robustness of drone target detection in complex environments are solved, and efficient and accurate drone target recognition is achieved.

CN120236067APending Publication Date: 2025-07-01HANGZHOU DIANZI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510648641.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In complex environments, the accuracy and robustness of drone target detection are insufficient. Especially when the background is complex, the lighting changes are large, and the contrast between the drone and the background is low, it is difficult for the prior art to achieve high accuracy detection.

Method used

Using the improved YOLOv9 model, the improved YOLOv9 model is constructed by setting the backbone network as a RepViT network and inserting the double-layer routing attention mechanism network afterwards, combined with the Slide Loss loss function, and data augmentation and training of drone images are carried out.

Benefits of technology

It improves the accuracy and robustness of drone target detection, enhances the model's detection ability of small targets, reduces computing costs, and is suitable for drone detection tasks in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236067A_ABST
    Figure CN120236067A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle target detection method and device based on an improved YOLOv9 model, equipment, a medium and a product, and relates to the technical field of target detection, and the method comprises the steps: obtaining unmanned aerial vehicle images under different environment conditions, carrying out the marking, and carrying out the data enhancement to form a training data set; setting a backbone network of the improved YOLOv9 model as a RepViT network, and inserting a double-layer routing attention mechanism network behind the RepViT network to construct an improved model; training the improved model by using the training data set to obtain a trained improved model; and inputting a to-be-detected unmanned aerial vehicle image into the trained improved model to obtain a prediction bounding box and a prediction confidence coefficient. According to the invention, the RepViT network and the double-layer routing attention mechanism are introduced, so that the detection precision and robustness of the unmanned aerial vehicle target are improved, and the unmanned aerial vehicle detection performance in a complex environment is improved at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of object detection, and particularly to an unmanned aerial vehicle object detection method, device, equipment, medium and product based on an improved YOLOv9 model. Background Technique

[0002] With the rapid development and wide application of unmanned aerial vehicle (UAV) technology, UAVs play an increasingly important role in multiple fields. From agricultural monitoring, logistics distribution to environmental monitoring, film and television production, the application of UAVs has greatly improved work efficiency and quality. However, the wide application of UAVs has also brought new challenges, especially in terms of security monitoring and privacy protection. The miniaturization and concealment of UAVs make their detection in complex environments particularly difficult, which poses higher requirements on existing monitoring systems.

[0003] In the field of UAV object detection, although various methods have been proposed, such as object detection algorithms based on deep learning, multi-sensor fusion technology, etc., accurate and efficient UAV detection in complex environments is still a challenge. Especially in the case of complex backgrounds, large changes in illumination, and low contrast between UAVs and backgrounds, related technologies are difficult to achieve high-accuracy detection.

[0004] Based on this, the present application proposes an unmanned aerial vehicle object detection method based on an improved YOLOv9 model to improve the detection accuracy and robustness of UAVs in complex environments. Summary of the Invention

[0005] The purpose of the present application is to provide an unmanned aerial vehicle object detection method, device, equipment, medium and product based on an improved YOLOv9 model, which can improve the accuracy, robustness and computational efficiency of UAV object detection and is applicable to UAV detection tasks in complex environments.

[0006] To achieve the above object, the present application provides the following solutions:

[0007] In a first aspect, the present application provides an unmanned aerial vehicle object detection method based on an improved YOLOv9 model, including:

[0008] Obtaining UAV images under different environmental conditions, and annotating the true bounding boxes and true confidence levels of UAVs in the UAV images; the environmental conditions include daytime, night, sunny day, rainy day, foggy day;

[0009] Performing data augmentation on the UAV images, and jointly constituting a UAV image training data set with the UAV regions in the annotated UAV images;

[0010] Set the backbone network of the improved YOLOv9 model to the RepViT network, and insert a double-layer routing attention mechanism network after the RepViT network to construct the improved YOLOv9 model;

[0011] Use the drone image training dataset to train the improved YOLOv9 model to obtain the trained improved YOLOv9 model;

[0012] Input the drone image to be detected into the trained improved YOLOv9 model to obtain the predicted bounding box and predicted confidence of the drone to be detected.

[0013] In a second aspect, the present application provides a drone target detection device based on the improved YOLOv9 model, including:

[0014] An image acquisition and annotation module for acquiring drone images under different environmental conditions and annotating the true bounding box and true confidence of the drones in the drone images; the environmental conditions include day, night, sunny, rainy, and foggy days;

[0015] A dataset construction module for performing data augmentation on the drone images and jointly constructing a drone image training dataset with the drone regions in the annotated drone images;

[0016] A model improvement and construction module for setting the backbone network of the improved YOLOv9 model to the RepViT network and inserting a double-layer routing attention mechanism network after the RepViT network to construct the improved YOLOv9 model;

[0017] A model training module for using the drone image training dataset to train the improved YOLOv9 model to obtain the trained improved YOLOv9 model;

[0018] A drone prediction module for inputting the drone image to be detected into the trained improved YOLOv9 model to obtain the predicted bounding box and predicted confidence of the drone to be detected.

[0019] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the steps of the drone target detection method based on the improved YOLOv9 model described in any one of the above.

[0020] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the drone target detection method based on the improved YOLOv9 model described in any one of the above are implemented.

[0021] In a fifth aspect, the present application provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the method for detecting drone targets based on the improved YOLOv9 model described in any one of the above.

[0022] According to the specific embodiments provided by the present application, the present application has the following technical effects:

[0023] The present application provides a method, device, equipment, medium and product for detecting drone targets based on an improved YOLOv9 model. By acquiring drone images under different environmental conditions and performing annotation, the problem of sample diversity in drone detection in diverse environments is solved, and accurate recognition of drone targets in multiple complex environments is achieved; by performing data augmentation on the drone images and jointly constructing a training data set with the annotated drone regions, the problem of sample imbalance is solved, and the recognition ability of the model for minority classes is improved; by setting the backbone network of the improved YOLOv9 model as a RepViT network and inserting a double-layer routing attention mechanism network after the RepViT network to construct the improved YOLOv9 model, the limitations of the traditional YOLOv9 model in feature extraction and attention focusing are solved, more efficient feature extraction and more accurate target localization are achieved, and the accuracy, robustness and computational efficiency of drone target detection are improved. Description of the Drawings

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0025] Figure 1 It is an application environment diagram of a method for detecting drone targets based on an improved YOLOv9 model in an embodiment of the present application;

[0026] Figure 2 It is a flowchart of a method for detecting drone targets based on an improved YOLOv9 model provided by an embodiment of the present application;

[0027] Figure 3 It is a network structure diagram of an improved YOLOv9 model provided by an embodiment of the present application;

[0028] Figure 4 It is a schematic diagram of a double-layer routing attention mechanism network provided by an embodiment of the present application;

[0029] Figure 5The experimental effect diagram of drone detection for the improved YOLOv9 model provided by an embodiment of this application;

[0030] Figure 6 The schematic diagram of the functional modules of a drone target detection device based on the improved YOLOv9 model provided by an embodiment of this application;

[0031] Figure 7 The schematic diagram of the structure of a computer device provided by an embodiment of this application. Detailed implementation manners

[0032] First, the technical terms involved in this application are introduced.

[0033] With the rapid development of deep learning, the target detection technology has evolved from the method based on artificial feature extraction to the end-to-end detection method based on deep neural networks. The deep neural network can automatically learn image features and improve the accuracy and robustness of target detection. According to different detection frameworks, target detection algorithms can be divided into two-stage detection algorithms and one-stage detection algorithms.

[0034] The two-stage detection algorithm (Two-stage Detector) usually generates regions of interest through candidate region proposals (Region Proposal), and performs target classification and location regression on this basis, such as R-CNN (Region- based ConvolutionalNeuralNetwork, Region-based Convolutional Neural Network), FastR-CNN (Fast Region-based Convolutional NeuralNetwork, fast region convolutional neural network), Faster R-CNN (Faster Region-based Convolutional NeuralNetwork, faster region convolutional neural network), Mask R-CNN (Mask Region-based Convolutional NeuralNetwork, mask region convolutional neural network), etc.

[0035] The one-stage detection algorithm (One-stage Detector) directly performs target detection on the entire image, regarding the detection task as a regression problem. Representative algorithms include YOLO (You Only Look Once), SSD (SingleShotMultiBox, single-shot multi-box detector), RetinaNet, etc. The one-stage detection method has a wide application prospect in the field of drone detection due to its fast detection speed and suitability for real-time detection tasks.

[0036] For drone target detection, the traditional YOLO series algorithms still have the following problems:

[0037] 1. Insufficient small target detection ability: When the YOLO series of algorithms detect small targets, due to the low resolution, features are easily lost, resulting in a decrease in detection accuracy.

[0038] 2. Background interference problem: Drones usually operate in complex backgrounds, such as environments with vegetation, sky, water bodies, etc., making target detection easily affected by interference.

[0039] 3. Computational complexity and efficiency problem: High-precision target detection algorithms often have a large amount of computation and are difficult to run efficiently on resource-constrained devices.

[0040] To solve the above problems, researchers have proposed various improvement strategies, such as optimizing the backbone network, introducing attention mechanisms, improving loss functions, etc. In recent years, RepViT (Re-parameterized Vision Transformer) as a lightweight and efficient backbone network, which introduces re-parameterization technology, improves the inference speed and accuracy of the network, and is especially suitable for embedded devices and edge computing environments. In addition, BiFormer (Bilateral Routing Transformer) as a two-level routing attention mechanism can establish effective connections between local and global features, improving the model's detection ability for small targets.

[0041] In the drone detection task, since drones usually fly in complex environments, its detection task faces the problem of sample imbalance, that is, the target area is usually small and the number is small, while the background area occupies most of the space, which may cause the loss function in related technologies to be biased towards the background category during training, reducing the detection accuracy. Slide Loss, as a loss function for solving the sample imbalance problem, can dynamically adjust the loss weight, enhance the detection ability for small targets, and improve the detection robustness.

[0042] In summary, this application proposes a drone target detection method based on an improved YOLOv9 model, combined with the Slide Loss loss function, to improve the accuracy, robustness, and computational efficiency of drone target detection, and is applicable to drone detection tasks in complex environments.

[0043] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0044] To make the above objects, features, and advantages of the present application more apparent and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0045] The drone target detection method based on the improved YOLOv9 model provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send drone images under different environmental conditions to the server 104. After receiving the drone images under different environmental conditions, the server 104 annotates the true bounding boxes and true confidence levels of the drones in the drone images; the environmental conditions include daytime, night, sunny, rainy, and foggy; perform data augmentation on the drone images, and jointly form a drone image training dataset with the drone regions in the annotated drone images; set the backbone network of the improved YOLOv9 model as the RepViT network, and insert a double-layer routing attention mechanism network after the RepViT network to construct the improved YOLOv9 model; use the drone image training dataset to train the improved YOLOv9 model to obtain a trained improved YOLOv9 model; input the drone image to be detected into the trained improved YOLOv9 model to obtain the predicted bounding box and predicted confidence level of the drone to be detected. The server 104 can feedback the predicted bounding box and predicted confidence level of the drone to be detected to the terminal 102. In addition, in some embodiments, the drone target detection method based on the improved YOLOv9 model can also be implemented by the server 104 or the terminal 102 alone. For example, the terminal 102 can directly perform model training and optimization processing on drone images under different environmental conditions, or the server 104 can obtain drone images under different environmental conditions from the data storage system and perform model training and optimization processing on these data.

[0046] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smartphones, tablets, Internet of Things devices, and portable wearable devices. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.

[0047] In an exemplary embodiment, as Figure 2 shown, a drone target detection method based on the improved YOLOv9 model is provided. This method is executed by a computer device, and can be specifically executed by a computer device such as a terminal or a server alone, or jointly executed by a terminal and a server. In the embodiments of the present application, this method is applied toFigure 1 Taking the server 104 in

[0048] Step 201: Obtain drone images under different environmental conditions, and label the true bounding boxes and true confidence levels of the drones in the drone images; the environmental conditions include day, night, sunny, rainy, and foggy days.

[0049] Obtain drone target image data under different environments, weather conditions, lighting changes, and complex backgrounds, including various scenarios such as day, night, sunny, rainy, and foggy days, to ensure the diversity and generalization ability of the data. Label the collected data, including true bounding boxes, true confidence levels, class labels (only for possible subsequent extensions), and target size information, to establish a high-quality data set.

[0050] In another exemplary embodiment of the present application, the calculation formula for the confidence level is:

[0051] p i = σ[FC(F grid,i )].

[0052] Wherein, p i represents the true confidence level of the i-th drone image; σ represents the Sigmoid function, and F grid,i is the grid cell feature of the drone area in the i-th drone image, and FC represents the fully connected layer.

[0053] Step 202: Perform data augmentation on the drone images, and jointly form a drone image training data set with the labeled drone areas in the drone images.

[0054] In another exemplary embodiment of the present application, performing data augmentation on the drone images specifically includes:

[0055] Perform data augmentation on the drone images through rotation, scaling, brightness adjustment, adding Gaussian noise, the MixUp method, or the Mosaic method. Data augmentation methods can be used to expand the data set, such as rotation, scaling, brightness adjustment, Gaussian noise, etc., to further improve the model's adaptability to real scenarios.

[0056] Step 203: Set the backbone network of the improved YOLOv9 model to the RepViT network, and insert a double-layer routing attention mechanism network after the RepViT network to construct an improved YOLOv9 model.

[0057] In another exemplary embodiment of the present application, the loss function of the improved YOLOv9 model in Step 203 is SlideLoss, specifically including:

[0058]

[0059] Among them, n represents the total number of UAV images in the UAV image training dataset; λ i represents the weight of the i-th UAV image in the UAV image training dataset; Loss i is the loss value of the i-th UAV image in the UAV image training dataset; λ′ represents the balance coefficient; α represents the positive sample weight; γ represents the negative sample weight; p i represents the predicted confidence of the i-th UAV image; IoU i represents the intersection over union of the predicted bounding box and the ground truth bounding box in the i-th UAV image; c i represents the length of the diagonal of the smallest closed bounding box containing the predicted bounding box and the ground truth bounding box in the i-th UAV image; ρ i represents the Euclidean distance from the center point of the predicted bounding box to the center point of the ground truth bounding box in the i-th UAV image; w i represents the width of the predicted bounding box in the i-th UAV image; h i represents the height of the predicted bounding box in the i-th UAV image.

[0060] Step 204: Use the UAV image training dataset to train the improved YOLOv9 model to obtain the trained improved YOLOv9 model.

[0061] In another exemplary embodiment of the present application, the training process of the improved YOLOv9 model specifically includes:

[0062] Use the UAV images in the UAV image training dataset as the input of the improved YOLOv9 model, output the predicted bounding box and predicted confidence of the UAV image, and train the improved YOLOv9 model until the difference between the predicted bounding box in the UAV image and the corresponding ground truth bounding box is less than the first preset threshold, and the difference between the predicted confidence in the UAV image and the corresponding ground truth confidence is less than the second preset threshold, then stop training.

[0063] In another exemplary embodiment of the present application, the progressive training method is used to train the improved YOLOv9 model, which specifically includes:

[0064] In the first preset period, freeze the RepViT backbone network of the improved YOLOv9 model, and only train the detection head module and FPN module of the improved YOLOv9 model.

[0065] In the second preset period, unfreeze all network parameters and train all modules of the improved YOLOv9 model, including the RepViT backbone network, detection head module and FPN module.

[0066] During the third preset period, the difficult sample mining method is adopted to increase the weights of negative samples with prediction confidence less than the third preset threshold, and all modules of the improved YOLOv9 model, including the RepViT backbone network, the detection head module, and the FPN module, are trained.

[0067] As an alternative implementation, when training the improved YOLOv9 model, an adaptive optimization algorithm (such as AdamW (Adaptive Moment Estimation with Weight Decay, adaptive moment estimation optimization algorithm with weight decay) or SGD (Stochastic Gradient Descent)) can be used to optimize network parameters, including biases, normalization means, and variances. During the training process, data augmentation strategies, such as MixUp augmentation and Mosaic augmentation, are combined to further improve the generalization ability of the model. At the same time, Slide Loss is used to calculate the loss and optimize model parameters, such as the learning rate, Biformer local and global routing parameters, etc., to make the detection accuracy and robustness reach the optimal state.

[0068] Step 205: Input the drone image to be detected into the trained improved YOLOv9 model to obtain the predicted bounding box and predicted confidence of the drone to be detected.

[0069] In practical applications, the trained improved YOLOv9 model is used for inference. After inputting the drone image to be detected, through feature extraction, attention enhancement, multi-scale feature fusion, and final detection head prediction, the target detection box and confidence information are output to achieve accurate detection of the drone target. To improve the real-time performance of detection, inference acceleration tools such as TensorRT and OpenVINO can be used for deployment.

[0070] By implementing the above steps 201 to 205, the present application can accurately identify and locate drone targets under complex backgrounds and different lighting conditions, providing technical support for drone monitoring and management. In addition, while reducing the computational cost, the present application enhances the adaptability and detection accuracy of the model, and is applicable to various application scenarios such as low-altitude security and drone supervision.

[0071] In another exemplary embodiment of the present application, the network structure of the improved YOLOv9 model is as Figure 3As shown in the figure, the original backbone network of the YOLOv9 model is replaced with RepViT to utilize its lightweight convolutional module to improve the inference speed and enhance the feature extraction ability at the same time. RepViT combines the structure re-parameterization technology and adopts a more optimized computational graph structure in the inference stage. It uses a multi-branch structure and equivalently converts multiple parallel convolutional layers into a single Conv2D (2-Dimensional Convolutional Layer), which improves the detection speed and reduces the computational resource requirements.

[0072] As an optional implementation, the RepViT backbone network adopts the RepVGG (Re-parameterized Visual Geometry Group) structure for structure re-parameterization to improve the inference efficiency of the model and maintain a high detection accuracy at the same time. Compared with traditional backbone networks such as ResNet (Residual Network), RepViT reduces the computational amount during inference, making the deployment more lightweight and efficient.

[0073] The BiFormer attention mechanism is introduced during the feature extraction process to combine local and global information and improve the detection ability for small targets in complex backgrounds. BiFormer combines the Bi-level Routing Attention mechanism, making the feature extraction more accurate and helping to distinguish the UAV targets from complex backgrounds. The BiFormer attention mechanism effectively enhances the feature expression ability and improves the detection performance for small UAVs through the combination of local routing and global routing. BiFormer adopts the self-attention mechanism in the Transformer architecture and combines the bi-level routing mechanism, reducing the computational amount while maintaining a strong feature extraction ability. The schematic diagram of the structure of the Biformer attention mechanism network is as Figure 4 shown. The Biformer attention mechanism network structure is used to efficiently process data such as UAV images to enhance the model's feature extraction and representation ability. It consists of an input module, an attention calculation module, and an output module. The input module receives a UAV image with a width of W, a height of H, and a channel number of C, and divides it into multiple windows with a window size of S as the basis for subsequent calculations. The attention calculation module first performs a linear transformation on the input features to generate a query matrix Q, and performs aggregation operations on the key matrix K and the value matrix V respectively to obtain the global key matrix K g and the global value matrix V g , where Q is used to query information, K g is used to calculate the similarity with Q to determine the position correlation, and V gFor weighted summation. Then, by calculating the dot product of Q and K g and normalizing it through the softmax function, the attention matrix A representing the attention weights between positions is obtained; finally, A and V g are subjected to matrix multiplication (mm operation) to obtain the weighted feature representation. The output module takes the result processed by the attention calculation module as the output feature O, which can be used for further processing in subsequent network layers. Where W and H are the width and height of the UAV image respectively, S is the window size, C is the number of channels, O is the output feature, A is the attention matrix, Q is the query matrix, V g is the global value matrix, K g is the global Key matrix, and mm is matrix multiplication.

[0074] Feature Pyramid Networks (FPN) is adopted for multi-scale feature fusion to enhance the feature expression ability of small targets. FPN combines deep semantic information and shallow detail information by fusing feature information at different levels, enhances target features, and improves the detection effect. By introducing additional skip connections, the detection accuracy can be further improved.

[0075] The Slide Loss function is used to optimize the detection head to alleviate the impact of positive and negative sample imbalance on detection performance. Slide Loss dynamically adjusts the weights of positive and negative samples, not only reducing the interference of excessive negative samples on the detection effect, improving the recall rate of small target UAVs, but also reducing the impact of extreme sample distributions on detection accuracy, and improving the adaptability of the improved YOLOv9 model to UAVs in small targets and complex backgrounds.

[0076] In another exemplary embodiment of the present application, the UAV target detection method based on the improved YOLOv9 model is implemented in the Ubuntu 18.04 environment, and the NVIDIA RTX 2080Ti graphics card is used to accelerate training. The experimental data includes images of UAV targets in different weather, lighting, and complex backgrounds.

[0077] The UAV target detection method based on the improved YOLOv9 model includes:

[0078] Step S1. Dataset construction and enhancement.

[0079] Data collection covers multiple scenarios, and complex backgrounds such as day / night, sunny / rainy / foggy, city / forest / mountain are collected through the visible light camera (such as SONY IMX477) carried by the UAV. The Labelimg tool is used to annotate the target bounding boxes, and the annotation category is "UAV". Data verification ensures annotation consistency through cross-validation, and blurry, duplicate, or mislabeled images are removed. Finally, about 440,000 valid samples are retained.

[0080] Data preprocessing includes geometric transformations (rotation ±15°, scaling 0.8 - 1.2 times, horizontal / vertical flipping) and photometric adjustments (brightness adjustment ±20%, contrast adjustment ±15%, adding Gaussian noise σ = 0.1). Advanced enhancement methods include Mosaic enhancement (randomly selecting 4 images for stitching) and MixUp enhancement (mixing two images and their labels with a weight of λ = 0.5). The final dataset is divided into a training set (about 220,000 images), a validation set (90,000 images), and a test set (about 130,000 images) in a ratio of 5:2:3.

[0081] Step S1. Improve the construction of the YOLOv9 model.

[0082] (1) RepViT backbone network

[0083] In this application, the original CSPDarkNet of the YOLOv9 model is replaced with RepViT. During the training phase, it adopts a multi-branch structure (parallel 3×3 convolution, 1×1 convolution, Identity branch), and improves feature diversity through the structural reparameterization technique. During the inference phase, the multi-branches are merged into a single-branch 3×3 convolution, reducing the computational graph complexity and increasing the inference speed by more than 30%. The RepViT-M1.5 version has only 5.6M parameters and 1.8G FLOPs, reducing the computational volume by 45% compared to ResNet50, and is suitable for edge device deployment.

[0084] (2) BiFormer attention mechanism

[0085] Embed the BiFormer module on the feature map output by RepViT to improve the object detection ability with a two-level routing mechanism. Local routing uses a sliding window (7×7) to divide the feature map and calculates the self-attention within the window to capture details such as drone propellers and airframe edges. Global routing calculates the cross-window attention through sparse sampling (Top-K = 16) to enhance the correlation between the background and the object. In addition, dynamic routing mechanism is adopted for computational optimization, and global attention calculation is only performed on important regions, reducing the memory occupancy by about 60% compared to the standard Transformer.

[0086] (3) Multi-scale feature fusion

[0087] Based on the original PANet of YOLOv9, add cross-layer skip connections. For shallow feature fusion, the feature maps of the C3 layer (high resolution, low semantics) and the C5 layer (low resolution, high semantics) are aligned in channels through 1×1 convolution and then stitched together, and the Feature Pyramid Network (FPN) is applied to extract multi-scale receptive field information using different scales of max pooling (5×5, 9×9, 13×13, 17×17).

[0088] (4) Slide Loss Loss function

[0089] Dynamic weight adjustment balances the loss terms by calculating the positive and negative sample ratios.

[0090] The positive sample weight formula is: α = N neg / (N pos + N neg ).

[0091] The negative sample weight formula is: γ = N pos / (N pos + N neg ).

[0092] Among them, α is the positive sample weight; N neg is the number of negative samples; N pos is the number of positive samples; γ is the negative sample weight.

[0093] Loss calculation includes classification loss (using Focal Loss, positive sample weight α = 0.25, negative sample weight γ = 2.0) and regression loss (using CIoU Loss), and the balance coefficient λ = 0.5. To ensure the optimal balance between accuracy and convergence speed.

[0094] Step S3. Model training and optimization.

[0095] The optimizer uses AdamW (initial learning rate is 3e -4 , weight decay is 0.05), the training period is 300 epochs, and the model weights are saved every 50 epochs. The Batch Size is set to 32 (single-card training), and the input image size is 640×640.

[0096] The key training strategies include progressive training. In the first 50 epochs, the RepViT backbone network is frozen, and only the detection head and FPN module are trained. From 50 to 200 epochs, all network parameters are unfrozen and the model is finely tuned. From 200 to 300 epochs, the hard sample mining strategy is enabled to strengthen the weights of negative samples with classification confidence <0.3. In addition, regularization uses Dropout (random inactivation) (ratio 0.2) and Label Smoothing (label smoothing) (smoothing parameter ε = 0.1) to prevent overfitting.

[0097] Step S4. Model deployment and inference acceleration.

[0098] The model is exported in ONNX format and quantized to FP16 through TensorRT, reducing the model size by 50%. Real-time detection with FPS≥30 is achieved on NVIDIA Jetson AGX Xavier. Memory access times during inference are reduced through layer fusion optimization (merging convolutional layers and BatchNorm layers (Batch Normalization Layer)).

[0099] The experimental effect diagram of the improved YOLOv9 model for drone detection is as Figure 5 shown. According to Figure 5 it can be seen that in complex backgrounds, the present application can accurately distinguish drones from windows and billboards with similar shapes in urban building complexes.

[0100] As an alternative implementation, for drones moving at high speeds (speed > 15m / s), stable tracking can be achieved through sliding window prediction and Kalman filter post-processing.

[0101] The present application also provides an application scenario that applies the above-mentioned drone target detection method based on the improved YOLOv9 model. Specifically: The drone target detection method based on the improved YOLOv9 model provided in this embodiment can be applied in low-altitude security monitoring scenarios. The low-altitude security monitoring scenario includes a data collection link, a target detection link, and an alarm and decision-making link. Image data enters the target detection link from the data collection link, undergoes detection and processing by the improved YOLOv9 model to obtain detection results, and enters the alarm and decision-making link. The drone target detection method based on the improved YOLOv9 model provided in this embodiment belongs to the target detection link. Specifically in the target detection link, the collected image data is input into the trained improved YOLOv9 model to detect and identify the target drones in the image, and after outputting the predicted bounding box and predicted confidence, the detection results enter the alarm and decision-making link. If an unauthorized drone is detected entering the no-fly zone, the system will immediately trigger an alarm and feedback the detected target information (such as position, flight direction, speed, etc.) to the monitoring center. The monitoring center conducts further analysis and decision-making based on this information, such as notifying security personnel or relevant regulatory departments to take measures.

[0102] Through the above process, the drone target detection method based on the improved YOLOv9 model provided in this embodiment plays an important role in the low-altitude security monitoring scenario, significantly improving the detection accuracy and robustness of small target drones in complex environments, while reducing the computational cost and enhancing the adaptability and real-time performance of the model.

[0103] The beneficial effects of the present application are as follows:

[0104] This application combines the RepViT backbone network, the BiFormer attention mechanism, and the Slide Loss function, improving the accuracy and generalization ability of UAV target detection. Especially in complex environments and small target detection tasks, it can effectively enhance the detection effect and reduce missed detections and false detections. Through the lightweight RepViT network structure, the computational cost is reduced and the inference speed is increased; through the BiFormer attention mechanism, the feature extraction ability is enhanced and the detection accuracy is improved; through the Slide Loss optimized loss function, the problem of sample imbalance is alleviated and the recall rate is increased. Finally, this application realizes an efficient, accurate, and robust UAV target detection method, which can operate stably under different environmental conditions and is applicable to multiple application scenarios such as low-altitude security and UAV supervision.

[0105] Based on the same inventive concept, an embodiment of this application also provides a UAV target detection device based on an improved YOLOv9 model for implementing the above-mentioned UAV target detection method involving the improved YOLOv9 model. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the UAV target detection device based on the improved YOLOv9 model provided below can refer to the limitations on the UAV target detection method based on the improved YOLOv9 model in the above text and will not be elaborated here.

[0106] In an exemplary embodiment, as Figure 6 shown, a UAV target detection device based on an improved YOLOv9 model is provided, including:

[0107] An image acquisition and annotation module 301, configured to obtain UAV images under different environmental conditions and annotate the true bounding boxes and true confidence levels of the UAVs in the UAV images; the environmental conditions include daytime, nighttime, sunny, rainy, and foggy days.

[0108] A dataset construction module 302, configured to perform data augmentation on the UAV images and jointly form a UAV image training dataset with the UAV regions in the annotated UAV images.

[0109] A model improvement and construction module 303, configured to set the backbone network of the improved YOLOv9 model as the RepViT network and insert a double-layer routing attention mechanism network after the RepViT network to construct the improved YOLOv9 model.

[0110] A model training module 304, configured to train the improved YOLOv9 model using the UAV image training dataset to obtain a trained improved YOLOv9 model.

[0111] The UAV prediction module 305 is configured to input the UAV image to be detected into the trained improved YOLOv9 model to obtain the predicted bounding box and predicted confidence of the UAV to be detected.

[0112] In another exemplary embodiment of the present application, another UAV target detection device based on the improved YOLOv9 model is provided, including a data management module, a model training module, and an inference deployment module.

[0113] The data management module supports batch uploading of images or video streams, is compatible with formats such as JPEG, PNG, and MP4, and automatically parses metadata such as shooting time and GPS (Global Positioning System) coordinates. At the same time, the device is built with a visualization tool for annotation verification to check whether there are missing or out-of-bounds problems with the annotation boxes and supports one-key correction. In addition, this module provides a GUI (Graphics User Interface), and the user can select different data augmentation strategies such as Mosaic and MixUp and preview the augmentation effect in real time.

[0114] The model training module has a visualization monitoring function, which can display key indicators such as loss curve, mAP (mean Average Precision), Recall, and Precision during the training process in real time, and supports comparative analysis of historical training records. The device also provides an automated hyperparameter search function, such as the optimization of learning rate and Batch Size, and recommends the optimal configuration based on the Bayesian optimization algorithm to improve the convergence speed and performance of the model.

[0115] The inference deployment module supports loading models in ONNX (Open Neural Network Exchange) and TensorRT formats and can automatically detect hardware compatibility such as GPU model and memory capacity. The system provides a RESTful API (Representational State Transfer API), which can receive image streams or single images input in the RTSP protocol (Real Time Streaming Protocol) and return detection results in JSON format, including target coordinates, confidence, and size classification. In addition, the device supports an alarm linkage mechanism. When an illegal UAV is detected, it can automatically trigger an audible and visual alarm or link to defense equipment such as a jammer gun.

[0116] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be asFigure 7 As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store UAV image processing data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a UAV target detection method based on an improved YOLOv9 model.

[0117] Those skilled in the art can understand that Figure 7 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout. In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0118] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0119] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0120] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0121] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, Resistive Random Access Memory (ReRAM), Magnetoresistive Random Access Memory (MRAM), Ferroelectric Random Access Memory (FRAM), Phase Change Memory (PCM), graphene memory, etc. Volatile memory can include Random Access Memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM), etc.

[0122] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0123] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0124] In this text, specific examples are used to elaborate on the principles and implementation manners of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application. At the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A drone target detection method based on an improved YOLOv9 model, characterized in that: The UAV target detection method based on the improved YOLOv9 model includes: Acquire drone images under different environmental conditions, and annotate the real bounding box and real confidence of the drone in the drone image; the environmental conditions include daytime, nighttime, sunny day, rainy day, and foggy day; Performing data enhancement on the drone image, and forming a drone image training data set together with the drone area in the labeled drone image; Set the backbone network of the improved YOLOv9 model to the RepViT network, and insert a two-layer routing attention mechanism network after the RepViT network to build the improved YOLOv9 model; Use the drone image training dataset to train the improved YOLOv9 model and obtain the trained improved YOLOv9 model; The image of the drone to be detected is input into the trained improved YOLOv9 model to obtain the predicted bounding box and prediction confidence of the drone to be detected.

2. The unmanned aerial vehicle target detection method based on the improved YOLOv9 model according to claim 1, characterized in that: The loss function of the improved YOLOv9 model is Slide Loss, which includes: Where n represents the total number of drone images in the drone image training dataset; λ i Represents the weight of the i-th drone image in the drone image training dataset; Loss i is the loss value of the i-th drone image in the drone image training dataset; λ′ represents the balance coefficient; α represents the positive sample weight; γ represents the negative sample weight; p i Represents the prediction confidence of the i-th drone image; IoU i represents the intersection-over-union ratio of the predicted bounding box and the true bounding box in the i-th drone image; c i The minimum diagonal length of the enclosing box containing the predicted bounding box and the true bounding box in the i-th drone image; ρ i The Euclidean distance from the center point of the predicted bounding box to the center point of the true bounding box in the i-th drone image; i represents the width of the predicted bounding box in the i-th drone image; h i represents the height of the predicted bounding box in the i-th drone image.

3. The unmanned aerial vehicle target detection method based on the improved YOLOv9 model according to claim 1, characterized in that: The calculation formula for the true confidence is: p i =σ[FC(F grid,i )]; Among them, p i represents the true confidence of the i-th drone image; σ represents the Sigmoid function, F grid,i is the grid cell feature of the drone area in the i-th drone image, and FC represents the fully connected layer.

4. The unmanned aerial vehicle target detection method based on the improved YOLOv9 model according to claim 1, characterized in that: Performing data enhancement on the drone image, specifically including: The drone image is subjected to data enhancement by rotating, scaling, adjusting brightness, adding Gaussian noise, using the MixUp method or the Mosaic method.

5. The unmanned aerial vehicle target detection method based on the improved YOLOv9 model according to claim 1, characterized in that: Improve the training process of the YOLOv9 model, including: The drone images in the drone image training data set are used as input of the improved YOLOv9 model, and the predicted bounding box and predicted confidence of the drone image are output. The improved YOLOv9 model is trained until the difference between the predicted bounding box in the drone image and the corresponding true bounding box is less than a first preset threshold, and the difference between the predicted confidence in the drone image and the corresponding true confidence is less than a second preset threshold, and the training is stopped.

6. The unmanned aerial vehicle target detection method based on the improved YOLOv9 model according to claim 1, characterized in that: The YOLOv9 model is trained and improved using a progressive training method, including: In the first preset cycle, the RepViT backbone network of the improved YOLOv9 model is frozen, and only the detection head module and FPN module of the improved YOLOv9 model are trained; In the second preset cycle, all network parameters are unfrozen, and the RepViT backbone network, detection head module, and FPN module of the improved YOLOv9 model are trained; During the third preset period, the difficult sample mining method is used to increase the weight of negative samples whose prediction confidence is less than the third preset threshold, and the RepViT backbone network, detection head module and FPN module of the improved YOLOv9 model are trained.

7. A drone target detection device based on an improved YOLOv9 model, characterized in that: The drone target detection device based on the improved YOLOv9 model includes: An image acquisition and annotation module, used to acquire drone images under different environmental conditions and annotate the real bounding box and real confidence of the drone in the drone image; the environmental conditions include daytime, nighttime, sunny, rainy, and foggy; A data set construction module, used to perform data enhancement on the drone image, and to form a drone image training data set together with the drone area in the annotated drone image; The model improvement and construction module is used to set the backbone network of the improved YOLOv9 model to the RepViT network, and insert a two-layer routing attention mechanism network after the RepViT network to build the improved YOLOv9 model; The model training module is used to train the improved YOLOv9 model using the drone image training dataset to obtain a trained improved YOLOv9 model; The drone prediction module is used to input the drone image to be detected into the trained improved YOLOv9 model to obtain the predicted bounding box and prediction confidence of the drone to be detected.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the drone target detection method based on the improved YOLOv9 model described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the unmanned aerial vehicle target detection method based on the improved YOLOv9 model described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the unmanned aerial vehicle target detection method based on the improved YOLOv9 model described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Optimization method and device of target detection model, equipment and medium

    CN121190911A

  • Smoke identification method and system suitable for comprehensive pipe gallery fire

    CN121259962A