Target detection method under view angle of unmanned aerial vehicle based on OSD-YOLO

By improving the dual small object detection structure, dynamic upsampling and attention mechanism of the YOLO network model, the false detection and missed detection problems of small object detection from the perspective of the drone are solved, and efficient drone target detection is achieved.

CN120355894APending Publication Date: 2025-07-22HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510403396.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Traditional drone target detection methods are difficult to effectively detect small targets in complex environments, and there are problems of missed detection and missed detection.

Method used

Using the UAV perspective object detection method based on OSD-YOLO, by building a dual small object detection structure, introducing dynamic upsampling Dysample module and DFMA attention mechanism, replacing the C2f module in the backbone network as the SPCC module, optimizing the model training parameters, and achieving efficient object detection.

Benefits of technology

It significantly improves the detection accuracy and efficiency of small targets from the perspective of the drone, reduces false detection and missed detection, and is suitable for real-time object detection of drone platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355894A_ABST
    Figure CN120355894A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle visual angle target detection method based on OSD-YOLO. The method is innovatively improved based on a YOLOv10 network model. Firstly, the small target detection capability is improved by constructing a dual small target detection structure; thirdly, a dynamic up-sampling Dysample module is introduced into the neck network, and the feature fusion effect is enhanced; then, a DFMA attention mechanism module is inserted, and the feature recognition capability in a complex environment is remarkably improved; and finally, a light-weight SPCC module is adopted to replace a C2f module, so that the calculation complexity is reduced while the precision is ensured. In a specific implementation process, a multi-scene unmanned aerial vehicle image data set is constructed, and professional preprocessing is performed; then end-to-end training and parameter optimization are carried out on the OSD-YOLO model; and finally deploying to an unmanned aerial vehicle system to realize real-time detection. Compared with the prior art, the method is especially suitable for small target detection of an unmanned aerial vehicle platform, effectively solves the problems of small target false detection, missing detection and low detection rate in a complex environment, and has important practical significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and target detection, and particularly relates to a target detection method from the perspective of an unmanned aerial vehicle (UAV) based on OSD-YOLO. Background Art

[0002] With the rapid development of UAV technology, its applications in fields such as aerial photography, monitoring, rescue, and agriculture are becoming increasingly widespread. As the core link of UAV applications, UAV aerial photography target detection technology can achieve automatic recognition and positioning of ground targets through image processing and artificial intelligence algorithms. This technology not only improves the working efficiency of UAVs but also provides reliable support for task execution in complex environments. For example, in disaster rescue, UAVs can quickly identify trapped people or damaged facilities; in agricultural monitoring, they can accurately identify pests and diseases to help farmers take targeted measures. However, traditional target detection methods face many challenges in UAV aerial photography scenarios. Images captured by UAVs usually have characteristics such as complex backgrounds, variable lighting conditions, and small target sizes, which make it difficult for traditional algorithms to meet the actual requirements in terms of detection accuracy and efficiency. Therefore, developing efficient and accurate UAV aerial photography target detection methods has become an urgent need for current technological development. Summary of the Invention

[0003] Object of the Invention: Aiming at the problems mentioned in the background art, the present invention proposes a target detection method from the perspective of an unmanned aerial vehicle (UAV) based on OSD-YOLO, which is suitable for small target detection on UAV platforms, solves the problems of misdetection and missed detection of small targets in the target detection task from the perspective of UAVs, and is suitable for deployment on UAV airborne hardware to carry out real-time target detection.

[0004] Technical Solution: The present invention proposes a target detection method from the perspective of an unmanned aerial vehicle (UAV) based on OSD-YOLO, including the following steps:

[0005] Step 1: Obtain an image dataset from the perspective of a UAV and perform preprocessing to divide it into a training set, a validation set, and a test set;

[0006] Step 2: Based on the YOLOv10 network model as the basic model, construct the OSD-YOLO network model; the OSD-YOLO network model adopts a dynamic upsampling Dysample module in the neck network and inserts a DFMA attention mechanism module, and replaces the C2f module in the backbone network and the neck network with an SPCC module; use the pre-divided training set and validation set to train the constructed OSD-YOLO network model and set validation parameters for training and validation;

[0007] Step 3: Use the trained OSD-YOLO network model to perform target detection on the image to be detected.

[0008] Furthermore, the OSD-YOLO network model also uses a dual small object detection structure for detection. The dual small object detection structure introduces a smaller object detection layer based on the original YOLOv10 model. First, a shallow feature map that has been downsampled by a factor of 4 is extracted from the backbone network, and its size is adjusted through a Conv downsampling operation to match the feature map sizes of other layers in the network. Then, the 80×80 feature map output by the Neck part is upsampled to generate a high-resolution feature map of 160×160, which is concatenated with the downsampled shallow feature map to form the P2 layer rich in small object information. Subsequently, the feature map output by the network is downsampled, fused with the P2 layer features, and then dimensionally reduced through Conv to make its size consistent with those of the P3, P4, and P5 layers. Finally, the feature maps of the P2, P3, P4, and P5 layers are jointly input into the detection head to achieve multi-scale object detection.

[0009] Furthermore, the dynamic upsampling Dysample module replaces the upsampling module, and the steps are as follows:

[0010] First, a linear transformation is performed on the input feature map X to generate an offset map H with the same size as the input. Subsequently, the offset map H is reshaped into O through a Pixel Shuffle operation. Then, the original sampling grid G is added to the reshaped offset O to obtain the final sampling set S. Finally, the input feature map X is resampled using the sampling set S to generate a high-resolution feature map X'.

[0011] Furthermore, the DFMA attention mechanism module inserted in the neck network is located after the SPPF module in the backbone network, and the feature map passing through the SPPF module is input into the DFMA attention mechanism module. The DFMA attention mechanism module first divides the feature map after 1×1 convolution into two parts, namely a and b. Among them, part b is processed sequentially through a feed-forward neural network FFN layer, a mixed local channel attention MLCA layer, and a feed-forward neural network FFN layer. Subsequently, parts a and b are concatenated, and finally feature fusion is performed through 1×1 convolution.

[0012] Furthermore, the feed-forward neural network FFN layer and the mixed local channel attention MLCA layer are specifically as follows:

[0013] The MLCA module includes two core components: one is the local and global information capture module, and the other is the channel and spatial information fusion module; the MLCA module extracts feature vectors containing both local details and global structures from the output feature map of the convolutional layer through local average pooling (LAP) and global average pooling (GAP) operations; subsequently, 1D convolutional operations are used to further learn the correlations between different channels, and local and global information is fused to generate importance scores for each channel and local region.

[0014] The feed-forward neural network (FFN) enhances the non-linear representation ability of the features and further processes and optimizes the features output by the MLCA module.

[0015] Furthermore, the SPCC module redesigned the Bottleneck layer of the C2f module. By introducing partial convolution (PConv), convolution calculations are only performed on some channels of the input feature map; at the same time, the SPCC module combines the ShuffleAttention mechanism, applies channel attention and spatial attention to the feature map output by PConv respectively, and mixes the feature information of different groups through channel rearrangement operations.

[0016] Furthermore, during the model training in step 2, the number of model training epochs is set to 200, the number of pictures input for one training is 16, the training log is observed in real time through Wandb during the training process, and the training results are saved after the training ends; the specific training parameters are set as follows: iterate 200 times, uniformly adjust the input image size to 640×640, set the initial learning rate to 0.01, the minimum learning rate to 0.001, the batch size to 16, and the optimizer to stochastic gradient descent (SGD).

[0017] Beneficial effects:

[0018] 1. In order to enhance the model's detection ability for small targets, the present invention adopts a dual small target detection structure in the Neck. The feature map of the P2 layer generated by this structure has more feature information of small targets, thus significantly improving the detection accuracy of the network model for small-sized targets.

[0019] 2. In order to reduce the loss of small target features during the upsampling process, the present invention adopts a more efficient dynamic upsampling module, DySample, which improves the feature fusion ability of the neck network from the perspective of point sampling.

[0020] 3. To better capture the key features in images, the present invention introduces an efficient attention mechanism (DFMA) in the Neck. DFMA is a comprehensive structure of MLCA and FFN, which can effectively extract the feature information of each channel and multi-scale spatial information. Therefore, this innovation can pay more attention to important regions more effectively and enhance the feature recognition ability of the model.

[0021] 4. To reduce the number of parameters and computational complexity of the model, the present invention redesigned the Bottleneck layer part based on the C2f module, and integrated a new lightweight module SPCC by leveraging the advantages of PConv and Shuffle Attention. This innovation not only reduces the computational complexity of the model, but also slightly improves the detection accuracy, realizing the lightweight of the model, which is suitable for deploying on UAV terminal devices for real-time target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is the flowchart of the method of the present invention;

[0023] Figure 2 is the network structure diagram of OSD-YOLO of the present invention;

[0024] Figure 3 is the flowchart of the Dysample module of the present invention;

[0025] Figure 4 is the schematic diagram of the DFMA module of the present invention;

[0026] Figure 5 is the schematic diagram of the SPCC module introduced by the present invention;

[0027] Figure 6 is the detection effect diagram of the benchmark model of the present invention;

[0028] Figure 7 is the detection effect diagram after the model improvement of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention fall within the scope defined by the appended claims of this application.

[0030] This embodiment proposes a target detection method from the perspective of a UAV based on the OSD-YOLO network. Refer to Figure 1 , and the specific implementation steps are as follows:

[0031] Step 1: Construction and preprocessing of the image dataset from the perspective of a UAV:

[0032] The image dataset obtained by the present invention from the perspective of the drone contains 10,209 pictures, which are taken by the drone from different angles and cover different scenes, weather and lighting conditions in 14 different cities in China. It covers 10 target categories such as pedestrians, cars, trucks and motorcycles. Among them, the training set contains 6,471 pictures, the validation set contains 548 pictures, and the test set contains 1,610 pictures. The complexity and diversity of this dataset provide an ideal test platform for evaluating and optimizing object detection algorithms.

[0033] Step 2: Dataset configuration of the OSD-YOLO network model

[0034] Set the name of the preprocessed dataset and construct the corresponding dataset.yaml configuration file. Write out the corresponding paths of the training set and the validation set. Set the number of dataset categories to 10. Uniformly adjust the input image size to 640×640. Set the number of training epochs to 200, the batch size to 16, the initial learning rate to 0.01, the minimum learning rate to 0.001, and the optimizer to Stochastic Gradient Descent (SGD).

[0035] Step 3. Based on the original model YOLOv10 as the base network, construct the OSD-YOLO network model. As Figure 1 shown. The following main improvements are made:

[0036] (1) In drone images, the resolution of the target to be detected is low and the feature information is scarce. Deep low-resolution features will cause the loss of small target features and missed detection phenomena. To address the problem that it is difficult to obtain small target information in the deep network, a dual small target detection structure is constructed. As shown in the appendix Figure 2 shown. First, this structure extracts the shallow feature map that has been downsampled by 4 times from the Backbone, and adjusts its size through the Conv downsampling operation to match the size of the feature maps of other layers in the network. Then, the 80×80 feature map output by the Neck part is upsampled to generate a 160×160 high-resolution feature map, which is concatenated with the downsampled shallow feature map to form the P2 layer rich in small target information. Subsequently, the feature map output by the network is downsampled, fused with the P2 layer features, and then dimensionally reduced through the Conv module to make its size consistent with the P3, P4, and P5 layers. Finally, the four feature maps of P2, P3, P4, and P5 are jointly input into the detection head to achieve multi-scale object detection. By fusing shallow detail information and deep semantic information, the detection performance of the model for small targets is effectively enhanced.

[0037] (2) The present invention introduces the DySample dynamic upsampling module into the Neck part, as shown in the appendix Figure 3As shown in the figure. Based on the dual small object detection structure, the neck network realizes more accurate multi-scale feature map alignment by dynamically adjusting the sampling position, thus significantly improving the feature fusion ability of the neck network.

[0038] The specific implementation process of the DySample module is shown in the following formulas (1)-(4):

[0039] Generate offset: H = linear(X) (1)

[0040] Offset reshaping: O = PixelShuffle(H, s) (2)

[0041] Calculate the sampling set: S = G + O (3)

[0042] Resampling: X' = grid_sample(X, S) (4)

[0043] In formula (1), linear represents a linear transformation of the input feature map X, and H is the generated offset map. In formula (2), PixelShuffle represents the feature map upsampling technique, s is the target upsampling multiple, and O is the reshaped offset map. In formula (3), G represents the position of the original pixels, and S is the final sampling set. In formula (4), grid_sample represents the resampling technique, and X' is the generated high-resolution feature map.

[0044] The DySample module dynamically determines the best sampling position for each output pixel through a learning mechanism, achieving content-aware upsampling. This method not only improves the quality and flexibility of upsampling, but also can more accurately restore small object and edge information, thus further enhancing the feature fusion ability of the neck network.

[0045] (3) The present invention introduces a DFMA attention mechanism in the Neck part, as shown in the appendix Figure 4 As shown. The DFMA attention mechanism module consists of two feed-forward neural networks (FFNs) and a mixed local channel attention module (MLCA). The specific implementation process is as follows: First, the feature map after 1×1 convolution is evenly divided into two parts (a and b), where part b is processed through the FFN layer, the MLCA layer, and the FFN layer in sequence, then parts a and b are concatenated, and finally feature fusion is performed through 1×1 convolution. The MLCA module is located between the two FFNs and is used to enhance the attention mechanism of the feature map, enabling the model to focus more on key regions and channel information.

[0046] Specifically, the Mixed Local Channel Attention (MLCA) consists of a local and global information capture module and a channel and spatial information fusion module. First, MLCA uses local average pooling (LAP) and global average pooling (GAP) operations to extract feature vectors containing both local details and global structures from the output feature map of the convolutional layer. Subsequently, through 1D convolution (Conv1d) operations, the correlations between different channels are further learned, and local and global information is fused to generate importance scores for each channel and local region. These scores are multiplied by the original feature map as weights to achieve recalibration of the features, thereby enhancing the attention to task-critical features while suppressing irrelevant features. The FFN is a feed-forward neural network that usually contains two linear layers (fully connected layers). The input feature map X first undergoes the first linear transformation, then a non-linear mapping is performed through the activation function σ, and then the second linear transformation is carried out, thus enhancing the feature expression ability and model adaptability. The MLCA module is embedded between two FFNs, and by considering the importance of both the spatial and channel dimensions simultaneously, the feature selection and fusion process is further optimized. Finally, the DFMA attention mechanism significantly improves the feature expression ability and attention focusing effect while maintaining high computational efficiency.

[0047] (4) In the present invention, the C2f module in the model is replaced with a new lightweight module SPCC, and the Bottleneck layer in the original C2f module is redesigned, as shown in the appendix Figure 5As shown. The newly designed bottleneck layer SPCC_Bottleneck integrates the partial convolution (PConv) and the Shuffle Attention mechanism. The core idea of the partial convolution (PConv) is to utilize the redundancy of the feature map and apply the conventional convolution only to a part of the input channels for spatial feature extraction, while the remaining channels remain unchanged. Finally, the processed feature map is concatenated with the unprocessed feature map to form the output feature map. In this way, PConv reduces the computational amount while retaining the feature expression ability. When the partial rate is set to 1 / 4, the computational amount of PConv is only 1 / 16 of that of the conventional convolution. Therefore, the present invention replaces the conventional convolution in the original bottleneck layer with PConv, effectively avoiding the problems of feature redundancy and gradient disappearance caused by the single-branch structure of the conventional convolution. The Shuffle Attention module is an efficient and complex attention mechanism that focuses on capturing multi-scale details. Its core idea is to group the channels of the input feature map and perform channel attention and spatial attention processing on each sub-feature group respectively. Specifically, Shuffle Attention first divides the channels of the feature map into multiple sub-feature groups. Each sub-feature group generates a channel attention map through global average pooling and a spatial attention map through group normalization. Subsequently, the channel and spatial attention maps are respectively applied to the corresponding sub-feature groups to generate weighted sub-feature maps. Finally, through the "channel shuffle" operation, information interaction between different sub-feature groups is realized and re-aggregated into a complete feature map. This design makes Shuffle Attention significantly superior to the traditional full attention mechanism in terms of computational efficiency.

[0048] By integrating the PConv and Shuffle Attention mechanisms, SPCC_Bottleneck performs convolution operations on some channels of the input feature map, then applies a 1×1 standard convolution, and combines an efficient attention mechanism, significantly reducing the model complexity. At the same time, this module improves the backbone network through extended functions, providing more valuable multi-scale feature information within limited computational capacity, thus achieving a balance between performance and efficiency.

[0049] Step 4. Set parameters to train the improved model

[0050] Preprocess the constructed dataset, divide it into a training set of 6471 images and a validation set of 548 images, set the number of training rounds of the model to 200, the number of images input for one training to 16, and then train with the modified model. During the training process, observe the training log in real time through Wandb and save the training results after the training ends.

[0051] Step 5. Use the constructed OSD-YOLO model for object detection from the perspective of a drone

[0052] Input the image to be detected, and perform target detection from the perspective of an unmanned aerial vehicle through the trained and verified OSD-YOLO network model.

[0053] In the implementation of the present invention, precision, recall, mean average precision (mAP), number of parameters (Params), and computational complexity (GFLOPs) are used as evaluation indicators. The formulas for each evaluation indicator are as follows:

[0054] Precision (P):

[0055]

[0056] Recall (R):

[0057]

[0058] Mean average precision (mAP50):

[0059]

[0060] Among them, TP refers to the number of samples that are positive classes and are predicted as positive classes, FP refers to the number of samples that are negative classes and are predicted as positive classes, FN refers to the number of samples that are positive classes and are predicted as negative classes, n is the number of detected target categories, and AP i is the AP of the i-th target category.

[0061] To prove the effectiveness of the proposed OSD-YOLO network model, we conducted ablation experiments based on the baseline network on the constructed dataset. Since the experiments of deep learning are all random, the experimental results in this paper adopt the average value of multiple experimental results to improve the feasibility of the experimental results. The experiment first tests the performance of each module in the baseline network one by one, and then adopts a stacking method to gradually add modules to the baseline network to compare the effectiveness of each module. The experimental results are shown in Table 1:

[0062] Table 1 Comparison of experimental results of different network models

[0063]

[0064] Note: √ indicates that the module is used

[0065] Compared with the baseline model, the OSD-YOLO network model of the present invention shows significant performance improvement. Specifically, after the model A generates the P2 feature map using a four-layer feature fusion structure, compared with the baseline model YOLOv10n, the mAP 50It has increased by 2.8%, but the floating-point FLOPs have also increased by 3.7 G. Then, the Dysample module is introduced for improvement based on Model A. Compared with Model A, the mAP of this Model B 50 has increased by 0.5%, but the floating-point numbers have not increased and have even decreased by 0.1 G. Then, the DFMA attention mechanism is introduced for improvement based on Model B. Compared with the baseline model, the mAP of this Model C 50 has increased by 4.1%, but the floating-point numbers have increased by 3.6 G. Therefore, a new type of lightweight module SPCC is introduced based on Model C in the present invention, and the final improved model is proposed. Compared with the baseline model, the P of the improved model has increased by 3.3%, and the mAP 50 has increased by 3.9%, the number of parameters has decreased by 0.4 M, and the floating-point numbers have only increased by 0.4 G.

[0066] The present invention aims at the problems existing in the vehicle target detection task from the perspective of drones and has better detection effects compared with the baseline model. As shown in the appendix Figure 6 and Figure 7 , the detections of the improved model Figure 7 and the original model Figure 6 are compared. In the detection of small targets, the original model 6 has the situations of misdetecting and missing pedestrians, and the occluded tricycles and motorcycles are not detected. While Figure 7 for the improved model, there are no misdetections and missing detections, and it can accurately detect each target. By comparing Figure 6 and Figure 7 , it can be seen that the detection performance of the improved model is better than that of the original model in the backgrounds of small target detection, target occlusion, etc.

[0067] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.

Claims

1. A target detection method from the perspective of an unmanned aerial vehicle based on OSD-YOLO, characterized in that, It includes the following steps: Step 1: Obtain the image dataset from the perspective of the drone, perform preprocessing, and divide it into a training set, a validation set, and a test set; Step 2: Based on the YOLOv10 network model, construct the OSD-YOLO network model; in the neck network of the OSD-YOLO network model, the dynamic upsampling Dysample module is adopted and the DFMA attention mechanism module is inserted, and the C2f module in the backbone network and the neck network is replaced with the SPCC module; use the pre-divided training set and validation set to train the constructed OSD-YOLO network model and set validation parameters for training and validation; Step 3: Use the trained OSD-YOLO network model to perform object detection on the image to be detected.

2. The method for object detection from a drone perspective based on OSD-YOLO according to claim 1, wherein The OSD-YOLO network model also uses a dual small object detection structure for detection. The dual small object detection structure introduces a smaller object detection layer on the basis of the original YOLOv10 model. First, extract the shallow feature map that has been downsampled by 4 times from the backbone network, and adjust the size through the Conv downsampling operation to match the feature map sizes of other layers in the network; then, upsample the 80×80 feature map output by the Neck part to generate a high-resolution feature map of 160×160, and splice it with the downsampled shallow feature map to form the P2 layer rich in small object information; subsequently, downsample the feature map output by the network, fuse it with the P2 layer feature, and then reduce the dimension through Conv to make its size consistent with the P3, P4, and P5 layers; finally, input the P2, P3, P4, and P5 four-layer feature maps into the detection head to achieve multi-scale object detection.

3. The method for object detection from the perspective of a drone based on OSD-YOLO according to claim 1, characterized in that, The steps for the dynamic upsampling Dysample module to replace the upsampling module are as follows: First, perform a linear transformation on the input feature map X to generate an offset map H with the same size as the input; then, reshape the offset map H into O through the Pixel Shuffle operation; next, add the original sampling grid G and the reshaped offset O to obtain the final sampling set S; finally, use the sampling set S to resample the input feature map X to generate a high-resolution feature map X'.

4. The method for target detection from a drone perspective based on OSD-YOLO according to claim 1, characterized in that, The DFMA attention mechanism module inserted in the neck network is located after the SPPF module of the backbone network, and the feature map passing through the SPPF module is input into the DFMA attention mechanism module; the DFMA attention mechanism module first divides the feature map after 1×1 convolution into two parts, namely a and b. Among them, part b is processed sequentially through the feed-forward neural network FFN layer, the mixed local channel attention MLCA layer, and the feed-forward neural network FFN layer, then splices the a and b parts, and finally performs feature fusion through 1×1 convolution.

5. The method for object detection from an unmanned aerial vehicle perspective based on OSD-YOLO according to claim 4, wherein The specific details of the feed-forward neural network FFN layer and the mixed local channel attention MLCA layer are as follows: The MLCA module includes two core components: one is the local and global information capture module, and the other is the channel and spatial information fusion module; the MLCA module extracts feature vectors containing both local details and global structures from the output feature map of the convolutional layer through local average pooling (LAP) and global average pooling (GAP) operations; subsequently, 1D convolutional operations are used to further learn the correlations between different channels, and local and global information is fused to generate importance scores for each channel and local region. The feed-forward neural network (FFN) enhances the non-linear expression ability of the features and further processes and optimizes the features output by the MLCA module.

6. The method for object detection from the perspective of an unmanned aerial vehicle based on OSD-YOLO according to claim 1, wherein The SPCC module redesigned the Bottleneck layer of the C2f module. By introducing partial convolution (PConv), convolution calculations are only performed on some channels of the input feature map; at the same time, the SPCC module combines the Shuffle Attention mechanism, applies channel attention and spatial attention to the feature map output by PConv respectively, and mixes the feature information of different groups through channel rearrangement operations.

7. The method for object detection from the perspective of an unmanned aerial vehicle based on OSD-YOLO according to claim 1, characterized in that, When training the model in step 2, the number of training epochs of the model is set to 200, the number of images input for one training is 16, the training log is observed in real time through Wandb during the training process, and the training results are saved after the training is completed. The specific training parameters are set as follows: the number of iterations is 200, the input image size is uniformly adjusted to 640×640, the initial learning rate is set to 0.01, the minimum learning rate is set to 0.001, the batch size is set to 16, and the optimizer is set to stochastic gradient descent (SGD).

Citation Information

Cited By

  • Improved YOLOv7-based steel marking area detection model and training method thereof

    CN120833474A

  • A steel marking area detection model based on improved YOLOv7 and a training method thereof

    CN120833474B

  • Small target detection method for unmanned aerial vehicle data

    CN120953859A

  • Text-fused multi-scale edge information multi-target detection method

    CN121582546A