Unmanned aerial vehicle aerial photography ground target detection method based on ST-YOLOv8n model
By improving the C2f module and loss function of the YOLOv8n model and combining it with the MD-Head multidimensional target detection head, the problems of diversity, complex background and real-time performance in aerial image ground target detection are solved, and high-precision ground target detection with low computational cost is achieved.
Patent Information
- Application Number
- CN202511222635.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-11-14
AI Technical Summary
Existing aerial image ground target detection technologies have significant limitations in terms of target scale diversity, complex backgrounds and environmental interference, real-time requirements and lightweight requirements, and geometric deformation positioning errors, making it difficult to meet the practical application needs of intelligent transportation and drone inspection.
The ST-YOLOv8n model is adopted, replacing the C2f module with the ST-C2f module and the CIoU loss function with the SDIoU loss function. The MD-Head multidimensional object detection head is integrated, and feature extraction and object localization are optimized through a dual-branch heterogeneous structure, dynamic gating fusion, multi-scale perception, semantic guidance and adaptive weighting mechanism.
It improves the detection accuracy of small targets, reduces the false negative rate and the false positive rate, enhances the robustness and positioning accuracy of the model in complex backgrounds, and meets the real-time and lightweight requirements of embedded terminals.
Smart Images

Figure CN120953852A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of image target detection, and in particular to a method for detecting ground targets in UAV aerial photography based on the ST-YOLOv8n model. Background Technology
[0002] With the rapid development of intelligent transportation, infrastructure maintenance, and drone inspection, ground target detection (such as pedestrians, vehicles, road cracks, potholes, etc.) from aerial images has become one of the core tasks for the practical application of computer vision technology. This technology uses aerial images collected by drones or vehicle-mounted equipment to automatically identify and locate ground targets. It can be widely applied in scenarios such as traffic flow monitoring, road defect investigation, and emergency rescue target search, and is of great significance for ensuring driving safety, improving infrastructure maintenance efficiency, and reducing the cost of manual inspections.
[0003] However, ground target detection in aerial images faces numerous technical challenges. Existing detection schemes (including the YOLOv8n-based basic model and other mainstream detection algorithms) still have significant limitations in practical applications, specifically in the following four aspects:
[0004] 1. The diversity of target scales leads to insufficient feature extraction;
[0005] Aerial photography scenes exhibit a wide range of ground targets, ranging from large-scale objects like buildings and trucks to small-scale targets such as pedestrians, bicycles, and narrow road cracks. Traditional detection models (such as the C2f module of the original YOLOv8n) employ a single feature integration strategy, achieving feature fusion only through basic convolutions and skip connections. This approach struggles to simultaneously capture the pixel-level details of small targets and the global contextual information of large targets. Features of small targets are easily lost during multiple downsampling iterations, resulting in a persistently high false negative rate (existing lightweight models generally have a false negative rate exceeding 25% for small targets like cracks). Large targets, on the other hand, suffer from insufficient receptive fields, failing to fully capture the correlation features between the background and the target, thus impacting localization accuracy.
[0006] 2. Complex backgrounds and environmental interference lead to insufficient detection robustness;
[0007] Aerial images are significantly affected by lighting conditions (such as strong light reflections and shadow occlusion), weather factors (rain and snow cover, dirt and pollution), and scene complexity (tree obstruction, densely built-up areas), and target features are easily submerged by complex background noise. Existing technologies (such as multi-scale feature fusion schemes) attempt to improve anti-interference capabilities by increasing feature layer interactions, but they suffer from high computational redundancy—model inference speeds are typically below 10 FPS, which cannot meet the real-time requirements of vehicle-mounted devices or drone terminals. While lightweight design schemes (such as simplified convolutional structures) can improve speed, they further weaken feature extraction capabilities, leading to an increased false detection rate, especially in the detection of small targets (such as pedestrians) in tree shadows and building shadows, where the false detection rate is more than 30% higher than in conventional scenes.
[0008] 3. There is a conflict between the requirements for real-time performance and lightweight design;
[0009] Ground target detection needs to be deployed on embedded terminals (such as UAV flight control systems and vehicle-mounted edge devices). These devices have limited computing resources (computing power and memory) and stringent requirements for lightweight and real-time performance of models (typically requiring inference speed ≥30FPS and computation <20GFLOPs). Existing high-precision detection models (such as YOLOv8 l / x and CNN-Transformer hybrid models) can improve detection accuracy, but their computational cost generally exceeds 100GFLOPs, and their large number of parameters makes them unsuitable for embedded terminals. While the original YOLOv8n is the lightest version in the YOLOv8 series (computational cost 8.1GFLOPs), it is not optimized for aerial photography scenarios and cannot meet the actual needs in terms of accuracy for small target detection and robustness against complex backgrounds, creating a contradiction between "accuracy, speed, and lightweight".
[0010] 4. The positioning error problem of geometrically deformed targets;
[0011] In aerial images, targets are prone to geometric deformation due to shooting angle (e.g., tilted aerial photography) and changes in target attitude (e.g., deformed vehicles, curved cracks). Traditional bounding box regression loss functions (such as CIoU) rely on center point distance and aspect ratio for error constraints, making them sensitive to such geometric deformations. For example, when calculating positioning error, CIoU only uses the target center point distance as the core penalty term, ignoring the deviation of the deformed target's diagonal position, resulting in a 30% higher positioning error for targets such as curved cracks and tilted vehicles compared to conventional targets. Subsequent improved loss functions such as GIoU and DIoU either rely on minimum bounding box area (GIoU) or only optimize center point distance (DIoU), still failing to construct an effective constraint mechanism for the geometric deformation characteristics of aerial targets, resulting in insufficient positioning stability.
[0012] Therefore, there is an urgent need for an improved solution that balances computational accuracy, computational speed, and lightweight design, while also being adapted to the characteristics of aerial photography scenarios, in order to meet the practical application needs of fields such as intelligent transportation and drone inspection. Summary of the Invention
[0013] The purpose of this invention is to overcome the shortcomings of the prior art and provide a ground target detection method based on the ST-YOLOv8n model that takes into account accuracy, speed and lightweight design, and is adapted to the characteristics of aerial photography scenarios.
[0014] To achieve the above objectives, the technical solution provided by this invention is as follows:
[0015] A UAV aerial ground target detection method based on the ST-YOLOv8n model is proposed. The original C2f module and CIoU loss function of the YOLOv8n model are replaced with the ST-C2f module and SDIoU loss function, respectively. The MD-Head multidimensional target detection head is integrated into the detection head part of the YOLOv8n model, thus obtaining the improved YOLOv8n model, namely the ST-YOLOv8n model.
[0016] The ST-C2f module features a dual-branch heterogeneous structure and a dynamic gated fusion unit. The dual-branch heterogeneous structure includes a local branch and a global branch. The local branch sequentially employs a 3×3 depthwise separable convolution to extract pixel-level local detail features and a 7×7 dilated convolution to expand the receptive field and integrate complex background features. The global branch introduces a SwingTransformer module, which calculates the long-range dependencies of the captured feature maps through window self-attention and combines a window translation mechanism to achieve global contextual information interaction. The dynamic gated fusion unit adaptively fuses the local features output from the local branch and the global features output from the global branch using a learnable weight matrix to obtain a fused feature map.
[0017] The SDIoU loss function is constructed by introducing the squared diagonal distance between the predicted box and the ground truth box and the average area ratio between the predicted box and the ground truth box, in order to optimize the target localization robustness and reduce the localization error of geometrically deformed targets.
[0018] The MD-Head multidimensional target detection head includes a multi-scale perception module, a semantic-guided fusion module, and an adaptive weighting module. The multi-scale perception module extracts target features at different scales by combining different pooling kernels. The semantic-guided fusion module aligns high-level semantic features with basic features and generates gating weights through semantic relevance calculation to achieve the filtering and fusion of semantic information related to small targets. The adaptive weighting module compresses the feature space dimension through global average pooling and generates channel weights to enhance the features of key target regions (finally outputting the ground target detection results of aerial images).
[0019] Train the ST-YOLOv8n model to obtain the trained ST-YOLOv8n model;
[0020] The trained ST-YOLOv8n model is used to detect ground targets.
[0021] Furthermore, the ST-C2f module processes the input feature map as follows:
[0022] First, perform dimensionality reduction on the input feature map X∈R using convolution. H×W×C Channel compression is performed to reduce subsequent computation and retain core features; then the compressed feature map is segmented by channel, and the segmented feature maps are input into the local branch and the global branch respectively;
[0023] C′=C / 2
[0024] X′=Conv(X)
[0025]
[0026] Where the dimension is R H×W×C H is the height of the input feature map, W is the width of the input feature map; C is the initial number of channels of the input feature map; C′ is the target number of channels after channel compression; X′ is the intermediate feature map after channel compression; X′ L This is a local feature branch feature map after channel segmentation; X′ G This is the global feature branch feature map after channel segmentation.
[0027] Furthermore, the working process of the SwinTransformer module in the global branch includes:
[0028] The segmented global feature branch feature map X′ is obtained by using a window function. G Divide into non-overlapping local windows
[0029] X=SplitWin(X)=[X1,X2,…,X n ]
[0030] SplitWin() is the function for dividing the window. M is the size of the window;
[0031] Calculate window self-attention within each local window using the following formula:
[0032] WinAtten(X i ) = SelfAtten(X i P Q ,X i P K ,Xi P V )
[0033] The window self-attention mechanism is as follows:
[0034]
[0035] In the formula, P Q P K P V is the projection matrix, Q, K, and V are three vectors generated by the input features through the projection matrix, corresponding to the query matrix, key matrix, and value matrix, respectively; d is the number of channels, and Softmax() is a function used to transform the dot product of Q and K into a probability distribution, thereby normalizing the attention weights;
[0036] The window translation mechanism is implemented to periodically move the window position and enable information interaction between adjacent windows;
[0037] ShiftWinAtt(Q,K,V)=WinAtten(Shift(Q),Shift(K),Shift(V))
[0038] In this context, ShiftWinAtt() represents the window translation mechanism, and Shift() represents the translation transformation.
[0039] Furthermore, the formula for the SDIoU loss function is as follows:
[0040]
[0041] Where IoU is the intersection-union ratio between the predicted bounding box B and the ground truth bounding box A:
[0042]
[0043] The squared Euclidean distance between the top-left corner of the ground truth bounding box A and the top-left corner of the predicted bounding box B is:
[0044]
[0045] The squared Euclidean distance between the bottom right corner of the ground truth bounding box A and the bottom right corner of the predicted bounding box B is:
[0046]
[0047] S t and S p These are the areas of the ground truth bounding box A and the predicted bounding box B, respectively.
[0048] S represents the shape similarity index:
[0049]
[0050] These are the coordinates of the top left and bottom right corners of the ground truth bounding box A and the predicted bounding box B, respectively.
[0051] α is a weighted correction coefficient:
[0052]
[0053] v is a quantitative indicator of shape difference:
[0054]
[0055] eps is a non-zero constant value used to avoid division by zero errors during calculations; w A and h A These are the width and height of the actual bounding box A, respectively. B and h B These are the width and height of the prediction box B, respectively.
[0056] Furthermore, in the MD-Head multidimensional target detection head, the multi-scale perception module, semantic-guided fusion module, and adaptive weighting module are connected in a serial manner, specifically in the order of multi-scale perception module → semantic-guided fusion module → adaptive weighting module; output through the following formula:
[0057] F out (X)=F3(F2(F1(X)))
[0058] Among them, F out ( ) represents the output of the MD-Head multidimensional target detection head, and F1(), F2(), and F3() represent the multi-scale perception module function, the semantic-guided fusion module function, and the adaptive weighting module function, respectively.
[0059] Furthermore, the multi-scale perception module includes a 3×3 pooling kernel, a 7×7 pooling kernel, and an 11×11 pooling kernel. These pooling kernels use parallel computing to extract small-scale target detail features, medium-scale target contour features, and large-scale target global features, respectively. Then, the three scale features are integrated into a unified scale multi-scale fusion feature through convolution fusion operation.
[0060] The formula used is as follows:
[0061]
[0062] Where F1(X) is the output feature of the multi-scale perception module; F N F is a nonlinear transformation function; Conv1 For convolution operations, P 3×3 (X) represents the output of a 3×3 convolution operation on input X; P7×7 (X) represents the output of a 7×7 convolution operation on input X; P 11×11 (X) is the output of an 11×11 convolution operation on input X.
[0063] Furthermore, the working process of the semantic guidance fusion module includes:
[0064] Upsampling is performed on high-level semantic features to make the resolution of high-level semantic features consistent with that of the basic features, thereby achieving feature alignment;
[0065] The semantic similarity between the aligned high-level semantic features and the basic features is calculated using matrix multiplication to generate the selection weights.
[0066] Sim(x,y,c)=X′(x,y,c)·S(x,y,c)
[0067] Sim(x,y,c) is a semantic similarity value, which represents the degree of semantic association between basic feature X′ and high-level semantic feature S at spatial location (x,y) and channel c.
[0068] Next, the selection weights are normalized, and semantically guided feature selection and fusion are performed:
[0069] F2(X)=BN(ReLU(F Conv (X′⊙σ(Sim))))
[0070] Where BN() and σ() represent normalization, ReLU() represents the activation function, and X′⊙σ(Sim) means applying the normalized screening matrix to X′.
[0071] Furthermore, the operation of the adaptive weighting module includes:
[0072] Perform global average pooling on the input feature map X″ to obtain global statistics for each channel; the global average pooling operation satisfies:
[0073]
[0074] Where x(i,j,:) represents the feature vector of all channels at position (i,j); H1 and W1 are the height and width of the input feature map X″, respectively;
[0075] Next, a nonlinear transformation is performed on the global features to generate a spatial weight matrix that matches the size of the input features;
[0076] Finally, the generated weight matrix is multiplied element-wise with the input features to enhance key regions and suppress noisy regions.
[0077] W = σ(UpSample(F)Conv3 (F 全局平均池化 )))
[0078] F3(X″) = X″⊙W
[0079] W is the spatial weight matrix; σ is the activation function; UpSample is the upsampling operation; F Conv3 For convolution operation, channel and spatial features are fused on the upsampled features to extract weight information, making the weight matrix more in line with semantic requirements and enhancing the discriminative power of the weights; F3(X″) is the final output feature; ⊙ is the element-wise multiplication operation, which multiplies each element of the spatial weight matrix W with the corresponding element of the input feature X″, adjusts the feature value according to the weight, and completes the dynamic weight allocation.
[0080] Furthermore, to achieve the above objectives, the present invention also provides an electronic device comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to achieve the above-mentioned [specific implementation].
[0081] Ground target detection method of ST-YOLOv8n model.
[0082] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing computer instructions, which, when executed by the computer, implement the above-described ground target detection method based on the ST-YOLOv8n model.
[0083] Compared with existing technologies, the principles and advantages of this technical solution are as follows:
[0084] 1. The ST-C2f module features a dual-branch heterogeneous structure and a dynamic gating fusion unit;
[0085] 1) Parallel extraction of multi-scale features via dual branches:
[0086] The local branch employs a combination of 3×3 depthwise separable convolution and 7×7 dilated convolution: the 3×3 depthwise separable convolution reduces computation while accurately capturing pixel-level details of small targets (such as pedestrians or narrow cracks); the 7×7 dilated convolution expands the receptive field, integrating background and contextual information of large targets (such as buildings or trucks), thus preventing the localization accuracy of large targets from being limited by insufficient receptive field. The global branch introduces the SwinTransformer module: through a "window self-attention + window translation mechanism," attention is calculated within a local window (reducing the high computational cost of global self-attention), while global information interaction is achieved through window translation, solving the problem of the original C2f module's lack of global feature extraction capabilities, while maintaining controllable computational complexity and ensuring lightweight characteristics.
[0087] 2) Dynamic gating fusion ensures the effectiveness of features:
[0088] By adaptively fusing detailed features of local branches with long-range dependency features of global branches using a learnable weight matrix, feature redundancy is avoided. At the same time, the module performs a "convolutional dimensionality reduction + channel segmentation" step beforehand to further reduce the amount of subsequent computation, ultimately achieving "no loss of details in small targets and no omission of global targets in large targets". Moreover, the number of model parameters only increases by 15% compared to the original YOLOv8n, and the inference speed can be adapted to embedded terminals (meeting the requirement of ≥30FPS), balancing accuracy and lightweight design.
[0089] 2. The MD-Head multidimensional target detection head constructs an anti-interference detection mechanism through multi-scale perception, semantic-guided gating fusion, and adaptive weighting;
[0090] 1) Multi-scale perception covers targets in complex scenes:
[0091] The feature extraction is carried out in parallel using "3×3 pooling + 7×7 pooling + 11×11 pooling". 3×3 pooling focuses on the details of small targets, while 7×7 and 11×11 pooling cover medium and large targets and complex background areas, avoiding the problem that a single pooling scale cannot adapt to diverse interference scenarios.
[0092] 2) Semantic-guided selection of target features:
[0093] By aligning high-level semantic features (including key semantic information such as target outline and mobility) with basic features, and generating gating weights through semantic relevance calculation, the relevant semantic features of small targets are accurately labeled, and the interference of background noise (such as tree shadows and road stains) is suppressed. This mechanism significantly improves the detection rate of small targets such as pedestrians, and reduces the background false detection rate compared with ordinary feature splicing schemes.
[0094] 3) Adaptive weighted enhancement of key regions:
[0095] By compressing the spatial dimension through global average pooling, channel weights are generated, and the features of key target regions (such as the location of small targets) are dynamically enhanced. This further reduces the impact of complex backgrounds on detection results and improves the robustness of the model under varying lighting conditions and occlusion scenarios.
[0096] 3. Construct the SDIoU loss function by combining the squared diagonal distance and the average area ratio to optimize positioning accuracy:
[0097] 1) Diagonal distance instead of center point distance:
[0098] Using the squared distance between the predicted bounding box and the ground truth bounding box as the core constraint, rather than the center point distance of CIoU, the center point deviation of geometrically deformed targets (such as tilted vehicles) may be small, but the diagonal point deviation is significant. SDIoU uses the diagonal distance constraint to more accurately capture the bounding box offset caused by deformation. Experiments have verified that the localization error in vehicle deformation scenes is reduced.
[0099] 2) Average area ratio supplementary constraint:
[0100] By introducing the average area ratio between the predicted bounding box and the ground truth bounding box, the stability of bounding box regression is further optimized, avoiding the "area deviation neglect" problem that may be caused by relying solely on diagonal distance. This makes the loss calculation more consistent with the geometric characteristics of the aerial deformed target, ultimately improving the localization robustness. Attached Figure Description
[0101] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the services required in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0102] Figure 1 This is a flowchart illustrating the principle of a UAV aerial ground target detection method based on the ST-YOLOv8n model, according to an embodiment of the present invention.
[0103] Figure 2 Here is a block diagram of the ST-C2f module;
[0104] Figure 3 This is a block diagram of the MD-Head multidimensional target detection head;
[0105] Figure 4 Performance curves for the YOLOv8n model;
[0106] Figure 5 Performance curves for the ST-YOLOv8n model;
[0107] Figure 6 The image shows the detection results of the YOLOv8n model.
[0108] Figure 7 This is a graph showing the detection results of the ST-YOLOv8n model.
[0109] ( Figure 4 and Figure 5In this context, Precision-Recall Curve represents the precision-recall curve; Recall-Confidence Curve represents the recall-confidence curve; pedestrian represents a pedestrian; people represents a crowd; bicycle represents a bicycle; car represents a car; van represents a van; truck represents a truck; tricycle represents a tricycle; awning-tricycle represents a covered tricycle; bus represents a bus; motor represents a motorcycle; all classes 0.348mAP@0.5 means that the mAP@0.5 for all classes is 0.348; Confidence represents confidence; Recall represents recall. Detailed Implementation
[0110] The present invention will be further described below with reference to specific embodiments:
[0111] like Figures 1 to 3 As shown in this embodiment, a UAV aerial photography ground target detection method based on the ST-YOLOv8n model includes the following steps:
[0112] S1. Replace the original C2f module and CIoU loss function of the YOLOv8n model with the ST-C2f module and SDIoU loss function respectively; and integrate the MD-Head multidimensional object detection head into the detection head part of the YOLOv8n model to obtain the improved YOLOv8n model, namely the ST-YOLOv8n model.
[0113] The ST-C2f module features a dual-branch heterogeneous structure and a dynamic gated fusion unit. The dual-branch heterogeneous structure includes a local branch and a global branch. The local branch sequentially employs a 3×3 depthwise separable convolution to extract pixel-level local detail features and a 7×7 dilated convolution to expand the receptive field and integrate complex background features. The global branch introduces a SwingTransformer module, which calculates the long-range dependencies of the captured feature maps through window self-attention and combines a window translation mechanism to achieve global contextual information interaction. The dynamic gated fusion unit adaptively fuses the local features output from the local branch and the global features output from the global branch using a learnable weight matrix to obtain a fused feature map.
[0114] The SDIoU loss function is constructed by incorporating the squared diagonal distance between the predicted box and the ground truth box, and the average area ratio between the predicted box and the ground truth box, in order to optimize the robustness of target localization and reduce the localization error of geometrically deformed targets.
[0115] The MD-Head multidimensional target detection head includes a multi-scale perception module, a semantic-guided fusion module, and an adaptive weighting module. The multi-scale perception module extracts target features at different scales by combining different pooling kernels. The semantic-guided fusion module aligns high-level semantic features with basic features and generates gating weights through semantic relevance calculation to achieve the filtering and fusion of semantic information related to small targets. The adaptive weighting module compresses the feature space dimension through global average pooling and generates channel weights to enhance the features of key target regions, ultimately outputting the ground target detection results of the aerial image.
[0116] S2. Train the ST-YOLOv8n model to obtain the trained ST-YOLOv8n model;
[0117] S3. Use the trained ST-YOLOv8n model to detect ground targets.
[0118] Specifically, in this embodiment, the ST-C2f module processes the input feature map as follows:
[0119] First, perform dimensionality reduction on the input feature map X∈R using convolution. H×W×C Channel compression is performed to reduce subsequent computation and retain core features; then the compressed feature map is segmented by channel, and the segmented feature maps are input into the local branch and the global branch respectively;
[0120] C′=C / 2
[0121] X′=Conv)X_
[0122]
[0123] Where the dimension is R H×W×C H is the height of the input feature map, W is the width of the input feature map; C is the initial number of channels of the input feature map; C′ is the target number of channels after channel compression; X′ is the intermediate feature map after channel compression; X′ L This is a local feature branch feature map after channel segmentation; X′ G This is the global feature branch feature map after channel segmentation.
[0124] The working process of the SwinTransformer module in the global branch includes:
[0125] The segmented global feature branch feature map X′ is obtained by using a window function. G Divide into non-overlapping local windows
[0126] X=SplitWin(X)=[X1,X2,…,X n ]
[0127] SplitWin() is the function for dividing the window. M is the size of the window;
[0128] Calculate window self-attention within each local window using the following formula:
[0129] WinAtten(X i ) = SelfAtten(X i P Q ,X i P K ,X i P V )
[0130] The window self-attention mechanism is as follows:
[0131]
[0132] In the formula, P Q P K P V is the projection matrix, Q, K, and V are three vectors generated by the input features through the projection matrix, corresponding to the query matrix, key matrix, and value matrix, respectively; d is the number of channels, and Softmax() is a function used to transform the dot product of Q and K into a probability distribution, thereby normalizing the attention weights;
[0133] The window translation mechanism is implemented to periodically move the window position and enable information interaction between adjacent windows;
[0134] ShiftWinAtt(Q,K,V)=WinAtten(Shift(Q),Shift(K),Shift(V))
[0135] In this context, ShiftWinAtt() represents the window translation mechanism, and Shift() represents the translation transformation.
[0136] Specifically, in this embodiment, the formula for the SDIoU loss function is as follows:
[0137]
[0138] Where IoU is the intersection-union ratio between the predicted bounding box B and the ground truth bounding box A:
[0139]
[0140] The squared Euclidean distance between the top-left corner of the ground truth bounding box A and the top-left corner of the predicted bounding box B is:
[0141]
[0142] The squared Euclidean distance between the bottom right corner of the ground truth bounding box A and the bottom right corner of the predicted bounding box B is:
[0143]
[0144] S t and S p These are the areas of the ground truth bounding box A and the predicted bounding box B, respectively.
[0145] S represents the shape similarity index:
[0146]
[0147] These are the coordinates of the top left and bottom right corners of the ground truth bounding box A and the predicted bounding box B, respectively.
[0148] α is a weighted correction coefficient:
[0149]
[0150] v is a quantitative indicator of shape difference:
[0151]
[0152] eps is a non-zero constant value used to avoid division by zero errors during calculations; w A and h A These are the width and height of the actual bounding box A, respectively. B and h B These are the width and height of the prediction box B, respectively.
[0153] Specifically, in this embodiment, the multi-scale perception module, semantic-guided fusion module, and adaptive weighting module in the MD-Head multi-dimensional target detection head are connected in a serial manner, specifically in the order of multi-scale perception module → semantic-guided fusion module → adaptive weighting module; output by the following formula:
[0154] F out (X)=F3(F2(F1(X)))
[0155] Among them, F out ( ) represents the output of the MD-Head multidimensional target detection head, and F1(), F2(), and F3() represent the multi-scale perception module function, the semantic-guided fusion module function, and the adaptive weighting module function, respectively.
[0156] More specifically, the multi-scale perception module includes 3×3 pooling kernels, 7×7 pooling kernels, and 11×11 pooling kernels. These pooling kernels use parallel computing to extract small-scale target detail features, medium-scale target contour features, and large-scale target global features, respectively. Then, through convolution fusion operation, the three scale features are integrated into a unified scale multi-scale fusion feature.
[0157] The formula used is as follows:
[0158]
[0159] Where F1(X) is the output feature of the multi-scale perception module; F N F is a nonlinear transformation function; Conv1 For convolution operations, P 3×3 (X) represents the output of a 3×3 convolution operation on input X; P 7×7 (X) represents the output of a 7×7 convolution operation on input X; P 11×11 (X) is the output of an 11×11 convolution operation on input X.
[0160] More specifically, the working process of the multi-semantic guidance fusion module includes:
[0161] Upsampling is performed on high-level semantic features to make the resolution of high-level semantic features consistent with that of the basic features, thereby achieving feature alignment;
[0162] The semantic similarity between the aligned high-level semantic features and the basic features is calculated using matrix multiplication to generate the selection weights.
[0163] Sim(x,y,c)=X′(x,y,c)·S(x,y,c)
[0164] Sim(x,y,c) is a semantic similarity value, which represents the degree of semantic association between basic feature X′ and high-level semantic feature S at spatial location (x,y) and channel c.
[0165] Next, the selection weights are normalized, and semantically guided feature selection and fusion are performed:
[0166] F2(X)=BN(ReLU(F Conv (X′⊙σ(Sim))))
[0167] Where BN() and σ() represent normalization, ReLU() represents the activation function, and X′⊙σ(Sim) represents applying the normalized screening matrix to X′.
[0168] More specifically, the adaptive weighting module works by:
[0169] Perform global average pooling on the input feature map X″ to obtain global statistics for each channel; the global average pooling operation satisfies:
[0170]
[0171] Where x(i,j,:) represents the feature vector of all channels at position (i,j); H1 and W1 are the height and width of the input feature map X″, respectively;
[0172] Next, a nonlinear transformation is performed on the global features to generate a spatial weight matrix that matches the size of the input features;
[0173] Finally, the generated weight matrix is multiplied element-wise with the input features to enhance key regions and suppress noisy regions.
[0174] W = σ(UpSample(F) Conv3 (F 全局平均池化 )))
[0175] F3(X″) = X″⊙W
[0176] W is the spatial weight matrix; σ is the activation function; UpSample is the upsampling operation; F Conv3 For convolution operation, channel and spatial features are fused on the upsampled features to extract weight information, making the weight matrix more in line with semantic requirements and enhancing the discriminative power of the weights; F3(X″) is the final output feature; ⊙ is the element-wise multiplication operation, which multiplies each element of the spatial weight matrix W with the corresponding element of the input feature X″, adjusts the feature value according to the weight, and completes the dynamic weight allocation.
[0177] To demonstrate the effectiveness and superiority of the method described in this invention, the following experiments were conducted:
[0178] 1) Dataset
[0179] The dataset used in this experiment is VisDrone2019, which is specifically designed for UAV vision tasks. It contains 8599 images, divided into training, validation, and test sets. The dataset covers 10 different target categories, including pedestrians, bicycles, cars, trucks, and buses. Each image is accompanied by detailed annotations, including the target's location coordinates, size, and morphological features, providing a standardized basis for the quantitative evaluation of the algorithm's detection results.
[0180] Furthermore, the VisDrone2019 dataset exhibits significant diversity and complexity, with its image samples covering scenes under different weather conditions, light intensities, shooting angles, and flight altitudes. It can accurately simulate various complex real-world environments, thus providing comprehensive support for verifying the effectiveness and robustness of the algorithm.
[0181] 2) Ablation experiment
[0182] To verify the effectiveness of the three improvements in the method described in this invention, this experiment selected YOLOv8n as the baseline model and used the VisDrone2019 dataset. The three improvements were added one by one to evaluate the impact of each improvement on the model's performance in detecting small targets. The experimental results are shown in Table 1.
[0183] Table 1
[0184] Group Model mAP@50 mAP@50:95 GFLOPs A YOLOv8n 0.348 0.204 8.1 B A+ST-C2f 0.359 0.211 9.7 C A+SDIoU 0.355 0.206 8.1 D A+ Small Target Detection Head 0.391 0.233 14.7 E D+ST-C2f+SDIoU 0.403 0.241 16.1
[0185] Experimental results show that the improved model using ST-C2f achieves mAP@50 and mAP@50:95 of 0.359 and 0.211, respectively, both superior to the original model's 0.348 and 0.204. When SDIoU is applied to the model, its related metrics also improve, without increasing model complexity. The most significant improvement is achieved when the small target detection head is applied, increasing mAP@50 to 0.391 (a 4.3 percentage point improvement) and mAP@50:95 to 0.233 (a 2.9 percentage point improvement). When all four modules are used simultaneously, the model's mAP@50 and mAP@50:95 reach 0.403 and 0.241, respectively, representing improvements of 5.5 and 3.7 percentage points, demonstrating the best detection performance. However, the model's complexity increases significantly.
[0186] 3) Performance curve comparison
[0187] Figure 4 and Figure 5 Performance evaluation curves for the YOLOv8n model and the improved model of this invention (ST-YOLOv8n model) are shown separately. The PR curve (left) contains 11 curves of different colors, corresponding to the 10 target categories and the overall mAP@50 metric. It can be seen that every curve of the improved model is superior to the original model. The RC curve (right) reflects the change in the model's recall rate with confidence level. The comparison shows that, compared to the original model, the improved model can effectively reduce the number of false positives and false negatives.
[0188] 4) Visual comparison
[0189] Figure 6 and Figure 7 The detection results of the YOLOv8n model and the improved model are shown separately. The comparison reveals that the improved model achieves higher confidence in predicting multiple targets and detects more targets.
[0190] Finally, the present invention also includes an electronic device and a computer-readable storage medium.
[0191] The electronic device includes a memory and a processor, which are interconnected. The memory stores computer instructions, and the processor executes these computer instructions to implement the aforementioned ground target detection method based on the ST-YOLOv8n model.
[0192] The computer-readable storage medium stores computer instructions, which, when executed by a computer, implement the aforementioned ground target detection method based on the ST-YOLOv8n model.
[0193] The above-described embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Therefore, any changes made in accordance with the shape and principle of the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for detecting ground targets in UAV aerial photography based on the ST-YOLOv8n model, characterized in that, include: The original C2f module and CIoU loss function of the YOLOv8n model are replaced with the ST-C2f module and SDIoU loss function, respectively; and the MD-Head multidimensional object detection head is integrated into the detection head part of the YOLOv8n model, thus obtaining the improved YOLOv8n model, namely the ST-YOLOv8n model. The ST-C2f module features a dual-branch heterogeneous structure and a dynamic gated fusion unit. The dual-branch heterogeneous structure includes a local branch and a global branch. The local branch sequentially employs a 3×3 depthwise separable convolution to extract pixel-level local detail features and a 7×7 dilated convolution to expand the receptive field and integrate complex background features. The global branch introduces a SwingTransformer module, which calculates the long-range dependencies of the captured feature maps through window self-attention and combines a window translation mechanism to achieve global contextual information interaction. The dynamic gated fusion unit adaptively fuses the local features output from the local branch and the global features output from the global branch using a learnable weight matrix to obtain a fused feature map. The SDIoU loss function is constructed by introducing the squared diagonal distance between the predicted box and the ground truth box and the average area ratio between the predicted box and the ground truth box, in order to optimize the target localization robustness and reduce the localization error of geometrically deformed targets. The MD-Head multidimensional target detection head includes a multi-scale perception module, a semantic-guided fusion module, and an adaptive weighting module. The multi-scale perception module extracts target features at different scales by combining different pooling kernels. The semantic-guided fusion module aligns high-level semantic features with basic features and generates gating weights by calculating semantic relevance, thereby achieving the filtering and fusion of semantic information related to small targets. The adaptive weighting module compresses the feature space dimension through global average pooling and generates channel weights to enhance the features of key target regions. Train the ST-YOLOv8n model to obtain the trained ST-YOLOv8n model; The trained ST-YOLOv8n model is used to detect ground targets.
2. The UAV aerial ground target detection method based on the ST-YOLOv8n model according to claim 1, characterized in that, The ST-C2f module processes the input feature map as follows: First, perform dimensionality reduction on the input feature map X∈R using convolution. H×W×C Channel compression is performed to reduce subsequent computation and retain core features; then the compressed feature map is segmented by channel, and the segmented feature maps are input into the local branch and the global branch respectively; C′=C / 2 X′=Conv(X) Where the dimension is R H×W×C H is the height of the input feature map, W is the width of the input feature map; C is the initial number of channels of the input feature map; C′ is the target number of channels after channel compression; X′ is the intermediate feature map after channel compression; X′ L This is a local feature branch feature map after channel segmentation; X′ G This is the global feature branch feature map after channel segmentation.
3. The UAV aerial ground target detection method based on the ST-YOLOv8n model according to claim 2, characterized in that, The working process of the SwinTransformer module in the global branch includes: The segmented global feature branch feature map X′ is obtained by using a window function. G Divide into non-overlapping local windows X=SplitWin(X)=[X1,X2,…,X n ] SplitWin() is the function for dividing the window. M is the size of the window; Calculate window self-attention within each local window using the following formula: WinAtten(X i )=SelfAtten(X i P Q ,X i P K ,X i P V ) The window self-attention mechanism is as follows: In the formula, P Q P K P V is the projection matrix, Q, K, and V are three vectors generated by the input features through the projection matrix, corresponding to the query matrix, key matrix, and value matrix, respectively; d is the number of channels; Softmax() is a function used to transform the dot product of Q and K into a probability distribution, thereby normalizing the attention weights. The window translation mechanism is implemented to periodically move the window position and enable information interaction between adjacent windows; ShiftWinAtt(Q,K,V)=WinAtten(Shift(Q),Shift(K),Shift(V)) ShiftWinAtt() represents the window translation mechanism, and Shift() represents the translation transformation.
4. The UAV aerial ground target detection method based on the ST-YOLOv8n model according to claim 1, characterized in that, The formula for the SDIoU loss function is as follows: Where IoU is the intersection-union ratio between the predicted bounding box B and the ground truth bounding box A: The squared Euclidean distance between the top-left corner of the ground truth bounding box A and the top-left corner of the predicted bounding box B is: The squared Euclidean distance between the bottom right corner of the ground truth bounding box A and the bottom right corner of the predicted bounding box B is: S t and S p These are the areas of the ground truth bounding box A and the predicted bounding box B, respectively. S represents the shape similarity index: These are the coordinates of the top left and bottom right corners of the ground truth bounding box A and the predicted bounding box B, respectively. α is a weighted correction coefficient: v is a quantitative indicator of shape difference: eps is a non-zero constant value used to avoid division by zero errors during calculations; w A and h A These are the width and height of the actual bounding box A, respectively. B and h B These are the width and height of the prediction box B, respectively.
5. The UAV aerial ground target detection method based on the ST-YOLOv8n model according to claim 1, characterized in that, In the MD-Head multidimensional target detection head, the multi-scale perception module, semantic-guided fusion module, and adaptive weighting module are connected in a serial manner, specifically in the order of multi-scale perception module → semantic-guided fusion module → adaptive weighting module; the output is shown by the following formula: F out (X)=F3(F2(F1(X))) Among them, F out ( ) represents the output of the MD-Head multidimensional target detection head, and F1(), F2(), and F3() represent the multi-scale perception module function, the semantic-guided fusion module function, and the adaptive weighting module function, respectively.
6. The UAV aerial ground target detection method based on the ST-YOLOv8n model according to claim 5, characterized in that, The multi-scale perception module includes a 3×3 pooling kernel, a 7×7 pooling kernel, and an 11×11 pooling kernel. These pooling kernels use parallel computing to extract small-scale target detail features, medium-scale target contour features, and large-scale target global features, respectively. Then, the three scale features are integrated into a unified scale multi-scale fusion feature through convolution fusion operation. The formula used is as follows: Where F1(X) is the output feature of the multi-scale sensing module; F N F is a nonlinear transformation function; Conv1 For convolution operations, P 3×3 (X) represents the output of a 3×3 convolution operation on input X; P 7×7 (X) represents the output of a 7×7 convolution operation on input X; P 11×11 (X) is the output of an 11×11 convolution operation on input X.
7. The UAV aerial ground target detection method based on the ST-YOLOv8n model according to claim 5, characterized in that, The working process of the semantic guidance fusion module includes: Upsampling is performed on high-level semantic features to make the resolution of high-level semantic features consistent with that of the basic features, thereby achieving feature alignment; The semantic similarity between the aligned high-level semantic features and the basic features is calculated using matrix multiplication to generate the selection weights. Sim(x,y,c)=X′(x,y,c)·S(x,y,c) Sim(x,y,c) is a semantic similarity value, which represents the degree of semantic association between basic feature X′ and high-level semantic feature S at spatial location (x,y) and channel c. Next, the selection weights are normalized, and semantically guided feature selection and fusion are performed: F2(X)=BN(ReLU(F Conv (X′⊙σ(Sim)))) Where BN() and σ() represent normalization, ReLU() represents the activation function, and X′⊙σ(Sim) means applying the normalized screening matrix to X′.
8. The UAV aerial ground target detection method based on the ST-YOLOv8n model according to claim 5, characterized in that, The working process of the adaptive weighting module includes: Perform global average pooling on the input feature map X″ to obtain global statistics for each channel; the global average pooling operation satisfies: Where x(i,j,:) represents the feature vector of all channels at position (i,j); H1 and W1 are the height and width of the input feature map X″, respectively; Next, a nonlinear transformation is performed on the global features to generate a spatial weight matrix that matches the size of the input features; Finally, the generated weight matrix is multiplied element-wise with the input features to enhance key regions and suppress noisy regions. W=σ(UpSample(F Conv3 (F 全局平均池化 ))) F3(X″) = X″⊙W W is the spatial weight matrix; σ is the activation function; UpSample is the upsampling operation; F Conv3 For convolution operation, channel and spatial features are fused on the upsampled features to extract weight information, making the weight matrix more in line with semantic requirements and enhancing the discriminative power of the weights; F3(X″) is the final output feature; ⊙ is the element-wise multiplication operation, which multiplies each element of the spatial weight matrix W with the corresponding element of the input feature X″, adjusts the feature value according to the weight, and completes the dynamic weight allocation.
9. An electronic device, characterized in that, include: The system includes a memory and a processor, which are interconnected. The memory stores computer instructions, and the processor executes these computer instructions to implement the ground target detection method based on the ST-YOLOv8n model as described in any one of claims 1-8.
10. A computer-readable storage medium storing computer instructions, characterized in that: When the computer instructions are executed by the computer, they implement the ground target detection method based on the ST-YOLOv8n model as described in any one of claims 1-8.
Citation Information
Cited By
Unmanned aerial vehicle aerial photography road defect detection method and system based on improved YOLOv11
CN121353964A
A Method and System for Road Defect Detection Based on Improved YOLOv11 Aerial Photography
CN121353964B
Parking space detection method based on YOLOv8n vehicle-mounted unmanned aerial vehicle data
CN121708445A