Unmanned aerial vehicle video image recognition method and system based on YOLOV13
By replacing the DS-C3k2 module in the YOLOV13 model with a multi-scale grouped dilated convolution module and the A2C2f module with a global Transformer module, an improved YOLOV13 model is constructed. This solves the problems of insufficient efficiency in simultaneous multi-scale target detection and insufficient global feature extraction, and achieves accurate and efficient identification and multi-scale feature representation of hidden danger targets at construction sites.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-04-03
AI Technical Summary
In the current UAV image hazard detection at highway construction sites, there are problems such as insufficient efficiency in simultaneous detection of multi-scale targets, difficulty in balancing the efficiency and accuracy of hazard identification at different scales, and lack of ability to extract and express multi-scale features across the entire domain, making it difficult to accurately and efficiently identify hazard targets with significant scale differences and complex types.
The DS-C3k2 module in the original YOLOV13 model was replaced with a multi-scale grouped dilated convolution module, and the A2C2f module was replaced with a global Transformer module to construct an improved YOLOV13 model. The multi-scale grouped dilated convolution module improves the performance of simultaneous multi-scale target detection, and the global Transformer module fills the gap in the ability to extract and express global multi-scale features.
It enables accurate and efficient identification of potential hazards at construction sites, meets actual detection needs, improves the accuracy and computational efficiency of multi-scale target detection, and adapts to multi-scale feature representation capabilities in complex scenarios.
Smart Images

Figure CN121789084A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and system for UAV video image recognition based on YOLOv13. Background Technology
[0002] In modern construction management of highway projects, drones have become a core technological means for identifying potential hazards at construction sites due to their flexibility, wide coverage, and high operational efficiency. By collecting high-definition video images of the construction site, drones can quickly identify and issue early warnings for various risk targets such as violations, equipment malfunctions, and lack of safety protection. This provides data support for project safety management, significantly reduces the cost and safety risks of manual inspections, and ensures the orderly progress of the construction process.
[0003] Currently, deep learning-based target detection algorithms are the core technology supporting hazard identification in UAV images. Among them, mainstream solutions, represented by the YOLO series algorithms, are widely used in various visual inspection scenarios due to their efficient real-time inference performance. Through continuous iterative optimization of the backbone network, feature fusion structure, and detection head design, the YOLO series algorithms have gradually improved the accuracy and speed of target detection, enabling them to meet the multi-target detection needs in common scenarios and providing a feasible technical path for hazard identification in UAV images.
[0004] However, in the existing scenarios of drone image hazard detection at highway construction sites, there are problems such as insufficient efficiency in simultaneous detection of multi-scale targets and difficulty in balancing the efficiency and accuracy of hazard identification at different scales. In addition, there are shortcomings such as a lack of ability to extract and express multi-scale features across the entire domain and an inability to fully explore and utilize information across the entire domain. Ultimately, this makes it difficult to accurately and efficiently identify hazard targets with significant scale differences and complex types, and thus fails to meet actual detection needs. Summary of the Invention
[0005] To address the challenges in existing UAV image-based hazard detection scenarios at highway construction sites, which suffer from insufficient efficiency in simultaneous multi-scale target detection and difficulty in balancing the efficiency and accuracy of hazard identification at different scales, as well as a lack of global multi-scale feature extraction and expression capabilities and an inability to fully utilize global information, ultimately leading to the inability to accurately and efficiently identify hazard targets with significant scale differences and diverse types, thus failing to meet actual detection needs, this invention provides a UAV video image recognition method and system based on YOLOv13.
[0006] The technical solutions provided by the embodiments of the present invention are as follows: The first aspect of this invention provides a UAV video image recognition method based on YOLOv13, comprising: S1: Acquire drone video images of the construction site of the highway construction project to be identified.
[0007] S2: Replace the DS-C3k2 module in the original YOLOV13 model with a multi-scale grouped dilated convolution module.
[0008] S3: Replace the A2C2f module in the original YOLOV13 model with the global Transformer module.
[0009] S4: Construct an improved YOLOv13 model based on a multi-scale grouped dilated convolution module and a global Transformer module.
[0010] S5: Input the UAV video images into the improved YOLOV13 model for hazard identification, and output the hazard identification results of the construction site of the highway construction project to be identified.
[0011] The second aspect of this invention provides a UAV video image recognition system based on YOLOv13, comprising: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the UAV video image recognition method based on YOLOv13 as described in the first aspect.
[0012] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the UAV video image recognition method based on YOLOv13 as described in the first aspect.
[0013] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, the DS-C3k2 module in the original YOLOV13 model is replaced with a multi-scale grouped dilated convolution module, which can effectively improve the synchronous detection efficiency of multi-scale targets, while taking into account the recognition efficiency and accuracy of hidden danger targets at different scales. The A2C2f module in the original YOLOV13 model is replaced with a global Transformer module, which fills the gap in the ability to extract and express global multi-scale features, and realizes the full mining and utilization of global information. The UAV video images are input into the improved YOLOV13 model for hidden danger recognition, which can realize accurate and efficient recognition of hidden danger targets at construction sites and meet the actual detection needs. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart illustrating a UAV video image recognition method based on YOLOv13, provided as an embodiment of the present invention.
[0016] Figure 2 This is an architecture diagram of a multi-scale grouped dilated convolution module provided in an embodiment of the present invention.
[0017] Figure 3 This is an architecture diagram of a global Transformer module provided in an embodiment of the present invention.
[0018] Figure 4 This is an architecture diagram of an improved YOLOV13 model provided in an embodiment of the present invention.
[0019] Figure 5 This is a schematic diagram of the structure of a UAV video image recognition system based on YOLOv13, provided as an embodiment of the present invention. Detailed Implementation
[0020] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0021] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0022] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0023] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0024] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0025] Reference manual attached Figure 1 The diagram shows a flowchart of a UAV video image recognition method based on YOLOv13 provided by an embodiment of the present invention.
[0026] This invention provides a UAV video image recognition method based on YOLOv13. This method can be implemented by a YOLOv13-based UAV video image recognition device, which can be a terminal or a server. The processing flow of the YOLOv13-based UAV video image recognition method may include the following steps:
[0027] S1: Acquire drone video images of the construction site of the highway construction project to be identified.
[0028] Reference manual attached Figure 2 The diagram shows the structure of a multi-scale grouped dilated convolution module for a UAV video image recognition method based on YOLOv13 provided by an embodiment of the present invention.
[0029] It should be noted that the Multi-Scale Grouped Dilated Convolutional Module (MSGDC) specifically includes: Bi (binarization component), Bi-3×3 grouped convolutional layer and RPRELU activation layer.
[0030] Among them, Bi (binarization component) is used to binarize the input features and branch output features, compressing feature data while improving representation capability.
[0031] The Bi-3×3 grouped convolutional layer consists of three parallel grouped convolutional units (Group=n, all with an expansion rate of 1), with the number of input channels n as the number of groups, to achieve local feature extraction under multi-path conditions.
[0032] The RPRELU activation layer corresponds to each group convolutional layer and is used to introduce non-linear transformations to enhance the expressive power of features.
[0033] Furthermore, the Multi-Scale Grouped Dilated Convolutional Module (MSGDC) starts with input features of n channels, first performs Bi binarization, and then simultaneously connects to three parallel branches. Each branch contains a structure of "Bi-3×3 grouped convolution (Group=n, dilation rate 1) + RPRELU activation layer". The three branches complete feature extraction and activation in parallel. Finally, the outputs of the three branches are summed and subjected to Bi binarization again to form the complete feature processing flow of the MSGDC module, realizing feature capture and fusion under multiple paths.
[0034] S2: Replace the DS-C3k2 module in the original YOLOV13 model with a multi-scale grouped dilated convolution module.
[0035] The original YOLOv13 model is designed around "lightweight and high performance," abandoning the traditional linear computation paradigm and adopting an innovative architecture to achieve a dual improvement in detection accuracy and computational efficiency. Its backbone network extracts multi-scale feature maps from B1 to B5 through specific modules, and incorporates a hypergraph association enhancement mechanism (HyperACE) and a full-process aggregation-distribution paradigm (FullPAD) in the neck section to optimize feature fusion and gradient propagation efficiency. The detection head is adapted to multi-scale feature output. In the MSCOCO benchmark test, the lightweight version has significant optimizations in key indicators such as mAP, inference speed, and memory usage compared to previous generation models (such as YOLOv11-N and YOLOv12-N), and is suitable for deployment in multiple scenarios such as CPU, GPU, and mobile devices.
[0036] The DS-C3k2 module is a core lightweight component of the original YOLOv13 model's backbone network, integrating depthwise separable convolution (DSConv) with the C3k2 residual structure design. Its core function is to efficiently extract multi-scale features while minimizing the number of model parameters and computational complexity (FLOPs). Compared to traditional large-kernel convolutions, this module reduces redundant computations while maintaining good feature extraction capabilities, providing crucial support for the lightweight deployment and efficient inference of the original YOLOv13 model. It is an important foundation for achieving a "precision-efficiency balance" in the model.
[0037] The Multi-Scale Grouped Dilated Convolutional Module (MSGDC) is a feature extraction module designed for multi-scale target detection. Its core consists of three groups of grouped convolutional layers with different dilation rates, coupled with an RPRRELU activation layer and Bi binarization. It reduces computational overhead through grouped convolution, expands the receptive field by using different dilation rates, and achieves efficient capture and fusion of feature information at different scales. Simultaneously, it enhances feature representation capabilities through binarization and specific activation mechanisms. Compared to conventional convolution and self-attention modules, this module significantly reduces model size while enhancing feature adaptability for multi-scale targets, making it particularly suitable for detecting potentially hazardous targets with significant scale differences in UAV images.
[0038] Optionally, the multi-scale grouped dilated convolution module specifically includes: The system consists of a first feature extraction branch, a second feature extraction branch, and a third feature extraction branch. Each feature extraction branch includes a binarized 3×3 grouped dilated convolutional layer and an RPReLU activation layer.
[0039] It should be noted that the Multi-Scale Grouped Dilated Convolutional Module (MSGDC) comprises four stages in the feature pyramid, each with different feature sizes in both spatial and channel dimensions. Given an input image, convolution-based patch embedding layers are used to segment the image and project it onto the feature sequence.
[0040] Furthermore, the pyramid structure can extract multi-scale features and improve feature performance by increasing the channel dimension of the output features. However, this also results in features with large spatial resolution in the early stages of the model, significantly increasing the complexity of the attention module. MSGDC effectively solves this problem. MSGDC consists of three 3×3 grouped dilated convolutional layers, each with a different dilation rate to achieve multi-scale feature fusion. Simultaneously, compared to ordinary convolution and self-attention modules, grouped convolution significantly reduces the model's parameters and computational complexity.
[0041] The multi-scale grouped dilated convolution module is specifically used for: The input feature maps of the multi-scale grouped dilated convolution module are respectively input to the first feature extraction branch, the second feature extraction branch, and the third feature extraction branch, and the first feature map, the second feature map, and the third feature map are output.
[0042] Optionally, the input feature map of the multi-scale grouped dilated convolution module is specifically as follows: ; ; in, X 0 indicates the input feature map of the multi-scale grouped dilated convolution module. P e This represents a learnable embedding location. H 0 Represents the initial feature map, GELU represents gelu Activation function bn Indicates the normalization layer. Cov Indicates a convolutional layer. I This represents the output image of the previous layer.
[0043] It should be noted that, X 0 represents the output image of its preceding module (i.e., the "upper layer" in the formula) before the MSGDC module begins performing feature extraction operations. I The features obtained after preprocessing are the initial input features of the MSGDC module feature processing flow.
[0044] Optionally, the calculation formula for the feature extraction branch is as follows: ; ; ; in, Indicates the first l -1st floor n The output feature maps of each feature extraction branch n =1,2,3 RPReLU This represents the RPReLU activation function. Indicates expansion rate dil for 2n Binarized 3×3 grouped dilated convolution with -1 , B a The binary function parameter is represented by `sign`, which indicates the sign function. x i express RPReLU The function in the first i Input on each channel, γ i and All indicate the first i Learnable displacements distributed across each channel β i Indicates the first i Learnable coefficients controlling the negative slope on each channel. X Indicates input features, b Indicates the scaling factor. a This indicates a bias towards science departments.
[0045] The first, second, and third feature maps are summed element-wise, and the sums are then batch-normalized to obtain the batch-normalized feature maps. ; in, H l-1 Indicates the first l Batch-normalized feature maps output from layer -1 Indicates the first l The first feature map output from layer -1 Indicates the first l The second feature map output from layer -1 Indicates the first l The third feature map output from layer -1.
[0046] The input feature map and the batch-normalized feature map of the multi-scale grouped dilated convolutional module are added element-wise to obtain the output feature map of the multi-scale grouped dilated convolutional module: ; in, In the multi-scale grouped dilated convolution module, the first... lThe output feature map of layer -1, where MSGDC represents the multi-scale grouped dilated convolutional module. X l-1 In the multi-scale grouped dilated convolution module, the first... l Input feature map of layer -1.
[0047] It should be noted that, X l-1 It is the input feature map of a specific level within the MSGDC module. It is the initial input of the level before the core operation is performed. It may be the output feature of the previous level within the module, or it may be the initial input of a certain stage when the module is processed in stages. But in essence, it is the pre-operation input of a certain level within the module.
[0048] In this embodiment of the invention, by replacing the core DS-C3k2 module in the backbone network of the original YOLOV13 model with a Multi-Scale Grouped Dilated Convolutional Module (MSGDC), the original model's core advantages of lightweight and efficient inference are maintained, while the key pain points of multi-scale target detection and feature extraction are specifically addressed. MSGDC, with its three groups of 3×3 grouped dilated convolutions with different dilation rates, can efficiently capture and fuse feature information at different scales. Combined with binarization operations and RPRRELU activation layers to enhance feature representation capabilities, it significantly improves the adaptability to targets with significant scale differences in UAV images. Simultaneously, through grouped convolution and a reasonable structural design, it reduces the number of model parameters and computational complexity, avoids a surge in attention module complexity, and ensures the integrity and stability of feature extraction through a complete process of learnable embedding positions, multi-branch feature extraction, element-wise addition, and batch normalization. Ultimately, it achieves dual optimization in multi-scale target detection accuracy and computational efficiency, making it more suitable for the actual needs of UAV image hazard detection at highway construction sites.
[0049] Reference manual attached Figure 3 The diagram shows the global Transformer module framework of a UAV video image recognition method based on YOLOv13 provided by an embodiment of the present invention.
[0050] It should be noted that the global Transformer module specifically includes: a region self-attention module, a global self-attention module, and auxiliary components.
[0051] The region self-attention module starts with the input features (H×W×C), first divides them into windows (such as L×L×C) and then reshapes them (R) to generate query (Q), key (K), and value (V) vectors. Through attention calculation and reshaping operations, it outputs the result of fusing local / regional features, and then adds it to the original feature residual to enhance the capture of local details.
[0052] Among them, the global self-attention module is used to receive the output features of the region self-attention. After reshaping to generate Q / K / V, it models the whole graph dependency through dot product attention and combines residual connection and reconstruction operations to achieve the fusion of global context information.
[0053] The auxiliary components include window partitioning, reshaping, and residual connections, which are used for feature dimension adjustment and information preservation to ensure the collaborative operation of the dual attention modules.
[0054] Furthermore, the input features (H×W×C) first enter the region self-attention module, where they are windowed (split into multiple L×L×C sub-features) and reshaped to generate Q / K / V vectors. The region features are obtained through attention calculation and reshaping operations, and then added to the residual of the original input features. Subsequently, this fused feature is input into the global self-attention module, where it is reshaped to generate new Q / K / V vectors. After completing the attention calculation across the entire image, the final features (H×W×C) are output by combining residual connections and reconstruction operations. The entire process achieves the coordinated operation of region self-attention and global self-attention through residual connections and feature propagation, completing the "local-global" feature fusion.
[0055] S3: Replace the A2C2f module in the original YOLOV13 model with the global Transformer module.
[0056] The A2C2f module is the core feature fusion module in the neck region of the original YOLOv13 model. Optimized from the classic C2f module, it significantly enhances cross-scale feature interaction capabilities. Through hierarchical residual connections and multi-path feature propagation, it efficiently integrates features from different levels of the backbone network output. This preserves detailed information from low-level features while incorporating semantic information from high-level features. Simultaneously, it optimizes gradient propagation efficiency and reduces information loss during feature fusion, providing more discriminative fused features for subsequent detection heads. It is a crucial supporting component for the original YOLOv13 model to achieve high-precision object detection.
[0057] The Global Transformer module is an innovative module designed for global feature extraction and multi-scale fusion. Its core integrates wavelet encoding / decoding structures and dual attention mechanisms (Regional Self-Attention RSA and Global Self-Attention GSA). It first decomposes the high- and low-frequency components of features through wavelet feature downsampling (WFD), which are then processed specifically by the Transformer and residual blocks. Wavelet feature upscaling (WFU) then achieves cross-scale feature alignment and fusion. Simultaneously, RSA focuses on capturing local and regional feature details, while GSA is responsible for mining global contextual relationships. The two work together to achieve a deep collaborative expression of local, regional, and global features, effectively compensating for the shortcomings of traditional algorithms in global information mining and significantly improving the feature representation capabilities of multi-scale targets in complex scenes.
[0058] Optionally, the global Transformer module specifically includes: Regional self-attention mechanism submodule and global self-attention mechanism submodule.
[0059] The region self-attention mechanism submodule is an attention component focused on capturing local and region-level dependencies of input features. By dividing the attention computation range into fixed-size windows or key regions, it effectively controls computational complexity while accurately capturing local details (such as target edges and textures). It first splits the input features into multiple independent window feature blocks, then enhances local contextual information through reshaping, depthwise separable convolution, and pointwise convolution. It models feature relationships within the region using learnable parameters and replaces traditional large-size attention maps with simplified region attention maps to improve efficiency. Finally, it outputs an enhanced representation that fuses local and region features, adapting to scenarios requiring enhanced local detail capture, such as small targets or occluded targets.
[0060] The global self-attention mechanism submodule overcomes the limitations of local perspectives and models attention components that address long-distance dependencies across the entire feature map, enabling the discovery of global semantic associations across regions and scales. It first extracts and optimizes global features through pointwise convolution and depthwise separable convolution. Then, it subdivides the feature channels into multiple attention heads to learn feature associations from multiple perspectives. Finally, it generates attention weights across the entire image by calculating the dot product of feature vectors, achieving efficient integration and transmission of global information. Without limiting the computational scope, it fully utilizes global contextual information, compensating for the limitations of local attention, and balances computational efficiency through a reasonable structural design, making it suitable for complex scenarios requiring the association of global information.
[0061] It's important to note that effectively utilizing local, regional, and global features is crucial for improving recognition accuracy. Local regions, containing multiple pixels, are best modeled using small kernels (1×1 or 3×3) to capture typical features such as local details. Regional features, containing dozens of pixels, are modeled using large kernel convolutions or window-based Transformers. Most methods focus only on utilizing local and global features, or local and regional features, resulting in poor recognition performance. By constructing a global Transformer as the main module for feature extraction, it consists of two parts: Regional Self-Attention (RSA) focuses on extracting local and regional features, while Global Self-Attention (GSA) is responsible for extracting both local and global features.
[0062] In one possible implementation, the global Transformer module is specifically used for: Wavelet downsampling is performed on the output feature map of the depthwise separable convolutional layer to obtain the input features of the global Transformer module and the input features of the first residual block.
[0063] Among them, depthwise separable convolutional layers are the core components supporting lightweight models and efficient inference. They are widely integrated into key structures such as the DS-C3k2 module of the backbone network. They replace traditional standard convolutions with a split design of "depthwise convolution (DWConv) and pointwise convolution (PWConv)". First, depthwise convolutions independently extract spatial features for each channel of the input features, capturing only local spatial information within a single channel. Then, 1×1 pointwise convolutions are used to complete feature fusion and dimensionality optimization between different channels. While retaining the model's ability to effectively capture target features, this significantly reduces the number of parameters and computational complexity, effectively reducing redundant operations.
[0064] It should be noted that introducing wavelet downsampling can effectively solve the problem of irreversible image distortion and unclear edges caused by traditional downsampling.
[0065] In one possible implementation, the method for determining the input features of the global Transformer module specifically includes: Extract shallow features from the output feature map of depth-separable convolutional layers.
[0066] It should be noted that the shallow feature extraction method is as follows: it is directly obtained from the feature map (including spatial resolution and corresponding number of channels) output by the depth separable convolutional layer, without complex deep transformation or multiple rounds of high-level feature fusion. It retains the basic texture, edge and low-level semantic information of the input image, and can be directly used for wavelet feature downsampling processing to separate the low-frequency part and the high-frequency part in each direction, which can be adapted to the feature input requirements of the global Transformer module and the residual block respectively.
[0067] By using wavelet transform, shallow features are processed in a hierarchical manner to obtain multiple wavelet bands: ; in, This represents the low-frequency component of shallow features. This represents the high-frequency component of shallow layer features in the horizontal direction. This represents the high-frequency component of shallow layer features in the vertical direction. This represents the high-frequency component along the diagonal direction of shallow features. WT Represents wavelet transform, F 1 indicates shallow features.
[0068] The low-frequency components are used as the input features of the global Transformer module; the high-frequency components in each direction are used as the input features of the first residual block. ; in, F high Indicates the enhanced high-frequency portion, R Represents the residual block. Represents the high-frequency component in the horizontal direction. This represents the high-frequency component in the vertical direction. This indicates the high-frequency component along the diagonal direction.
[0069] The input features of the global Transformer module are fed into the regional self-attention mechanism submodule, which outputs a regional attention map.
[0070] In one possible implementation, the method for determining the region attention map specifically includes: The region self-attention mechanism submodule extracts the input features of the global Transformer module to obtain window features: ; in, X m Indicates the first m A window feature block, m =1,2,…, n , n Indicates the number of window features. Split ( ) represents the partitioning function. X Indicates the window size.
[0071] By combining pointwise convolution and depthwise separable convolution, the window features are projected onto the query vector, key vector, and value vector of the region self-attention mechanism submodule: ; in, Q i Indicates the first i A query vector, K i Indicates the first i A key vector, V i Indicates the first i A vector of values RS Represents the reshaping operator, D This indicates a depthwise separable convolutional layer. P This indicates a pointwise convolutional layer. X i Indicates the first i A window feature block.
[0072] Based on the query vector, key vector, and value vector, determine the region attention map: ; in, The low-frequency portion of the enhanced shallow features is represented by the region attention map. Attention Indicates regional self-attention. Q i Indicates the first i A query vector, K i Indicates the first i A key vector, V Represents a value vector. V i Indicates the first i A vector of values ReLU This represents the activation function. Indicates the first i Transpose of a key vector α This represents the learnable parameters.
[0073] The input features of the region attention map and the first residual block are added element by element, and the result of the element-by-element addition is upgraded with wavelet features to obtain the input feature map of the global self-attention mechanism submodule and the input features of the second residual block.
[0074] It should be noted that since feature resolutions differ at different scales, upsampling is required to align them to the same resolution before feature fusion. However, direct fusion is not optimal because it introduces high- and low-frequency aliasing to some extent. Introducing wavelet feature upscaling for image scale transformation can effectively utilize features at different scales.
[0075] In one possible implementation, the method for determining the input feature map of the global self-attention mechanism submodule specifically includes: The input features of the region attention map and the first residual block are added element by element to obtain the fused feature input map.
[0076] Perform wavelet transform on the fused feature input map to obtain wavelet subbands: ; in, This represents the low-frequency portion of the fused feature input map. This represents the high-frequency component in the horizontal direction of the fused feature input map. This represents the high-frequency component in the vertical direction of the fused feature input map. This represents the high-frequency component along the diagonal direction of the fused feature input map. WT Represents wavelet transform, F s This represents the fused feature input map that is larger than a preset scale.
[0077] By processing the wavelet subbands using residual blocks and inverse wavelet transform, the input feature map of the global self-attention mechanism submodule is obtained: ; in, This represents the input feature map of the global self-attention mechanism submodule. IWT This represents the inverse wavelet transform. C Indicates splicing, This represents the low-frequency portion of the fused feature input map. F s+1 This represents the fused feature input map smaller than a preset scale. R Represents the residual block. This represents the high-frequency component in the horizontal direction of the fused feature input map. This represents the high-frequency component in the vertical direction of the fused feature input map. This represents the high-frequency component along the diagonal direction of the fused feature input map.
[0078] The input feature map of the global self-attention mechanism submodule is input into the global self-attention mechanism submodule, and the output is the global attention map.
[0079] It should be noted that for the input feature map of the global self-attention mechanism submodule, local information is first extracted from the input feature map using 1×1 pointwise convolution and 3×3 depthwise separable convolution to ensure accurate recovery of image feature details. Then, the channels are subdivided into multiple heads. h Simultaneously, it learns different self-attention maps. Specifically, it enhances the overall image features after local details are rendered based on the generated query vector, key vector, and value vector. Finally, a global attention map is created by calculating the dot product of the query vector and the key vector.
[0080] The global attention map and the input features of the second residual block are added element-wise to obtain the output feature map of the global Transformer module.
[0081] In this embodiment of the invention, by replacing the A2C2f module in the neck of the original YOLOV13 model with the global Transformer (FDT) module, a significant improvement in feature extraction and fusion capabilities is achieved while maintaining the cross-scale feature interaction advantages of the original module. The FDT module, supported by a wavelet encoding / decoding structure, effectively avoids the image distortion problem of traditional downsampling through wavelet feature downsampling, accurately splits and specifically processes high- and low-frequency components of features, and then solves the high- and low-frequency aliasing problem in cross-scale feature fusion through wavelet feature upsampling, ensuring the effectiveness of feature alignment and fusion. Simultaneously, its integrated regional self-attention and global self-attention dual mechanisms can accurately capture local details, regional features, and global contextual relationships, respectively, overcoming the shortcomings of traditional methods that only focus on some hierarchical features, and achieving deep collaborative expression of local, regional, and global features. With the help of a complete feature processing workflow, the ability to mine information across the entire domain and represent features at multiple scales has been greatly enhanced. This effectively solves the problem of insufficient feature extraction in complex scenarios by the original model, enabling the model to more accurately identify potential hazards with significant scale differences and complex types in UAV images of highway construction sites, and further improve detection accuracy and scene adaptability.
[0082] Reference manual attached Figure 4 The diagram shows an improved YOLOv13 model framework diagram of a UAV video image recognition method based on YOLOv13 provided by an embodiment of the present invention.
[0083] It should be noted that the instruction manual includes... Figure 4 In this context, Conv represents a convolutional layer, MSGDC3K2 represents a multi-scale grouped dilated convolutional module, DSConv represents a depthwise separable convolutional layer, FDT represents a global Transformer module, Upsample represents an upsampling layer, Concat represents a concatenation layer, FullPAD Tunnel represents a feature tunnel, and Detect-P5 represents a detection head.
[0084] S4: Construct an improved YOLOv13 model based on a multi-scale grouped dilated convolution module and a global Transformer module.
[0085] Optionally, the improved YOLOV13 model specifically includes: a backbone network, a neck network, and a detection head network.
[0086] The backbone network specifically includes: a first convolutional layer, a second convolutional layer, a first multi-scale grouped dilated convolutional module, a third convolutional layer, a second multi-scale grouped dilated convolutional module, a first depthwise separable convolutional layer, a first global Transformer module, a second depthwise separable convolutional layer, and a second global Transformer module.
[0087] It should be noted that in the backbone network, the first convolutional layer, the second convolutional layer, the first multi-scale grouped dilated convolutional module, the third convolutional layer, the second multi-scale grouped dilated convolutional module, the first depthwise separable convolutional layer, the first global Transformer module, the second depthwise separable convolutional layer, and the second global Transformer module are connected sequentially from top to bottom.
[0088] The neck network specifically includes: HyperACE module, first feature tunnel, first upsampling layer, first splicing layer, third multi-scale grouped dilated convolution module, second upsampling layer, second splicing layer, fourth multi-scale grouped dilated convolution module, second feature tunnel, fourth convolutional layer, third splicing layer, fifth multi-scale grouped dilated convolution module, fifth convolutional layer, fourth splicing layer, sixth multi-scale grouped dilated convolution module, and third feature tunnel.
[0089] It should be noted that the first feature tunnel includes a first feature fusion layer, a second feature fusion layer, and a third feature fusion layer from bottom to top; the second feature tunnel includes a fourth feature fusion layer and a fifth feature fusion layer from bottom to top; and the third feature tunnel includes a sixth feature fusion layer and a seventh feature fusion layer from bottom to top. Each feature fusion layer is used to add the two input features element by element.
[0090] Furthermore, the second multi-scale grouped dilated convolutional module, the first global Transformer module, and the second global Transformer module are all connected to the HyperACE module. The second global Transformer module and the HyperACE module (where the output of the HyperACE module is H5) are both connected to the first feature fusion layer. The first global Transformer module and the HyperACE module (where the output of the HyperACE module is H4) are both connected to the second feature fusion layer. The second multi-scale grouped dilated convolutional module and the HyperACE module (where the output of the HyperACE module is H3) are both connected to the third feature fusion layer. The third multi-scale grouped dilated convolutional module and the HyperACE module (where the output of the HyperACE module is H4) are both connected to the fourth feature fusion layer. The fourth multi-scale grouped dilated convolutional module and the HyperACE module (where the output of the HyperACE module is H3) are both connected to the fifth feature fusion layer. The sixth multi-scale grouped dilated convolutional module and the HyperACE module (the output of the HyperACE module here is H5) are both connected to the sixth feature fusion layer. The fifth multi-scale grouped dilated convolutional module and the HyperACE module (the output of the HyperACE module here is H4) are both connected to the seventh feature fusion layer. The first feature fusion layer is connected to the first upsampling layer. The second feature fusion layer and the first upsampling layer are both connected to the first concatenation layer. The first concatenation layer is connected to the third multi-scale grouped dilated convolutional module. The third multi-scale grouped dilated convolutional module is connected to the second upsampling layer. The third feature fusion layer and the second upsampling layer are both connected to the second concatenation layer. The second concatenation layer is connected to the fourth multi-scale grouped dilated convolutional module. The fifth feature fusion layer is connected to the fourth convolutional layer. Both the fourth feature fusion layer and the fourth convolutional layer are connected to the third splicing layer. The third splicing layer is connected to the fifth multi-scale grouped dilated convolutional module. The fifth multi-scale grouped dilated convolutional module is connected to the fifth convolutional layer. Both the first feature fusion layer and the fifth convolutional layer are connected to the fourth splicing layer. The fourth splicing layer is connected to the sixth multi-scale grouped dilated convolutional module.
[0091] It should be noted that H5 refers to low-resolution enhancement features, H4 refers to medium-resolution enhancement features, and H3 refers to high-resolution enhancement features.
[0092] The detection head network specifically includes: a first detection head, a second detection head, and a third detection head.
[0093] It should be noted that the fifth feature fusion layer is connected to the first detection head, the seventh feature fusion layer is connected to the second detection head, and the sixth feature fusion layer is connected to the third detection head.
[0094] It should be noted that MSGDC employs three grouped convolutional layers with different dilation rates to effectively fuse information at different scales, enhancing the representational power of binary activations and improving the model's representational capabilities, especially when processing large-scale data. Simultaneously, compared to conventional convolutional and self-attention modules, it reduces the model size, significantly improving computational cost and efficiency. The Global Attention Feature Fusion (FAT) module primarily strengthens feature extraction and enhances feature expressiveness. Specifically, it interacts with global information across features at different scales, effectively fusing local detail information, regional feature information, and global shape and position information in a weighted manner, ensuring the model can capture features across different ranges, resulting in better model recognition performance. Therefore, by complementing the advantages of the MSGDC and FAT algorithms, the YOLOv13 algorithm framework is optimized.
[0095] In this embodiment of the invention, an improved YOLOv13 model is constructed by integrating a multi-scale grouped dilated convolutional module (MSGDC) and a global Transformer module (FDT). This achieves synergistic optimization of the backbone network, neck network, and detection head network, fully leveraging the complementary advantages of the two modules. The MSGDC module, with its grouped convolutional design using different dilation rates, reduces model size and improves computational efficiency while efficiently integrating multi-scale features, enhancing adaptability to potential targets at different scales. The global Transformer module, through wavelet encoding / decoding and a dual attention mechanism, deeply mines global information and collaboratively expresses local, regional, and global features, compensating for the shortcomings of traditional models in global feature extraction. By combining HyperACE modules, feature tunnels, upsampling and stitching layers, and other components, the improved model forms a closed loop in the entire process of backbone network feature extraction, neck network feature fusion and accurate output of the detection head. It not only continues the lightweight and efficient inference characteristics of the original YOLOV13, but also significantly improves the accuracy of multi-scale target detection and adaptability to complex scenes. It can more accurately and efficiently meet the needs of detecting hidden dangers with significant scale differences and complex types in UAV images of highway construction sites.
[0096] S5: Input the UAV video images into the improved YOLOV13 model for hazard identification, and output the hazard identification results of the construction site of the highway construction project to be identified.
[0097] Reference manual attached Figure 5 The diagram shows a schematic of the structure of a UAV video image recognition system based on YOLOv13 provided by the present invention.
[0098] The present invention also provides a UAV video image recognition system 20 based on YOLOv13, applied to the aforementioned UAV video image recognition method based on YOLOv13, comprising: Processor 201.
[0099] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201, they implement the UAV video image recognition method based on YOLOv13 as described in the method embodiment.
[0100] The UAV video image recognition system 20 based on YOLOv13 provided by this invention can execute the above-mentioned UAV video image recognition method based on YOLOv13 and achieve the same or similar technical effects. To avoid repetition, this invention will not elaborate further.
[0101] It should be understood that the processor in the embodiments of the present invention can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0102] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0103] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0104] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0105] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0106] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0107] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0108] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0109] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0110] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0111] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0112] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0113] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the UAV video image recognition method based on YOLOv13 as described in the method embodiment.
[0114] The present invention provides a computer-readable storage medium that can implement the steps and effects of the UAV video image recognition method based on YOLOv13 in the above-described method embodiments. To avoid repetition, the present invention will not repeat them.
[0115] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0116] The following points need to be explained: (1) The accompanying drawings of the embodiments of the present invention only involve the structures involved in the embodiments of the present invention. Other structures can refer to the general design.
[0117] (2) For clarity, the thickness of layers or regions is enlarged or reduced in the drawings used to describe embodiments of the invention, i.e., these drawings are not drawn to scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element or there may be intermediate elements.
[0118] (3) Where there is no conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0119] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A UAV video image recognition method based on YOLOv13, characterized in that, include: S1: Acquire drone video images of the construction site of the highway construction project to be identified; S2: Replace the DS-C3k2 module in the original YOLOV13 model with a multi-scale grouped dilated convolution module; S3: Replace the A2C2f module in the original YOLOV13 model with a global Transformer module; S4: Construct an improved YOLOv13 model based on the multi-scale grouped dilated convolution module and the global Transformer module; S5: Input the UAV video images into the improved YOLOV13 model for hazard identification, and output the hazard identification results of the construction site of the highway construction project to be identified.
2. The UAV video image recognition method based on YOLOv13 according to claim 1, characterized in that, The improved YOLOV13 model specifically includes: a backbone network, a neck network, and a detection head network; The backbone network specifically includes: a first convolutional layer, a second convolutional layer, a first multi-scale grouped dilated convolutional module, a third convolutional layer, a second multi-scale grouped dilated convolutional module, a first depthwise separable convolutional layer, a first global Transformer module, a second depthwise separable convolutional layer, and a second global Transformer module. The neck network specifically includes: a HyperACE module, a first feature tunnel, a first upsampling layer, a first splicing layer, a third multi-scale grouped dilated convolutional module, a second upsampling layer, a second splicing layer, a fourth multi-scale grouped dilated convolutional module, a second feature tunnel, a fourth convolutional layer, a third splicing layer, a fifth multi-scale grouped dilated convolutional module, a fifth convolutional layer, a fourth splicing layer, a sixth multi-scale grouped dilated convolutional module, and a third feature tunnel; The detection head network specifically includes: a first detection head, a second detection head, and a third detection head.
3. The UAV video image recognition method based on YOLOv13 according to claim 1, characterized in that, The multi-scale grouped dilated convolution module specifically includes: The first feature extraction branch, the second feature extraction branch, and the third feature extraction branch, each of the feature extraction branches includes a binarized 3×3 grouped dilated convolutional layer and an RPReLU activation layer; The multi-scale grouped dilated convolution module is specifically used for: The input feature maps of the multi-scale grouped dilated convolution module are respectively input to the first feature extraction branch, the second feature extraction branch, and the third feature extraction branch, and the first feature map, the second feature map, and the third feature map are output. The first feature map, the second feature map, and the third feature map are added element-wise, and the result of the element-wise addition is batch normalized to obtain the batch normalized feature map: ; in, H l-1 Indicates the first l Batch-normalized feature maps output from layer -1 Indicates the first l The first feature map output from layer -1 Indicates the first l The second feature map output from layer -1 Indicates the first l -1 layer output third feature map; The input feature map of the multi-scale grouped dilated convolutional module and the batch normalized feature map are added element-wise to obtain the output feature map of the multi-scale grouped dilated convolutional module: ; in, In the multi-scale grouped dilated convolution module, the first... l The output feature map of layer -1, where MSGDC represents the multi-scale grouped dilated convolutional module. X l-1 In the multi-scale grouped dilated convolution module, the first... l Input feature map of layer -1.
4. The UAV video image recognition method based on YOLOv13 according to claim 3, characterized in that, The input feature map of the multi-scale grouped dilated convolution module is specifically as follows: ; ; in, X 0 indicates the input feature map of the multi-scale grouped dilated convolution module. P e This represents a learnable embedding location. H 0 Represents the initial feature map, GELU represents gelu Activation function bn Indicates the normalization layer. Cov Indicates a convolutional layer. I This represents the output image of the previous layer; The specific calculation formula for the feature extraction branch is as follows: ; ; ; in, Indicates the first l -1st floor n The output feature maps of each feature extraction branch n =1,2,3 RPReLU This represents the RPReLU activation function. Indicates expansion rate dil for 2n Binarized 3×3 grouped dilated convolution with -1 , B a The binary function parameter is represented by `sign`, which indicates the sign function. x i express RPReLU The function in the first i Input on each channel, γ i and All indicate the first i Learnable displacements distributed across each channel β i Indicates the first i Learnable coefficients controlling the negative slope on each channel. X Indicates input features, b Indicates the scaling factor. a This indicates a bias towards science departments.
5. The UAV video image recognition method based on YOLOv13 according to claim 1, characterized in that, The global Transformer module specifically includes: Regional self-attention mechanism submodule and global self-attention mechanism submodule.
6. The UAV video image recognition method based on YOLOv13 according to claim 5, characterized in that, The global Transformer module is specifically used for: Wavelet downsampling is performed on the output feature map of the depthwise separable convolutional layer to obtain the input features of the global Transformer module and the input features of the first residual block; The input features of the global Transformer module are input into the region self-attention mechanism submodule, and the region attention map is output. The input features of the region attention map and the first residual block are added element by element, and the result of the element-by-element addition is upgraded with wavelet features to obtain the input feature map of the global self-attention mechanism submodule and the input features of the second residual block. The input feature map of the global self-attention mechanism submodule is input into the global self-attention mechanism submodule, and the global attention map is output. The global attention map and the input features of the second residual block are added element-wise to obtain the output feature map of the global Transformer module.
7. The UAV video image recognition method based on YOLOv13 according to claim 6, characterized in that, The specific methods for determining the input features of the global Transformer module include: Extract shallow features from the output feature map of the depth-separable convolutional layer; By performing wavelet transform on the shallow features in a hierarchical manner, multiple wavelet bands are obtained: ; in, This represents the low-frequency component of shallow features. This represents the high-frequency component of shallow layer features in the horizontal direction. This represents the high-frequency component of shallow layer features in the vertical direction. This represents the high-frequency component along the diagonal direction of shallow features. WT Represents wavelet transform, F 1 indicates shallow features; The low-frequency components are used as input features of the global Transformer module; the high-frequency components in each direction are used as input features of the first residual block. ; in, F high Indicates the enhanced high-frequency portion, R Represents the residual block. Represents the high-frequency component in the horizontal direction. This represents the high-frequency component in the vertical direction. Indicates the high-frequency component in the diagonal direction; The method for determining the region attention map specifically includes: The region self-attention mechanism submodule extracts the input features of the global Transformer module to obtain window features: ; in, X m Indicates the first m A window feature block, m =1,2,…, n , n Indicates the number of window features. Split ( ) represents the partitioning function. X Indicates window size; By combining pointwise convolution and depthwise separable convolution, the window features are projected onto the query vector, key vector, and value vector of the region self-attention mechanism submodule: ; in, Q i Indicates the first i A query vector, K i Indicates the first i A key vector, V i Indicates the first i A vector of values RS Represents the reshaping operator, D This indicates a depthwise separable convolutional layer. P This indicates a pointwise convolutional layer. X i Indicates the first i Each window feature block; The region attention map is determined based on the query vector, the key vector, and the value vector: ; in, The low-frequency portion of the enhanced shallow features is represented by the region attention map. Attention Indicates regional self-attention. Q i Indicates the first i A query vector, K i Indicates the first i A key vector, V Represents a value vector. V i Indicates the first i A vector of values ReLU This represents the activation function. Indicates the first i Transpose of a key vector α This represents the learnable parameters.
8. The UAV video image recognition method based on YOLOv13 according to claim 6, characterized in that, The specific method for determining the input feature map of the global self-attention mechanism submodule includes: The input features of the region attention map and the first residual block are added element by element to obtain the fused feature input map; Perform wavelet transform on the fused feature input map to obtain wavelet sub-bands: ; in, This represents the low-frequency portion of the fused feature input map. This represents the high-frequency component in the horizontal direction of the fused feature input map. This represents the high-frequency component in the vertical direction of the fused feature input map. This represents the high-frequency component along the diagonal direction of the fused feature input map. WT Represents wavelet transform, F s This represents a fused feature input map that is larger than a preset scale; The wavelet subband is processed using residual blocks and inverse wavelet transform to obtain the input feature map of the global self-attention mechanism submodule: ; in, This represents the input feature map of the global self-attention mechanism submodule. IWT This represents the inverse wavelet transform. C Indicates splicing, This represents the low-frequency portion of the fused feature input map. F s+1 This represents the fused feature input map smaller than a preset scale. R Represents the residual block. This represents the high-frequency component in the horizontal direction of the fused feature input map. This represents the high-frequency component in the vertical direction of the fused feature input map. This represents the high-frequency component along the diagonal direction of the fused feature input map.
9. A UAV video image recognition system based on YOLOv13, characterized in that, include: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the UAV video image recognition method based on YOLOv13 as described in any one of claims 1 to 8.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the UAV video image recognition method based on YOLOv13 as described in any one of claims 1 to 8.