Highway construction project construction site hidden danger identification method and system

By replacing the modules in the YOLOV13 model with multi-scale grouped dilated convolution, global Transformer, and frequency modulation modules, the problem of efficiency and accuracy imbalance in UAV image hazard identification was solved, improving the identification accuracy and adaptability to interference factors, and meeting the detection needs of highway construction sites.

CN121789085APending Publication Date: 2026-04-03CHINA ACAD OF TRANSPORTATION SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing UAV image hazard identification technology struggles to balance efficiency and accuracy in multi-scale target synchronous detection, resulting in identification imbalances, an inability to fully extract information from the entire field, and poor adaptability to interference factors such as noise, fog, and low light, leading to a decrease in identification accuracy and making it difficult to meet the actual needs of highway construction sites.

Method used

The DS-C3k2 module in the original YOLOV13 model was replaced with a multi-scale grouped dilated convolution module, the A2C2f module was replaced with a global Transformer module, and some feature tunnels were replaced with frequency modulation modules to construct an improved YOLOV13 model, thereby enhancing feature extraction and recognition capabilities.

Benefits of technology

It achieves a balance between efficiency and accuracy in identifying potential hazards at highway construction sites, fully leverages information across the entire area, enhances adaptability to interference factors, improves identification accuracy, and meets actual detection needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789085A_ABST
    Figure CN121789085A_ABST
Patent Text Reader

Abstract

The invention provides a road construction project construction site hidden danger identification method and system, and relates to the technical field of data processing, and the method comprises the steps: obtaining an unmanned plane video image of a to-be-identified road construction project construction site; all DS-C3k2 modules in the original YOLOV13 model are replaced by a multi-scale grouping expansion convolution module; all A2C2f modules in the original YOLOV13 model are replaced by a global Transform module, and the A2C2f modules in the original YOLOV13 model are replaced by a global Transform module; replacing a part of feature tunnels in the original YOLOV13 model with a frequency modulation module; the method comprises the following steps: constructing an improved YOLOV13 model based on a multi-scale grouping expansion convolution module, a global Transform module and a frequency modulation module; and inputting the unmanned aerial vehicle video image into the improved YOLOV13 model for hidden danger recognition, and outputting a hidden danger recognition result of the to-be-recognized road construction project construction site.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for identifying potential hazards at highway construction sites. Background Technology

[0002] At highway construction sites, hazard identification is a crucial step in ensuring construction safety and project quality. With the widespread adoption of drone technology, using drones equipped with cameras to collect video images for hazard identification at construction sites has become an efficient and convenient method.

[0003] Currently, in the field of UAV image hazard identification, target detection algorithms, represented by the YOLO series, are the mainstream technology choice. These algorithms are widely used in various target detection scenarios due to their high detection speed and certain recognition accuracy. The YOLO series algorithms, by constructing a specific network framework, achieve feature extraction and classification detection of targets in images. Their core idea is to transform the target detection task into a regression problem, directly predicting the target's location and category on the image. They possess the advantage of end-to-end detection and can meet the basic needs of real-time data acquisition and analysis by UAVs.

[0004] However, existing technologies struggle to balance efficiency and accuracy in multi-scale target synchronous detection, easily leading to identification imbalances. Furthermore, they do not fully mine information across the entire domain, failing to achieve accurate identification of hazards based on full-domain features. At the same time, they are poorly adaptable to interference factors such as noise, fog, and low light, resulting in decreased identification accuracy and making it difficult to meet the actual hazard detection needs of highway construction sites. Summary of the Invention

[0005] To address the challenges of existing technologies in balancing efficiency and accuracy during multi-scale target synchronous detection, which often leads to identification imbalances and insufficient mining of global information, making it impossible to accurately identify hazards based on global features; furthermore, these technologies are poorly adapted to interference factors such as noise, fog, and low light, resulting in decreased identification accuracy and making it difficult to meet the actual hazard detection needs of highway construction sites, this invention provides a method and system for identifying hazards at highway construction sites.

[0006] The technical solutions provided by the embodiments of the present invention are as follows: The first aspect of this invention provides a method for identifying potential hazards at highway construction sites, comprising: S1: Acquire drone video images of the construction site of the highway construction project to be identified.

[0007] S2: Replace all DS-C3k2 modules in the original YOLOV13 model with multi-scale grouped dilated convolutional modules.

[0008] S3: Replace all A2C2f modules in the original YOLOV13 model with global Transformer modules.

[0009] S4: Replace some feature tunnels in the original YOLOV13 model with frequency modulation modules.

[0010] S5: Construct an improved YOLOv13 model based on a multi-scale grouped dilated convolution module, a global Transformer module, and a frequency modulation module.

[0011] S6: Input the UAV video images into the improved YOLOV13 model for hazard identification, and output the hazard identification results of the construction site of the highway construction project to be identified.

[0012] A second aspect of the present invention provides a hazard identification system for highway construction sites, comprising: processor; The memory stores computer-readable instructions, which, when executed by the processor, implement the method for identifying potential hazards at highway construction sites as described in the first aspect.

[0013] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, all DS-C3k2 modules in the original YOLOV13 model are replaced with multi-scale grouped dilated convolution modules, which can balance efficiency and accuracy and prevent identification imbalance. All A2C2f modules are replaced with global Transformer modules, which can fully explore global information and achieve accurate identification of hidden dangers. At the same time, replacing some feature tunnels with frequency modulation modules can enhance adaptability to interference factors, improve identification accuracy, and fully match the actual hidden danger detection needs of highway construction sites. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a flowchart illustrating a method for identifying potential hazards at a highway construction site, as provided in an embodiment of the present invention.

[0016] Figure 2 This is an architecture diagram of a multi-scale grouped dilated convolution module provided in an embodiment of the present invention.

[0017] Figure 3 This is an architecture diagram of a global Transformer module provided in an embodiment of the present invention.

[0018] Figure 4 This is an architecture diagram of a frequency modulation module provided in an embodiment of the present invention.

[0019] Figure 5 This is an architecture diagram of an improved YOLOV13 model provided in an embodiment of the present invention.

[0020] Figure 6 This is a structural schematic diagram of a hazard identification system for highway construction projects, provided as an embodiment of the present invention. Detailed Implementation

[0021] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0022] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0023] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0024] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0025] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0026] Reference manual attached Figure 1 The diagram shows a flowchart of a method for identifying potential hazards at a highway construction site, as provided in an embodiment of the present invention.

[0027] This invention provides a method for identifying potential hazards at highway construction sites. This method can be implemented using a hazard identification device for highway construction sites, which can be a terminal or a server. The processing flow of the hazard identification method for highway construction sites may include the following steps:

[0028] S1: Acquire drone video images of the construction site of the highway construction project to be identified.

[0029] Reference manual attached Figure 2 The diagram shows an architecture of a multi-scale grouped dilated convolution module provided in an embodiment of the present invention.

[0030] It should be noted that the Multi-Scale Grouped Dilated Convolutional Module (MSGDC) specifically includes: Bi (binarization component), Bi-3×3 grouped convolutional layer and RPRELU activation layer.

[0031] Among them, Bi (binarization component) is used to binarize the input features and branch output features, compressing feature data while improving representation capability.

[0032] The Bi-3×3 grouped convolutional layer consists of three parallel grouped convolutional units (Group=n, all with an expansion rate of 1), with the number of input channels n as the number of groups, to achieve local feature extraction under multi-path conditions.

[0033] The RPRELU activation layer corresponds to each group convolutional layer and is used to introduce non-linear transformations to enhance the expressive power of features.

[0034] Furthermore, the Multi-Scale Grouped Dilated Convolutional Module (MSGDC) starts with input features of n channels, first performs Bi binarization, and then simultaneously connects to three parallel branches. Each branch contains a structure of "Bi-3×3 grouped convolution (Group=n, dilation rate 1) + RPRELU activation layer". The three branches complete feature extraction and activation in parallel. Finally, the outputs of the three branches are summed and subjected to Bi binarization again to form the complete feature processing flow of the MSGDC module, realizing feature capture and fusion under multiple paths.

[0035] S2: Replace all DS-C3k2 modules in the original YOLOV13 model with multi-scale grouped dilated convolutional modules.

[0036] The original YOLOV13 model, designed with lightweight and high accuracy as its core objectives, adopts the classic YOLO architecture paradigm of "Backbone - HyperACE - FullPAD - Detect". The backbone network progressively extracts multi-scale basic features of the image through modules such as Conv, DS-C3k2, DSConv, and FDT, outputting feature maps of different levels, B3, B4, and B5. After receiving the backbone features, the HyperACE module generates enhanced features H3, H4, and H5 through hypergraph modeling and branch fusion (aligned with the resolution and channels of B3, B4, and B5, respectively). The FullPAD module is responsible for distributing the enhanced features throughout the entire process to the backbone-neck junction, each layer of the neck, and the neck-detector junction. It completes cross-scale feature fusion through operations such as stitching, upsampling, and addition to enhance semantic consistency. Finally, the Detect module classifies and locates the targets based on the fused features, achieving efficient detection of targets at different scales.

[0037] The DS-C3k2 module is the core lightweight feature extraction unit in the backbone network and Neck module of the YOLOv13 model. Based on the CSP paradigm and incorporating the design of depthwise separable convolution (DSConv), it first reduces the dimensionality of the input features and unifies the channels through 1×1 convolution. Then, following the CSP approach, it is divided into a direct shortcut branch (preserving shallow features) and a main branch containing multiple cascaded DS-Bottlenecks. Finally, the two feature channels are concatenated and fused through 1×1 convolution to restore the channels. This significantly reduces the number of parameters and computation while taking into account the receptive field and multi-scale feature expression capabilities, supporting the model's real-time and efficient detection.

[0038] Among them, the multi-scale grouped dilated convolution module uses the number of channels as n The input is characterized by binarization (Bi) and then three groups are connected in parallel. n The convolutional branches are binarized and grouped with a 3×3 kernel and an expansion rate of 1. Each branch extracts features separately, and then the nonlinear expression is enhanced by the RPRelu activation function. Finally, the multi-branch outputs are fused to capture multi-scale feature information in a lightweight manner, thereby improving the efficiency and richness of feature extraction.

[0039] It should be noted that the Multi-Scale Grouped Dilated Convolutional Module (MSGDC) comprises four stages in the feature pyramid, each with different feature sizes in both spatial and channel dimensions. Given an input image, convolution-based patch embedding layers are used to segment the image and project it onto the feature sequence.

[0040] Furthermore, the pyramid structure can extract multi-scale features and improve feature performance by increasing the channel dimension of the output features. However, this also results in features with large spatial resolution in the early stages of the model, significantly increasing the complexity of the attention module. MSGDC effectively solves this problem. MSGDC consists of three 3×3 grouped dilated convolutional layers, each with a different dilation rate to achieve multi-scale feature fusion. Simultaneously, compared to ordinary convolution and self-attention modules, grouped convolution significantly reduces the model's parameters and computational complexity.

[0041] Optionally, the original YOLOV13 model includes a first feature tunnel, a second feature tunnel, and a third feature tunnel.

[0042] Optionally, the multi-scale grouped dilated convolution module specifically includes: The system consists of a first feature extraction branch, a second feature extraction branch, and a third feature extraction branch. Each feature extraction branch includes a binarized 3×3 grouped dilated convolutional layer and an RPReLU activation layer.

[0043] Optionally, the multi-scale grouped dilated convolution module is specifically used for: The input feature maps of the multi-scale grouped dilated convolution module are respectively input to the first feature extraction branch, the second feature extraction branch, and the third feature extraction branch, and the first feature map, the second feature map, and the third feature map are output.

[0044] Optionally, the input feature map of the multi-scale grouped dilated convolution module is specifically as follows: ; ; in, X 0 indicates the input feature map of the multi-scale grouped dilated convolution module. P e This represents a learnable embedding location. H 0 Represents the initial feature map, GELU represents gelu Activation function bn Indicates the normalization layer. Cov Indicates a convolutional layer. I This represents the output image of the previous layer.

[0045] Optionally, the calculation formula for the feature extraction branch is as follows: ; ; ; in, Indicates the first l -1st floorn The output feature maps of each feature extraction branch n =1,2,3 RPReLU This represents the RPReLU activation function. Indicates expansion rate dil for 2n Binarized 3×3 grouped dilated convolution with -1 , B a The binary function parameter is represented by `sign`, which indicates the sign function. x i express RPReLU The function in the first i Input on each channel, γ i and All indicate the first i Learnable displacements distributed across each channel β i Indicates the first i Learnable coefficients controlling the negative slope on each channel. X Indicates input features, b Indicates the scaling factor. a This indicates a bias towards science departments.

[0046] The first, second, and third feature maps are summed element-wise, and the sums are then batch-normalized to obtain the batch-normalized feature maps. ; in, H l-1 Indicates the first l Batch-normalized feature maps output from layer -1 Indicates the first l The first feature map output from layer -1 Indicates the first l The second feature map output from layer -1 Indicates the first l The third feature map output from layer -1.

[0047] The input feature map and the batch-normalized feature map of the multi-scale grouped dilated convolutional module are added element-wise to obtain the output feature map of the multi-scale grouped dilated convolutional module: ; in, In the multi-scale grouped dilated convolution module, the first... l The output feature map of layer -1, where MSGDC represents the multi-scale grouped dilated convolutional module. X l-1 In the multi-scale grouped dilated convolution module, the first... lInput feature map of layer -1.

[0048] In this embodiment of the invention, by replacing the backbone network of the original YOLOv13 model and the DS-C3k2 module in the Neck module with the Multi-Scale Grouped Dilated Convolution (MSGDC) module, the advantages of the original model's lightweight deployment are retained. At the same time, the multi-scale feature capture capability and nonlinear expression are enhanced simultaneously without significantly increasing the number of parameters and computational complexity by using binarized grouped dilated convolution with three branches and different dilation rates, RPRelu activation and residual connection design. Meanwhile, the computational efficiency is further optimized by using binarization processing and grouped convolution, which solves the module complexity problem caused by high-resolution features in the early stage of the feature pyramid. By enhancing the residuals of input features and multi-branch fused features, the richness and robustness of feature extraction are improved, and finally, the model's detection accuracy and inference efficiency are synergistically optimized.

[0049] Reference manual attached Figure 3 The diagram illustrates the architecture of a global Transformer module provided in an embodiment of the present invention.

[0050] It should be noted that the global Transformer module specifically includes: a region self-attention module, a global self-attention module, and auxiliary components.

[0051] The region self-attention module starts with the input features (H×W×C), first divides them into windows (such as L×L×C) and then reshapes them (R) to generate query (Q), key (K), and value (V) vectors. Through attention calculation and reshaping operations, it outputs the result of fusing local / regional features, and then adds it to the original feature residual to enhance the capture of local details.

[0052] Among them, the global self-attention module is used to receive the output features of the region self-attention. After reshaping to generate Q / K / V, it models the whole graph dependency through dot product attention and combines residual connection and reconstruction operations to achieve the fusion of global context information.

[0053] The auxiliary components include window partitioning, reshaping, and residual connections, which are used for feature dimension adjustment and information preservation to ensure the collaborative operation of the dual attention modules.

[0054] Furthermore, the input features (H×W×C) first enter the region self-attention module, where they are windowed (split into multiple L×L×C sub-features) and reshaped to generate Q / K / V vectors. The region features are obtained through attention calculation and reshaping operations, and then added to the residual of the original input features. Subsequently, this fused feature is input into the global self-attention module, where it is reshaped to generate new Q / K / V vectors. After completing the attention calculation across the entire image, the final features (H×W×C) are output by combining residual connections and reconstruction operations. The entire process achieves the coordinated operation of region self-attention and global self-attention through residual connections and feature propagation, completing the "local-global" feature fusion.

[0055] S3: Replace all A2C2f modules in the original YOLOV13 model with global Transformer modules.

[0056] The core of the A2C2f module is to integrate region attention with cross-stage features. It is achieved through 1×1 convolution dimensionality reduction / upgrading, ABlock (including region attention and MLP) and residual connections. It can flexibly switch attention on / off (when enabled, attention is calculated independently by region to reduce complexity and preserve receptive field; when disabled, it is replaced by conventional convolution bottleneck). It is widely deployed in the backbone and neck, which can effectively enhance multi-scale feature expression and discrimination, and balance accuracy and efficiency. It is suitable for dense prediction tasks such as object detection and segmentation.

[0057] The Global Transformer module is a highly efficient global context modeling unit for vision / sequence tasks. Its core is to overcome window / local constraints and directly model long-distance dependencies of the entire image / sequence through multi-head self-attention. It is often paired with positional encoding, residual connections, and layer normalization, supplemented by 1×1 convolutions or feedforward networks to complete feature transformation and nonlinear enhancement. To balance computational power and performance, optimization strategies such as parallel global attention and local convolution / window attention, sparse attention, or linear attention are often adopted in engineering to reduce computational complexity while maintaining global perception capabilities. It can be used as a core module of the backbone / neck and is widely used in image classification, object detection, semantic segmentation, and multimodal tasks, effectively improving the ability to model large-scale targets and complex contexts.

[0058] It's important to note that effectively utilizing local, regional, and global features is crucial for improving recognition accuracy. Local regions, containing multiple pixels, are best modeled using small kernels (1×1 or 3×3) to capture typical features such as local details. Regional features, containing dozens of pixels, are modeled using large kernel convolutions or window-based Transformers. Most methods focus only on utilizing local and global features, or local and regional features, resulting in poor recognition performance. By constructing a global Transformer as the main module for feature extraction, it consists of two parts: Regional Self-Attention (RSA) focuses on extracting local and regional features, while Global Self-Attention (GSA) is responsible for extracting both local and global features.

[0059] Optionally, the global Transformer module specifically includes: Regional self-attention mechanism submodule and global self-attention mechanism submodule.

[0060] The region self-attention mechanism submodule is a highly efficient attention unit for real-time vision tasks. It divides the input feature map into several continuous regions along the horizontal / vertical direction, and independently calculates multi-head self-attention within each region, replacing global self-attention to significantly reduce computational complexity while retaining a large receptive field. The input features are normalized by layers and linearly projected to obtain Q, K, and V. After being reshaped by region, attention weights and weighted sums are calculated in parallel. The region features are then stitched back to their original size. Combined with MLP, residual connections, and activation functions, feature transformation and nonlinear enhancement are completed. It is often used in conjunction with 1×1 convolutional dimensionality reduction / upgrading and embedded in ABlock modules such as A2C2f. It is widely used in backbone and neck regions, balancing accuracy and inference speed.

[0061] The global self-attention mechanism submodule is a core component of the global Transformer-like model. It allows each position to directly focus on all positions in the entire input sequence / feature map to model long-distance dependencies. After the input is normalized by layers, Q, K, and V are obtained through linear projection. Similarity is calculated by scaling dot product and then normalized by softmax to obtain attention weights. These weights are then used to sum V in a weighted manner. A multi-head mechanism is commonly used to learn complementary attention patterns in parallel from different subspaces. Subsequently, it is combined with feedforward networks, residual connections, and activation functions to complete feature transformation and nonlinear enhancement.

[0062] The global Transformer module is specifically used for: Wavelet downsampling is performed on the output feature map of the depthwise separable convolutional layer to obtain the input features of the global Transformer module and the input features of the first residual block.

[0063] Among them, depthwise separable convolutional layers decompose standard convolution into two steps: depthwise convolution and pointwise convolution. This significantly reduces the number of parameters and computational cost (FLOPs) while maintaining similar feature extraction capabilities, making it a core component of efficient models for mobile / edge devices. First, depthwise convolution is used to independently convolve each channel of the input feature map using a single convolutional kernel (capturing only intra-channel spatial information). Then, 1×1 pointwise convolution is used to linearly combine the output channels of the depthwise convolution (integrating cross-channel information), ultimately outputting a feature map with the same dimensions as standard convolution.

[0064] In one possible implementation, the method for determining the input features of the global Transformer module specifically includes: Extract shallow features from the output feature map of depth-separable convolutional layers.

[0065] It should be noted that the shallow feature extraction method is as follows: it is directly obtained from the feature map (including spatial resolution and corresponding number of channels) output by the depth separable convolutional layer, without complex deep transformation or multiple rounds of high-level feature fusion. It retains the basic texture, edge and low-level semantic information of the input image, and can be directly used for wavelet feature downsampling processing to separate the low-frequency part and the high-frequency part in each direction, which can be adapted to the feature input requirements of the global Transformer module and the residual block respectively.

[0066] By using wavelet transform, shallow features are processed in a hierarchical manner to obtain multiple wavelet bands: ; in, This represents the low-frequency component of shallow features. This represents the high-frequency component of shallow layer features in the horizontal direction. This represents the high-frequency component of shallow layer features in the vertical direction. This represents the high-frequency component along the diagonal direction of shallow features. WT Represents wavelet transform, F 1 indicates shallow features.

[0067] The low-frequency components are used as the input features of the global Transformer module; the high-frequency components in each direction are used as the input features of the first residual block. ; in, F high Indicates the enhanced high-frequency portion, R Represents the residual block. Represents the high-frequency component in the horizontal direction. This represents the high-frequency component in the vertical direction. This indicates the high-frequency component along the diagonal direction.

[0068] The input features of the global Transformer module are fed into the regional self-attention mechanism submodule, which outputs a regional attention map.

[0069] In one possible implementation, the method for determining the region attention map specifically includes: The region self-attention mechanism submodule extracts the input features of the global Transformer module to obtain window features: ; in, X m Indicates the first m A window feature block, m =1,2,…, n , n Indicates the number of window features. Split ( ) represents the partitioning function. X Indicates the window size.

[0070] By combining pointwise convolution and depthwise separable convolution, the window features are projected onto the query vector, key vector, and value vector of the region self-attention mechanism submodule: ; in, Q i Indicates the first i A query vector, K i Indicates the first i A key vector, V i Indicates the first i A vector of values RS Represents the reshaping operator, D This indicates a depthwise separable convolutional layer. P This indicates a pointwise convolutional layer. X i Indicates the first i A window feature block.

[0071] Based on the query vector, key vector, and value vector, determine the region attention map: ; in, The low-frequency portion of the enhanced shallow features is represented by the region attention map. Attention Indicates regional self-attention. Q i Indicates the first i A query vector, Ki Indicates the first i A key vector, V Represents a value vector. V i Indicates the first i A vector of values ReLU This represents the activation function. Indicates the first i Transpose of a key vector α This represents the learnable parameters.

[0072] The input features of the region attention map and the first residual block are added element by element, and the result of the element-by-element addition is upgraded with wavelet features to obtain the input feature map of the global self-attention mechanism submodule and the input features of the second residual block.

[0073] In one possible implementation, the method for determining the input feature map of the global self-attention mechanism submodule specifically includes: The input features of the region attention map and the first residual block are added element by element to obtain the fused feature input map.

[0074] Perform wavelet transform on the fused feature input map to obtain wavelet subbands: ; in, This represents the low-frequency portion of the fused feature input map. This represents the high-frequency component in the horizontal direction of the fused feature input map. This represents the high-frequency component in the vertical direction of the fused feature input map. This represents the high-frequency component along the diagonal direction of the fused feature input map. WT Represents wavelet transform, F s This represents the fused feature input map that is larger than a preset scale.

[0075] By processing the wavelet subbands using residual blocks and inverse wavelet transform, the input feature map of the global self-attention mechanism submodule is obtained: ; in, This represents the input feature map of the global self-attention mechanism submodule. IWT This represents the inverse wavelet transform. C Indicates splicing, This represents the low-frequency portion of the fused feature input map. F s+1 This represents the fused feature input map smaller than a preset scale. R Represents the residual block. This represents the high-frequency component in the horizontal direction of the fused feature input map. This represents the high-frequency component in the vertical direction of the fused feature input map. This represents the high-frequency component along the diagonal of the fused feature input map.

[0076] The input feature map of the global self-attention mechanism submodule is input into the global self-attention mechanism submodule, and the output is the global attention map.

[0077] The global attention map and the input features of the second residual block are added element-wise to obtain the output feature map of the global Transformer module.

[0078] In this embodiment of the invention, the A2C2f module in YOLOv13 is replaced with a global Transformer module, breaking through the local constraints of the original module's regional attention. Through the collaborative design of regional self-attention (RSA) and global self-attention (GSA), it accurately captures multi-level features at the local, regional, and global levels, solving the recognition limitations caused by most methods focusing only on some feature levels. By leveraging wavelet downsampling and residual fusion of high-frequency features, the feature details and contextual relationships are enhanced. Furthermore, through optimization strategies such as region partitioning for attention calculation and depthwise separable convolutional projection, the computational complexity of global attention is balanced while improving the modeling of large-scale targets and the feature discrimination of complex scenes, further enhancing the comprehensiveness and robustness of the model's feature representation and contributing to improved detection accuracy.

[0079] Reference manual attached Figure 4 The diagram shows an architecture diagram of a frequency modulation module provided in an embodiment of the present invention.

[0080] It should be noted that after the high-frequency and low-frequency features are input, the low-frequency features generate an attention descriptor through the corresponding branch, which is then multiplied element-wise with the high-frequency features to complete the modulation of the high-frequency from the low-frequency to the high-frequency. At the same time, the high-frequency features generate an attention descriptor through the corresponding branch, which is then multiplied element-wise with the low-frequency features to complete the modulation of the low-frequency from the high-frequency to the low-frequency. After that, the two modulated features are fused by element-wise addition, and then the channels are adjusted by 1×1 convolution to finally output the modulated features.

[0081] S4: Replace some feature tunnels in the original YOLOV13 model with frequency modulation modules.

[0082] It should be noted that the frequency modulation module, which allows information exchange between different frequency bands to calibrate low-frequency and high-frequency features, can effectively handle various types of image noise problems. Its goal is to ensure that one type of mined feature complements another. For example, high-frequency features contain edge and surprise texture details, so an ultra-lightweight spatial attention unit is used to enrich low-frequency mined features. Similarly, global information in low-frequency features is passed to the high-frequency branch through a channel attention unit.

[0083] In one possible implementation, S4 specifically involves replacing the second and third feature tunnels with a frequency modulation module.

[0084] Optionally, the frequency modulation module specifically includes: an ultra-lightweight spatial attention unit, a global information through-channel attention unit, and a 1×1 convolutional layer.

[0085] Among them, the ultra-lightweight spatial attention unit is a plug-and-play component for compact visual networks on mobile / edge devices. It replaces heavy matrix multiplication with lightweight operations such as depthwise separable convolution, 1×1 pointwise convolution and pooling, generating spatial weight maps with extremely low parameter and computational cost. It then weights the input feature map element by element, strengthening the target region, suppressing background redundancy, and improving spatial localization and context aggregation capabilities.

[0086] In this process, global information is first processed by channel attention units to compress the spatial dimension of the input feature map through operations such as global average pooling and global max pooling, and aggregated into channel-level descriptors carrying global spatial statistics. Then, nonlinear transformation and dimensionality reduction / restoration are performed through a lightweight multilayer perceptron or convolution. Channel-wise weights are generated by activation functions and finally multiplied with the original feature map channel-wise to complete channel recalibration guided by global information, which strengthens key channels, suppresses redundancy, and effectively models inter-channel dependencies and long-range context.

[0087] The frequency modulation module is specifically used for: The feature images from the multi-scale grouped dilated convolution module are input into the ultra-lightweight spatial attention unit to generate a spatial attention map. ; in, A H-L Representing a spatial attention map, δ This represents the sigmoid activation function. This indicates the 6th 7×7 convolutional layer. GAP c This indicates channel-level global average pooling. X h This represents the feature image of the multi-scale grouped dilated convolution module. GMP c [ , ] indicates channel-level max pooling, and [ , ] indicates concatenation operation.

[0088] Furthermore, the ultra-lightweight spatial attention unit computes a spatial attention map from high-frequency mined features, which is then used to supplement features in low-frequency branches. This unit utilizes two different channel pooling techniques in parallel to generate two single-channel spatial feature maps, each with a size of [size missing]. H × W×1. Then, these feature maps are concatenated along the channel dimension. The concatenated features are further refined by 7×7 convolution, followed by a sigmoid operation to generate the final spatial attention map.

[0089] Based on the spatial attention map, the first modulation feature is obtained through element-wise multiplication: ; in, Indicates the first modulation feature, X l ⊙ represents the low-frequency features discovered, and ⊙ represents element-wise multiplication.

[0090] Element-wise multiplication involves multiplying two tensors with identical dimensions element-wise at corresponding positions, outputting a new tensor with unchanged dimensions. This enables "weight modulation" or "feature enhancement," precisely strengthening the feature responses of key regions / channels and suppressing redundant information. It is also commonly used for complementary fusion of multi-branch features, allowing effective information to be superimposed element-wise from features from different sources. This operation is highly computationally efficient, introduces no additional parameters, and is a crucial fundamental operation for balancing model performance and inference speed.

[0091] The mined low-frequency features are input into the global information through the channel attention unit to generate an attention description map: ; in, A L-H This represents an attention description diagram. This indicates the 8th 1×1 convolutional layer. γ Represents the ReLU activation function. This indicates the 7th 1×1 convolutional layer. GAP s This represents global average pooling along the spatial dimension. This indicates the 10th 1×1 convolutional layer. This indicates the 9th 1×1 convolutional layer. GMP s This represents max pooling along the spatial dimension.

[0092] Furthermore, the global information is processed through a channel attention unit, a two-branch module, which handles the incoming low-frequency mined features, generating a feature descriptor that is then used to process high-frequency mined features. This unit applies global average pooling along the spatial dimension to the upper branch of the mined low-frequency features, resulting in a 1×1× CThe feature vectors are processed, then passed through two convolutional layers and a ReLU activation function; the lower branches of this unit also adopt the same structure, but the head uses max pooling; finally, the results of the two branches are summed. Based on this, the sigmoid function is applied to generate the final attention description map.

[0093] Based on the attention description map, the second modulation feature is obtained through element-wise multiplication: ; in, This indicates the second modulation feature.

[0094] The modulation feature is obtained by adding the first modulation feature and the second modulation feature element by element.

[0095] By using cross-attention units and 1×1 convolutional layers, the modulation features and feature images from the multi-scale grouped dilated convolutional module are fused to obtain the output image of the frequency modulation module.

[0096] The cross-attention unit is a cross-sequence / cross-modality / cross-branch attention component used to handle queries from one source and keys and values ​​from another different source. It generates weights through similarity calculation and weights the values, achieving cross-source information alignment and fusion, unlike query, key, and value-based attention. The query, key, and value are mapped to the same dimension through projection layers, and similarity is calculated using scaled dot product or additive methods. Attention weights are obtained through softmax normalization and then weighted and summed with the values. A multi-head design is commonly used to decompose the computation into multiple subspaces in parallel, and finally, the results are concatenated and projected for output, balancing multi-granularity associations and expressive power.

[0097] In this embodiment of the invention, the second and third feature tunnels of YOLOv13 are replaced with frequency modulation modules. Through bidirectional modulation of ultra-lightweight spatial attention units and global information channel attention units, complementary calibration of high-frequency features (including edge and texture details) and low-frequency features (including global context) is achieved, effectively handling image noise problems. Key features are precisely enhanced and redundancy is suppressed by element-level multiplication. Then, the modulation features and multi-scale grouped dilated convolution features are deeply fused through cross-attention units. This balances computational efficiency with lightweight design, enriches feature hierarchy through cross-frequency information interaction, and optimizes the feature transfer quality of feature tunnels, thereby improving the robustness and accuracy of the model in detecting targets of different scales in complex scenes.

[0098] Reference manual attached Figure 5 The diagram shows an architecture diagram of an improved YOLOV13 model provided by an embodiment of the present invention.

[0099] It should be noted that the instruction manual includes... Figure 4In this context, Conv represents a convolutional layer, MSGDC3K2 represents a multi-scale grouped dilated convolutional module, DSConv represents a depthwise separable convolutional layer, FDT represents a global Transformer module, Upsample represents an upsampling layer, Concat represents a stitching layer, FullPAD Tunnel represents a feature tunnel, FMoM represents a frequency modulation module, and Detect-P5 represents a detection head.

[0100] S5: Construct an improved YOLOv13 model based on a multi-scale grouped dilated convolution module, a global Transformer module, and a frequency modulation module.

[0101] It should be noted that MSGDC employs three grouped convolutional layers with different dilation rates to effectively fuse information at different scales, enhancing the representational power of binary activations and improving the model's representational capabilities, especially when processing large-scale data. Simultaneously, compared to conventional convolutions and self-attention modules, it reduces the model size, significantly improving computational cost and efficiency. FAT primarily strengthens feature extraction and enhances feature expressiveness. Specifically, it interacts with global information across features at different scales, effectively fusing local detail information, regional feature information, and global shape and position information in a weighted manner, ensuring the model can capture features across different ranges, resulting in better model recognition performance. Therefore, by complementing the advantages of MSGDC, FAT, and FMoM algorithms, a systematic optimization is performed on the backbone network and neck structure of the YOLO13 algorithm framework.

[0102] Optionally, the improved YOLOV13 model specifically includes: a backbone network, a neck network, and a detection head network.

[0103] The backbone network specifically includes: a first convolutional layer, a second convolutional layer, a first multi-scale grouped dilated convolutional module, a third convolutional layer, a second multi-scale grouped dilated convolutional module, a first depthwise separable convolutional layer, a first global Transformer module, a second depthwise separable convolutional layer, and a second global Transformer module.

[0104] It should be noted that in the backbone network, the first convolutional layer, the second convolutional layer, the first multi-scale grouped dilated convolutional module, the third convolutional layer, the second multi-scale grouped dilated convolutional module, the first depthwise separable convolutional layer, the first global Transformer module, the second depthwise separable convolutional layer, and the second global Transformer module are connected sequentially from top to bottom.

[0105] The neck network specifically includes: HyperACE module, first feature tunnel, first upsampling layer, first splicing layer, third multi-scale grouped dilated convolution module, second upsampling layer, second splicing layer, fourth multi-scale grouped dilated convolution module, first frequency modulation module, fourth convolutional layer, third splicing layer, fifth multi-scale grouped dilated convolution module, fifth convolutional layer, fourth splicing layer, sixth multi-scale grouped dilated convolution module, and second frequency modulation module.

[0106] It should be noted that the first feature tunnel includes a first feature fusion layer, a second feature fusion layer, and a third feature fusion layer from bottom to top; the first frequency modulation module includes a fourth feature fusion layer and a fifth feature fusion layer from bottom to top; and the second frequency modulation module includes a sixth feature fusion layer and a seventh feature fusion layer from bottom to top. Each feature fusion layer is used to add the two input features element by element.

[0107] Furthermore, the second multi-scale grouped dilated convolutional module, the first global Transformer module, and the second global Transformer module are all connected to the HyperACE module. The second global Transformer module and the HyperACE module (where the output of the HyperACE module is H5) are both connected to the first feature fusion layer. The first global Transformer module and the HyperACE module (where the output of the HyperACE module is H4) are both connected to the second feature fusion layer. The second multi-scale grouped dilated convolutional module and the HyperACE module (where the output of the HyperACE module is H3) are both connected to the third feature fusion layer. The third multi-scale grouped dilated convolutional module and the HyperACE module (where the output of the HyperACE module is H4) are both connected to the fourth feature fusion layer. The fourth multi-scale grouped dilated convolutional module and the HyperACE module (where the output of the HyperACE module is H3) are both connected to the fifth feature fusion layer. The sixth multi-scale grouped dilated convolutional module and the HyperACE module (the output of the HyperACE module here is H5) are both connected to the sixth feature fusion layer. The fifth multi-scale grouped dilated convolutional module and the HyperACE module (the output of the HyperACE module here is H4) are both connected to the seventh feature fusion layer. The first feature fusion layer is connected to the first upsampling layer. The second feature fusion layer and the first upsampling layer are both connected to the first concatenation layer. The first concatenation layer is connected to the third multi-scale grouped dilated convolutional module. The third multi-scale grouped dilated convolutional module is connected to the second upsampling layer. The third feature fusion layer and the second upsampling layer are both connected to the second concatenation layer. The second concatenation layer is connected to the fourth multi-scale grouped dilated convolutional module. The fifth feature fusion layer is connected to the fourth convolutional layer. Both the fourth feature fusion layer and the fourth convolutional layer are connected to the third splicing layer. The third splicing layer is connected to the fifth multi-scale grouped dilated convolutional module. The fifth multi-scale grouped dilated convolutional module is connected to the fifth convolutional layer. Both the first feature fusion layer and the fifth convolutional layer are connected to the fourth splicing layer. The fourth splicing layer is connected to the sixth multi-scale grouped dilated convolutional module.

[0108] It should be noted that H5 refers to low-resolution enhancement features, H4 refers to medium-resolution enhancement features, and H3 refers to high-resolution enhancement features.

[0109] The detection head network specifically includes: a first detection head, a second detection head, and a third detection head.

[0110] It should be noted that the fifth feature fusion layer is connected to the first detection head, the seventh feature fusion layer is connected to the second detection head, and the sixth feature fusion layer is connected to the third detection head.

[0111] In this embodiment of the invention, the constructed improved YOLOv13 model integrates three core modules—Multi-Scale Grouped Dilated Convolution (MSGDC), Global Transformer (FAT), and Frequency Modulation (FMoM)—to collaboratively optimize the backbone and neck structure of the original model. The backbone network consists of convolutional layers, MSGDC modules, and FAT modules connected in series. MSGDC's multi-dilation rate grouped convolutions are used to capture multi-scale features and improve binary activation representation capabilities. The FAT module is used to fuse local, regional, and global features to enhance long-range contextual modeling. The neck network, based on the enhanced features of the HyperACE module, forms a cross-scale feature interaction link through feature tunneling, frequency modulation modules, multi-scale grouped dilated convolution modules, upsampling, and concatenation operations. The frequency modulation module is used to achieve complementary calibration of high and low frequency features and noise suppression. Then, the feature fusion layer completes the accurate fusion of enhanced features (H3, H4, H5) at different resolutions. Finally, the detection head connects to the output of each feature fusion layer to achieve efficient classification and localization of multi-scale targets. The three modules complement each other, ensuring computational efficiency through lightweight design, and enhancing the richness and robustness of feature representation from three dimensions: multi-scale capture, global context modeling, and cross-band feature calibration, thus comprehensively optimizing the model's detection performance in complex scenarios.

[0112] S6: Input the UAV video images into the improved YOLOV13 model for hazard identification, and output the hazard identification results of the construction site of the highway construction project to be identified.

[0113] Reference manual attached Figure 6 The diagram shows a structural schematic of a hazard identification system for highway construction sites provided by the present invention.

[0114] The present invention also provides a hazard identification system 20 for highway construction sites, applied to the aforementioned hazard identification method for highway construction sites, comprising: Processor 201.

[0115] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201, the method for identifying hidden dangers at the construction site of a highway construction project, as described in the method embodiment, is implemented.

[0116] The highway construction site hazard identification system 20 provided by the present invention can execute the above-mentioned highway construction site hazard identification method and achieve the same or similar technical effects. To avoid duplication, the present invention will not elaborate further.

[0117] It should be understood that the processor in the embodiments of the present invention can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0118] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0119] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0120] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0121] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.

[0122] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0123] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0124] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0125] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0126] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0127] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0128] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0129] This invention provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method for identifying hidden dangers at highway construction sites as described in the method embodiments.

[0130] The present invention provides a computer-readable storage medium that can implement the steps and effects of the method for identifying hidden dangers at highway construction sites as described in the above-described method embodiments. To avoid repetition, the present invention will not elaborate further.

[0131] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0132] The following points need to be explained: (1) The accompanying drawings of the embodiments of the present invention only involve the structures involved in the embodiments of the present invention. Other structures can refer to the general design.

[0133] (2) For clarity, the thickness of layers or regions is enlarged or reduced in the drawings used to describe embodiments of the invention, i.e., these drawings are not drawn to scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element or there may be intermediate elements.

[0134] (3) Where there is no conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.

[0135] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for identifying potential hazards at highway construction sites, characterized in that, include: S1: Acquire drone video images of the construction site of the highway construction project to be identified; S2: Replace all DS-C3k2 modules in the original YOLOV13 model with multi-scale grouped dilated convolutional modules; S3: Replace all A2C2f modules in the original YOLOV13 model with global Transformer modules; S4: Replace some feature tunnels in the original YOLOV13 model with frequency modulation modules; S5: Construct an improved YOLOv13 model based on the multi-scale grouped dilated convolution module, the global Transformer module, and the frequency modulation module; S6: Input the UAV video images into the improved YOLOV13 model for hazard identification, and output the hazard identification results of the construction site of the highway construction project to be identified.

2. The method for identifying potential hazards at highway construction sites according to claim 1, characterized in that, The original YOLOV13 model includes a first feature tunnel, a second feature tunnel, and a third feature tunnel; Specifically, S4 is: Replace the second feature tunnel and the third feature tunnel with the frequency modulation module.

3. The method for identifying potential hazards at highway construction sites according to claim 2, characterized in that, The improved YOLOV13 model specifically includes: a backbone network, a neck network, and a detection head network; The backbone network specifically includes: a first convolutional layer, a second convolutional layer, a first multi-scale grouped dilated convolutional module, a third convolutional layer, a second multi-scale grouped dilated convolutional module, a first depthwise separable convolutional layer, a first global Transformer module, a second depthwise separable convolutional layer, and a second global Transformer module. The neck network specifically includes: a HyperACE module, a first feature tunnel, a first upsampling layer, a first splicing layer, a third multi-scale grouped dilated convolution module, a second upsampling layer, a second splicing layer, a fourth multi-scale grouped dilated convolution module, a first frequency modulation module, a fourth convolutional layer, a third splicing layer, a fifth multi-scale grouped dilated convolution module, a fifth convolutional layer, a fourth splicing layer, a sixth multi-scale grouped dilated convolution module, and a second frequency modulation module; The detection head network specifically includes: a first detection head, a second detection head, and a third detection head.

4. The method for identifying potential hazards at highway construction sites according to claim 1, characterized in that, The multi-scale grouped dilated convolution module specifically includes: The first feature extraction branch, the second feature extraction branch, and the third feature extraction branch, each of the feature extraction branches includes a binarized 3×3 grouped dilated convolutional layer and an RPReLU activation layer; The multi-scale grouped dilated convolution module is specifically used for: The input feature maps of the multi-scale grouped dilated convolution module are respectively input to the first feature extraction branch, the second feature extraction branch, and the third feature extraction branch, and the first feature map, the second feature map, and the third feature map are output. The first feature map, the second feature map, and the third feature map are added element-wise, and the result of the element-wise addition is batch normalized to obtain the batch normalized feature map: ; in, H l-1 Indicates the first l Batch-normalized feature maps output from layer -1 Indicates the first l The first feature map output from layer -1 Indicates the first l The second feature map output from layer -1 Indicates the first l -1 layer output third feature map; The input feature map of the multi-scale grouped dilated convolutional module and the batch normalized feature map are added element-wise to obtain the output feature map of the multi-scale grouped dilated convolutional module: ; in, In the multi-scale grouped dilated convolution module, the first... l The output feature map of layer -1, where MSGDC represents the multi-scale grouped dilated convolutional module. X l-1 In the multi-scale grouped dilated convolution module, the first... l Input feature map of layer -1.

5. The method for identifying potential hazards at highway construction sites according to claim 4, characterized in that, The input feature map of the multi-scale grouped dilated convolution module is specifically as follows: ; ; in, X 0 indicates the input feature map of the multi-scale grouped dilated convolution module. P e This represents a learnable embedding location. H 0 Represents the initial feature map, GELU represents gelu Activation function bn Indicates the normalization layer. Cov Indicates a convolutional layer. I This represents the output image of the previous layer; The specific calculation formula for the feature extraction branch is as follows: ; ; ; in, Indicates the first l -1st floor n The output feature maps of each feature extraction branch n =1,2,3 RPReLU This represents the RPReLU activation function. Indicates expansion rate dil for 2n Binarized 3×3 grouped dilated convolution with -1 , B a The binary function parameter is represented by `sign`, which indicates the sign function. x i express RPReLU The function in the first i Input on each channel, γ i and All indicate the first i Learnable displacements distributed across each channel β i Indicates the first i Learnable coefficients controlling the negative slope on each channel. X Indicates input features, b Indicates the scaling factor. a This indicates a bias towards science departments.

6. The method for identifying potential hazards at highway construction sites according to claim 1, characterized in that, The global Transformer module specifically includes: Region self-attention mechanism submodule and global self-attention mechanism submodule; The global Transformer module is specifically used for: Wavelet downsampling is performed on the output feature map of the depthwise separable convolutional layer to obtain the input features of the global Transformer module and the input features of the first residual block; The input features of the global Transformer module are input into the region self-attention mechanism submodule, and the region attention map is output. The input features of the region attention map and the first residual block are added element by element, and the result of the element-by-element addition is upgraded with wavelet features to obtain the input feature map of the global self-attention mechanism submodule and the input features of the second residual block. The input feature map of the global self-attention mechanism submodule is input into the global self-attention mechanism submodule, and the global attention map is output. The global attention map and the input features of the second residual block are added element-wise to obtain the output feature map of the global Transformer module.

7. The method for identifying potential hazards at highway construction sites according to claim 6, characterized in that, The specific methods for determining the input features of the global Transformer module include: Extract shallow features from the output feature map of the depth-separable convolutional layer; By performing wavelet transform on the shallow features in a hierarchical manner, multiple wavelet bands are obtained: ; in, This represents the low-frequency component of shallow features. This represents the high-frequency component of shallow layer features in the horizontal direction. This represents the high-frequency component of shallow layer features in the vertical direction. This represents the high-frequency component along the diagonal direction of shallow features. WT Represents wavelet transform, F 1 indicates shallow features; The low-frequency components are used as input features of the global Transformer module; the high-frequency components in each direction are used as input features of the first residual block. ; in, F high Indicates the enhanced high-frequency portion, R Represents the residual block. Represents the high-frequency component in the horizontal direction. This represents the high-frequency component in the vertical direction. Indicates the high-frequency component in the diagonal direction; The method for determining the region attention map specifically includes: The region self-attention mechanism submodule extracts the input features of the global Transformer module to obtain window features: ; in, X m Indicates the first m A window feature block, m =1,2,…, n , n Indicates the number of window features. Split ( ) represents the partitioning function. X Indicates window size; By combining pointwise convolution and depthwise separable convolution, the window features are projected onto the query vector, key vector, and value vector of the region self-attention mechanism submodule: ; in, Q i Indicates the first i A query vector, K i Indicates the first i A key vector, V i Indicates the first i A vector of values RS Represents the reshaping operator, D This indicates a depthwise separable convolutional layer. P This indicates a pointwise convolutional layer. X i Indicates the first i Each window feature block; The region attention map is determined based on the query vector, the key vector, and the value vector: ; in, The low-frequency portion of the enhanced shallow features is represented by the region attention map. Attention Indicates regional self-attention. Q i Indicates the first i A query vector, K i Indicates the first i A key vector, V Represents a value vector. V i Indicates the first i A vector of values ReLU This represents the activation function. Indicates the first i Transpose of a key vector α This represents the learnable parameters.

8. The method for identifying potential hazards at highway construction sites according to claim 6, characterized in that, The specific method for determining the input feature map of the global self-attention mechanism submodule includes: The input features of the region attention map and the first residual block are added element by element to obtain the fused feature input map; Perform wavelet transform on the fused feature input map to obtain wavelet sub-bands: ; in, This represents the low-frequency portion of the fused feature input map. This represents the high-frequency component in the horizontal direction of the fused feature input map. This represents the high-frequency component in the vertical direction of the fused feature input map. This represents the high-frequency component along the diagonal direction of the fused feature input map. WT Represents wavelet transform, F s This represents a fused feature input map that is larger than a preset scale; The wavelet subband is processed using residual blocks and inverse wavelet transform to obtain the input feature map of the global self-attention mechanism submodule: ; in, This represents the input feature map of the global self-attention mechanism submodule. IWT This represents the inverse wavelet transform. C Indicates splicing, This represents the low-frequency portion of the fused feature input map. F s+1 This represents the fused feature input map smaller than a preset scale. R Represents the residual block. This represents the high-frequency component in the horizontal direction of the fused feature input map. This represents the high-frequency component in the vertical direction of the fused feature input map. This represents the high-frequency component along the diagonal direction of the fused feature input map.

9. The method for identifying potential hazards at highway construction sites according to claim 1, characterized in that, The frequency modulation module specifically includes: an ultra-lightweight spatial attention unit, a global information transit-channel attention unit, and a 1×1 convolutional layer; The frequency modulation module is specifically used for: The high-frequency features in the output feature map of the third multi-scale grouped dilated convolution module are input into the ultra-lightweight spatial attention unit to generate a spatial attention map: ; in, A H-L Representing a spatial attention map, δ This represents the sigmoid activation function. This indicates the 6th 7×7 convolutional layer. GAP c This indicates channel-level global average pooling. X h Indicates high-frequency characteristics, GMP c represents channel-level max pooling, and [ , ] represents concatenation operation; Based on the spatial attention map, the first modulation feature is obtained through element-wise multiplication: ; in, Indicates the first modulation feature, X l ⊙ represents low-frequency characteristics; The low-frequency features in the output feature map of the fourth multi-scale grouped dilated convolution module are input into the global information channel attention unit to generate an attention description map: ; in, A L-H This represents an attention description diagram. This indicates the 8th 1×1 convolutional layer. γ Represents the ReLU activation function. This indicates the 7th 1×1 convolutional layer. GAP s This represents global average pooling along the spatial dimension. This indicates the 10th 1×1 convolutional layer. This indicates the 9th 1×1 convolutional layer. GMP s This represents max pooling along the spatial dimension; Based on the attention description map, the second modulation feature is obtained through the element-wise multiplication: ; in, Indicates the second modulation feature; The first modulation feature and the second modulation feature are added element by element to obtain the modulation feature; The frequency modulation module output image is obtained by fusing the modulation features and the feature image of the multi-scale grouped dilated convolution module through the cross-attention unit and the 1×1 convolutional layer.

10. A hazard identification system for highway construction sites, characterized in that, include: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the method for identifying hidden dangers at highway construction sites as described in any one of claims 1 to 9.