A severe weather traffic sign detection method and system based on multi-stage optimization and meteorological classification
By employing multi-level optimization and meteorological classification methods, the problems of poor module coordination and accuracy-efficiency imbalance in traffic sign detection under severe weather conditions are solved, achieving high-precision and high-efficiency traffic sign detection, which is suitable for autonomous driving environmental perception.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN INST OF TECH AT WEIHAI
- Filing Date
- 2026-04-23
- Publication Date
- 2026-07-24
AI Technical Summary
Existing traffic sign detection methods suffer from poor module coordination under adverse weather conditions, loss of key sign details or noise amplification due to the separate design of image enhancement and detection models, weak dynamic adaptability, and an imbalance between accuracy and efficiency.
A method for detecting traffic signs in severe weather based on multi-level optimization and meteorological classification is constructed. By systematically screening severe weather images and integrating refined data augmentation techniques, a MobileNetV3 model is used to enhance multi-scale feature representation. A hierarchical attention enhancement mechanism is designed, a dynamic feature enhancement network is constructed, and improved modules such as DSPPF, GDW-PSA, and EMA high-efficiency multi-scale attention modules are embedded to achieve adaptive image restoration and feature fusion.
Achieving high accuracy (mAP@0.5≥81.7%) and high efficiency (94.6 FPS) in traffic sign detection under adverse weather conditions overcomes the limitations of traditional methods and meets the real-time perception requirements of autonomous driving systems in complex weather environments.
Smart Images

Figure CN122454534A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental perception technology for autonomous driving, and more specifically, to a method and system for detecting traffic signs in severe weather based on multi-level optimization and meteorological classification. Background Technology
[0002] Traffic sign detection in adverse weather conditions is a core challenge in autonomous driving and intelligent transportation systems. This is because severe weather fundamentally disrupts the "visibility" and "discrimination" of images, directly challenging the basic assumptions of existing detection algorithms. Haze reduces image contrast through light scattering, causing red prohibition signs to fade into gray-white dots, resulting in the loss of key shape and color features and causing significant false negatives. Rain and snow introduce dynamic, semi-transparent stripe noise, partially obscuring signs and creating false textures, making it easy for algorithms to misidentify dense rain lines as structures, leading to numerous false positives. Furthermore, overexposure due to road surface reflections or rapid shadow transitions caused by cloud movement can completely obscure the sign pattern in overly bright or dark areas, further exacerbating detection instability.
[0003] The core challenge of this problem lies in the significant gap between the "laboratory" and the "real world." First, data acquisition is extremely difficult—collecting large-scale, pixel-level labeled datasets under real, severe weather conditions is costly and dangerous, while synthetic data often differs from real-world distributions, potentially leading to model overfitting. Second, model design must strike a difficult balance between accuracy and real-time performance: complex image preprocessing (such as dehazing and deraining) improves image quality but increases computational burden and reduces frames per second (FPS); while end-to-end model improvement is more efficient, it places extremely high demands on the diversity of training data and the robustness of the network structure. Therefore, effectively resisting multiple weather-related interferences while maintaining real-time detection speed has become a critical bottleneck that urgently needs to be overcome in this field.
[0004] Existing invention patent 1, a traffic sign detection method and system based on YOLOv5 (application number: CN202310831023.6), expands the dataset through online and offline hybrid data augmentation; designs a C2fGhost module to reconstruct the CSPDarkNet feature extraction network, retaining rich gradient flow information while being lightweight; uses shallow feature P2 to replace deep feature P5 to improve small target detection capability; introduces an Efficient-RepGFPN feature fusion network to enhance multi-scale detection capability; adopts Wise-IoU V3 loss function to replace CIoU loss function to optimize the training effect of low-quality data; and enhances inference performance through confidence scaling distillation. Existing invention patent 2, "A Road Traffic Sign Target Detection Method and System Based on YOLOv8" (application number: CN202411837047.3), addresses the challenge of detecting small road targets in occluded and unevenly lit scenarios. It constructs a CDFF-YOLO model, enhances feature fusion capabilities by embedding an MPA module, achieves multi-scale information fusion by constructing a lightweight MSF module integrating GSConv and CARAFE operators, and simultaneously enhances frequency and spatial domain features using a DFF module with two-dimensional Fourier transform, thereby improving the detection accuracy of small targets. However, both of these patents have inherent limitations under severe weather conditions such as heavy rain and dense fog, as their static feature extraction mechanisms do not model weather interference.
[0005] Existing invention patent 3, "A Method for Detecting Traffic Signs in Fog" (application number: CN202411192629.0), addresses the detection failure issues caused by image whitening distortion and blurred traffic signs in foggy weather. It constructs a lightweight network model, suppresses noise interference through an adaptive convolutional structure based on fog density, enhances sign contour features using a multi-scale edge enhancement module, and embeds real-time defogging residual units to jointly optimize image clarity, thereby improving the detection accuracy and robustness of traffic signs in foggy conditions. However, this patented traffic sign detection method only applies to foggy conditions and cannot handle adverse weather conditions such as rain or low light, limiting its applicability.
[0006] In summary, existing traffic sign detection methods suffer from three major technical bottlenecks under severe weather conditions: poor module coordination, separation of image enhancement and detection models leading to loss of key sign details or noise amplification, weak dynamic adaptability (static feature fusion mechanisms cannot respond to mixed weather changes, such as MSDA-YOLO performance fluctuations of ±15%), and an imbalance between accuracy and efficiency (high-precision models have an inference speed of less than 3 FPS, and lightweight models experience an accuracy drop of over 20% under extreme weather conditions). Summary of the Invention
[0007] The technical problem to be solved by this invention is:
[0008] To address the problems of poor module coordination, separation of image enhancement and detection model design leading to loss of key sign details or noise amplification, weak dynamic adaptability, and accuracy-efficiency imbalance in existing traffic sign detection methods under adverse weather conditions.
[0009] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:
[0010] This invention provides a method for detecting traffic signs in severe weather based on multi-level optimization and meteorological classification, comprising the following steps:
[0011] S100: Based on a general traffic sign dataset, a raw dataset with weather scene segmentation is constructed by systematically filtering its original severe weather images and integrating refined data augmentation techniques.
[0012] S200. Perform meteorological classification on the original dataset from step S100. Based on the MobileNetV3 model, design a multi-scale feature representation enhancement and propose a hierarchical attention enhancement mechanism. This includes deploying HDPA hybrid dual-path attention in shallow features to enhance the perception of local texture details, using HDPA hybrid dual-path attention in mid-level features to enhance cross-channel interaction to suppress weather noise, and retaining SE standard compression excitation in deep features to maintain the robustness of global semantic representation.
[0013] S300: Construct a dynamic feature enhancement network. Based on the multi-scale features extracted by the weather classification module, achieve adaptive image restoration for different scenarios through hierarchical feature reconstruction and gating fusion mechanism, including multi-scale feature adaptation, weather perception processing branch and dynamic gating fusion strategy.
[0014] S400. Based on the YOLOv11 backbone network, a triple improvement module is proposed and embedded between the end of the backbone network and the detection head. The triple improvement module includes the DSPPF dynamic deformable pooling module, the GDW-PSA grouped dynamic weighted feature aggregation module, and the EMA efficient multi-scale spatial attention mechanism. With the collaborative work of the triple improvement module, multi-scale feature enhancement, adaptive receptive field adjustment, context information fusion and key region focusing are used to improve the feature representation ability and detection robustness of the model under complex degradation conditions.
[0015] Further, in step S100, the traffic sign dataset CCTSDB-2021 undergoes multi-level cleaning, including:
[0016] Image histogram analysis was used to select an initial pool of severe weather samples; physical-guided degradation modeling was implemented using the imgaug enhancement library: a directional rain streak layer was added to rain samples, randomly rotating snowflake particles were superimposed on snow samples, and nighttime imaging degradation was simulated for low-light samples through gamma correction and random channel noise.
[0017] To evaluate the robustness of the model under complex weather combinations, an additional 3,000 mixed weather test sets were constructed. The performance degradation patterns of the model under specific weather conditions were analyzed using the controlled variable method on the single weather test sets.
[0018] Further, in step S200, the following are included:
[0019] S210. For the channel attention path, the input feature map Perform global average pooling, where H, W, and C are the height, width, and number of channels of the input feature map, respectively, to generate channel description vectors. ReLU activation learns the nonlinear relationship between channels, through Convolution followed by Hard sigmoid activation Output the channel weight vector :
[0020]
[0021] Wherein, GAP represents global average pooling. It is the ReLU activation function. For Hard sigmoid function; for Convolution; F is the convolution kernel;
[0022] S220. In the spatial attention path, spatial features are extracted through a lightweight convolutional sequence. The first convolutional layer extracts local texture features, and the second convolutional layer establishes spatial context associations.
[0023]
[0024] in, This indicates two concatenated depthwise separable convolutions; Spatial features extracted after two lightweight convolutional sequences;
[0025] S230, Feature fusion strategy, to achieve dual-path attention calibration of channel attention results and spatial attention results:
[0026]
[0027] in, It represents the Hadamah accumulation. For spatial dimension broadcasting operations, This is the feature map output after dual-path attention calibration.
[0028] Further, in step S300, the following are included:
[0029] S310, Multi-scale Features Adaptation and multi-level features have differentiated representation capabilities, employing different levels of features to process different harsh environments; among them, shallow high-resolution features Contains high-frequency detail information, suitable for localizing local noise; mid-level features During receptive field expansion, local texture and global contextual information are fused to provide robust characterization of contrast attenuation under mixed interference of haze and snow; deep low-resolution features High-dimensional semantic cues related to atmospheric transmittance and light intensity are screened through channel attention mechanism to support parameter estimation of physical degradation models;
[0030] S320, weather perception and processing branch, including RRM rain removal module, SRM snow removal module, DHM defogging module and LEM low light enhancement module;
[0031] S330, dynamic gating fusion strategy, based on the adverse weather classification probability in the classification module. Normalized branch weights are generated using Gumbel-Softmax. :
[0032]
[0033] in, For the noise perturbation term that follows a Gumbel distribution, Temperature coefficient; These are the classification probabilities of the i-th and j-th types of adverse weather, respectively;
[0034] Output characteristics of rain, snow, fog, and low light enhancement branches Perform feature alignment and weighted aggregation:
[0035]
[0036]
[0037] in, The output results of each module, This is the intermediate feature map after the Align feature alignment operation. This represents the final weighted aggregated output enhanced feature map.
[0038] Further, in step S320, the following is included:
[0039] S321. In the RRM rain removal module, for deep features... Discrete cosine transform (DCT) is performed to separate high-frequency noise from low-frequency content, and the high-frequency rain line region is located to suppress the rain line. Finally, spatial features are reconstructed using inverse discrete cosine transform (IDCT), and low-frequency components are superimposed to avoid blurring of traffic sign outlines.
[0040]
[0041]
[0042]
[0043] in, For discrete cosine transformation, For depthwise separable convolution, Hard sigmoid function; AvgPool is average pooling; This refers to the frequency domain features obtained after performing a Discrete Cosine Transform (DCT) on the deep features. This represents a rain line probability mask generated from high-frequency components using depthwise separable convolution and a Hard Sigmoid function. This represents the rain removal enhancement feature output after reconstruction by IDCT inverse discrete cosine transform and superposition of low-frequency components;
[0044] S322. In the SRM snow removal module, multi-scale snow suppression is used to target mid-layer features. Instance normalization and convolution operations generate a snow cover probability map. Through snow cover probability map Guided feature reconstruction preserves the edge structure of traffic signs and employs dilated convolution to expand the receptive field and suppress snow noise.
[0045]
[0046]
[0047] in, For instance normalization, For Hadama accumulation, The snow removal enhancement feature output after suppressing snow noise;
[0048] S323. In the DHM dehazing module, the transmittance is predicted based on deep features, and the brightest region is located through a spatial attention mechanism. Image dehazing is then achieved by combining a physical scattering model.
[0049]
[0050]
[0051]
[0052] in, This is a transmittance map; Maxpool is the maximum pooling method. This represents the global atmospheric light value. For the input features of the foggy image, Features of the clear image output after dehazing;
[0053] S324. In the LEM low-light enhancement module, based on feature scale Predict channel gain and estimate noise components using a light quantum network to simultaneously boost brightness and suppress noise:
[0054]
[0055]
[0056]
[0057] in, For channel gain, For noise components, Low-light enhancement features are used to improve brightness and suppress noise in the output.
[0058] Further, in step S400, the following is included:
[0059] S410. The DSPPF dynamic deformable pooling module includes three parallel branches: Fixed Grid Pooling, Deformable Pooling, and Global Context. The Fixed Grid Pooling branch uses 5×5 max pooling to extract local features at a fixed scale. The Deformable Pooling branch adjusts the sampling position of the 3×3 pooling kernel based on dynamic offsets to enhance adaptability to irregular targets. The Global Context branch captures scene-level semantic information through global average pooling. The three branch outputs are concatenated and compressed to generate an efficient representation that integrates multi-scale features.
[0060] Deformable pooling dynamically adjusts the receptive field by predicting the sampling grid offset:
[0061]
[0062] in, The learnable weight parameters for the first 1×1 convolutional layer, with dimension 1. , The learnable weight parameters for the second 1×1 convolutional layer have a dimension of . , For pooling core size, This represents the dynamic offset of the deformable pooling sampling grid. Input feature map of deformable pooling branch
[0063] The three branch outputs are spliced along the channel and then compressed using a 1×1 convolution:
[0064]
[0065] in, For weighting; The output is the fused multi-scale feature. D is an abbreviation for Deformable, which is a 3×3 deformable max pooling with dynamic offset.
[0066] S420, in the GDW-PSA grouped dynamic weighted feature aggregation module, a three-stage design of grouped feature decoupling, channel recalibration, and dynamic weight fusion is used to achieve multi-scale feature extraction and adaptive information aggregation while maintaining lightweight computational efficiency; including,
[0067] The input features are divided into multiple subgroups along the channel dimension, and each subgroup is independently subjected to depthwise separable convolution and channel attention recalibration:
[0068]
[0069]
[0070] in, For learnable weights, Representing the Intermediate feature maps extracted from each subgroup after passing through DWConv depthwise separable convolution;
[0071] A lightweight gating network is introduced to dynamically predict the fusion weights of each subgroup, and the weight distribution is generated through global average pooling and fully connected layers:
[0072]
[0073]
[0074] in, , To compress dimensions, Here, represents the gated network parameters, and w represents the dynamic fusion weight distribution of each feature subgroup. This represents a one-dimensional vector obtained by flattening the input features after global average pooling, where g represents group / subgroup fusion gating;
[0075] The recalibrated subgroup features are then weighted and concatenated according to dynamic weights:
[0076]
[0077] S430. An EMA high-efficiency multi-scale attention module is embedded in front of the YOLOv11 detector head. The EMA high-efficiency multi-scale attention module performs global channel and local space modeling respectively, thereby achieving cross-space fusion.
[0078] A severe weather traffic sign detection system based on multi-level optimization and meteorological classification is provided. The system has program modules corresponding to the above steps and executes the steps in the above-mentioned severe weather traffic sign detection method based on multi-level optimization and meteorological classification when running.
[0079] A computer-readable storage medium storing a computer program configured to, when invoked by a processor, implement steps of a severe weather traffic sign detection method based on multi-level optimization and meteorological classification.
[0080] Compared with the prior art, the beneficial effects of the present invention are:
[0081] This invention proposes an AWEN-YOLO three-level optimization architecture based on meteorological perception. Through a meteorological classification-guided adaptive enhancement and dynamic feature fusion mechanism, it achieves high-precision and high-efficiency traffic sign detection under adverse weather conditions, overcoming the limitations of traditional methods in monitoring adverse weather. Experiments have demonstrated that this invention, through the meteorological classification-guided adaptive enhancement and dynamic feature fusion mechanism, achieves high-precision (mAP@0.5≥81.7%) and high-efficiency (94.6 FPS) traffic sign detection under adverse weather conditions, meeting the real-time perception requirements of autonomous driving systems in complex meteorological environments.
[0082] This invention proposes a three-level optimization architecture, AWEN YOLO, based on meteorological perception. Through a three-level collaborative approach of weather classification, dynamic enhancement, and detection optimization, it overcomes the problems of low detection accuracy and poor adaptability under severe weather conditions. The meteorological classification module employing hybrid dual-path attention (HDPA) effectively suppresses rain and snow noise, improving mAP@0.5 by 1.2%. The dynamic gating feature enhancement network constructs four restoration branches for rain, snow, fog, and low light conditions and adaptively fuses them, further improving mAP@0.5 by 2.8% under mixed weather conditions. The triple improvement modules (DSPPF+GDW PSA+EMA) respectively enhance the adaptability to irregular deformation, grouped dynamic feature aggregation, and small target spatial localization; introducing each module individually can improve mAP@0.5 by 1.2%~3.4%. The overall model achieved an mAP@0.5 of 81.7% on the severe weather test set, a 5.8% improvement over the benchmark YOLOv11s, with an inference speed of 94.6 FPS and only 11.1M parameters, achieving an excellent balance between accuracy and efficiency and meeting the real-time perception requirements of autonomous driving systems for complex weather environments. Attached Figure Description
[0083] Figure 1 These are simulation diagrams of different weather conditions in embodiments of the present invention;
[0084] Figure 2 This is a network structure diagram of the HDPA module in an embodiment of the present invention;
[0085] Figure 3 This is a structural diagram of the DSPPF module in an embodiment of the present invention;
[0086] Figure 4 This is a diagram of the GDW-PSA network structure in an embodiment of the present invention;
[0087] Figure 5 The figures shown are experimental comparison results of the present invention and existing marker detection models in the embodiments of the present invention. Among them, (a) is a comparison figure of the mAP@0.5 evaluation index, (b) is a comparison figure of the mAP@0.5:0.95 evaluation index, (c) is a comparison figure of the Precision evaluation index, and (d) is a comparison figure of the Recall evaluation index.
[0088] Figure 6 This is a flowchart of a severe weather traffic sign detection method based on multi-level optimization and meteorological classification in an embodiment of the present invention;
[0089] Figure 7 This is a diagram of the improved YOLO11 framework in an embodiment of the present invention. Detailed Implementation
[0090] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0091] Specific Implementation Plan 1: Combining Figures 1 to 4 , Figure 6 and Figure 7 As shown, this invention provides a method for detecting traffic signs in severe weather based on multi-level optimization and meteorological classification, including a data processing module, a meteorological classification module, a dynamic feature enhancement module, and an improved YOLO detection network module.
[0092] S100. In the data processing module, in response to the severe lack of severe weather samples in the existing traffic sign dataset, a new dataset with strict weather scene division is constructed based on the traffic sign dataset CCTSDB-2021 by systematically screening its original severe weather images and integrating refined data augmentation techniques.
[0093] During the data construction phase, the original dataset underwent multi-level cleaning: image histogram analysis (HSV spatial brightness variance ≤30 indicates low light) was used to select an initial severe weather sample pool; subsequently, the imgaug enhancement library was used to implement physically guided degradation modeling—a directionally controllable rain streak layer (parameters: raindrop length 10-25 pixels, tilt angle ±30°, density 0.1-0.3) was added to rainy day samples, and randomly rotated snowflake particles (diameter 5-15 pixels, transparency 0.4-0.8) were superimposed on snowy day samples; low light samples were simulated for nighttime imaging degradation through gamma correction (γ=1.5-2.5) and random channel noise (σ=0.05-0.1).
[0094] To evaluate the model's robustness under complex weather combinations, an additional 3,000 mixed weather test images (each image containing ≥2 types of weather interference) were constructed. The performance degradation patterns of the model under specific weather conditions were analyzed using the controlled variable method on the single-weather test images. The results are as follows: Figure 1 As shown;
[0095] S200. In the meteorological classification module, to address the feature degradation caused by interference from rain, fog, and snow, this invention, based on MobileNetV3, implements a multi-scale feature representation enhancement design and proposes a hierarchical attention enhancement mechanism: Hybrid Dual-Path Attention (HDPA) is deployed in shallow features to enhance local texture detail perception; HDPA is used in mid-layer features to enhance cross-channel interaction and suppress weather noise; and the standard compression excitation (SE) module is retained in deep features to maintain the robustness of global semantic representation, thereby improving feature discrimination under complex meteorological conditions. The attention focus direction is adaptively adjusted to comprehensively enhance feature discrimination and classification capabilities under complex and variable meteorological conditions. HPDA, for example... Figure 2 As shown, the specific process is as follows:
[0096] S210, Channel Attention Path:
[0097] For the input feature map Perform global average pooling, where H, W, and C are the height, width, and number of channels of the input feature map, respectively, to generate channel description vectors. ReLU activation learns the nonlinear relationship between channels, through Convolution followed by Hard sigmoid activation Output the channel weight vector :
[0098]
[0099] Wherein, GAP represents global average pooling. It is the ReLU activation function. For Hard sigmoid function; for Convolution; F is the convolution kernel;
[0100] S220, Spatial Attention Path:
[0101] Spatial features are extracted using lightweight convolutional sequences. The first convolutional layer extracts local texture features, and the second convolutional layer establishes spatial context relationships.
[0102]
[0103] in, This indicates two concatenated depthwise separable convolutions; Spatial features extracted after two lightweight convolutional sequences;
[0104] S230, Feature Fusion Strategy:
[0105] Dual-path attention calibration is performed on both channel attention and spatial attention results:
[0106]
[0107] In the formula, It represents the Hadamah accumulation. Broadcasting operations for spatial dimensions; This represents the feature map output after dual-path attention calibration.
[0108] S300. In the feature enhancement module, this invention proposes a dynamic feature enhancement network, which is based on multi-scale features extracted by the weather classification module. This module achieves adaptive image restoration for different scenarios through hierarchical feature reconstruction and gating fusion mechanisms. It mainly realizes image enhancement under severe weather conditions through the following three-stage processing:
[0109] S310, multi-scale feature adaptation
[0110] Multi-level features possess differentiated representation capabilities, employing different levels of features to process different harsh environments; among them, shallow high-resolution features... Contains high-frequency detail information, suitable for locating localized noise such as rain lines and snowflakes; mid-level features During receptive field expansion, local texture and global contextual information are fused to provide robust characterization of contrast attenuation under mixed interference of haze and snow; deep low-resolution features High-dimensional semantic cues related to atmospheric transmittance and light intensity can be screened through channel attention mechanisms to support parameter estimation of physical degradation models;
[0111] S320, Weather Sensing and Processing Branch
[0112] S321. Rain Removal Module (RRM): For deep features Discrete cosine transform is performed to separate high-frequency noise from low-frequency content, and the high-frequency rain line region is located to suppress the rain line. Finally, spatial features are reconstructed through inverse discrete cosine transform (IDCT), and low-frequency components are superimposed to avoid blurring of the traffic sign outline. The specific formula is as follows:
[0113]
[0114]
[0115]
[0116] in, For discrete cosine transformation, For depthwise separable convolution, Hard sigmoid function; AvgPool is average pooling; This represents the frequency domain features obtained after performing a Discrete Cosine Transform (DCT) on the deep features. This represents a rain line probability mask generated from high-frequency components using depthwise separable convolution and a Hard Sigmoid function. This represents the rain removal enhancement feature of the output after reconstruction by inverse discrete cosine transform (IDCT) and superposition of low-frequency components;
[0117] S322, Snow Removal Module (SRM): This module suppresses snow accumulation at multiple scales, affecting mid-level features. Instance normalization and convolution operations generate a snow cover probability map. Through snow cover probability map Guided feature reconstruction preserves the edge structure of traffic signs and employs dilated convolution to expand the receptive field and suppress snow noise.
[0118]
[0119]
[0120] in, For instance normalization, For Hadamah accumulation; The snow removal enhancement feature output after suppressing snow noise;
[0121] S323, Dehazing Module (DHM): This module predicts transmittance based on deep features and locates the brightest region through a spatial attention mechanism, thus avoiding estimation bias caused by global averaging. Finally, it combines a physical scattering model to achieve image dehazing.
[0122]
[0123]
[0124]
[0125] in, This is a transmittance map; Maxpool is the maximum pooling method. It is the global atmospheric light value. This represents the features of the input foggy image. This indicates the features of the clear image output after dehazing;
[0126] S324, Low-light Enhancement Module (LEM): Based on feature scale Predict channel gain and estimate noise components using a light quantum network to simultaneously boost brightness and suppress noise:
[0127]
[0128]
[0129]
[0130] in, For channel gain, This is a noise component; To enhance low-light output characteristics while suppressing noise; S330, dynamic gating fusion strategy,
[0131] Based on the probability of adverse weather classification in the classification module (Rain, snow, fog, low light), normalized branch weights are generated using Gumbel-Softmax. :
[0132]
[0133] in, For the noise perturbation term that follows a Gumbel distribution, Temperature coefficient; These are the classification probabilities of the i-th and j-th types of adverse weather, respectively;
[0134] Output characteristics of rain, snow, fog, and low light enhancement branches Perform feature alignment and weighted aggregation:
[0135]
[0136]
[0137] in, The output of each module; representing the feature map output after dual-path attention calibration; This represents the intermediate feature map after the feature alignment operation. This represents the final weighted aggregation output enhanced feature map;
[0138] S400. In the YOLO detection module, this invention proposes a triple core improvement module based on the efficient and real-time YOLOv11 backbone network, and embeds it between the end of the backbone network and the detection head to build a more powerful feature pyramid: Dynamic Deformable Pooling Module (DSPPF), Grouped Dynamic Weighted Feature Aggregation Module (GDW-PSA), and Efficient Multi-Scale Spatial Attention Mechanism (EMA). These modules work together to significantly improve the model's feature representation ability and detection robustness under complex degradation conditions through multi-scale feature enhancement, adaptive receptive field adjustment, contextual information fusion, and key region focusing.
[0139] S410, Dynamically Deformable Pooling Module (DSPPF).
[0140] This invention designs a dynamically deformable pooling module (Deformable SPPF, DSPPF); this module includes three parallel branches: fixed mesh pooling, deformable pooling, and global context, combined with... Figure 3 As shown, the FixedGrid Pooling branch uses 5×5 max pooling to extract local features at a fixed scale; the Deformable Pooling branch adjusts the sampling position of the 3×3 pooling kernel based on dynamic offsets to enhance adaptability to irregular targets; the Global Context branch captures scene-level semantic information through Global Average Pooling (GAP); the three-branch outputs are concatenated and compressed to generate an efficient representation that integrates multi-scale features.
[0141] Deformable pooling branches dynamically adjust the receptive field by predicting the sampling grid offset:
[0142]
[0143] in, The learnable weight parameters for the first 1×1 convolutional layer, with dimension 1. , The learnable weight parameters for the second 1×1 convolutional layer have a dimension of . For pooling core size, This represents the dynamic offset of the deformable pooling sampling grid. The input feature map of the deformable pooling branch is the original input feature of the entire DSPPF dynamic deformable pooling module.
[0144] The three branch outputs are spliced along the channel and then compressed using a 1×1 convolution:
[0145]
[0146] in, For weighting; The output is the fused multi-scale feature; D is an abbreviation for Deformable, which is 3×3 deformable max pooling with dynamic offset;
[0147] S420, Grouped Dynamic Weighted Feature Aggregation Module (GDW-PSA).
[0148] Severe weather conditions cause contrast degradation and detail blurring in traffic sign images. Traditional detection methods suffer from poor feature extraction robustness and loss of multi-scale information, leading to a high false negative rate. Existing Pyramid Squeeze Attention (PSA) modules are limited by the redundancy of static weight fusion mechanisms and global channel interactions, making it difficult to effectively cope with complex and ever-changing degradation patterns. To address this, this invention proposes a Grouping dynamically weighted PyramidSqueeze Attention (GDW-PSA) module. This module, through a three-stage design of grouped feature decoupling, channel recalibration, and dynamic weight fusion, can achieve multi-scale feature extraction and adaptive information aggregation while maintaining lightweight computational efficiency.
[0149] This module first divides the input features into multiple subgroups along the channel dimension, and each subgroup is independently subjected to depthwise separable convolution and channel attention recalibration:
[0150]
[0151]
[0152] in, These are learnable weights; Representing the Intermediate feature maps extracted from each subgroup after passing through depthwise separable convolution (DWConv);
[0153] This section innovatively introduces a lightweight gating network to dynamically predict the fusion weights of each subgroup, generating the weight distribution through global average pooling and fully connected layers:
[0154]
[0155]
[0156] in, , To compress dimensions, Here, represents the gated network parameters, and w represents the dynamic fusion weight distribution of each feature subgroup. This represents a one-dimensional vector obtained by flattening the input features after global average pooling (GAP), where g represents group / subgroup fusion gating.
[0157] Finally, the recalibrated subgroup features are weighted and concatenated according to dynamic weights:
[0158]
[0159] S430, EMA attention mechanism module
[0160] To address the issue of traffic sign feature attenuation under adverse conditions, this invention embeds a high-efficiency multi-scale attention module (EMA) in front of the YOLOv11 detection head. The EMA module performs global channel and local spatial modeling respectively, thereby achieving cross-space fusion. This process retains cross-scale nonlinear correlation while reducing computational load through parameter sharing mechanism. The EMA module improves feature signal-to-noise ratio by suppressing rain and fog noise, and its multi-scale characteristics are particularly suitable for small target detection scenarios.
[0161] Specific Implementation Scheme 2: The present invention provides a severe weather traffic sign detection system based on multi-level optimization and meteorological classification. This system has program modules corresponding to the above steps, and executes the steps in the above-mentioned severe weather traffic sign detection method based on multi-level optimization and meteorological classification when running.
[0162] The other combinations and connections in this implementation scheme are the same as in Specific Implementation Scheme 1.
[0163] Specific Implementation Scheme 3: The present invention provides a computer-readable storage medium storing a computer program configured to implement, when called by a processor, the steps of a severe weather traffic sign detection method based on multi-level optimization and meteorological classification.
[0164] The other combinations and connections in this implementation scheme are the same as in Specific Implementation Scheme 1.
[0165] Example
[0166] To explore the advantages of this invention compared with currently popular traffic sign detection models, we compared it with currently popular traffic sign detection models.
[0167] Evaluation metrics: Precision, recall, mAP@0.5 (mean precision at an intersection-union ratio threshold of 0.5), mAP@0.5:0.95 (mean precision at intersection-union ratio thresholds of 0.5 to 0.95), parameters, FPS (frames per second), and FLOPs(G) (number of floating-point operations).
[0168] Table 1 Comparison with other classic algorithms
[0169]
[0170] Experimental results: Combining Figure 5 As shown, the AWEN-YOLO model designed in this invention outperforms all comparative models with mAP@0.5 of 81.7% and mAP@0.95 of 47.4%, respectively, and improves upon the closest performing YOLOv11s by 5.8% and 2.9%, respectively. Its inference speed reaches 94.6 FPS, demonstrating a significant advantage in the balance between accuracy and efficiency. Thanks to the collaborative design of the weather classification module and Dynamic Deformable Pooling (DSPPF), the model achieves significantly improved accuracy compared to YOLOv8s in complex scenarios such as rain and fog, and is 2.25 times faster than the traditional YOLOv4's 42 FPS. Compared to traditional models, the AWEN-YOLO of this invention can complete traffic sign detection in adverse weather conditions with better accuracy and efficiency.
[0171] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A method for detecting traffic signs in severe weather based on multi-level optimization and meteorological classification, characterized in that, Includes the following steps: S100: Based on a general traffic sign dataset, a raw dataset with weather scene segmentation is constructed by systematically filtering its original severe weather images and integrating refined data augmentation techniques. S200. Perform meteorological classification on the original dataset from step S100. Based on the MobileNetV3 model, design a multi-scale feature representation enhancement and propose a hierarchical attention enhancement mechanism. This includes deploying HDPA hybrid dual-path attention in shallow features to enhance the perception of local texture details, using HDPA hybrid dual-path attention in mid-level features to enhance cross-channel interaction to suppress weather noise, and retaining SE standard compression excitation in deep features to maintain the robustness of global semantic representation. S300: Construct a dynamic feature enhancement network. Based on the multi-scale features extracted by the weather classification module, achieve adaptive image restoration for different scenarios through hierarchical feature reconstruction and gating fusion mechanism, including multi-scale feature adaptation, weather perception processing branch and dynamic gating fusion strategy. S400. Based on the YOLOv11 backbone network, a triple improvement module is proposed and embedded between the end of the backbone network and the detection head. The triple improvement module includes the DSPPF dynamic deformable pooling module, the GDW-PSA grouped dynamic weighted feature aggregation module, and the EMA efficient multi-scale spatial attention mechanism. Through the collaborative work of the three improvement modules, multi-scale feature enhancement, adaptive receptive field adjustment, contextual information fusion, and key region focusing are used to improve the model's feature representation ability and detection robustness under complex degradation conditions.
2. The method for detecting traffic signs in severe weather based on multi-level optimization and meteorological classification according to claim 1, characterized in that: In step S100, the traffic sign dataset CCTSDB-2021 undergoes multi-level cleaning, including: Image histogram analysis was used to select an initial pool of severe weather samples; physical-guided degradation modeling was implemented using the imgaug enhancement library: a directional rain streak layer was added to rain samples, randomly rotating snowflake particles were superimposed on snow samples, and nighttime imaging degradation was simulated for low-light samples through gamma correction and random channel noise. To evaluate the robustness of the model under complex weather combinations, an additional 3,000 mixed weather test sets were constructed. The performance degradation patterns of the model under specific weather conditions were analyzed using the controlled variable method on the single weather test sets.
3. The method for detecting traffic signs in severe weather based on multi-level optimization and meteorological classification according to claim 2, characterized in that: Step S200 includes, S210. For the channel attention path, the input feature map Perform global average pooling, where H, W, and C are the height, width, and number of channels of the input feature map, respectively, to generate channel description vectors. ReLU activation learns the nonlinear relationship between channels, through Convolution followed by Hard sigmoid activation Output the channel weight vector : Wherein, GAP represents global average pooling. It is the ReLU activation function. For Hard sigmoid function; for Convolution; F is the convolution kernel; S220. In the spatial attention path, spatial features are extracted through a lightweight convolutional sequence. The first convolutional layer extracts local texture features, and the second convolutional layer establishes spatial context associations. in, This indicates two concatenated depthwise separable convolutions; Spatial features extracted after two lightweight convolutional sequences; S230, Feature fusion strategy, to achieve dual-path attention calibration of channel attention results and spatial attention results: in, It represents the Hadamah accumulation. For spatial dimension broadcasting operations, This is the feature map output after dual-path attention calibration.
4. The method for detecting traffic signs in severe weather based on multi-level optimization and meteorological classification according to claim 3, characterized in that: Step S300 includes, S310, Multi-scale Features Adaptation and multi-level features have differentiated representation capabilities, employing different levels of features to process different harsh environments; among them, shallow high-resolution features Contains high-frequency detail information, suitable for localizing local noise; mid-level features During receptive field expansion, local texture and global contextual information are fused to provide robust characterization of contrast attenuation under mixed interference of haze and snow; deep low-resolution features High-dimensional semantic cues related to atmospheric transmittance and light intensity are screened through channel attention mechanism to support parameter estimation of physical degradation models; S320, weather perception and processing branch, including RRM rain removal module, SRM snow removal module, DHM defogging module and LEM low light enhancement module; S330, dynamic gating fusion strategy, based on the adverse weather classification probability in the classification module. Normalized branch weights are generated using Gumbel-Softmax. : in, For the noise perturbation term that follows a Gumbel distribution, Temperature coefficient; These are the classification probabilities of the i-th and j-th types of adverse weather, respectively; Output characteristics of rain, snow, fog, and low light enhancement branches Perform feature alignment and weighted aggregation: in, The output results of each module, This is the intermediate feature map after the Align feature alignment operation. This represents the final weighted aggregated output enhanced feature map.
5. The method for detecting traffic signs in severe weather based on multi-level optimization and meteorological classification according to claim 4, characterized in that: Step S320 includes, S321. In the RRM rain removal module, for deep features... Discrete cosine transform (DCT) is performed to separate high-frequency noise from low-frequency content, and the high-frequency rain line region is located to suppress the rain line. Finally, spatial features are reconstructed using inverse discrete cosine transform (IDCT), and low-frequency components are superimposed to avoid blurring of traffic sign outlines. in, For discrete cosine transformation, For depthwise separable convolution, Hard sigmoid function; AvgPool is average pooling; This refers to the frequency domain features obtained after performing a Discrete Cosine Transform (DCT) on the deep features. This represents a rain line probability mask generated from high-frequency components using depthwise separable convolution and a Hard Sigmoid function. This represents the rain removal enhancement feature output after reconstruction by IDCT inverse discrete cosine transform and superposition of low-frequency components; S322. In the SRM snow removal module, multi-scale snow suppression is used to target mid-layer features. Instance normalization and convolution operations generate a snow cover probability map. Through snow cover probability map Guided feature reconstruction preserves the edge structure of traffic signs and employs dilated convolution to expand the receptive field and suppress snow noise. in, For instance normalization, For Hadama accumulation, The snow removal enhancement feature output after suppressing snow noise; S323. In the DHM dehazing module, the transmittance is predicted based on deep features, and the brightest region is located through a spatial attention mechanism. Image dehazing is then achieved by combining a physical scattering model. in, This is a transmittance map; Maxpool is the maximum pooling method. This represents the global atmospheric light value. Given the features of the foggy image as input. Features of the clear image output after dehazing; S324. In the LEM low-light enhancement module, based on feature scale Predict channel gain and estimate noise components using a light quantum network to simultaneously boost brightness and suppress noise: in, For channel gain, For noise components, Low-light enhancement features are used to improve brightness and suppress noise in the output.
6. The method for detecting traffic signs in severe weather based on multi-level optimization and meteorological classification according to claim 5, characterized in that: Step S400 includes, S410. The DSPPF dynamic deformable pooling module includes three parallel branches: Fixed Grid Pooling, Deformable Pooling, and Global Context. The Fixed Grid Pooling branch uses 5×5 max pooling to extract local features at a fixed scale. The Deformable Pooling branch adjusts the sampling position of the 3×3 pooling kernel based on dynamic offsets to enhance adaptability to irregular targets. The Global Context branch captures scene-level semantic information through global average pooling. The three branch outputs are concatenated and compressed to generate an efficient representation that integrates multi-scale features. Deformable pooling dynamically adjusts the receptive field by predicting the sampling grid offset: in, The learnable weight parameters for the first 1×1 convolutional layer, with dimension 1. , The learnable weight parameters for the second 1×1 convolutional layer have a dimension of . , For pooling core size, This represents the dynamic offset of the deformable pooling sampling grid. Input feature map of deformable pooling branch; The three branch outputs are spliced along the channel and then compressed using a 1×1 convolution: in, For fusion weights; The output is the fused multi-scale feature. D is an abbreviation for Deformable, which is a 3×3 deformable max pooling with dynamic offset. S420, in the GDW-PSA grouped dynamic weighted feature aggregation module, a three-stage design of grouped feature decoupling, channel recalibration, and dynamic weight fusion is used to achieve multi-scale feature extraction and adaptive information aggregation while maintaining lightweight computational efficiency; including, The input features are divided into multiple subgroups along the channel dimension, and each subgroup is independently subjected to depthwise separable convolution and channel attention recalibration: in, For learnable weights, Representing the Intermediate feature maps extracted from each subgroup after passing through DWConv depthwise separable convolution; A lightweight gating network is introduced to dynamically predict the fusion weights of each subgroup, and the weight distribution is generated through global average pooling and fully connected layers: in, , To compress dimensions, Here, represents the gated network parameters, and w represents the dynamic fusion weight distribution of each feature subgroup. This represents a one-dimensional vector obtained by flattening the input features after global average pooling, where g represents group / subgroup fusion gating; The recalibrated subgroup features are then weighted and concatenated according to dynamic weights: S430. An EMA high-efficiency multi-scale attention module is embedded in front of the YOLOv11 detector head. The EMA high-efficiency multi-scale attention module performs global channel and local space modeling respectively, thereby achieving cross-space fusion.
7. A severe weather traffic sign detection system based on multi-level optimization and meteorological classification, characterized in that: The system has a program module corresponding to the steps of any one of the claims 1-6 above, and executes the steps in the above-described method for detecting traffic signs in severe weather based on multi-level optimization and meteorological classification when it is run.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program configured to, when invoked by a processor, implement the steps of any one of claims 1-6: a method for detecting severe weather traffic signs based on multi-level optimization and meteorological classification.
Citation Information
Patent Citations
Traffic sign detection method and system based on YOLOv5
CN116977976A
A traffic sign detection method under foggy weather
CN119181076B
Road traffic sign target detection method based on YOLOv8
CN119785312A