Traffic sign defogging detection method based on improved MSR-YOLO
Through the improved MSR-YOLO model, the multi-scale adaptive network and feature enhancement technology are used to solve the problem of traffic sign recognition in foggy environments, and the accuracy and effect of traffic sign detection in foggy days are improved.
Patent Information
- Application Number
- CN202510447145.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-08
AI Technical Summary
In foggy environments, existing image processing algorithms are difficult to accurately identify traffic signs, resulting in increased driving risks.
The improved MSR-YOLO model is adopted to perform defog pre-treatment through the multi-scale adaptive network module MS, combined with dark channel prior defog filter, Gamma filter and sharpening filter for image enhancement, and the C3K2-RFAConv module is introduced into the YOLOv11 network to improve feature extraction capabilities and add a small object detection layer to improve detection accuracy.
Effectively reduce the impact of foggy environment on traffic sign recognition, and improve the accuracy and detection effect of foggy traffic sign detection.
Smart Images

Figure CN120278916A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and particularly to a traffic sign dehazing detection method based on improved MSR-YOLO. Background Art
[0002] Efficient and accurate identification of traffic signs is crucial for the safety and reliability of active driving assistance and driverless vehicles. However, accurately detecting traffic signs in extreme situations remains challenging.
[0003] In foggy weather, particulate matter and water vapor in the atmosphere cause a sharp drop in visibility, shortening the recognition distance of traffic signs. Drivers need to be closer to accurately identify the signs. This increases driving risks because drivers may only notice the signs when they are close, leaving them unable to react in time.
[0004] Existing image processing algorithms can accurately identify traffic signs in most cases. However, in foggy weather, these algorithms may be affected by factors such as lighting conditions and sign clarity, resulting in poor recognition effects. Summary of the Invention
[0005] Object of the Invention: Aiming at the problems existing in the prior art, the present invention provides a traffic sign dehazing detection method based on improved MSR-YOLO, which can effectively reduce the impact of foggy weather on traffic sign recognition, thereby improving the accuracy of traffic sign detection in foggy weather.
[0006] Technical Solution: The present invention provides a traffic sign dehazing detection method based on improved MSR-YOLO, which is characterized by including the following steps:
[0007] S001: Use on-vehicle monitoring to collect various traffic sign pictures in foggy weather to obtain a dataset;
[0008] S002: Use the multi-scale adaptive network module MS for dehazing preprocessing. The multi-scale adaptive network module MS improves the image quality and detail performance through three different-scale convolutional kernels, and then performs dehazing processing through a dark channel prior dehazing filter, a Gamma filter, and a sharpening filter in sequence;
[0009] S003: Based on the improvement of the YOLOv11 network, introduce the C3K2-RFAConv module to replace the C3K2 module in both the Backbone part and the Neck part. Replace the traditional convolutional module with RFAConv in the Backbone part to obtain a traffic sign detection network model and train it using the dataset;
[0010] S004: Use the trained traffic sign detection network model to perform object detection on the dehazed preprocessed image.
[0011] Furthermore, the specific process of the multi-scale adaptive network module MS in S002 is as follows:
[0012] Input the foggy image, process it through three convolutional kernels of different scales, which consists of three parallel branches. The first branch uses five layers of convolution with 7×7 convolutions to process images with complex backgrounds and extract large-scale global information; the second branch uses five layers of 5×5 convolutions to extract medium-scale features; the third branch directly inputs the original foggy image to retain initial information and details; then add the outputs of the three branches to achieve multi-scale feature fusion, and finally use five layers of 3×3 convolutions to further extract details; then perform dehazing processing through three methods: dark channel prior dehazing filter, Gamma filter, and sharpening filter, and finally output the processed image to the YOLOv11 network model for detection.
[0013] Furthermore, perform dehazing processing through the dark channel prior dehazing filter, Gamma filter, and sharpening filter in sequence, as follows:
[0014] S00211: The dark channel prior dehazing filter estimates the dark channel before the image to eliminate the fog effect and restore the clarity and contrast of the image. First, construct a fog image based on the atmospheric scattering model:
[0015] Ix = Jxtx + A(1 - tx)
[0016] where Ix represents the image to be dehazed, Jx is the restored image, A is the global atmospheric light component, and tx is the transmittance;
[0017] tx = e -β d(x)
[0018] where β represents the atmospheric scattering coefficient and d(x) represents the scene depth; select the top 0.1% brightest pixels in the dark channel image and use the corresponding pixels in the original image to estimate the atmospheric light A. For each pixel in the input image, select the minimum value from all color channels within the local window to form the dark channel image:
[0019] I dark x = min min I c y y∈Ω(x) c∈Ω(x)
[0020] where I dark represents the dark channel image, I c represents each channel of the color image, and Ω(x) represents a window centered on pixel x. Then perform the final transformation according to the above formula:
[0021] tx = 1 - ωI dark x / A
[0022] where ω is a hyperparameter;
[0023] S00212: The Gamma filter improves the visual quality of an image by adjusting the brightness and contrast of the image, and is defined as follows:
[0024]
[0025] where r, g, b represent the RGB channel pixels, and G represents the hyperparameter to be optimized;
[0026] S00213: The sharpening filter emphasizes the edges and fine structures in an image:
[0027] Fx,λ = Ix + λIx - Gaulx
[0028] where Ix represents the input image, Gaul represents the Gaussian filter, and λ represents a proportionality factor.
[0029] Furthermore, the C3K2 - RFAConv module improves the Bottleneck in the C3K2 module, and replaces some of the standard convolutions Conv in the Bottleneck with RFAConv.
[0030] Furthermore, the RFAConv combines the receptive field attention RFA and the coordinated attention CA mechanism to enhance the feature extraction ability.
[0031] Furthermore, a small target detection layer for tiny feature extraction is added to the detection head part, the original 4 C3K2 modules are increased to 6 C3K2 - RFAConv modules, an upsampling and Concat module is set before the third C3K2 - RFAConv module, a Conv module and a Concat module are added before the fourth C3K2 - RFAConv module, and the outputs of the third, fourth, fifth, and sixth C3K2 - RFAConv modules are respectively output to the detection head part.
[0032] Beneficial effects:
[0033] The designed MSR-YOLO model of the present invention, on the one hand, introduces MS (multi-scale adaptive network module) for defogging preprocessing. With the three different-scale convolutional kernels (7×7, 5×5, 3×3) of the MS module, it effectively improves the image quality and detail performance. It performs defogging processing through three methods: dark channel prior defogging filter, Gamma filter, and sharpening filter, which can effectively reduce the impact of foggy environment on traffic sign recognition. On the other hand, the YOLOv11 network model is improved by introducing C3K2-RFAConv to overcome the defect of reduced performance of standard convolution caused by convolutional kernel parameter sharing, thereby enhancing the multi-scale information processing ability of the model. The present invention also adds a small target detection layer for extracting tiny features on the basis of the YOLOv11 network model. The small target detection layer is located after the last few convolutional blocks of the backbone network and contains several additional convolutional layers for extracting more detailed features. This detection layer can output a feature map with a size of 160 pixels × 160 pixels and can detect targets larger than 4 pixels × 4 pixels. Description of the Drawings
[0034] Figure 1 It is a diagram of the MSR-YOLO model;
[0035] Figure 2 It is a flowchart of the MS module;
[0036] Figure 3 It is a diagram of the C3K2-RFAConv module. Detailed Implementation Manner
[0037] The following further clarifies the present invention in combination with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications made by those skilled in the art to the present invention all fall within the scope defined by the appended claims of this application.
[0038] This article provides a traffic sign defogging detection method based on improved MSR-YOLO, including:
[0039] S001 Use on-vehicle monitoring to collect various traffic sign pictures. Since it is difficult to obtain a foggy traffic sign dataset, for the purpose of experiments, the present invention performs fogging processing on the conventionally obtained traffic sign pictures using a fogging algorithm to obtain a foggy traffic sign dataset.
[0040] S002 Improve its network framework based on the YOLOv11 model to obtain the improved MSR-YOLO model, as follows Figure 1 , The improvement of the MSR-YOLO model is as follows:
[0041] S0021: Use MS (Multi-Scale Adaptive Network Module) for defogging preprocessing; first, with the help of three different-scale convolutional kernels (7×7, 5×5, 3×3) of the MS module, effectively improve the image quality and detail performance. Then, perform defogging processing through three methods: dark channel prior defogging filter, Gamma filter, and sharpening filter. Finally, output the processed image to YOLOv11 for detection.
[0042] Use MS (Multi-Scale Adaptive Network Module) for defogging preprocessing. First, with the help of three different-scale convolutional kernels (7×7, 5×5, 3×3) of the MS module, effectively improve the image quality and detail performance. The specific operations are as follows: Input the foggy image. The architecture consists of three parallel branches. The first branch uses five layers of 7×7 convolutions, which can effectively process images with complex backgrounds and extract large-scale global information through multiple layers of 7×7 convolutions. The second branch uses five layers of 5×5 convolutions to balance the receptive field size and computational complexity. This branch enriches the overall feature representation by extracting medium-scale features through multiple layers of 5×5 convolutions. The third branch directly inputs the original foggy image to retain the initial information and details. Then, add the outputs of the three branches to achieve multi-scale feature fusion. Then, through five layers of 3×3 convolutions to further extract details. Then, perform defogging processing through three methods: dark channel prior defogging filter, Gamma filter, and sharpening filter. Finally, output the processed image to YOLOv11 for detection. The flowchart of the MS module is as follows Figure 2 .
[0043] The dark channel prior defogging filter estimates the dark channel before the image to eliminate the fog effect, thereby restoring the clarity and contrast of the image. First, construct the fog image based on the atmospheric scattering model:
[0044] Ix = Jx tx + A(1 - tx)
[0045] Ix represents the image to be defogged, Jx is the restored image, A is the global atmospheric light component, and tx is the transmittance.
[0046] tx = e -β d(x)
[0047] β represents the atmospheric scattering coefficient, and d(x) represents the scene depth. Select the top 0.1% brightest pixels in the dark channel image and use the corresponding pixels in the original image to estimate the atmospheric light A. For each pixel in the input image, select the minimum value from all color channels within the local window. This forms the dark channel image.
[0048] I dark x = min min I c y y∈Ω(x) c∈Ω(x)
[0049] I dark represents the dark channel image, and then performs the final transformation according to the above formula:
[0050] tx = 1 - ωI dark x / A
[0051] ω is a hyperparameter that is optimized through backpropagation of the multi-scale adaptive network to improve the performance of the dehazing filter for foggy image detection.
[0052] As a commonly used non-linear processing method in image processing, the Gamma filter improves the visual quality of the image by adjusting the brightness and contrast of the image. It is defined as follows:
[0053]
[0054] G represents the hyperparameter to be optimized.
[0055] As an image processing tool, the sharpening filter is used to enhance the details and edges in the image to make it look clearer. By emphasizing the edges and fine structures in the image, the sharpening filter can make the visual effect more vivid and dynamic.
[0056] Fx,λ = Ix + λIx - Gaulx
[0057] Ix represents the input image, Gaul represents the Gaussian filter, and λ represents a proportionality factor.
[0058] S0022: In the YOLOv11 model, C3K2-RFAConv is introduced to overcome the defect of reduced performance of the standard convolution caused by the sharing of convolution kernel parameters, thereby improving the multi-scale information processing ability of the model; the traditional convolution module is replaced by RFAConv in the Backbone part.
[0059] The C3K2-RFAConv module combines the receptive field attention (RFA) and coordinated attention (CA) mechanisms to enhance the feature extraction ability. The RFA mechanism dynamically adjusts the weights of feature processing in different local regions through fine control of each convolution operation, achieving intelligent weighting of features within the receptive field. This process not only enhances the network's ability to identify key regions in the image but also enables the network to focus more on these important features, thus maintaining a low computational complexity and parameter increment while improving performance. The CA module calculates attention in both the channel and spatial dimensions, better capturing the long-term dependencies between features while retaining precise position information. By coordinating spatial attention and channel attention, important feature information is integrated, improving the network's feature representation ability. Compared with the traditional self-attention mechanism, RFAConv improves the effectiveness of the convolution process. The C3K2-RFAConv module is as follows Figure 3as shown
[0060] S0023: Add a small target detection layer to improve detection accuracy.
[0061] Since the targets in traffic sign detection lack obvious feature information, the downsampling multiples of the 3 detection layers of the original YOLOV11 algorithm are relatively large, making it difficult to capture the features of some tiny targets. To address this issue, the neck and detection head parts were improved by adding 1 small target detection layer for extracting tiny features. The small target detection layer is located after the last few convolutional blocks of the backbone network and contains several additional convolutional layers for extracting more detailed features. This detection layer can output a feature map with a size of 160 pixels × 160 pixels and can detect targets larger than 4 pixels × 4 pixels. This method is more suitable for detecting small targets such as traffic signs. That is, the original 4 C3K2 modules were increased to 6 C3K2-RFAConv modules. An upsampling and Concat module was set before the third C3K2-RFAConv module, and a Conv module and a Concat module were added before the fourth C3K2-RFAConv module. The outputs of the third, fourth, fifth, and sixth C3K2-RFAConv modules were sent to the detection head part respectively.
[0062] The experimental environment for this time is as follows: The operating system of the computer is Windows 11 Professional Edition 64-bit, the GPU is NVIDIA GeForce RTX4060Ti, the video memory size is 16G, the CUDA version is 12.1, the deep learning framework is Pytorch 2.1.0, and Python is 3.11. The experimental training cycle (epoch) is 300, the number of threads (workers) is 8, the batch size (batchsize) is 16, the input photo size (imgsz) is 640, and the initial learning rate (lr0) is 0.01.
[0063] The samples of the object detection results can be roughly divided into three categories. True positive (TP) represents the correctly detected object, false negative (FN) represents the undetected object, and false positive (FP) represents the wrongly detected object. This experiment mainly uses mAP (mean average precision) to evaluate the performance.
[0064] mAP is a commonly used metric for evaluating object detectors. It quantifies the precision and recall of an object detector at different intersections above a joint threshold to measure its accuracy. IoU measures the overlap between the predicted bounding box and the actual ground truth box.
[0065]
[0066] In addition, we also measure the speed and efficiency of the target detector by the number of model parameters and GFLOPs. The lower the values of the first two metrics, the higher the model efficiency.
[0067] To verify the effectiveness of the model, the detection structures of various mainstream object detection networks were reproduced on the self-built dataset, and the improved model was compared with YOLOv8 and YOLOv11. The comparison results are shown in the following table.
[0068]
[0069] Among them, when YOLOv11 (modified) is combined with the dehazing model MS for object detection, its accuracy is greatly improved compared with other models.
[0070] The above embodiments are only for illustrating the technical concept and features of the present invention, aiming to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and shall not be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention shall be covered within the protection scope of the present invention.
Claims
1. A traffic sign defogging detection method based on improved MSR-YOLO, characterized in that The steps are as follows: S001: Use in-vehicle monitoring to collect various traffic sign pictures in foggy weather to obtain a dataset; S002: Use the multi-scale adaptive network module MS for defogging preprocessing. The multi-scale adaptive network module MS improves the image quality and detail performance through three different-scale convolutional kernels, and then performs defogging processing through a dark channel prior defogging filter, a Gamma filter, and a sharpening filter in sequence; S003: Based on the improvement of the YOLOv11 network, introduce the C3K2-RFAConv module to replace the C3K2 module in both the Backbone part and the Neck part. In the Backbone part, RFAConv replaces the traditional convolutional module to obtain a traffic sign detection network model and train it using the dataset; S004: Use the trained traffic sign detection network model to perform object detection on the defogged preprocessed image.
2. The traffic sign defogging detection method based on the improved MSR-YOLO according to claim 1, wherein The specific process of the multi-scale adaptive network module MS in S002 is as follows: Input the foggy image, process it through three different-scale convolutional kernels, which consists of three parallel branches. The first branch uses five layers of convolution with 7×7 convolutions to process images with complex backgrounds and extract large-scale global information; the second branch uses five layers of 5×5 convolutions to extract medium-scale features; the third branch directly inputs the original foggy image to retain the initial information and details; then add the outputs of the three branches to achieve multi-scale feature fusion, and finally use five layers of 3×3 convolutions to further extract details; then perform defogging processing through three methods: a dark channel prior defogging filter, a Gamma filter, and a sharpening filter, and finally output the processed picture to the YOLOv11 network model for detection.
3. The traffic sign defogging detection method based on the improved MSR-YOLO according to claim 2, wherein Perform defogging processing through a dark channel prior defogging filter, a Gamma filter, and a sharpening filter in sequence, as follows: S00211: The dark channel prior defogging filter estimates the dark channel in front of the image to eliminate the fog effect and restore the clarity and contrast of the image. First, construct a fog image based on the atmospheric scattering model: Ix = Jxtx + A(1 - tx) where Ix represents the image to be defogged, Jx is the restored image, A is the global atmospheric light component, and tx is the transmittance; tx = e -β d(x) where β represents the atmospheric scattering coefficient, and d(x) represents the scene depth; select the top 0.1% brightest pixels in the dark channel image and use the corresponding pixels in the original image to estimate the atmospheric light A. For each pixel in the input image, select the minimum value from all color channels within the local window to form a dark channel image: I dark x = min min I c y y ∈ Ω(x) c ∈ Ω(x) Among them, I dark represents the dark channel image, I c represents each channel of the color image, Ω(x) represents a window centered on pixel x, and then the following final transformation is performed according to the above formula: tx = 1 - ωI dark x / A where ω is a hyperparameter; S00212: The Gamma filter improves the visual quality of the image by adjusting the brightness and contrast of the image, and is defined as follows: where r, g, b represent the RGB channel pixels, and G represents the hyperparameter to be optimized; S00213: The sharpening filter emphasizes the edges and fine structures in the image: Fx,λ = Ix + λIx - Gaul x where Ix represents the input image, Gaul represents the Gaussian filter, and λ represents a proportionality factor.
4. The traffic sign defogging detection method based on the improved MSR-YOLO according to claim 1, wherein The C3K2-RFAConv module improves the Bottleneck in the C3K2 module, replacing some of the standard convolutions Conv in the Bottleneck with RFAConv.
5. The traffic sign defogging detection method based on the improved MSR - YOLO according to claim 4, wherein, The RFAConv combines the receptive field attention RFA and the coordinate attention CA mechanism to enhance the feature extraction ability.
6. The traffic sign defogging detection method based on the improved MSR-YOLO according to claim 1, characterized in that Add a small object detection layer for tiny feature extraction in the detection head part, increase the original 4 C3K2 modules to 6 C3K2-RFAConv modules, set an upsampling and Concat module before the third C3K2-RFAConv module, add a Conv module and a Concat module before the fourth C3K2-RFAConv module, and output the third, fourth, fifth, and sixth C3K2-RFAConv modules to the detection head part respectively.
Citation Information
Cited By
Remote sensing image haze removal detection method and device and electronic equipment
CN120931629A
Automobile central control screen small target detection method based on YOLOv11 improvement
CN121415218A