Night road waterlogging detection method and device based on improved yo11

By improving the YOLO11 model and introducing a reflectance sensing feature enhancement and fusion module, the accuracy problem of nighttime road water accumulation detection was solved, and the detection accuracy and robustness were improved, especially the water accumulation detection effect in complex reflective environments.

CN120894531APending Publication Date: 2025-11-04WUXI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511080246.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify road flooding at night, especially given the complex road textures and reflective features, which increases the risk of traffic accidents such as loss of vehicle control.

Method used

An improved YOLO11 model is adopted, incorporating Mosaic data augmentation, adaptive anchor box calculation mechanism, and image adaptive scaling strategy. Combined with the Reflectance Sensing Feature Enhancement Module (RAFE) and the Reflectance Sensing Feature Fusion Module (RAFF), the extraction and fusion of image feature information are optimized to improve detection accuracy.

Benefits of technology

It significantly improves the accuracy of water accumulation detection in complex reflective environments, enhances the ability to detect small-sized water accumulation targets, and improves the robustness and accuracy of the detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894531A_ABST
    Figure CN120894531A_ABST
Patent Text Reader

Abstract

The invention discloses a night road waterlogging detection method and device based on improved yo11, and the method comprises the steps: obtaining historical night road waterlogging scene data, carrying out the preprocessing of the data, and building a night road waterlogging detection data set; training and generating a night road waterlogging detection model based on an improved yolo11 model by using the established night road waterlogging detection data set; and inputting an image of a to-be-detected ponding target in the night road scene by using the generated night road ponding detection model, and obtaining ponding information of the target in the to-be-detected image. The RAFE module is added in the backbone network of the original yolo11 model, and the concat connection of the yolo11 is replaced by the designed RAFF module, so that the extraction capability of the model on the global position information of the ponding point in the shallow network is enhanced, and the distribution and fusion of the weight of the semantic information such as the reflective feature and the texture feature in the deep network are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of road water detection, and particularly relates to a night road water detection method and device based on improved yolo11. BACKGROUND

[0002] Road water is a common traffic hazard that can cause vehicles to lose control, leading to various accidents from minor fender bending to serious collisions, posing a serious threat to road safety. Due to complex road textures and different colors affected by reflective features, especially at night, the light is strong and weak at times, and existing technologies are difficult to accurately identify road water. In recent years, the principles of artificial intelligence technology and its application in various fields have made great breakthroughs, and have gradually become an important issue in daily life and scientific and technological research. Among them, image recognition technology in artificial intelligence technology is very important, and with the continuous development of the information age, image intelligent recognition technology has gradually emerged and been widely applied.

[0003] Deep learning is a kind of machine learning, which is based on the structure of artificial neural network, and extracts features through learning a large amount of data, so as to realize the recognition and understanding of complex patterns. In the field of target detection, deep learning can automatically learn image features and effectively deal with the uncertainty of detection targets. At present, convolutional neural network (CNN) is one of the most widely used algorithms in deep learning, which has excellent performance in image classification, target detection and other fields.

[0004] Yolo11 is a kind of target detection algorithm based on deep learning, which can effectively deal with the problem of target size change through multi-scale feature fusion strategy. However, in order to solve the above challenges, we propose a night road detection method based on improved yolo11, which is used for active detection of road water and improvement of traffic safety. SUMMARY

[0005] The present application aims to at least solve the technical problems existing in the related art to some extent.

[0006] The present application aims to provide a detection method that can effectively deal with the size change of road water targets.

[0007] In order to achieve the above-mentioned purpose, the present application provides a night road waterlogging detection method based on improved yolo11 in one aspect, comprising the following steps: S1, obtaining historical night road waterlogging scene data, and preprocessing the data to establish a night road waterlogging detection dataset; S2, based on the improved yolo11 model, a night road waterlogging detection model is trained and generated with the established night road waterlogging detection dataset; S21, a night road waterlogging detection model is constructed, which takes the improved yolo11 model as the benchmark model and includes an input module, a backbone module, a neck module and a head module: the input module is based on Mosaic data enhancement, adaptive anchor frame calculation mechanism and image adaptive scaling strategy, and is used for data set scale expansion, anchor frame parameter optimization and image preprocessing filling; the backbone module introduces a reflection perception feature enhancement module RAFE for enhancing and extracting image feature information; the image features are divided into shallow features and deep features; the neck module introduces a reflection perception feature fusion module RAFF to replace the concat module, which is used to fuse the geometric information of the shallow network and the semantic information of the deep network, and improve the detection performance of the model; the head module uses C3 module to adjust the channel of the feature map output by the neck module, and then generates road waterlogging information through convolution operation; S22, based on the extracted night road waterlogging features, the constructed road waterlogging detection model is trained to obtain the optimal road waterlogging detection model; S3, using the generated night road waterlogging detection model, input the image of the waterlogging target to be detected in the night road scene, and obtain the waterlogging information of the target in the image to be detected.

[0008] Preferably, the preprocessing in step S1 is specifically: using Robotflow tool to label the picture and setting the data set format to txt format to make the night road waterlogging detection data set.

[0009] As preferred, the reflection-aware feature enhancement module RAFE enhances the features affected by reflection in the water accumulation area, and the specific processing process is as follows: step 1, based on the feature map output by the Conv module, extract reflection features of different feature scales through three parallel reflection feature branches; step 2, based on the reflection features of different feature scales, generate enhanced reflection features based on the Concat splicing and the Conv module; step 3, based on the feature map output by the Conv module, generate noise weights through the noise branch noise_branch; step 4, based on the feature map output by the Conv module, generate a global noise ratio through the noise weight predictor noise_weight_predictor; step 5, based on the noise weights and the global noise ratio, suppress reflection noise and enhance water accumulation features, and generate features fused with context information through a multi-head attention mechanism; step 6, based on the generated enhanced reflection features and the features fused with context information, perform residual connection and output through convolution.

[0010] As preferred, the three parallel reflection feature branches in step 1 use 1x1, 3x3, and 5x5 convolution kernels respectively, and each reflection branch is compressed through grouped convolution, followed by normalization, ReLU activation, and dropout.

[0011] As preferred, the reflection-aware feature fusion module RAFF is used to optimize the fusion of shallow and deep features, enhance the boundary and depth information of water accumulation detection, and the specific processing process is as follows: step one, receive shallow features shallow and deep features deep, and upsample the deep features to the shallow resolution through bilinear interpolation; at the same time, initialize the parameters dynamically according to the input channel number; step two, generate the weights of shallow and deep features through 1x1 convolution and Sigmoid activation, and apply them to shallow and deep features to generate weighted_shallow and weighted_deep respectively; step three, based on the generated weighted_deep, generate reflection attention weights through 3x3 convolution and 1x1 convolution to suppress reflection interference; step four, based on the reflection attention weights, generate dense weights through the dense_weight module to enhance feature expression; step five, based on the projection of the deep features in step S24 to the shallow channel number, add the weighted shallow features, and output the fused features through 3x3 convolution refine.

[0012] In another aspect, the application provides a night road water detection device based on improved YOLO11, which comprises: a preprocessing module: enhancing and labeling the night water data set to generate a txt format data set; a feature processing module: used for extracting and fusing night road water features according to the improved YOLO11 model and introducing RAFE and RAFF modules; a model training module: used for training and optimizing the night road water detection model based on the improved YOLO11 according to the night road water features; and a water detection module: used for real-time road water detection by using the trained model.

[0013] Preferably, the feature processing module comprises: a feature extraction unit: used for extracting road water features by the RAFE module according to the data set; a feature fusion unit: used for fusing multi-scale features by the RAFF module according to the road water features; and a weight adjustment unit: used for suppressing reflection noise by adaptively adjusting feature weights.

[0014] Beneficial effects: The water detection method provided by the application significantly improves the performance of the water detection model in a complex reflection environment, especially the detection accuracy in the urban night road water scene, by introducing the reflection perception feature enhancement module and the reflection perception feature fusion module. The multi-scale feature enhancement mechanism is introduced to improve the feature expression ability of the model to the water area by fusing feature information of different scales, thereby enhancing the detection ability to small size water targets. At the same time, the reflection perception fusion technology is introduced to optimize the integration of shallow and deep features, thereby significantly improving the detection accuracy of the model to the water boundary and depth. BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1 FIG. 1 is a flowchart of a road water detection method based on an improved YOLO11 provided by the application.

[0016] Figure 2 FIG. 2 is a Yolo11 network structure diagram of a road water detection method based on an improved YOLO11 provided by the application.

[0017] Figure 3 FIG. 3 is a RAFE structure diagram provided by the application.

[0018] Figure 4 FIG. 4 is a RAFF structure diagram provided by the application.

[0019] Figure 5 FIG. 5 is a curve diagram of the recall rate changing with the confidence value provided by the application.

[0020] Figure 6 FIG. 6 is a curve diagram of the average precision changing with the recall rate provided by the application.

[0021] Figure 7The present application provides a far night road water accumulation identification schematic diagram.

[0022] Figure 8 The present application provides a near night road water accumulation identification schematic diagram. DETAILED DESCRIPTION

[0023] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments, and they should not be understood as limiting the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application. In the description of the present application, it should be understood that the terms used are only for the purpose of description, and should not be understood as indicating or implying relative importance.

[0024] Firstly, some nouns designed in the present application are explained:

[0025] At present, the mainstream target detection algorithm is roughly divided into two-stage and single-stage detection algorithms. The two-stage detection algorithm is represented by the RCNN series, and the single-stage target detection algorithm is represented by the YOLO and RT-DETR series. Among them, the yolo series has the advantages of fast detection speed, high precision, etc., and has become the mainstream target detection algorithm at present. In view of the technical challenges of early small area water accumulation warning detection under a large number of targets in the urban road water accumulation scene, the present application optimizes and improves the YOLO11 detection algorithm for specific application scenarios to meet the actual needs of the specific application scenarios.

[0026] YOLO11 introduces C3k2 module and C2PSA module and other innovative designs, which significantly improves the accuracy of target detection while maintaining real-time detection speed, especially good at handling small targets and occluded targets in complex scenes. C3k2 module is an optimized version of YOLOv11 for traditional CSP Bottleneck structure, and its core design improves feature extraction efficiency through parallel convolution branches and flexible parameter configuration. The input feature map is divided into two parts, one part is directly transmitted to retain the shallow features, and the other part processes the deep features through multiple Bottleneck or C3k modules (variable convolution kernel), and finally realizes multi-scale feature fusion through splicing fusion. C2PSA (Cross Stage Partial with Pyramid Squeeze Attention) module combines CSP structure and attention mechanism, and improves the attention ability to complex occluded objects and key regions through multi-scale convolution kernel and channel weighting. The feature map is divided into two parts, one part is directly transmitted, and the other part is processed through the PSA attention module, and finally spliced and fused.

[0027] From the perspective of model architecture, the YOLOv11 model is mainly composed of four core parts: input module (Input), backbone network (Backbone), neck module (Neck), and head module (Head). Each module closely cooperates to complete the target detection task. The specific architecture of the improved model is shown in FIG. 1. Figure 2

[0028] In the model input stage, we mainly adopted three innovative technologies: Mosaic data enhancement, adaptive anchor box calculation mechanism, and image adaptive scaling strategy, which respectively provide solutions for key problems such as data set scale expansion, anchor box parameter optimization, and image preprocessing padding. Specifically, the Mosaic data enhancement technology randomly selects four images from the training set, performs dynamic scaling, intelligent cropping, and random spatial arrangement operations, and then splices them into a synthetic image. This innovative data synthesis method not only effectively expands the diversity of training samples, but also enhances the model's detection ability for multi-target and multi-scale targets by simulating complex scenes, significantly improving the generalization performance and robustness of the target detection network. By combining the loss function CIOU to adaptively calculate the anchor boxes of different data sets, the anchor box with the highest recall rate is saved for model training, enabling the model to converge better.

[0029] The Backbone module is composed of CBS, C2f, and SPP structures to extract image feature information. The CBS structure is composed of convolution (Conv), batch normalization layer (BN), and SiLU activation function, which is used for feature extraction and down-sampling of images. The C2f module is composed of multiple CBS structures, which enhances the learning ability of the network and optimizes multi-scale feature representation by cross-layer connection and feature fusion while maintaining gradient flow. The SPP structure is divided into multiple branches, which uses a combination of Maxpool layer and CBS layer to extract features at different receptive fields. The upper branch uses Maxpool operations with different sizes (kernel size of 5x5, 9x9, 13x13, and step size of 1) to pool the feature map and preserve global information. The lower branch uses the CBS structure (convolution kernel size of 1x1 and step size of 1) to adjust the channel number of the feature map, and then fuses it with the pooling result through the cat operation, thereby improving the model's feature extraction ability for complex scenes. This design enables the Backbone module of YOLO11 to maintain high efficiency while significantly improving feature expression ability and robustness.

[0030] ​The Neck module employs a Path Aggregation Feature Pyramid Network (PAFPN) structure, which improves the model's detection performance by bidirectionally fusing geometric information from shallow networks and semantic information from deep networks. The Head module uses the C3 module to adjust the channels of the feature map output by PAFPN, and then generates object confidence, category, and bounding box information through a convolution operation with a kernel size of 1×1 and a stride of 1.

[0031] Example 1: As Figure 1 As shown, this embodiment of the invention provides a method for detecting nighttime road water accumulation based on an improved YOLOv11, including: S1: acquiring historical nighttime road water accumulation scene data, preprocessing the data, and establishing a nighttime road water accumulation detection dataset; collecting nighttime road water accumulation scene data, using Robotflow to annotate the images, setting the dataset format to txt format, and creating a nighttime road detection dataset.

[0032] Create a dataset for nighttime road detection, and divide the training set and validation set in a 7:3 ratio.

[0033] S2: Using the established nighttime road flooding detection dataset, a nighttime road flooding detection model is trained based on the improved YOLO11 model; S21: Construct the nighttime road flooding detection model, which uses the improved YOLO11 model as the baseline model and includes an input module, a backbone module, a neck module, and a head module: The input module, based on Mosaic data augmentation, adaptive anchor box calculation mechanism, and image adaptive scaling strategy, is used to expand the dataset size, optimize anchor box parameters, and preprocess and fill images; The backbone module introduces the Reflectance Awareness Feature Enhancement (RAFE) module... This paper focuses on enhancing and extracting image feature information. The image features are divided into shallow and deep features. The neck module introduces a Reflectance-Aware Feature Fusion (RAFF) module to replace the concat module, fusing geometric information from the shallow network and semantic information from the deep network to improve the model's detection performance. The head module uses the C3 module to adjust the channels of the feature map output by the neck module, and then generates object confidence, category, and bounding box information through convolution. An improved YOLOv11 nighttime road flooding detection model is built, and a Reflectance-Aware Feature Enhancement (RAFE) module and a Reflectance-Aware Feature Fusion (RAFF) module are designed, significantly improving the performance of the flooding detection model in complex reflection environments. The following is a detailed technical description of the two modules: The Reflectance-Aware Feature Enhancement (RAFE) module is designed to enhance features affected by reflection in flooded areas, such as... Figure 2As shown, RAFE is used in the Backbone part for the extraction of waterlogged features, which is mainly responsible for enhancing features affected by reflection, such as Figure 3 As shown, the first input feature map passes through three parallel reflective feature branches (ReflectiveBranch), in which three parallel ReflectiveBranch branches use 1x1, 3x3, and 5x5 convolution kernels, respectively, so that reflective features of different feature scales can be extracted. Each reflective branch is compressed through group convolution (groups=channels / / 4) and then connected with BatchNorm2d, ReLU activation, and dropout (default rate 0.15) to prevent overfitting and enhance robustness. The three parallel reflective branches are concatenated along the channel dimension, and then restored to the original channel through 1x1 convolution to generate enhanced reflective features (reflective_features). At the same time, the input feature also enters the noise branch (noise_branch), which generates noise weights denoted as noise_weight through 3x3 convolution (groups=channels / / 4), batch normalization, ReLU activation, 1x1 convolution, and Sigmoid activation. At the same time, the noise weight predictor (noise_weight_predictor) reduces the spatial dimension through AdaptiveAvgPool2d, and then generates a global noise scale denoted as noise_weight_scale through 1x1 convolution and Sigmoid. Finally, the reflective noise is suppressed and the waterlogged feature is enhanced through the formula reflective_features * (1-noise_weight_scale*noise_weight), denoted as enhanced_features. Subsequently, the multi-head attention mechanism (MultiheadAttention) is used to enhance the features and capture the spatial dependency relationship, and then the features fused with the context information are output. Finally, the reflective_features are connected in residual through convolution.

[0034] The reflective perception feature fusion module (RAFF) is used to optimize the fusion of shallow and deep features and enhance the boundary and depth information of water detection, as shown in Figure 4As shown, the RAFF module fuses shallow and deep features, optimizes boundary and depth information. Both shallow (shallow, high resolution) and deep (deep, low resolution) features are received, the deep features are upsampled to the shallow resolution by bilinear interpolation to ensure spatial consistency, then the parameters are dynamically initialized according to the input channel number, and the reduction is set to 16 by default to reduce the channel to reduce the computational complexity and reduce the risk of overfitting. Subsequently, after the shallow and deep features are spliced (Concat), the weight (weight_conv) module, that is, (1x1 convolution+BatchNorm2d+ReLU+1x1 convolution+Sigmoid), generates weights, which are divided into shallow weights weighted_shallow and deep weights weighted_deep along the channel dimension, and are applied to shallow and deep features, respectively. Subsequently, weighted_deep generates reflection weights (reflection_weight) through the reflection attention reflection_attention module, that is, (3x3 convolution, groups=deep_channels / / reduction+BatchNorm2d+ReLU+1x1 convolution+Sigmoid), and suppresses the interference of water reflection through weighted_deep*(1-reflection_weight), denoted as reflection_suppressed, and then generates dense weights (dense_weight) through the dense weight (dense_weight) module (similar to reflection_attention), and enhances feature expression through the formula reflection_suppressed*(1+dense_weight). Subsequently, the fusion features are generated by projecting through the 1x1 convolution layer to the shallow channel number and adding weighted_shallow. Finally, the fusion features are output through the refine module (3x3 convolution+BatchNorm2d+ReLU).

[0035] S22, based on the extracted night road water feature, training the constructed road water detection model to obtain an optimal road water detection model:

[0036] The training set in the night road waterlogging detection data set constructed in S1 is input into the improved yolo11 night road waterlogging detection model in S2 for training, network model training hyperparameters are set, the number of training rounds is set to 200 rounds, the batch size (bacth-size) is 32, SGD is used as the optimizer, which can make the convergence more smooth, the model training environment details are configured as shown in Table 1, after the training is completed, the verification set is used for verification, and the night road waterlogging detection model effect optimal weight is saved, here the model optimal weight refers to the parameter values (weights and biases) learned by each layer (convolutional layer, BN layer, classification head, etc.) of the neural network, which can make the waterlogging detection precision best, and the best.pt is named, the best.pt is loaded into the road waterlogging detection model, and the optimal road waterlogging detection model is generated, which is used for inference detection in the future.

[0037] S3, using the night road waterlogging detection model generated by training, inputting the image of the waterlogging target to be detected in the night road scene, obtaining the waterlogging information of the target to be detected in the image.

[0038] To verify the waterlogging detection effect of the embodiment, the same batch of images to be detected is detected for waterlogging by using the YOLO11 algorithm and the detection method of the present application. As shown in Table 1, the experimental environment of the improved yolo11 night road waterlogging detection method based on the present application is Windows10, PyTorch2.4.1 version, the graphics card is NVIDIA GeForce RTX 4060Ti, and the specific configuration is shown in Table 1. The network is trained in GPU mode, the network model parameters are set, the number of training rounds is set to 200 rounds, the bacth-size is 32, and SGD is used as the optimizer. After the training is completed, the night road waterlogging detection model effect optimal weight is saved, and is named as best.pt.

[0039] Parameter Configuration CPU 12th Gen Intel(R) Core(TM) i5-12400F GPU NVIDIA GeForce RTX 4060 Ti System environment Wondows10 CUDA version CUDA12.1 Edit language version Python3.10 Deep learning framework PyTorch2.4.1 The image of the waterlogging target to be detected in the night road scene is input into the improved yolo11 model, the optimal weight file best.pt is loaded into the detection model for inference detection, and whether the target to be detected in the image contains waterlogging information is obtained.

[0040] On this basis, the category and the detection accuracy of the detection are displayed in the output picture, and the detection accuracy value in the detection result is proportional to the detection effect. In order to evaluate the accuracy and stability of the model, the precision P, the recall rate R and the average precision value mAP are selected as evaluation indexes, and the calculation formula is as follows ​​​; wherein TP represents the number of positive samples predicted by the model, FP represents the number of negative samples predicted by the model, FN represents the number of positive samples predicted by the model, and n represents the class of the sample; is the average value of the i-th class sample;

[0041] A recall rate Recall and precision rate Precision graph generated after training 200 rounds of the improved yolo11 nighttime road water detection method are respectively as shown in Figure 5 , Figure 6 Table 2 is a model performance comparison table. Under the condition of keeping consistent experimental environment and parameters, the recall Recall value of the algorithm in the present application reaches 75.3%, which is improved by 2.6% compared with yolo11, and the average precision mAP value reaches 86%, which is improved by 2% compared with yolo11, proving that the improved yolo11 nighttime road water detection method proposed in the present application is feasible.

[0042] Algorithm name Recall / % mAP / % Yolo11 72.7% 84% Algorithm in this paper 75.3% 86% The water detection method proposed in the present application significantly improves the performance of the water detection model in a complex reflection environment, especially the detection accuracy in the urban nighttime road water scene, by introducing a reflection perception feature enhancement module and a reflection perception feature fusion module. By introducing a multi-scale feature enhancement mechanism, the feature expression ability of the model to the water area is improved by fusing feature information of different scales, and the detection ability to small size water targets is enhanced. At the same time, through the reflection perception fusion technology, the integration of shallow and deep features is optimized, and the detection accuracy of the model to the water boundary and depth is significantly improved. The test of the method of the present application on the collected private data set and part of the open source data set shows that compared with the original YOLO11 model, the detection accuracy of the present application is improved by about 2%, and the detection performance in the complex reflection environment is particularly obvious. Figure 7 、 Figure 8 The detection effects of the present application in different scenes are respectively shown, proving the superiority thereof. The water detection method proposed in the present application can be widely applied to the fields of urban drainage system monitoring, road water early warning, etc., has high practical value and popularization potential, can effectively improve the accuracy and reliability of water detection, and solves the detection of traditional methods in complex scenes.

[0043] Embodiment 2: An improved yolo11-based night road waterlogging detection device, comprising: a preprocessing module: enhancing and labeling a night waterlogging dataset to generate a txt format dataset; a feature processing module: used for extracting and fusing night road waterlogging features according to an improved YOLO11 model by introducing RAFE and RAFF modules; a model training module: used for training and optimizing an improved YOLO11-based night road waterlogging detection model according to the night road waterlogging features; and a waterlogging detection module: used for real-time road waterlogging detection by using the trained model.

[0044] As preferred, the feature processing module comprises: a feature extraction unit: used for extracting road waterlogging features by a RAFE module according to the dataset; a feature fusion unit: used for fusing multi-scale features by a RAFF module according to the road waterlogging features; and a weight adjustment unit: used for suppressing reflection noise by adaptively adjusting feature weights.

[0045] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for detecting nighttime road water accumulation based on an improved YOLOv11, characterized in that, The process includes the following steps: S1. Acquire historical nighttime road flooding scene data and preprocess the data to establish a nighttime road flooding detection dataset; S2. Using the established nighttime road flooding detection dataset, train and generate a nighttime road flooding detection model based on an improved YOLOv11 model; S21. Construct a nighttime road flooding detection model, which uses the improved YOLOv11 model as the baseline model and includes an input module, a backbone module, a neck module, and a head module: The input module, based on Mosaic data augmentation, adaptive anchor box calculation mechanism, and image adaptive scaling strategy, is used to expand the dataset size, optimize anchor box parameters, and fill in image preprocessing; The backbone module introduces a Reflectance Awareness Feature Enhancement (RAFE) module to enhance and extract image feature information; The image features are divided into shallow features and deep features; The neck module introduces a Reflectance Awareness Feature Fusion (RAFF) module to replace the concat module, which fuses the geometric information of the shallow network and the semantic information of the deep network to improve the detection performance of the model. The head module uses the C3 module to adjust the channels of the feature map output by the neck module, and then generates road water accumulation information through convolution. S22: Based on the extracted nighttime road water accumulation features, the constructed road water accumulation detection model is trained to obtain the optimal road water accumulation detection model. S3: Using the generated nighttime road water accumulation detection model, the image of the target water accumulation in the nighttime road scene is input to obtain the water accumulation information of the target in the image to be detected.

2. The nighttime road water accumulation detection method based on improved YOLOv11 according to claim 1, characterized in that, The preprocessing described in step S1 specifically involves: using the Robotflow tool to annotate the images, setting the dataset format to txt, and creating a nighttime road flooding detection dataset.

3. The nighttime road water accumulation detection method based on the improved YOLOv11 according to claim 2, characterized in that, The Reflection Sensing Feature Enhancement (RAFE) module enhances the features affected by reflection in the water accumulation area. Its specific processing steps are as follows: Step 1: Based on the feature map output by the Conv module, extract reflection features at different feature scales through three parallel reflection feature branches; Step 2: Based on the reflection features at different feature scales, generate enhanced reflection features using Concat concatenation and the Conv module; Step 3: Based on the feature map output by the Conv module, generate noise weights through the noise branch (noise_branch); Step 4: Based on the feature map output by the Conv module, generate a global noise ratio using the noise weight predictor (noise_weight_predictor); Step 5: Based on the noise weights and the global noise ratio, suppress reflection noise and enhance water accumulation features, generating features fused with contextual information through a multi-head attention mechanism; Step 6: Based on the generated enhanced reflection features and features fused with contextual information, perform residual connections and output after convolution.

4. The nighttime road water accumulation detection method based on the improved YOLOv11 according to claim 3, characterized in that, The three parallel reflection feature branches mentioned in step 1 use 1x1, 3x3, and 5x5 convolution kernels, respectively. Each reflection branch is followed by normalization, ReLU activation, and dropout after channel compression through grouped convolution.

5. The nighttime road water accumulation detection method based on the improved YOLOv11 according to claim 4, characterized in that, The Reflection Sensing Feature Fusion (RAFF) module is used to optimize the fusion of shallow and deep features, enhancing the boundary and depth information of water accumulation detection. Its specific processing steps are as follows: Step 1: Receive shallow and deep features, and upsample the deep features to the shallow resolution using bilinear interpolation; simultaneously, dynamically initialize parameters based on the number of input channels; Step 2: Generate weights for the shallow and deep features using 1x1 convolution and sigmoid activation, applying them to the shallow and deep features to generate weighted_shallow and weighted_deep, respectively; Step 3: Based on the generated weighted_deep, generate reflection attention weights using 3x3 and 1x1 convolutions to suppress reflection interference; Step 4: Based on the reflection attention weights, generate dense weights using the dense_weight module, thereby enhancing feature representation. Step 5: Based on the deep features described in step S24, project them to the number of shallow channels, add them to the weighted shallow features, and then refine the output fused features through a 3x3 convolution.

6. A nighttime road water accumulation detection device based on an improved YOLOv11, employing the nighttime road water accumulation detection method based on an improved YOLOv11 as described in any one of claims 1-5, characterized in that, The device includes: a preprocessing module for enhancing and labeling the nighttime water accumulation dataset to generate a txt format dataset; a feature processing module for extracting and fusing nighttime road water accumulation features by introducing RAFE and RAFF modules based on the improved YOLO11 model; a model training module for training and optimizing a nighttime road water accumulation detection model based on the nighttime road water accumulation features; and a water accumulation detection module for real-time road water accumulation detection using the trained model.

7. The nighttime road water accumulation detection device based on the improved YOLOv11 according to claim 6, characterized in that, The feature processing module includes: a feature extraction unit for extracting road water accumulation features using the RAFE module based on the dataset; a feature fusion unit for fusing multi-scale features using the RAFF module based on the road water accumulation features; and a weight adjustment unit for suppressing reflection noise by adaptively adjusting feature weights.

Citation Information

Cited By

  • Establishment method and application of ponding image day and night bidirectional conversion model based on reflection map consistency

    CN121961833A