Method for recognizing road surface conditions in severe weather through expressway monitoring video
By designing a hybrid perception and fusion network model for weather element characteristics, the problem of interference between multiple weather conditions in automatic highway identification is solved, high-precision identification of road conditions in bad weather is achieved, and technical support is provided for expressway management and traffic control.
Patent Information
- Application Number
- CN202510629146.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Automatically identify the weather conditions of the expressway, multiple weather conditions interfere with each other, and the quality of the monitoring video image is affected by the weather conditions, making it difficult to accurately identify the road conditions of the severe weather conditions on the expressway.
A hybrid perception fusion network model of weather element characteristics was designed. By obtaining highway monitoring video data, a bad weather image data set was constructed, and the model training was carried out using the SGD optimizer and cross-entropy loss function to achieve high-precision classification of five types of road conditions: dry, water accumulation, ice accumulation, snow accumulation and heavy fog.
Real-time automatic identification of bad weather on highways has been realized, the accuracy of road conditions affected by five types of bad weather has been improved, the identification efficiency has been enhanced, and technical support for the highway management and traffic control of the transportation department.
Smart Images

Figure CN120147982A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for identifying road conditions in bad weather, specifically a method for identifying road conditions in bad weather from highway surveillance videos, belonging to the technical field of traffic meteorological image processing. Background Art
[0002] In modern road traffic, bad weather has a great impact on transportation. Snow, fog, freezing rain, icing and other extreme weather conditions may lead to traffic jams and serious traffic accidents, greatly affecting people's lives. Bad weather such as rain, fog and snow will cause the visibility value to decrease, and the bad weather reduces the friction and roughness between the vehicle and the road surface, thus directly affecting the driver's line of sight and vehicle driving performance, making it difficult for the driver to observe the road conditions and unable to judge the traffic conditions ahead in a timely and accurate manner, and prone to vehicle accidents; bad weather such as road icing and strong winds will hinder the normal driving of vehicles and easily cause the vehicles to skid or deviate from the normal driving route. Bad weather greatly deteriorates the driving conditions on highways. Weather recognition methods need to learn more complex phenomena such as the illumination and reflection of objects and scene surfaces, and due to the diversity, variability and high mutual dependence of weather characteristics, the research on this topic has always been extremely challenging. Therefore, it is particularly important to monitor road weather conditions in real time.
[0003] In traditional weather forecasting, meteorological forecasting is achieved by constructing an atmospheric model based on the monitoring results of meteorological satellites and ground observation stations. This method can cover a wide range of weather trends, but the accuracy of real-time weather forecasting in a small area is relatively low. The prediction accuracy of highways can be improved by densely building professional automatic weather stations, but the construction cost of weather stations is relatively high and the maintenance cost is large, and the power supply equipment on highways is not fully supplied, making it difficult to implement. However, highway video surveillance has a comprehensive coverage, and obtaining visual images has the advantages of low cost and high efficiency. By combining all the traffic surveillance video cameras on highways with the outdoor images of highways obtained by deep learning recognition cameras, the road weather conditions can be covered more meticulously. Combining this result with traditional short-term weather forecasting methods can make weather forecasting more accurate. Therefore, the analysis of weather image conditions based on deep learning has significant advantages in terms of cost and efficiency.
[0004] Related patent literature: CN118469837A discloses a method and device for constructing a model to enhance the clarity of highway surveillance videos, including the following steps: constructing an architecture for enhancing the clarity of highway surveillance videos, which includes a feature extraction module, a frequency feature modulation module, and a decoding module; obtaining surveillance video images in a highway scenario, adding noise to them, and using them as training samples, and performing unsupervised training on the architecture for enhancing the clarity of highway surveillance videos with the training samples as the input to obtain a model for enhancing the clarity of highway surveillance videos. This solution enhances and modulates frequencies by converting perspectives and processing image features in the frequency domain using a multi-layer perceptron, thereby enabling real-time enhancement of the clarity of highway surveillance videos. CN113917564A discloses a multi-parameter analysis remote sensing type road surface weather condition detector and detection method. The detector uses auxiliary measurements of the near-road surface atmospheric temperature T and humidity Hum to analyze environmental parameters, analyzes the change in the dew point temperature T0, and judges the dew condensation condition; uses the remote sensing type infrared radiation temperature measurement principle to measure the change in the road surface temperature Tr parameter and judges and analyzes the road surface icing and snow accumulation conditions; comprehensively measures multiple parameters, designs an identification and judgment algorithm, and analyzes and judges road surface weather conditions such as dry, wet, waterlogged, icy, snow-covered, and ice-water mixture on the road surface; all components are controlled by a microcontroller and data is collected. The microcontroller has a multi-parameter analysis and judgment algorithm built-in, and after calculating and analyzing the multi-parameter data, the measured data is sent out through the RS232 / 485 serial port; all components are installed in an all-weather and durable housing, which can withstand harsh weather and provide accurate data under any weather conditions.
[0005] The above technologies do not solve the problems that in the automatic identification of highway weather conditions, multiple weather conditions interfere with each other, and the quality of surveillance video images is affected by weather conditions, making it difficult to accurately identify the road surface conditions in bad weather on highways. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for identifying road surface conditions in bad weather from highway surveillance videos, which can achieve real-time automatic identification of bad weather on highways, so as to solve the problems that in the automatic identification of highway weather conditions, multiple weather conditions interfere with each other, and the quality of surveillance video images is affected by weather conditions, making it difficult to accurately identify the road surface conditions in bad weather on highways, thereby providing technical support for highway management and traffic control by the transportation department.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A method for identifying road surface conditions in bad weather from highway surveillance videos, the technical solution thereof is that it includes the following steps: S1: Obtain the real monitoring video data of the highway, and process the video image data into a dataset of labels for bad weather on the highway road, mainly including five types of labels that affect the road, namely dry, waterlogged, icy, snowy, and foggy, so as to construct a dataset of images of bad weather in highway monitoring videos; specifically, obtain the video data in the highway monitoring network, and download the video data of five types of road surface conditions, namely dry, waterlogged, icy, snowy, and foggy, required for the research, providing basic data support for the research on the road surface conditions of the highway in this invention.
[0008] S2: Design a hybrid perception fusion network model for weather element features to extract the features of the label data of bad weather on the highway road, and output the probabilities of five types of road conditions, namely dry, waterlogged, icy, snowy, and foggy.
[0009] S3: Design the learning rate, optimizer, and loss function for the convergence of the network model.
[0010] S4: Use the dataset of images of bad weather in highway monitoring videos constructed in step S1, input the label dataset into the model, and adopt the SGD optimizer and cross-entropy loss function to realize the parameter optimization and update of the network model until the model finally converges; this step completes the training of the model and obtains the hyperparameters and optimal model weight parameters.
[0011] S5: Save the hyperparameters and optimal model weight parameters obtained in step S4, and load them into the model constructed in step S2 to obtain the optimal parameters and its neural network model structure for identifying five types of road conditions, namely dry, waterlogged, icy, snowy, and foggy, and construct a method for identifying bad weather road surface conditions in highway monitoring videos; S6: Input the highway monitoring video image data into the method for identifying bad weather road surface conditions in highway monitoring videos constructed in step S5 for identifying the features of five types of road conditions, namely dry, waterlogged, icy, snowy, and foggy, complete the accurate identification of five types of road conditions, namely dry, waterlogged, icy, snowy, and foggy, and finally output the recognition result of the road state picture; The above steps S1, S2, S3, and S4 are image training methods, and steps S5 and S6 are image recognition methods.
[0012] In the above technical solution, the preferred technical solution may be that step S1 specifically includes: S1.1: Obtain the real monitoring video data in bad weather from the highway monitoring network, including the road surface conditions of the highway in different seasons and different time periods. To ensure the image quality, one frame of image needs to be extracted from the obtained video data every 5 minutes, and the images with poor quality are excluded from the extracted images to obtain a dataset of images of highway road surface conditions without labels.
[0013] S1.2: Label the corresponding tags for the unlabeled highway pavement condition image data to form a dataset of five types of pavement condition tags: dry, waterlogging, icing, snow accumulation, and heavy fog. Then, screen the video data with rich pavement feature conditions for dry, waterlogging, icing, snow accumulation, and heavy fog.
[0014] The further operation steps of step S1 include: S101: Use the VideoCapture() function in the opencv-python library to read the surveillance video data and obtain the total number of frames, frame rate, image height, and image width information of the video data; S102: Extract one frame of image by setting a 5-minute time interval, and save each frame of the image in JPG format to complete the conversion of the highway surveillance video data into image data; S103: The images need to be further screened to remove those with large noise and poor image quality, calculate the pixel differences corresponding to the picture frames, sort them according to the difference size as needed, and save the pictures with large differences as valid images; S104: For the unlabeled image data obtained in step S103, manually label the image tags according to the proportion of liquid water and solid water on the road surface for dry, waterlogging, icing, and snow accumulation, and according to the road surface visibility for heavy fog. When the liquid and solid water accumulation on the road surface is both 0, there is no waterlogging on the road surface at this time, and the road surface condition is dry. When the water film on the road surface is greater than 1 kg / m 2 runoff will occur. Therefore, use the liquid water greater than 1 kg / m 2 as the numerical judgment for waterlogging. When solid water accumulation appears on the ground, that is, when the solid water accumulation is greater than 0, regardless of whether there is liquid water accumulation on the ground, the road surface condition is defined as ice / snow, including the ice (snow) water mixture state and the icing (snow accumulation) state. Fog is the suspension of a large number of small water droplets or ice crystal particles in the near-surface air, making people's line of sight blurred, and the horizontal visibility distance of the person drops below 1000 meters.
[0015] In the above technical solution, a preferred technical solution may further be that in step S2, a weather element feature hybrid perception fusion network model is designed based on deep learning, which can effectively focus on the differences in images of different weather phenomena in the highway monitoring scenario. The main feature extraction structure of the weather element feature hybrid perception fusion network model mainly includes a hybrid perception module and a deep context feature fusion module. The hybrid perception feature fusion module is mainly composed of Regional-to-Local Attention, Neighborhood Attention, and HaloAttention. Regional-to-Local Attention divides the image into multiple regions by introducing the idea of image segmentation, processes each region as input, divides the picture into multiple non-overlapping large blocks, and each large block can be further divided into small blocks. Then, the large blocks first perform self-attention to obtain global information, and then the updated large block information is exchanged with the small block information it belongs to, so that the small blocks can have global information. The purpose of this operation is to retain both the global and local information in the image, be able to process images of different scales and aspect ratios, and output regional features and local features. Neighborhood Attention is based on improving traditional self-attention. This mechanism limits the attention range of each element to a specific area near it. Neighborhood Attention not only reduces the computational cost but also can better capture short-distance dependencies while maintaining the ability to understand long-distance dependencies. HaloAttention uses self-attention to capture the correlation between pixels, effectively processes multi-scale feature maps, and perceives key image-dependent features. The input regional features are added to the neighborhood features processed by Neighborhood Attention, pixel fine-grained features, and the initial regional features. While learning the neighborhood features and pixel fine-grained features, the important information of the original features is retained. The weather element feature hybrid perception fusion network model is inspired by the RegionViT model. The weather element feature hybrid perception fusion network model first passes the feature information to the main feature extraction structure through two groups of tokens through the regional-to-local transformer encoder. The main feature extraction structure is composed of four stacked hybrid perception feature modules to form a pyramid structure to generate multi-scale features. Each layer of the pyramid structure will receive the regional features (RegionalFeature) and local features (Local Feature) of the previous stage, and use the downsampling process to halve the spatial resolution while doubling the channel size on the regional and local tokens, and then enter the next stage. And all local tokens are used in each stage to provide more fine-grained position information. The multi-scale features output by the pyramid structure are fused by the deep context feature fusion module for regional features and local features. The fused feature tokens are used as the final embedding features for classification, and the classification probability of each category is output. Finally, the highway pavement condition classification result is obtained.
[0016] In the above technical solution, a preferred technical solution may further be that in step S3, the learning rate (LearningRate), the training batch size (Batchsize), the SGD optimizer, the loss function, and the total number of iterations (Epoch) of the training model are set as hyperparameters of the experiment. The loss function is Cross Entropy Loss. A framework for the training model is constructed, and the weather element feature hybrid perception fusion network model is trained to converge to obtain the optimal weights and model. The main process is as follows: data preparation, model definition, optimizer definition, loss function definition, loop training, model evaluation and saving, and model adjustment. Step S3 specifically includes: S3.1: Data preparation: The image label data is input into the model and converted into structured data that meets the requirements, and the images and corresponding label information that the input model can adapt to are input.
[0017] S3.2: Constructing the training loop: When training the model, a training loop needs to be constructed to iteratively train the model. In each training step, the input data and labels need to be provided, and the loss function of the model is calculated.
[0018] S3.3: Optimizer selection: Select a suitable optimizer to update the weights of the model. Commonly used optimizers include Stochastic Gradient Descent (SGD) and Adam.
[0019] S3.4: Training the model: Use the training loop and optimizer to train the model. During the training process, the performance of the model can be improved by adjusting hyperparameters, using different activation functions, etc.
[0020] S3.5: Evaluating the model: After the training is completed, the performance of the model needs to be evaluated. The test dataset can be used to test indicators such as the accuracy, precision, and recall rate of the model.
[0021] To quantitatively evaluate the effectiveness of the algorithm, three commonly used evaluation indicators are the overall accuracy (Accuracy), precision (Precision), and recall rate (Recall).
[0022] Accuracy: It is an indicator to measure the global accuracy, indicating the percentage of correctly predicted road samples in the total number of samples. It is calculated using the following formula:
[0023] In the above formula, TP is the positive sample predicted as the positive class (True Positive), FP is the negative sample predicted as the positive class (False Positive), TN is the negative sample predicted as the negative class (True Negative), and FN is the positive sample predicted as the negative class (False Negative).
[0024] Precision: This evaluation metric is mainly used to reflect the percentage of accurate classifications in the classification results of different weather categories by the model. It represents the proportion of samples predicted as correctly classified as this weather phenomenon. It can be expressed in the following way:
[0025] Recall: This evaluation metric represents the percentage of correctly predicted weather classification results in the total number of this category. Recall can be expressed in the following way:
[0026] S3.6: Adjust the model: According to the evaluation results, the parameters or structure of the model can be adjusted to improve the performance of the model.
[0027] S3.7: Training the weather element feature hybrid perception fusion network model requires data preparation, defining the model, defining the optimizer, defining the loss function, loop training, model evaluation and saving, and adjusting the model. By continuously adjusting and optimizing, the performance of the model can be improved to better meet the requirements of practical applications.
[0028] In the above technical solutions, the preferred technical solution can also be that step S4 specifically includes: S4.1: Use the highway surveillance video severe weather image dataset constructed in step S1, and randomly divide this labeled dataset into a training set, a validation set, and a test set. The division ratio of the training set, the validation set, and the test set is 7:1:2.
[0029] S4.2: Input the training set and the validation set into the model, and use the SGD optimizer and the CrossEntropy Loss function to implement the parameter optimization and update of the network model until the model finally converges, and finally obtain the optimal model weight parameters.
[0030] S4.3: Use the test set that has not participated in model training to verify and evaluate the weather element feature hybrid perception fusion network model designed in step S2, obtain the classification accuracy result, and obtain the optimal model.
[0031] In step S4, according to the process of step S4, DataSet and DataLoader of the deep learning framework Pytorch are used for data preparation to handle data loading and batch processing; the model is defined mainly according to the model structure designed in step S3, inheriting the Module in the Pytorch framework, and defining the initial base class __init__() and the propagation layer forward() of the model; the optimizer is defined as SGD; the loss function is defined as the cross-entropy loss function Cross Entropy Loss; the loop training is carried out by setting the model to train() and continuously performing forward propagation and backward propagation to train the model, the model evaluation sets the model to eval() to evaluate the results of each round, and the model is saved by saving the parameter information of the model.
[0032] The further operation steps of step S4 include: S401: For data preparation, a custom data loading method is defined according to data characteristics, which is mainly divided into: Reading: Read the original data from disk or network; Preprocessing: including cleaning, conversion, normalization, etc.; Batch processing: Organize the data in batches for parallel processing; Loading: Load the data into memory and pass it to the model, and load the data from the dataset by creating an instance of DataLoader of the deep learning framework Pytorch; S402: Select a suitable optimizer to update the weights of the model. The optimizers used include Stochastic Gradient Descent (SGD) to reduce the gap between the model prediction and the actual result. SGD is a variant of the gradient descent algorithm. Its core principle is that in each iteration search, the algorithm randomly selects a sample or data point (or a small batch of samples), calculates the gradient of this sample, and then updates the model parameters with this gradient; S403: Define the model, inherit the Module class of the model, and it is necessary to initialize the initial base class __init__() and the propagation layer forward() of the model; S404: When training the model, a training loop needs to be constructed to iteratively train the model. In each training step, it is necessary to provide the input data and labels, and calculate the loss function of the model. The loss function selects the cross-entropy loss function Cross Entropy Loss: , where M is the number of categories; is the sign function (0 or 1). If the sample i 's true category is equal to c take 1, otherwise take 0; is the observed sample i belongs to the category cPrediction probability.
[0033] S405: Use the training loop and optimizer to train the model. During the training process, use the train() method to set the model to the training state. The train() method is used to enable dropout, batch normalization, and other training-specific operations when training a neural network. This method notifies the model to perform backpropagation and update the weights and biases of the model. By enabling these specific training operations, the model can better learn and adapt to the data during the training process, thereby improving the generalization ability of the model.
[0034] S406: After the training is completed, it is necessary to evaluate the performance of the model. Set the model to the validation state through the eval() method, stop the update of the parameters during the validation process, and use the test data set to test the accuracy, precision, recall, and other metrics of the model, and output the results of the model training in real time.
[0035] S407: According to the evaluation results, the parameters or structure of the model can be adjusted to improve the performance of the model.
[0036] In the above technical solution, the preferred technical solution may also be that in steps S5 and S6: Input the obtained real highway monitoring image into a method for identifying the road surface conditions in bad weather in highway monitoring videos constructed in step S5 to achieve high-precision classification of five types of road conditions on the highway, namely dry, waterlogged, icy, snowy, and foggy road surfaces, and output the recognition results of the road state pictures. In this way, through the monitoring videos arranged on the highway, the target images of the section to be recognized are obtained; the target images are input into the model trained in step S4, and the classification results are output.
[0037] The present invention provides a method for identifying the road surface conditions in bad weather in highway monitoring videos, which can realize the real-time automatic identification of bad weather on highways, proposes a weather element feature hybrid perception fusion network model, and improves the classification accuracy of road conditions affected by five types of bad weather. Through experiments, compared with the deep learning recognition method, the present invention improves the accuracy of the classification results, enhances the recognition efficiency, and effectively realizes the automatic and accurate classification of images of bad weather phenomena on highways (such as a certain province). The present invention solves the problem that in the automatic identification of highway weather conditions, multiple weather conditions interfere with each other, and the quality of the monitoring video images is affected by the weather conditions, making it difficult to accurately identify the road surface conditions in bad weather on highways, thereby providing technical support for the highway management and traffic control of the transportation department.
[0038] The present invention mainly uses a dataset of highway surveillance video images in bad weather, which is randomly divided into a training set, a validation set, and a test set according to the ratio of 7:1:2. In the experiments of the present invention, classic networks were selected from different perspectives for comparison of experimental effects. Mainly, a series of classic recognition network models of Transformer were compared, including Visual Transformer (Vit), RegionVit, and MaxVit. In addition, the classic convolutional networks ResNet101 and DesNet101 were also used. In addition, the lightweight networks EfficientNet and ShuffleNet were selected. To ensure the reliability of the experimental results, multiple experiments were conducted to calculate the average value. In this study, the dataset of highway surveillance video images in bad weather was used to test the invention model to evaluate the robustness and reliable engineering performance of the network model. Through comparative analysis of classic Transformer network models, compared with the best MaxVit model proposed by the present invention, the accuracy rate and recall rate were increased by 0.9% and 1.2% respectively. Compared with the benchmark network model RegionVit, the accuracy rate and recall rate classification were increased by 1.8% and 2.1% respectively. Compared with the convolutional network model EfficientNet with the best experimental results, the accuracy rate and recall rate were increased by 1.2% and 2.2% respectively. In addition, compared with the RegionVit model, the number of parameters of the model of the present invention increased by 10.46M, and the accuracy result was significantly improved. Compared with the lightweight network models EfficientNet and ShuffleNet in terms of the number of model parameters, the model of the present invention discarded the parameter load and improved the accuracy of the model in recognition. Compared with the experimental results of the traditional convolutional network ResNet101, the model of the present invention significantly reduced the parameter load, improved the accuracy, and reduced the complexity at the same time. The comparative experimental results fully prove the good characteristics of the model of the present invention in the recognition of highway bad weather. See Table 1.
[0039] Table 1
[0040] The experimental process ensured that the experimental environments of the comparison models were all kept consistent and were all carried out in the same server. The server device parameters were two Intel Xeon(R) Gold 5218 processors and one 4090 graphics card. The hyperparameter settings of the experiment were that the learning rate was 0.001, the training batch size was 8, the optimizer was SGD, and the total number of iterations of the experiment was 100 rounds. Description of the Drawings
[0041] Figure 1 It is a flowchart of the method for recognizing the road conditions in bad weather of highway surveillance video of the present invention.
[0042] Figure 2 This is the structural diagram of the weather element feature hybrid perception fusion network model provided by the present invention.
[0043] Figure 3 This is the training flowchart of the weather element feature hybrid perception fusion network model provided by the present invention. Detailed implementation manners
[0044] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Based on these embodiments, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of the present invention.
[0045] Embodiment 1: As Figure 1 、 Figure 2 、 Figure 3 shown, the method for identifying the road surface conditions in bad weather from the highway surveillance videos of the present invention includes steps S1, S2, S3, S4, S5, and S6. Among them, steps S1, S2, S3, and S4 are image training methods, and steps S5 and S6 are image recognition methods.
[0046] Specifically, the method for identifying the road surface conditions in bad weather from the highway surveillance videos of the present invention includes the following steps: S1: Obtain the real highway surveillance video data, and process the video image data into a highway road bad weather label data set, mainly including five types of labels that affect the road, namely dry, water accumulation, ice accumulation, snow accumulation, and heavy fog, so as to construct a highway surveillance video bad weather image data set; specifically, obtain the video data in the highway surveillance network, and download the video data of the five types of road surface conditions of dry, water accumulation, ice accumulation, snow accumulation, and heavy fog required for the research, providing basic data support for the research on the road surface conditions of the highway of the present invention. Step S1 specifically includes: S1.1: Obtain the real highway surveillance video data in bad weather from the highway surveillance network, including the road surface conditions of the highway at different seasons and different time periods. In order to ensure the image quality, one frame of image needs to be extracted from the obtained video data every 5 minutes, and the images with poor quality are removed from the extracted images to obtain an unlabeled highway road surface condition image data set.
[0047] S1.2: Label the corresponding labels for the unlabeled highway road surface condition image data to form a label data set for the five types of road surface conditions of dry, water accumulation, ice accumulation, snow accumulation, and heavy fog. Thus, the video data with rich road surface feature conditions of dry, water accumulation, ice accumulation, snow accumulation, and heavy fog is screened.
[0048] The further operation steps of step S1 include: S101: Use the VideoCapture() function in the opencv - python library to read the monitoring video data, and obtain the total number of frames, frame rate, image height, and image width information of the video data; S102: Extract one frame of image by setting a 5 - minute time interval, and save each frame of the image in JPG format, completing the conversion of highway monitoring video data into image data; S103: For the images to be further processed, filter out those with large noise and poor image quality, calculate the pixel differences corresponding to the image frames, sort them according to the difference size as needed, and save the valid images with large differences; S104: For the unlabeled image data obtained in step S103, manually label the image tags according to dry, water accumulation, ice accumulation, snow accumulation, which are based on the ratio of liquid water and solid water on the road surface, and fog is based on the road surface visibility. When both the liquid and solid water accumulation amounts on the road surface are 0, there is no water accumulation on the road surface at this time, and the road surface condition is dry. When the water film on the road surface is greater than 1 kg / m 2 will runoff occur. Therefore, use the liquid water greater than 1 kg / m 2 as the numerical judgment for water accumulation. When there is solid water accumulation on the ground, that is, when the solid water accumulation amount is greater than 0, regardless of whether there is liquid water accumulation on the ground, the road surface condition is defined as ice / snow, including the ice (snow) - water mixed state and the ice - forming (snow - forming) state. Fog is the suspension of a large number of small water droplets or ice crystal particles in the near - low - level air, making people's line of sight blurred, and the horizontal visibility distance of the person drops to less than 1000 meters.
[0049] S2: Design a hybrid perception fusion network model for weather element features to extract the features of the label data of bad weather on highway roads, and output the probabilities of five types of road conditions: dry, waterlogged, icy, snowy, and foggy. In step S2, based on deep learning, a hybrid perception fusion network model for weather element features is designed, which can effectively focus on the differences in images of different weather phenomena in the highway monitoring scenario. The main feature extraction structure of the hybrid perception fusion network model for weather element features mainly includes a hybrid perception module and a deep context feature fusion module. The hybrid perception feature fusion module is mainly composed of Regional-to-Local Attention, Neighborhood Attention, and Halo Attention. Regional-to-Local Attention introduces the idea of image segmentation, divides the image into multiple regions, processes each region as an input, divides the picture into multiple non-overlapping large blocks, and each large block can be further divided into small blocks. Then, the large blocks first perform self-attention to obtain global information, and then the updated large block information is exchanged with the small block information to which it belongs, so that the small blocks can have global information. The purpose of this operation is to simultaneously retain the global and local information in the image, and can process images of different scales and aspect ratios, and output regional features and local features. Neighborhood Attention is based on improving traditional self-attention. This mechanism restricts the attention range of each element to a specific area near it. Neighborhood Attention not only reduces the computational cost, but also can better capture short-distance dependencies while maintaining the ability to understand long-distance dependencies. Halo Attention uses self-attention to capture the correlation between pixels, effectively processes multi-scale feature maps, and perceives key image-dependent features. Add the neighborhood features processed by Neighborhood Attention, the pixel fine-grained features, and the initial regional features of the input regional features. While learning the neighborhood features and pixel fine-grained features, the important information of the original features is retained.The hybrid perception fusion network model for weather element feature is inspired by the RegionViT model. First, the hybrid perception fusion network model for weather element feature passes feature information to the main feature extraction structure through two groups of tokens via the region-to-local transformer encoder. The main feature extraction structure consists of four stacked hybrid perception feature modules, forming a pyramid structure to generate multi-scale features. Each layer of the pyramid structure will receive the regional feature and local feature from the previous stage, and use the downsampling process to halve the spatial resolution while doubling the channel size on the regional and local tokens, and then enter the next stage. Using all local tokens at each stage provides more fine-grained location information. The multi-scale features output by the pyramid structure are fused with the regional feature and local feature by the deep context feature fusion module, and the fused feature tokens are used as the final embedding features for classification, outputting the classification probability for each category, and finally obtaining the classification result of the highway pavement condition.
[0050] S3: Design the learning rate, optimizer, and loss function for the convergence of the network model. In step S3, set the learning rate, training batch size, SGD optimizer, loss function, and the total number of iterations (Epoch) of the training model as the hyperparameters of the experiment. The loss function is Cross Entropy Loss. Build the framework of the training model, and train the hybrid perception fusion network model for weather element feature to converge to obtain the optimal weights and model. The main processes are: data preparation, model definition, optimizer definition, loss function definition, loop training, model evaluation and saving, and model adjustment. Step S3 specifically includes: S3.1: Data preparation: Input the image label data into the model, convert it into structured data that meets the requirements, and input the images and corresponding label information that the model can adapt to; S3.2: Build the training loop: When training the model, a training loop needs to be built to iteratively train the model. In each training step, input data and labels need to be provided, and the loss function of the model needs to be calculated; S3.3: Optimizer selection: Select a suitable optimizer to update the weights of the model. Commonly used optimizers include Stochastic Gradient Descent (SGD), Adam; S3.4: Train the model: Use the training loop and optimizer to train the model. During the training process, the performance of the model can be improved by adjusting hyperparameters, using different activation functions, etc.; S3.5: Evaluate the model: After the training is completed, the performance of the model needs to be evaluated. The test dataset can be used to test indicators such as the accuracy, precision, and recall rate of the model; To quantitatively evaluate the effectiveness of the algorithm, three commonly used evaluation metrics are the overall accuracy (Accuracy), precision (Precision), and recall (Recall).
[0051] Accuracy: It is a metric to measure the global accuracy, representing the percentage of correctly predicted road samples in the total number of samples. It is calculated using the following formula:
[0052] In the above formula, TP is the positive sample predicted as the positive class (True Positive), FP is the negative sample predicted as the positive class (False Positive), TN is the negative sample predicted as the negative class (True Negative), and FN is the positive sample predicted as the negative class (False Negative).
[0053] Precision: This evaluation metric is mainly used to reflect the percentage of accurately classified results in the classification results of different weather categories of the model. It represents the proportion of samples predicted to be correctly classified as this weather phenomenon. It can be expressed in the following way:
[0054] Recall: This evaluation metric represents the percentage of correctly predicted weather classification results in the total number of this category. Recall can be expressed in the following way:
[0055] S3.6: Adjust the model: According to the evaluation results, the parameters or structure of the model can be adjusted to improve the performance of the model.
[0056] S3.7: Training the weather element feature hybrid perception fusion network model requires data preparation, defining the model, defining the optimizer, defining the loss function, loop training, model evaluation and saving, and adjusting the model. By continuously adjusting and optimizing, the performance of the model can be improved to better meet the needs of practical applications.
[0057] S4: Use the highway surveillance video adverse weather image dataset constructed in step S1. Input the label dataset into the model, and adopt the SGD optimizer and cross-entropy loss function to optimize and update the parameters of the network model until the model finally converges. This step completes the training of the model and obtains the hyperparameters and the optimal model weight parameters. Step S4 specifically includes: S4.1: Use the highway surveillance video adverse weather image dataset constructed in step S1. Randomly divide the label dataset into a training set, a validation set, and a test set, and the division ratio of the training set, the validation set, and the test set is 7:1:2.
[0058] S4.2: Input the training set and the validation set into the model, and use the SGD optimizer and the CrossEntropy Loss function to optimize and update the parameters of the network model until the model finally converges, and finally obtain the optimal model weight parameters.
[0059] S4.3: Use the test set that has not participated in model training to verify and evaluate the weather element feature hybrid perception fusion network model designed in step S2, obtain the classification accuracy result, and get the optimal model.
[0060] In step S4, according to the process of step S4, the DataSet and DataLoader of the deep learning framework Pytorch are used for data preparation to handle data loading and batch processing; the model is defined mainly according to the model structure designed in step S3, inheriting the Module in the Pytorch framework, and defining the initial base class __init__() and the propagation layer forward() of the model; the optimizer is defined as SGD; the loss function is defined as the Cross Entropy Loss function; the loop training is carried out by setting the model to train(), continuously performing forward propagation and backward propagation to train the model, the model evaluation sets the model to eval(), evaluates the results of each round, and the model is saved by saving the parameter information of the model.
[0061] The further operation steps of step S4 include: S401: Customize the data loading method according to the data characteristics for data preparation, mainly including: Reading: Read the original data from the disk or network; Preprocessing: including cleaning, conversion, normalization, etc.; Batch processing: Organize the data in batches for parallel processing; Loading: Load the data into the memory and pass it to the model, and load the data from the dataset by creating an instance of the DataLoader of the deep learning framework Pytorch; S402: Select a suitable optimizer to update the weights of the model. The optimizers used include Stochastic Gradient Descent (SGD) to reduce the gap between the model prediction and the actual result. SGD is a variant of the gradient descent algorithm. Its core principle is that in each iteration search, the algorithm randomly selects a sample or data point (or a small batch of samples), calculates the gradient of this sample, and then uses this gradient to update the model parameters; S403: Define the model, inherit the Module class of the model, and it is necessary to initialize the initial base class __init__() and the propagation layer forward() of the model; S404: When training the model, a training loop needs to be constructed to iteratively train the model. In each training step, input data and labels need to be provided, and the loss function of the model is calculated. The loss function selected is the Cross Entropy Loss: , where M is the number of classes; is the sign function (0 or 1). If the true class of the sample i is equal to c , take 1, otherwise take 0; is the observed sample i belonging to the class c 's predicted probability; S405: Use the training loop and optimizer to train the model. During the training process, use the train() method to set the model to the training state. The train() method is used to enable dropout, batch normalization, and other training-specific operations when training a neural network. This method will notify the model to perform backpropagation and update the weights and biases of the model. By enabling these specific training operations, the model can better learn and adapt to the data during the training process, thereby improving the generalization ability of the model; S406: After the training is completed, the performance of the model needs to be evaluated. Set the model to the validation state through the eval() method, stop the update of the parameters during the validation process, use the test dataset to test the accuracy, precision, recall, and other metrics of the model, and output the results of the model training in real time; S407: According to the evaluation results, the parameters or structure of the model can be adjusted to improve the performance of the model.
[0062] S5: Save the hyperparameters and the optimal model weight parameters obtained in step S4, and load them into the model constructed in step S2 to obtain the optimal parameters and its neural network model structure for identifying five types of road conditions: dry, waterlogged, icy, snowy, and foggy, and construct a method for identifying poor weather road conditions in highway surveillance videos; S6: Input the highway surveillance video image data into the method for identifying poor weather road conditions in highway surveillance videos for identifying the characteristics of five types of road conditions: dry, waterlogged, icy, snowy, and foggy in step S5, complete the accurate identification of five types of road conditions: dry, waterlogged, icy, snowy, and foggy, and finally output the road state picture recognition result; in step S5 : The obtained real highway monitoring images are input into a method for identifying road conditions in bad weather in highway monitoring videos constructed in step S5, achieving high-precision classification of five types of road conditions on highway pavements, namely dry, waterlogged, icy, snow-covered, and foggy, and outputting the recognition results of road condition pictures. In this way, through the monitoring videos deployed on the highway, the target images of the section to be recognized are obtained; the target images are input into the model trained in step S4, and the classification results are output.
[0063] The purpose of the present invention is to provide a method for identifying road conditions in bad weather in highway monitoring videos, which can realize real-time automatic identification of bad weather on highways, so as to solve the problem that in the automatic identification of highway weather conditions, multiple weather conditions interfere with each other, and the quality of monitoring video images is affected by weather conditions, making it difficult to accurately identify the road conditions in bad weather on highways, thereby providing technical support for the highway management and traffic control of the transportation department.
Claims
1. A method for identifying bad weather road conditions using highway monitoring video, characterized in that It includes the following steps: S1: Obtain real highway surveillance video data, process the video image data into a highway road severe weather label dataset, including five types of labels affecting roads: dryness, water accumulation, ice accumulation, snow accumulation, and fog, so as to construct a highway surveillance video severe weather image dataset; S2: Design a hybrid perception fusion network model of weather element features to extract the features of the label data of severe weather on highways and output the probabilities of five types of road conditions: dry, waterlogged, iced, snowed, and foggy. S3: Design learning rate, optimizer, and loss function for network model convergence; S4: using the highway surveillance video severe weather image dataset constructed in step S1, inputting the label dataset into the model, and using the SGD optimizer and the cross entropy loss function to optimize and update the parameters of the network model until the model finally converges; this step completes the training of the model and obtains the hyperparameters and the optimal model weight parameters; S5: Save the hyperparameters and optimal model weight parameters obtained in step S4, and load them into the model constructed in step S2, so as to obtain the optimal parameters and neural network model structure for identifying five types of road conditions: dry, water-logged, ice-logged, snow-logged and foggy, and construct a method for identifying bad weather road conditions from highway monitoring videos; S6: inputting the highway monitoring video image data into the method for identifying the road condition in bad weather by using the highway monitoring video in step S5, in which the five road condition characteristics of dryness, water accumulation, ice accumulation, snow accumulation and fog are identified, completing the accurate identification of the five road conditions of dryness, water accumulation, ice accumulation, snow accumulation and fog, and finally outputting the road state image recognition result; The above steps S1, S2, S3 and S4 are image training methods, and steps S5 and S6 are image recognition methods.
2. The method for identifying bad weather road conditions using highway monitoring video according to claim 1 is characterized in that Step S1 comprises: S1.1: Obtain real surveillance video data under bad weather conditions from the highway monitoring network, extract a frame of image every 5 minutes, remove the images with poor quality from the extracted images, and obtain an unlabeled highway road condition image dataset; S1.2: Label the unlabeled highway road condition image data with corresponding labels to form a label dataset of five types of road conditions: dry, waterlogged, iced, snowed, and foggy. This allows for the selection of video data with rich characteristic road conditions: dry, waterlogged, iced, snowed, and foggy.
3. The method for identifying bad weather road conditions using highway monitoring video according to claim 1 is characterized in that In step S2, the weather element feature hybrid perception fusion network model first passes the feature information to the main feature extraction structure through two groups of tokens through the regional to local converter encoder. The main feature extraction structure is composed of stacking four layers of hybrid perception feature modules to form a pyramid structure to generate multi-scale features. Each layer of the pyramid structure will accept the regional features and local features of the previous stage, and use the downsampling process to halve the spatial resolution, while doubling the channel size on the regional and local labels before entering the next stage. All local labels are used in each stage to provide more fine-grained location information. The multi-scale features output by the pyramid structure are fused with regional features and local features by the deep context feature fusion module. The fused feature labels are used as the final embedded features for classification, and the classification probability of each category is output. Finally, the classification results of the highway pavement conditions are obtained.
4. The method for identifying bad weather road conditions using highway monitoring video according to claim 1 is characterized in that In step S3, the learning rate, training batch size, SGD optimizer, loss function, and total number of iterations (Epoch) of the training model are set as the hyperparameters of the experiment. The loss function is Cross Entropy Loss. The framework of the training model is constructed, and the weather element feature hybrid perception fusion network model is trained to converge to obtain the optimal weights and model. The process is: data preparation, model definition, optimizer definition, loss function definition, cyclic training, model evaluation and preservation, and model adjustment.
5. The method for identifying bad weather road conditions using highway monitoring video according to claim 4 is characterized in that Step S3 includes: S3.1: Data preparation: Image label data is input into the model and converted into structured data that meets the requirements. The input model can adapt to the image and corresponding label information. S3.2: Construct a training loop: When training a model, you need to construct a training loop to iteratively train the model. In each training step, you need to provide input data and labels, and calculate the loss function of the model. S3.3: Optimizer selection: Select a suitable optimizer to update the weights of the model. Optimizers include stochastic gradient descent (SGD) and Adam. S3.4: Training model: Use training loops and optimizers to train the model. During the training process, adjust hyperparameters and use different activation functions to improve the performance of the model. S3.5: Evaluate the model: After training, you need to evaluate the performance of the model and use the test data set to test the accuracy, precision, and recall of the model. S3.6: Adjust the model: Based on the evaluation results, adjust the parameters or structure of the model to improve the performance of the model; S3.7: Training the hybrid perception fusion network model of weather element characteristics requires data preparation, model definition, optimizer definition, loss function definition, cyclic training, model evaluation and preservation, and model adjustment. Through continuous adjustment and optimization, the performance of the model can be improved to better meet the needs of practical applications.
6. The method for identifying bad weather road conditions using highway monitoring video according to claim 1 is characterized in that Step S4 includes: S4.1: Using the highway surveillance video severe weather image dataset constructed in step S1, the labeled dataset is randomly divided into a training set, a validation set, and a test set, and the division ratio of the training set, the validation set, and the test set is 7:1:2; S4.2: Input the training set and validation set into the model, use the SGD optimizer and the cross-entropy loss function to optimize and update the parameters of the network model until the model finally converges and finally obtains the optimal model weight parameters; S4.3: Use the test set that has not participated in the model training to verify the weather element feature hybrid perception fusion network model designed in the evaluation step S2, obtain the classification accuracy results, and obtain the optimal model.
7. The method for identifying bad weather road conditions using highway monitoring video according to claim 1 is characterized in that In step S4, according to the process of step S4, data preparation uses DataSet and DataLoader of the deep learning framework Pytorch to handle data loading and batch processing; the model is defined mainly according to the model structure designed in step S3, inheriting the Module in the Pytorch framework, and defining the model's initial base class __init__() and the model's propagation layer forward(); the optimizer is defined as SGD; the loss function is defined as the cross entropy loss function Cross Entropy Loss; the cyclic training sets the model to train(), continuously performs forward propagation and back propagation training on the model, and the model evaluation sets the model to eval() to evaluate the results of each round, and the model is saved by saving the model's parameter information.
8. The method for identifying bad weather road conditions using highway monitoring video according to claim 7 is characterized in that Step S4 further comprises the following steps: S401: Data preparation: Customize the data loading method based on data characteristics, which can be divided into: Read: read raw data from disk or network; Preprocessing: including cleaning, conversion, normalization, etc.; Batch processing: Organize data into batches for parallel processing; Loading: Load the data into memory and pass it to the model. Load the data from the dataset by creating a DataLoader instance of the deep learning framework Pytorch. S402: Select a suitable optimizer to update the weight of the model. The optimizer includes stochastic gradient descent (SGD) to reduce the gap between the model prediction and the actual result. SGD is a variant of the gradient descent algorithm. Its core principle is that in each iterative search, the algorithm randomly selects a sample or data point, calculates the gradient of the sample, and then uses this gradient to update the model parameters. S403: Define the model, inherit the Module class of the model, and initialize the initial base class __init__() and the propagation layer forward() of the model; S404: When training the model, you need to build a training loop to iteratively train the model. In each training step, you need to provide input data and labels, and calculate the loss function of the model. The loss function is the cross entropy loss function: , M is the number of categories; is a sign function (0 or 1), if the sample i The true category is equal to c Take 1, otherwise take 0; For the observed sample i Belongs to category c The predicted probability of S405: Use the training loop and optimizer to train the model. During the training process, use the train() method to set the model to the training state. The train() method is used to enable dropout, batchnormalization, and other training-specific operations when training a neural network. This method notifies the model to perform backpropagation and update the model's weights and biases. By enabling these specific training operations, the model can better learn and adapt to the data during the training process, thereby improving the model's generalization ability. S406: After the training is completed, the performance of the model needs to be evaluated. The model is set to the verification state through the eval() method. During the verification process, the parameter update is stopped. The accuracy, precision, recall and other indicators of the model are tested using the test data set, and the results of the model training are output in real time. S407: According to the evaluation results, adjust the parameters or structure of the model to improve the performance of the model.
9. The method for identifying bad weather road conditions using highway monitoring video according to claim 1, characterized in that In steps S5 and S6: The acquired real highway monitoring image is input into a method for highway monitoring video recognition of severe weather road conditions constructed in step S5, so as to achieve high-precision classification of five types of highway road conditions: dry road surface, water accumulation, ice accumulation, snow accumulation and heavy fog, and output the road status image recognition result.
Citation Information
Patent Citations
Multi-parameter analysis remote sensing type pavement meteorological condition detector and detection method
CN113917564A
Construction method and device of model for enhancing expressway monitoring video definition
CN118469837A
Expressway severe weather identification method based on artificial intelligence
CN110866593A
Expressway severe weather identification method based on multi-scale fusion network
CN113392818A
Road monitoring weather video recognition method based on deep learning
CN117351394A