Road condition monitoring and recognition method, device, equipment and medium based on machine vision

Through multi-source data fusion and intelligent processing, the data dependence and identification capabilities of the existing road condition monitoring system are solved, efficient perception and precise regulation of complex road conditions are achieved, and road maintenance costs are reduced.

CN119445503BActive Publication Date: 2025-07-04SHANGHAI ZZLD INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411487566.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2025-07-04
Estimated Expiration
2044-10-24

AI Technical Summary

Technical Problem

The existing road condition monitoring system relies on a single data source to be easily affected by weather and light, lacks the ability to fusion multiple sources of data, lacks recognition capabilities in complex scenarios, and lacks intelligent decision-making support, making it difficult to generate optimized traffic control strategies.

Method used

Multi-source data acquisition, dynamic time regularization and multi-modal feature fusion, combined with multi-scale pyramid decomposition and deep feature extraction, semantic analysis is performed using feature mapping and knowledge graph enhancement, multi-level risk quantification and space-time propagation simulation, and finally multi-objective trade-offs and causal chain reasoning are carried out to generate adaptive traffic regulation instructions.

Benefits of technology

It improves the comprehensiveness and reliability of road condition monitoring, enhances the perception of complex road conditions, improves the accuracy of feature extraction and the accuracy of semantic understanding, can predict risk evolution trends, and finds the optimal balance between safety, traffic efficiency and maintenance costs, and generates accurate regulatory strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119445503B_ABST
    Figure CN119445503B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and discloses a road condition monitoring and recognition method, device, equipment and medium based on machine vision. The method includes: collecting multi-source data of the road environment to obtain an original data set; performing dynamic time warping and multi-modal feature fusion processing on the original data set to obtain a spatio-temporal feature tensor; performing multi-scale pyramid decomposition and deep feature extraction on the spatio-temporal feature tensor to obtain a road condition feature vector; performing semantic parsing on the road condition feature vector through feature mapping and knowledge graph enhancement to obtain a road condition attribute set; performing multi-level risk quantification and spatio-temporal propagation simulation on the road condition attribute set to obtain a hierarchical warning index and a risk distribution map; performing multi-objective trade-off and causal chain reasoning on the hierarchical warning index and the risk distribution map to obtain an adaptive traffic control instruction set. The present application improves the efficiency and accuracy of road condition monitoring and recognition based on machine vision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and particularly to a road condition monitoring and recognition method, device, equipment and medium based on machine vision. Background Art

[0002] Existing road condition monitoring and recognition technologies mainly rely on the combination of fixed cameras and manual inspections. These systems usually adopt video image analysis technology to identify abnormal situations on the road, such as traffic congestion, accidents or road surface damage. Some advanced systems also integrate sensor networks to monitor environmental parameters such as road surface temperature and humidity in real time. At the same time, some regions have begun to try using unmanned aerial vehicles for road inspections to improve the flexibility and coverage of monitoring.

[0003] However, there are still some obvious deficiencies in the existing technologies. First, a single data source (such as relying only on visible light images) is easily affected by weather and lighting conditions, resulting in unstable recognition accuracy. Second, most systems lack effective multi-source data fusion capabilities and cannot fully utilize the advantages of different types of sensors. Moreover, traditional image processing methods perform poorly in complex scenarios, especially in identifying subtle road surface damage or predicting potential risks. Finally, existing systems generally lack intelligent decision-making support functions and it is difficult to automatically generate optimized traffic control strategies according to real-time road conditions. Summary of the Invention

[0004] This application provides a road condition monitoring and recognition method, device, equipment and medium based on machine vision, for improving the efficiency and accuracy of road condition monitoring and recognition based on machine vision.

[0005] In a first aspect, this application provides a road condition monitoring and recognition method based on machine vision. The road condition monitoring and recognition method based on machine vision includes: collecting multi-source data of the road environment to obtain an original data set, where the original data set includes: high-definition image stream, infrared thermal map, laser point cloud, millimeter wave radar data and environmental parameters; performing dynamic time warping and multi-modal feature fusion processing on the original data set to obtain a spatio-temporal feature tensor; performing multi-scale pyramid decomposition and depth feature extraction on the spatio-temporal feature tensor to obtain a road condition feature vector; performing semantic parsing on the road condition feature vector through feature mapping and knowledge graph enhancement to obtain a road condition attribute set, where the road condition attribute set includes road surface type, damage state, attachment distribution and traffic flow parameters; performing multi-level risk quantification and spatio-temporal propagation simulation on the road condition attribute set to obtain a hierarchical early warning index and a risk distribution map; performing multi-objective trade-off and causal chain reasoning on the hierarchical early warning index and the risk distribution map to obtain an adaptive traffic control instruction set.

[0006] Second aspect, the present application provides a road condition monitoring and recognition device based on machine vision, and the road condition monitoring and recognition device based on machine vision includes:

[0007] An acquisition module, configured to perform multi-source data acquisition on the road environment to obtain an original data set, where the original data set includes: a high-definition image stream, an infrared thermal map, a laser point cloud, millimeter-wave radar data, and environmental parameters;

[0008] A fusion module, configured to perform dynamic time warping and multi-modal feature fusion processing on the original data set to obtain a spatio-temporal feature tensor;

[0009] An extraction module, configured to perform multi-scale pyramid decomposition and depth feature extraction on the spatio-temporal feature tensor to obtain a road condition feature vector;

[0010] An analysis module, configured to perform semantic analysis on the road condition feature vector through feature mapping and knowledge graph enhancement to obtain a road condition attribute set, where the road condition attribute set includes road surface type, damage state, attachment distribution, and traffic flow parameters;

[0011] A simulation module, configured to perform multi-level risk quantification and spatio-temporal propagation simulation on the road condition attribute set to obtain a hierarchical warning index and a risk distribution map;

[0012] An inference module, configured to perform multi-objective trade-off and causal chain inference on the hierarchical warning index and the risk distribution map to obtain an adaptive traffic control instruction set.

[0013] The third aspect of the present application provides a computer device, where the memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through a bus. When the machine-readable instructions are executed by the processor, the steps of the above-mentioned road condition monitoring and recognition method based on machine vision are executed.

[0014] The fourth aspect of the present application provides a computer-readable storage medium, where instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is made to execute the above-mentioned road condition monitoring and recognition method based on machine vision.

[0015] In the technical solution provided by this application, multi-source data collection significantly improves the comprehensiveness and reliability of data, including the comprehensive utilization of high-definition image streams, infrared thermal maps, laser point clouds, millimeter-wave radar data, and environmental parameters, enabling road condition monitoring to no longer be limited to a single data source and greatly enhancing the system's perception ability of complex road conditions. The dynamic time warping and multi-modal feature fusion technologies effectively solve the spatio-temporal alignment problem of heterogeneous data, ensuring the consistency and comparability of data from different sources. The multi-scale pyramid decomposition and deep feature extraction methods can capture the multi-scale features of road conditions, fully representing both the macroscopic road network structure and the microscopic road surface details, improving the accuracy and robustness of feature extraction. The introduction of feature mapping and knowledge graph enhancement combines machine learning with domain expert knowledge, significantly improving the accuracy and interpretability of road condition semantic understanding. The multi-level risk quantification and spatio-temporal propagation simulation technologies can not only evaluate the current road condition risks but also predict the evolution trends of risks, providing a scientific basis for proactive prevention and timely intervention. Finally, the application of multi-objective trade-off and causal chain reasoning enables the system to find the optimal balance among multiple objectives such as safety, traffic efficiency, and maintenance cost, and formulate more accurate and effective control strategies through causal analysis. The road maintenance cost is reduced through precise maintenance strategies. Brief Description of the Drawings

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a schematic diagram of an embodiment of the road condition monitoring and recognition method based on machine vision in the embodiments of this application;

[0018] Figure 2 It is a schematic diagram of an embodiment of the road condition monitoring and recognition device based on machine vision in the embodiments of this application;

[0019] Figure 3 It is a schematic diagram of the structure of the computer device in the embodiments of this application. Detailed Embodiments

[0020] The embodiments of the present application provide a road condition monitoring and recognition method, device, equipment and medium based on machine vision. Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or equipment that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.

[0021] For ease of understanding, the specific process of the embodiments of the present application is described below. Please refer to Figure 1 One embodiment of the road condition monitoring and recognition method based on machine vision in the embodiments of the present application includes:

[0022] Step S101: Collect multi-source data of the road environment to obtain an original data set. The original data set includes: high-definition image stream, infrared thermal map, laser point cloud, millimeter-wave radar data and environmental parameters;

[0023] Step S102: Perform dynamic time warping and multi-modal feature fusion processing on the original data set to obtain a spatio-temporal feature tensor;

[0024] Step S103: Perform multi-scale pyramid decomposition and depth feature extraction on the spatio-temporal feature tensor to obtain a road condition feature vector;

[0025] Step S104: Perform semantic parsing on the road condition feature vector through feature mapping and knowledge graph enhancement to obtain a road condition attribute set. The road condition attribute set includes road surface type, damage status, attachment distribution and traffic flow parameters;

[0026] Step S105: Perform multi-level risk quantification and spatio-temporal propagation simulation on the road condition attribute set to obtain a hierarchical early warning index and a risk distribution map;

[0027] Step S106: Perform multi-objective trade-off and causal chain reasoning on the hierarchical early warning index and the risk distribution map to obtain an adaptive traffic control instruction set.

[0028] It can be understood that the execution subject of the present application can be a road condition monitoring and recognition device based on machine vision, or a terminal or a server. Specifically, it is not limited here. The embodiments of the present application are described by taking the server as the execution subject as an example.

[0029] Specifically, multi-source data collection is carried out on the road environment to obtain the original dataset. This involves using multiple sensor devices to simultaneously collect road environment information, including high-definition image streams captured by high-definition cameras, infrared thermal maps obtained by infrared thermal imagers, laser point clouds scanned by lidar, millimeter-wave radar data detected by millimeter-wave radars, and environmental parameters measured by various environmental sensors. The high-definition image stream provides visible light information of the road scene, the infrared thermal map can identify heat sources under low-light conditions, the laser point cloud precisely depicts the three-dimensional structure of the road, the millimeter-wave radar data can detect moving objects, and the environmental parameters include temperature, humidity, light intensity, etc. The original dataset is processed by dynamic time warping and multi-modal feature fusion to obtain a spatio-temporal feature tensor. Dynamic time warping aims to solve the problem of inconsistent data collection frequencies of different sensors, and aligns each data stream to a unified time scale through methods such as interpolation or downsampling. Multi-modal feature fusion uses deep learning techniques, such as multi-modal autoencoders or attention mechanisms, to fuse different types of data features into a unified representation. The output of this step is a high-dimensional tensor that contains information about time, space, and multiple data modalities.

[0030] Subsequently, multi-scale pyramid decomposition and deep feature extraction are performed on the spatio-temporal feature tensor to obtain a road condition feature vector. Multi-scale pyramid decomposition is an image processing technique that creates feature maps of different resolutions by successive downsampling, which helps to capture road condition information at different scales. Deep feature extraction uses deep learning models such as convolutional neural networks (CNNs) to extract high-level semantic features from the multi-scale feature maps. The output of this step is a low-dimensional road condition feature vector that condenses the key information in the original data. Semantic parsing is performed on the road condition feature vector through feature mapping and knowledge graph enhancement to obtain a road condition attribute set. Feature mapping uses a pre-trained deep learning model to map the feature vector to a semantic space. Knowledge graph enhancement introduces prior knowledge in the road domain, such as pavement material properties, traffic rules, etc., to assist in interpreting the semantic meaning of the features. Through this step, the system can identify road condition attributes such as road surface type (e.g., asphalt, concrete), damage status (e.g., cracks, potholes), attachment distribution (e.g., water accumulation, snow accumulation), and traffic flow parameters (e.g., vehicle flow, average speed).

[0031] Then, perform multi-level risk quantification and spatio-temporal propagation simulation on the road condition attribute set to obtain hierarchical warning indicators and risk distribution maps. The multi-level risk quantification considers the influence degree of different attributes on road safety. By methods such as the Analytic Hierarchy Process (AHP) or fuzzy comprehensive evaluation, the qualitative road condition attributes are converted into quantitative risk indicators. The spatio-temporal propagation simulation uses models such as cellular automata or graph neural networks to simulate the propagation process of risks in the road network and generate dynamic risk distribution maps. Finally, perform multi-objective trade-off and causal chain reasoning on the hierarchical warning indicators and risk distribution maps to obtain an adaptive traffic control instruction set. The multi-objective trade-off considers multiple objectives such as safety, traffic efficiency, and maintenance cost, and uses multi-objective optimization algorithms such as NSGA-II to find balanced solutions. The causal chain reasoning is based on Bayesian networks or structural equation models to analyze the impact of different control measures on road conditions and traffic flow, so as to generate the optimal control strategy.

[0032] For example, a certain section of the highway uses this method for road condition monitoring. First, environmental parameters such as high-definition video streams with a resolution of 1080p, infrared thermal maps of 320×240 pixels, laser point clouds of 1 million points per second, millimeter-wave radar data with an angular resolution of 0.1°, and temperature and humidity updated per second are collected. These data are aligned to a 1-second interval through the dynamic time window method and fused into a 512-dimensional spatio-temporal feature tensor using a multi-modal variational autoencoder. Perform a 4-layer Gaussian-Laplacian pyramid decomposition on this tensor to obtain multi-scale feature maps, and then extract a 2048-dimensional road condition feature vector through a ResNet50 network. The feature vector is mapped to a 300-dimensional semantic space through a pre-trained Word2Vec model, and combined with a road knowledge graph containing 10,000 concept nodes to parse out the road condition attribute set. The attribute set includes road surface type (asphalt, 95% confidence), damage status (minor cracks, area about 0.5 square meters), attachment distribution (ponding, depth 3 cm, coverage rate 2%), and traffic flow parameters (traffic volume 150 vehicles / hour, average speed 85 km / h). Based on these attributes, use a five-layer AHP model to quantify the risk and obtain a comprehensive risk index of 0.7 (full score 1). Through a spatio-temporal propagation model based on a graph convolutional network, it is predicted that the risk may spread to adjacent 5-kilometer sections within the next 2 hours. The multi-objective optimization balances safety and traffic efficiency, and combines Bayesian causal reasoning to finally generate control instructions: reduce the speed limit to 70 km / h, turn on the electronic warning sign, and arrange maintenance personnel to deal with ponding and cracks within 2 hours.

[0033] In the embodiments of the present application, multi-source data collection significantly improves the comprehensiveness and reliability of data, including the comprehensive utilization of high-definition image streams, infrared thermal maps, laser point clouds, millimeter-wave radar data, and environmental parameters, enabling road condition monitoring to no longer be limited to a single data source and greatly enhancing the system's perception ability of complex road conditions. The dynamic time warping and multi-modal feature fusion technologies effectively solve the spatio-temporal alignment problem of heterogeneous data, ensuring the consistency and comparability of data from different sources. The multi-scale pyramid decomposition and deep feature extraction methods can capture the multi-scale features of road conditions, and can be fully characterized from the macroscopic road network structure to the microscopic road surface details, improving the accuracy and robustness of feature extraction. The introduction of feature mapping and knowledge graph enhancement combines machine learning with domain expert knowledge, greatly improving the accuracy and interpretability of road condition semantic understanding. The multi-level risk quantification and spatio-temporal propagation simulation technologies can not only evaluate the current road condition risks, but also predict the evolution trend of risks, providing a scientific basis for proactive prevention and timely intervention. Finally, the application of multi-objective trade-off and causal chain reasoning enables the system to find the optimal balance among multiple objectives such as safety, traffic efficiency, and maintenance costs, and formulate more accurate and effective control strategies through causal analysis. Reduce road maintenance costs through precise maintenance strategies.

[0034] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0035] (1) Perform high-resolution scanning on the visible light scene in the road environment to obtain a high-definition image stream, and perform infrared imaging on the thermal radiation scene in the road environment to obtain an infrared thermal map;

[0036] (2) Perform three-dimensional laser scanning on the road environment to obtain a laser point cloud, and detect moving targets in the road environment through a millimeter-wave radar to obtain millimeter-wave radar data;

[0037] (3) Collect multi-parameters of the meteorological conditions in the road environment to obtain environmental parameters;

[0038] (4) Perform time synchronization processing on the high-definition image stream, infrared thermal map, laser point cloud, millimeter-wave radar data, and environmental parameters through a timestamp alignment algorithm to obtain time-aligned data, and perform coordinate transformation and spatial correspondence processing on the time-aligned data through a spatial registration algorithm to obtain spatially consistent data;

[0039] (5) Perform data cleaning and outlier detection on the spatially consistent data to obtain preprocessed data, and perform feature extraction on the preprocessed data through a multi-modal feature extraction algorithm to obtain an original data set.

[0040] Specifically, the base first performs a high-resolution scan on the visible light scene in the road environment to obtain a high-definition image stream. A high-resolution camera, such as a camera with 4K resolution (3840×2160 pixels), is used to continuously collect road scene images at a rate of 30 frames per second. At the same time, infrared imaging is performed on the thermal radiation scene in the road environment to obtain an infrared thermal image. Infrared imaging uses a long-wave infrared thermal imager with a working wavelength range of 8-14μm, a resolution of 640×480 pixels, and a temperature resolution of 0.05°C, which can effectively capture the road surface temperature distribution and thermal anomalies. Three-dimensional laser scanning is performed on the road environment to obtain a laser point cloud. The laser scanner uses a 64-line lidar with a scanning frequency of 10Hz, a ranging accuracy of ±2cm, a horizontal field of view of 360°, a vertical field of view of 26.9°, and can collect approximately 1.3 million points per second. At the same time, millimeter-wave radar is used to detect moving targets in the road environment to obtain millimeter-wave radar data. The millimeter-wave radar operates at a frequency of 77GHz, has a range resolution of 0.2m, a speed resolution of 0.1m / s, and can track up to 100 targets simultaneously.

[0041] Then, multi-parameter collection of the meteorological conditions in the road environment is carried out to obtain environmental parameters. The collection of meteorological parameters includes temperature (accuracy ±0.1°C), humidity (accuracy ±2%RH), wind speed (accuracy ±0.5m / s), precipitation (accuracy ±0.2mm), and light intensity (accuracy ±2%), etc., with a sampling frequency of 1Hz. Subsequently, time synchronization processing is performed on the high-definition image stream, infrared thermal image, laser point cloud, millimeter-wave radar data, and environmental parameters through a timestamp alignment algorithm to obtain time-aligned data. The timestamp alignment algorithm uses a linear interpolation method to align all data streams to a unified time axis with a time resolution of 0.1 seconds. Coordinate system conversion and spatial correspondence processing are performed on the time-aligned data through a spatial registration algorithm to obtain spatially consistent data. Spatial registration uses the ICP (Iterative Closest Point) algorithm based on feature point matching to convert different sensor data into a unified world coordinate system, and the registration accuracy is better than 5cm.

[0042] Finally, data cleaning and outlier detection are performed on the spatially consistent data to obtain preprocessed data. Data cleaning includes removing invalid data points, filling in missing values, and smoothing noise. Outlier detection uses an algorithm based on the Local Outlier Factor (LOF) to identify and mark data points that deviate from the normal range. Feature extraction is performed on the preprocessed data through a multi-modal feature extraction algorithm to obtain an original data set. Multi-modal feature extraction uses a deep learning model, such as a multi-modal autoencoder, to extract a unified feature representation from different types of data.

[0043] For example, at a monitoring point on a certain highway, the specific implementation process of this method is as follows: First, a 4K high-definition camera captures road images at a frame rate of 30fps, while an infrared thermal imager with a resolution of 640×480 captures the road surface temperature distribution. A 64-line lidar scans the road environment at a frequency of 10Hz, obtaining approximately 130,000 spatial points per scan. A 77GHz millimeter-wave radar outputs target detection results every 0.1 seconds, including distance, speed, and direction information. The weather station records environmental parameters once per second, including an air temperature of 22.5°C, a relative humidity of 65%, a wind speed of 3.2m / s, a precipitation of 0mm, and a light intensity of 15000lux. The timestamp alignment algorithm aligns this heterogeneous data to a time interval of 0.1 seconds. For example, the 30fps image stream is downsampled to 10fps to be consistent with the sampling rate of the lidar point cloud; linear interpolation is performed on the millimeter-wave radar data and environmental parameters to synchronize them with other data. During the spatial registration process, first, feature points (such as roadside lines, traffic signs, etc.) are extracted from the high-definition image and the lidar point cloud, and then the ICP algorithm is used to iteratively optimize the transformation matrix between different coordinate systems, and finally all data is unified into the world coordinate system based on the road center line.

[0044] In the data cleaning stage, invalid points in the lidar point cloud with a reflectivity lower than 5% (about 2% of the total number of points) are removed, local vacancies in the infrared image caused by occlusion are filled (about 1% of the image area), and Gaussian filtering is applied to the high-definition image to reduce noise. Outlier detection identifies an area of abnormally high temperature (10°C higher than the surrounding temperature), which may indicate potential road surface problems. Multimodal feature extraction inputs the cleaned data into a pre-trained deep neural network to extract a 1024-dimensional feature vector, which contains comprehensive information such as road surface texture, temperature distribution, three-dimensional structure, and dynamic targets.

[0045] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0046] (1) Perform time series analysis on each data stream in the original dataset to obtain a timestamp sequence, and perform alignment processing on the timestamp sequence to obtain a regularized time series;

[0047] (2) Perform interpolation processing on the regularized time series to obtain a data series with a unified sampling rate, and segment the data series with a unified sampling rate to obtain a time window dataset;

[0048] (3) Extract features from each data modality in the time window dataset to obtain a multimodal feature set, and perform scale adjustment on the multimodal feature set to obtain a normalized feature set;

[0049] (4) Cross-modal feature fusion is performed on the normalized feature set through a multi-modal feature fusion algorithm to obtain a fused feature vector, and temporal correlation analysis is performed on the fused feature vector to obtain a temporal feature matrix;

[0050] (5) Spatial relationship analysis is performed on the temporal feature matrix through a spatial correlation analysis algorithm to obtain a spatio-temporal feature tensor.

[0051] Specifically, time series analysis is performed on each data stream to obtain a time stamp sequence. This involves extracting the acquisition time of each data point to form a discrete time series. Subsequently, alignment processing is performed on these time stamp sequences to obtain a regularized time series. The alignment processing uses the dynamic time warping (DTW) algorithm, which finds the optimal time warping path to make the time axes of different data streams correspond. Then, interpolation processing is performed on the regularized time series to obtain a data sequence with a unified sampling rate. The interpolation processing uses cubic spline interpolation to create a smooth curve between the original data points, thereby obtaining data points at fixed time intervals. The unified sampling rate is usually set as the greatest common divisor of the highest sampling rates in all data streams to ensure no information loss. Then, the data sequence with a unified sampling rate is segmented to obtain a time window data set. The segmentation uses a sliding window technique, and the window size is set according to specific application requirements, usually 1 - 10 seconds, and the window overlap rate is 50%, to capture local features and change trends in time.

[0052] Subsequently, feature extraction is performed on each data modality in the time window dataset to obtain a multi-modal feature set. Different methods are used for feature extraction depending on the type of data: for image data, a convolutional neural network (CNN) is used to extract spatial features; for time series data, such as radar signals, a long short-term memory network (LSTM) is used to extract temporal features; for point cloud data, networks such as PointNet are used to extract three-dimensional structural features. The extracted features are usually high-dimensional vectors, with each modality possibly having hundreds to thousands of dimensions. Next, the multi-modal feature set is scaled to obtain a normalized feature set. Normalization uses the Z-score normalization method to adjust the mean of each feature to 0 and the standard deviation to 1, ensuring that features of different modalities are processed on the same scale. Then, cross-modal feature fusion is performed on the normalized feature set through a multi-modal feature fusion algorithm to obtain a fused feature vector. The fusion algorithm uses a multi-modal Transformer model with an attention mechanism, which can adaptively learn the correlation and importance weights between features of different modalities. The dimension of the fused feature vector is usually between 1000 - 2000 and contains the key information of each modality. Temporal correlation analysis is performed on the fused feature vector to obtain a temporal feature matrix. Temporal correlation analysis uses the autocorrelation function (ACF) and the partial autocorrelation function (PACF) to identify periodic patterns and time-dependent relationships in the feature sequence.

[0053] Finally, spatial relationship analysis is performed on the temporal feature matrix through a spatial correlation analysis algorithm to obtain a spatio-temporal feature tensor. Spatial correlation analysis uses spatial autocorrelation indices (such as Moran's I) and geographically weighted regression (GWR) methods to capture the distribution pattern and local correlation of features in the spatial dimension. The finally obtained spatio-temporal feature tensor is a three-dimensional data structure containing information on time, space, and features.

[0054] For example, at a monitoring point on a certain highway, the specific implementation process of this method is as follows: First, 4K video streams at 30fps, 64-line lidar point clouds at 10Hz, millimeter-wave radar data at 100Hz, and environmental parameter data at 1Hz are collected. These data streams are aligned through the DTW algorithm to obtain a regular time series at intervals of 0.01 seconds. All data is adjusted to a unified sampling rate of 100Hz using cubic spline interpolation, and then segmented with a window size of 5 seconds and an overlap of 2.5 seconds to obtain a time window dataset. For each 5-second window, 150 frames of images are extracted from the 4K video, and 2048-dimensional image features are extracted through a pre-trained ResNet50 network; for the lidar point cloud data, 1024-dimensional three-dimensional structural features are extracted using the PointNet++ network; the millimeter-wave radar data extracts 512-dimensional time-frequency features through 1D-CNN; the environmental parameters are directly used as 5-dimensional feature vectors. After these features are standardized by Z-score, they are input into a multi-modal Transformer for fusion to obtain a 1536-dimensional fused feature vector.

[0055] ACF and PACF analyses are performed on the fused feature vector, and it is identified that there is an obvious 15-minute periodicity in the feature sequence, which may correspond to the regular changes in traffic flow. Spatial correlation analysis shows that some features (such as road surface temperature) show a high positive correlation within a range of 2 kilometers (Moran's I = 0.85), while other features (such as traffic flow density) show a moderate negative correlation (Moran's I = -0.42). The finally generated spatio-temporal feature tensor has a dimension of 100×100×1536, where 100×100 represents the spatial grid of the monitoring area, and 1536 is the feature dimension of each spatio-temporal point. This tensor comprehensively captures the change characteristics of the road environment in the time and space dimensions, providing a rich information basis for subsequent road condition analysis and early warning. Through this multi-modal and multi-scale data processing and feature extraction method, the accuracy and comprehensiveness of road condition monitoring are greatly improved, and potential road safety hazards and traffic anomalies can be better identified.

[0056] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0057] (1) Perform multi-scale decomposition on the spatio-temporal feature tensor through the Gaussian pyramid algorithm to obtain a multi-scale feature pyramid, and perform differential processing on the multi-scale feature pyramid through the Laplacian pyramid algorithm to obtain multi-scale differential features;

[0058] (2) Perform feature enhancement on the multi-scale differential features through convolution operations to obtain an enhanced feature map, and perform dimensionality reduction processing on the enhanced feature map through pooling operations to obtain a compressed feature representation;

[0059] (3) Feature transformation is performed on the compressed feature representation through a non - linear activation function to obtain an activated feature map, and multi - layer feature fusion is performed on the activated feature map to obtain a fused feature representation;

[0060] (4) Key information extraction is performed on the fused feature representation through an attention mechanism to obtain attention - weighted features, and dimensionality reduction processing is performed on the attention - weighted features through a fully - connected layer to obtain a road condition feature vector.

[0061] Specifically, multi - scale decomposition is performed on it through the Gaussian pyramid algorithm to obtain a multi - scale feature pyramid. The Gaussian pyramid algorithm performs downsampling and smoothing on the original tensor in an iterative manner, and the resolution of each layer is half of the previous layer. Specifically, for the original tensor T, the Gaussian pyramid of the i - th layer is denoted as Gi(T), and the calculation process is to convolve Gi - 1(T) with a 5×5 Gaussian kernel and then perform 2×2 downsampling. Usually, a pyramid of 4 to 5 layers is constructed to capture features at different scales.

[0062] Subsequently, differential processing is performed on the multi - scale feature pyramid through the Laplacian pyramid algorithm to obtain multi - scale differential features. The Laplacian pyramid is the differential representation of the Gaussian pyramid, and the calculation method is to subtract two adjacent layers of the Gaussian pyramid. For the Laplacian representation Li(T) of the i - th layer, its calculation formula is Li(T)=Gi(T) - upsample(Gi + 1(T)), where upsample represents the upsampling operation. The Laplacian pyramid can highlight the feature edges and detail information at different scales. Then, feature enhancement is performed on the multi - scale differential features through convolution operations to obtain an enhanced feature map. In this step, multiple convolutional kernels are used to perform convolution operations on the differential features of each layer to extract richer feature representations. The size of the convolutional kernel is usually 3×3 or 5×5, and the number ranges from 32 to 128, depending on specific requirements. After the convolution operation, dimensionality reduction processing is performed on the enhanced feature map through a pooling operation to obtain a compressed feature representation. The pooling operation commonly uses max - pooling or average - pooling, with a pooling window size of 2×2 and a stride of 2, which can significantly reduce the data volume while retaining the main features.

[0063] Then, the compressed feature representation is subjected to feature transformation through a non-linear activation function to obtain an activated feature map. Commonly used non-linear activation functions include ReLU (Rectified Linear Unit) or LeakyReLU, which can introduce non-linear transformations and enhance the expressive power of the model. The activated feature map undergoes multi-level feature fusion to obtain a fused feature representation. Feature fusion adopts the idea of skip connection, combining feature maps at different levels through element-wise addition or concatenation to comprehensively utilize multi-scale information. Finally, key information extraction is performed on the fused feature representation through an attention mechanism to obtain attention-weighted features. The attention mechanism calculates the importance weights of each feature position, focusing on the regions and features that are most critical for road condition recognition. Specifically, the self-attention mechanism can be adopted, and the attention weights are calculated in the way of query-key-value. After obtaining the attention-weighted features, dimensionality reduction processing is performed on them through a fully connected layer to obtain the final road condition feature vector. The fully connected layer maps high-dimensional features to a lower-dimensional vector space, usually reduced to 256 or 512 dimensions, for subsequent classification or regression tasks.

[0064] For example, in a certain highway monitoring system, the specific implementation process of this method is as follows: First, a 5-layer Gaussian pyramid is constructed for a spatio-temporal feature tensor of 100×100×1536 to obtain multi-scale feature representations with resolutions of 100×100, 50×50, 25×25, 13×13, and 7×7 in sequence. Then, the Laplacian pyramid is calculated to obtain 4 layers of differential features, with each layer containing 1536 channels. 64 3×3 convolutional kernels are applied to each layer of differential features for feature enhancement to obtain 64 feature maps. Taking the first layer (100×100) as an example, the enhanced feature dimension is 100×100×64. Then, a 2×2 max pooling operation is performed to halve the size of the feature map, obtaining a compressed feature representation of 50×50×64. The ReLU activation function, f(x) = max(0, x), is applied to the compressed feature to introduce non-linear transformation. Then, the activated feature maps at different levels are upsampled and concatenated to obtain a fused feature representation of 100×100×256. The self-attention mechanism is applied to the fused feature to calculate the attention weight matrix A, with a dimension of 100×100, representing the importance of each spatial position. Multiply A by the original feature map to obtain the attention-weighted features.

[0065] Finally, a 100×100×256 feature map is mapped into a 512-dimensional road condition feature vector through two fully connected networks. The first fully connected layer reduces the input to 2048 dimensions, and the second layer further reduces it to 512 dimensions. The finally obtained 512-dimensional vector contains key information about road conditions, such as road surface type, damage degree, traffic flow, etc., providing a highly condensed feature representation for subsequent road condition analysis and decision-making. This multi-scale and multi-level feature extraction and fusion method can effectively capture the complex features of the road environment, and can comprehensively represent from the macroscopic road network structure to the microscopic road surface details. By introducing the attention mechanism, the recognition ability for key regions and features is further improved, greatly enhancing the accuracy and reliability of road condition monitoring.

[0066] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0067] (1) Perform semantic space projection on the road condition feature vector through a feature mapping algorithm to obtain a semantic feature representation, and perform feature grouping on the semantic feature representation to obtain preliminary semantic categories;

[0068] (2) Perform concept alignment on the preliminary semantic categories through a knowledge graph matching algorithm to obtain a semantic concept set, and construct a semantic relationship network for the semantic concept set;

[0069] (3) Perform road surface type recognition on the semantic relationship network to obtain the road surface type, and perform road surface damage analysis on the semantic relationship network to obtain the damage state;

[0070] (4) Perform road surface attachment recognition on the semantic relationship network to obtain the attachment distribution, and calculate traffic flow parameters for the semantic relationship network to obtain traffic flow parameters;

[0071] (5) Integrate the road surface type, damage state, attachment distribution and traffic flow parameters through an attribute aggregation algorithm to obtain a road condition attribute set.

[0072] Specifically, after obtaining the road condition feature vector, the road condition monitoring and recognition method based on machine vision first performs semantic space projection through a feature mapping algorithm to obtain a semantic feature representation. This process uses a pre-trained word embedding model, such as Word2Vec or GloVe, to map the high-dimensional feature vector into a low-dimensional semantic space. Assuming the road condition feature vector is 512-dimensional, it can be converted into a 300-dimensional semantic vector through feature mapping. The mapping process can be expressed as:

[0073] S = W·F + b

[0074] Among them, S is a 300-dimensional semantic vector, F is a 512-dimensional road condition feature vector, W is a weight matrix of 512×300, and b is a 300-dimensional bias vector. These parameters are obtained through pre-training and can convert the original features into representations with semantic meanings. Then, the semantic feature representations are grouped to obtain preliminary semantic categories. Feature grouping uses the K-means clustering algorithm to divide the vectors in the 300-dimensional semantic space into K clusters. The selection of cluster centers is based on prior knowledge in the road condition field, such as categories like road surface type and damage degree. Each cluster represents a preliminary semantic category, such as "asphalt road surface" and "slight damage".

[0075] Subsequently, the preliminary semantic categories are conceptually aligned through a knowledge graph matching algorithm to obtain a semantic concept set. The knowledge graph contains professional concepts and relationships in the road field, such as road surface materials, damage types, traffic parameters, etc. The concept alignment process uses cosine similarity to calculate the similarity between the preliminary semantic categories and the concepts in the knowledge graph, and selects the concept with the highest similarity as the matching result. For example, "asphalt road surface" may be aligned with the concept of "flexible road surface" in the knowledge graph. A semantic relationship network is constructed for the semantic concept set. Relationship construction is based on predefined relationship types in the knowledge graph, such as "is a kind of", "contains", "causes", etc., to connect the matched concepts into a network structure. Next, road surface type recognition is performed on the semantic relationship network to obtain the road surface type. Road surface type recognition uses a graph convolutional network (GCN) to process the semantic relationship network and classify using node features and topological structures. The output of the GCN passes through the softmax function to obtain the probability distribution of each road surface type, and the category with the highest probability is selected as the recognition result. At the same time, road surface damage analysis is performed on the semantic relationship network to obtain the damage status. Damage analysis uses a multi-label classification model, such as LSTM (long short-term memory network), considering the temporal features of damage types and outputting multiple damage labels and their severities.

[0076] Then, road surface attachment recognition is performed on the semantic relationship network to obtain the attachment distribution. Attachment recognition uses a variant of the region convolutional neural network (R-CNN), such as Mask R-CNN, which can perform object detection and semantic segmentation simultaneously to identify the type, location, and coverage area of the attachments. Traffic flow parameter calculation is performed on the semantic relationship network to obtain traffic flow parameters. Traffic flow parameter calculation uses time series analysis methods, such as ARIMA (autoregressive integrated moving average model), to predict indicators such as vehicle flow and average speed in the short term. Finally, the road surface type, damage status, attachment distribution, and traffic flow parameters are integrated through an attribute aggregation algorithm to obtain a road condition attribute set. Attribute aggregation uses a weighted summation method to assign weights according to the importance of each attribute to road condition evaluation. The aggregation formula is as follows:

[0077]

[0078] wherein, R is the comprehensive road condition index, A i is the quantization value of the i-th attribute, w i is the corresponding weight, and n is the total number of attributes. The determination of the weight is based on expert experience and historical data analysis.

[0079] For example, a certain highway monitoring system uses this method to analyze a 1-kilometer-long section of the road. First, the 512-dimensional road condition feature vector is mapped to a 300-dimensional semantic vector through a pre-trained Word2Vec model. K-means clustering (K = 10) divides the semantic vectors into 10 preliminary categories. Knowledge graph matching aligns these 10 categories with the road domain knowledge graph containing 5,000 nodes to obtain accurate semantic concepts. Semantic relationship construction forms a semantic relationship network containing 50 nodes and 200 edges. The GCN model processes this network and identifies that the road surface type is "asphalt concrete" (confidence 0.92). The LSTM model analyzes and obtains the damage states as "slight cracks" (area 0.5 square meters, depth 2 mm) and "rutting" (depth 5 mm, length 10 meters). Mask R-CNN identifies the attachments as "ponding" (coverage rate 2%, average depth 1 cm) and "gravel" (coverage rate 0.5%). The ARIMA model calculates that the current traffic flow is 1,200 vehicles per hour and the average speed is 85 km / h.

[0080] When aggregating attributes, the weight of the road surface type is 0.3, the weight of the damage state is 0.25, the weight of the attachment distribution is 0.25, and the weight of the traffic flow parameters is 0.2. After quantifying each attribute, the calculated result of the comprehensive road condition index is 0.82 (full score 1), indicating that the road condition is good but there are minor problems. This comprehensive index provides an intuitive road condition assessment result for the traffic management department, helping to timely discover potential problems and formulate maintenance plans.

[0081] In a specific embodiment, the process of executing step S105 may specifically include the following steps:

[0082] (1) Identify risk factors for the road condition attribute set to obtain a risk factor set, and sort the risk factor set by importance to obtain weighted risk factors;

[0083] (2) Perform hierarchical processing on the weighted risk factors through the analytic hierarchy process to obtain a risk hierarchy structure, and divide the risk hierarchy structure by risk level to obtain hierarchical risk indicators;

[0084] (3) Estimate the spatial distribution of the hierarchical risk indicators through a spatio-temporal interpolation algorithm to obtain a risk spatial distribution map, and perform dynamic evolution simulation on the risk spatial distribution map through a cellular automaton algorithm to obtain a risk propagation model;

[0085] (4) Conduct multi-scenario simulations on the risk propagation model to obtain the risk evolution sequence, and calculate the probability distribution of the risk evolution sequence to obtain the risk probability graph;

[0086] (5) Divide the risk regions of the risk probability graph to obtain the hierarchical warning indicators, and graphically represent the hierarchical warning indicators to obtain the risk distribution map.

[0087] Specifically, first identify its risk factors to obtain the risk factor set. A method combining principal component analysis (PCA) and random forest algorithm is adopted. PCA is used for dimensionality reduction and extraction of main features, while random forest is used to evaluate the importance of features. For n road condition attributes, PCA retains the principal components that explain 90% of the variance, usually reducing the dimension to m (m < n). Subsequently, the random forest algorithm scores the importance of these m principal components to obtain the importance weights of each factor. The formula for calculating the importance weight of a risk factor is as follows:

[0088]

[0089] Among them, W i is the importance weight of the i-th risk factor, and E i is the importance score of this factor in the random forest, and m is the number of principal components. Then, hierarchically process the weighted risk factors through the analytic hierarchy process (AHP) to obtain the risk hierarchy structure. The AHP method first constructs a hierarchical structure model, including the goal layer (overall risk), the criterion layer (such as road surface condition, traffic flow, environmental factors), and the scheme layer (specific risk factors). Then, pairwise comparisons are made to establish a judgment matrix. The element a ij of the judgment matrix A represents the importance of factor i relative to factor j, and its value range is usually an integer from 1 to 9. Calculate the eigenvector and the maximum eigenvalue, and conduct a consistency test. If the consistency test is passed, the eigenvector after normalization is the weight vector of each layer.

[0090] Conduct risk level division on the risk hierarchy structure to obtain the hierarchical risk indicators. The risk level division adopts the fuzzy comprehensive evaluation method to convert the quantitative values of risk factors into qualitative risk levels. First, determine the evaluation factor set and the comment set, and then establish a fuzzy relation matrix R. The formula for calculating the fuzzy comprehensive evaluation result B is:

[0091]

[0092] Among them, W is the weight vector, R is the fuzzy relation matrix, represents the fuzzy composition operation, b1, b2,..., b mThey are the m elements of vector B, and each element represents the membership degree of the final evaluation result corresponding to a certain comment level. Then, the spatial distribution of the classified risk indicators is estimated through a spatio-temporal interpolation algorithm to obtain a risk spatial distribution map. Spatio-temporal interpolation uses the Kriging method, which takes into account spatial autocorrelation and can provide the best linear unbiased estimate. The basic formula for Kriging interpolation is:

[0093]

[0094] Among them, Z*(x0) is the predicted value of the point to be estimated, Z(x i ) is the observed value of the known point, and λ i is the weight coefficient. The weight coefficient is solved through the variogram and the Kriging equation system. Next, the risk spatial distribution map is dynamically evolved and simulated through the cellular automaton algorithm to obtain a risk propagation model. The cellular automaton model defines a grid-like space, and each cell represents a road section. The state of each cell is updated according to the states of the surrounding cells and predefined rules within discrete time steps. The state update rule can be expressed as:

[0095] S t+1 (i, j) = f(S t (i, j), N t (i, j))

[0096] Among them, S t (i, j) is the state of the cell at the position (i, j) at time t, N t (i, j) is the state of its neighborhood, and f is the state transition function. Multi-scenario simulation is carried out on the risk propagation model to obtain a risk evolution sequence. Multi-scenario simulation considers different initial conditions and parameter settings, such as weather changes and traffic flow fluctuations. The cellular automaton model is run multiple times (such as 1000 times) under each scenario, and the risk evolution process of each simulation is recorded. Probability distribution calculation is carried out on the risk evolution sequence to obtain a risk probability map. The probability distribution calculation uses the kernel density estimation method to estimate the probability density of the risk values at each spatial position at different time points.

[0097] Finally, the risk probability map is divided into risk regions to obtain the hierarchical early warning indicators. The risk region division uses clustering algorithms such as DBSCAN (Density-Based Spatial Clustering of Applications with Noise) to group regions with similar risk levels together. The hierarchical early warning indicators are graphically represented to obtain the risk distribution map. The graphical representation combines a heat map and an isoline map to intuitively display the spatial distribution and intensity of risks. For example, a certain section of highway (50 kilometers in length) uses this method for risk assessment. First, 15 main risk factors are identified from the road condition attribute set, including pavement damage degree, traffic flow, weather conditions, etc. PCA analysis reduces the 15 factors to 8 principal components, explaining 92% of the total variance. The random forest algorithm scores the importance of these 8 principal components, obtaining a weight distribution of [0.25, 0.20, 0.15, 0.12, 0.10, 0.08, 0.06, 0.04]. The AHP method constructs a three-layer risk structure: overall risk (goal layer), pavement condition, traffic flow, environmental factors (criterion layer), and 8 specific risk factors (scheme layer). Through expert evaluation and consistency test, the criterion layer weights [0.5, 0.3, 0.2] are obtained. Fuzzy comprehensive evaluation divides the risk level into three levels: low, medium, and high, obtaining the risk level of each 500-meter section.

[0098] Based on the risk levels of these 500-meter sections, the Kriging interpolation method generates a continuous risk spatial distribution map. The cellular automaton model divides the 50-kilometer section into 100 cells, with each cell representing 500 meters. State transition rules are defined, such as when the average risk level of surrounding cells is higher than that of the current cell, there is an 80% probability that the risk level of the current cell will increase. 1000 multi-scenario simulations are performed, each simulating the risk evolution over 24 hours. Through kernel density estimation, the risk probability distribution of each cell at different time points is calculated. The DBSCAN algorithm (parameters: ε = 0.05, MinPts = 4) clusters the risk probability map, identifying 3 high-risk regions, 5 medium-risk regions, and 2 low-risk regions. The finally generated risk distribution heat map clearly shows the spatial distribution of risks, with high-risk regions presented in red, mainly concentrated at the 10-15 kilometer and 30-35 kilometer sections of the road.

[0099] In a specific embodiment, the process of performing step S106 may specifically include the following steps:

[0100] (1) Construct a multi-objective function for the hierarchical early warning indicators to obtain a multi-objective optimization problem, and set constraints for the multi-objective optimization problem to obtain an optimization problem with constraints;

[0101] (2) Use the Pareto optimization algorithm to perform multi-objective solution for the optimization problem with constraints to obtain a candidate solution set, and evaluate the candidate solution set to obtain weighted candidate solutions;

[0102] (3) Identify the causal relationships in the risk distribution map to obtain a causal map, and perform probability inference on the causal map to obtain the event probability distribution;

[0103] (4) Evaluate the intervention effect on the event probability distribution through the causal intervention analysis algorithm to obtain an intervention plan set, and match the intervention plan set with the weighted candidate plans to obtain a regulation strategy set;

[0104] (5) Make adaptive adjustments to the regulation strategy set to obtain an adaptive traffic regulation instruction set.

[0105] Specifically, construct a multi-objective function for the hierarchical warning indicators to obtain a multi-objective optimization problem. The multi-objective function usually includes aspects such as safety, traffic efficiency, and maintenance cost. For example, the safety objective can be expressed as minimizing the length of high-risk sections, the traffic efficiency objective can be expressed as maximizing the average vehicle speed, and the maintenance cost objective can be expressed as minimizing the road repair frequency. These objectives often conflict with each other and a balance needs to be found among them. Set constraints for the multi-objective optimization problem to obtain a constrained optimization problem. The constraint conditions include budget limitations, human resource limitations, time limitations, etc. For example, the annual road maintenance budget does not exceed 10 million yuan, the emergency repair response time does not exceed 2 hours, and traffic control measures should not cause the average vehicle speed to decrease by more than 20%. These constraint conditions ensure the feasibility of the optimization results in actual operations.

[0106] Next, perform multi-objective solution for the constrained optimization problem through the Pareto optimization algorithm to obtain a candidate solution set. The Pareto optimization algorithm, such as NSGA-II (Non-dominated Sorting Genetic Algorithm II), can find the optimal trade-off among multiple objectives. This algorithm generates a series of non-dominated solutions, i.e., the Pareto front, through evolutionary computing methods. Each solution represents a possible combination of regulation strategies, such as adjusting the speed limit, implementing traffic control, arranging road repairs, etc. Evaluate the candidate solution set to obtain weighted candidate plans. The evaluation process considers the performance of each plan on each objective and weights them according to the decision-maker's preferences. The evaluation indicators may include the degree of risk reduction, the improvement effect of traffic flow, the input-output ratio, etc. The weighting method uses the Analytic Hierarchy Process (AHP) or the fuzzy comprehensive evaluation method, considering both quantitative and qualitative factors.

[0107] Then, the risk distribution map is subjected to causal relationship identification to obtain a causal graph. Causal relationship identification uses structural equation models (SEM) or Bayesian networks to analyze the mutual influence between various risk factors. For example, it is identified how the slipperiness of the road surface affects the braking distance of the vehicle, which in turn affects the probability of an accident. The causal graph is represented in the form of a directed acyclic graph (DAG), where nodes represent risk factors and edges represent causal relationships. Probabilistic reasoning is performed on the causal graph to obtain the probability distribution of events. Probabilistic reasoning uses the Bayesian reasoning method to calculate the probability of each event occurring under given observation conditions. For example, based on the current weather conditions, traffic flow, and road conditions, the probability distribution of traffic accidents within the next 2 hours is calculated.

[0108] The intervention effect is evaluated on the event probability distribution through the causal intervention analysis algorithm to obtain the intervention plan set. Causal intervention analysis is based on the do-calculus theory and simulates the impact of different intervention measures on the system. For example, evaluate the effect of reducing the speed limit on reducing the probability of accidents, or analyze the effect of increasing the road friction coefficient on improving driving safety. The intervention plan set contains various possible control measures and their expected effects. The intervention plan set is matched with the weighted candidate plans to obtain the control strategy set. The matching process uses heuristic algorithms, such as simulated annealing or genetic algorithms, to find the best combination of intervention measures. Matching criteria include consistency between the intervention effect and the goals of the candidate plans, implementation difficulty, resource requirements, etc.

[0109] Finally, the control strategy set is adaptively adjusted to obtain the adaptive traffic control instruction set. The adaptive adjustment takes into account the real-time road condition changes and historical data feedback to dynamically optimize the control strategy. Reinforcement learning algorithms such as Q-learning or deep Q network (DQN) are used to continuously optimize the decision-making process and improve the control effect.

[0110] For example, a highway section (100 km long) uses this method for traffic control. First, a multi-objective function is constructed, including minimizing accident risk (f1), maximizing traffic efficiency (f2), and minimizing maintenance cost (f3). Constraints include an annual budget of 100 million yuan and an average speed of no less than 80 km / h. The NSGA-II algorithm generates 100 Pareto optimal solutions, each of which contains a different combination of control strategies. In the scheme evaluation stage, the AHP method is used to assign weights [0.5, 0.3, 0.2] to the three objectives. After evaluation, the top 10 weighted candidate schemes are selected. Causal relationship identification is analyzed through structural equation modeling, and a causal graph with 20 nodes and 35 edges is obtained. Bayesian reasoning calculates the accident probability distribution of each section in the next 4 hours based on real-time data (such as current traffic volume of 1500 vehicles / hour and rainfall of 5mm / h).

[0111] Causal intervention analysis evaluated various measures, such as reducing the speed limit by 20 km / h on high-risk sections (10 - 20 km and 60 - 70 km), which is expected to reduce the accident probability by 30%. The matching process selected the optimal combination of control strategies: set the speed limit to 100 km / h on the 10 - 20 km section, implement temporary traffic control on the 60 - 70 km section, and add 4 mobile speed measurement points throughout the line. Adaptive adjustment uses the DQN algorithm to dynamically adjust the strategy according to real-time traffic flow data and weather changes. For example, when the rainfall increases to 10 mm / h, automatically further reduce the speed limit to 90 km / h and activate the road surface drainage system. This adaptive control method can flexibly adjust the strategy according to the actual situation, effectively improving road safety and traffic efficiency, while optimizing resource utilization.

[0112] The above described the method for road condition monitoring and recognition based on machine vision in the embodiments of the present application. Next, the device for road condition monitoring and recognition based on machine vision in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the device for road condition monitoring and recognition based on machine vision in the embodiments of the present application includes:

[0113] The acquisition module 201 is used to collect multi-source data of the road environment to obtain an original data set, where the original data set includes: high-definition image stream, infrared thermal map, laser point cloud, millimeter-wave radar data, and environmental parameters;

[0114] The fusion module 202 is used to perform dynamic time warping and multi-modal feature fusion processing on the original data set to obtain a spatio-temporal feature tensor;

[0115] The extraction module 203 is used to perform multi-scale pyramid decomposition and deep feature extraction on the spatio-temporal feature tensor to obtain a road condition feature vector;

[0116] The parsing module 204 is used to perform semantic parsing on the road condition feature vector through feature mapping and knowledge graph enhancement to obtain a road condition attribute set, where the road condition attribute set includes road surface type, damage state, attachment distribution, and traffic flow parameters;

[0117] The simulation module 205 is used to perform multi-level risk quantification and spatio-temporal propagation simulation on the road condition attribute set to obtain a hierarchical warning index and a risk distribution map;

[0118] The inference module 206 is used to perform multi-objective trade-off and causal chain reasoning on the hierarchical warning index and the risk distribution map to obtain an adaptive traffic control instruction set.

[0119] Through the collaborative cooperation of the above-mentioned various components, multi-source data acquisition significantly improves the comprehensiveness and reliability of data, including the comprehensive utilization of high-definition image streams, infrared thermal maps, laser point clouds, millimeter-wave radar data, and environmental parameters, enabling road condition monitoring to no longer be limited to a single data source and greatly enhancing the system's perception ability of complex road conditions. The dynamic time warping and multi-modal feature fusion technologies effectively solve the spatio-temporal alignment problem of heterogeneous data, ensuring the consistency and comparability of data from different sources. The multi-scale pyramid decomposition and deep feature extraction methods can capture the multi-scale features of road conditions, fully representing both the macroscopic road network structure and the microscopic road surface details, improving the accuracy and robustness of feature extraction. The introduction of feature mapping and knowledge graph enhancement combines machine learning with domain expert knowledge, greatly improving the accuracy and interpretability of road condition semantic understanding. The multi-level risk quantification and spatio-temporal propagation simulation technologies can not only evaluate the current road condition risks but also predict the evolution trends of risks, providing a scientific basis for proactive prevention and timely intervention. Finally, the application of multi-objective trade-off and causal chain reasoning enables the system to find the optimal balance among multiple objectives such as safety, traffic efficiency, and maintenance costs, and formulate more accurate and effective control strategies through causal analysis. Reduce road maintenance costs through precise maintenance strategies.

[0120] Based on the same inventive concept, an embodiment of the present application further provides an electronic device. Referring to Figure 3 As shown, it is a schematic structural diagram of an electronic device 300 provided by an embodiment of the present application, including a processor 301, a memory 302, and a bus 303. Among them, the memory 302 is used to store execution instructions, including an internal memory 3021 and an external memory 3022; here, the internal memory 3021 is also called the main memory, which is used to temporarily store the operation data in the processor 301 and the data exchanged with the external memory 3022 such as a hard disk. The processor 301 exchanges data with the external memory 3022 through the internal memory 3021. When the electronic device 300 runs, communication between the processor 301 and the memory 302 is carried out through the bus 303.

[0121] The present application also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is made to execute the steps of the road condition monitoring and recognition method based on machine vision.

[0122] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0123] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0124] As described above, the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. A road condition monitoring and recognition method based on machine vision, characterized in that, The machine vision-based road condition monitoring and recognition method includes: Performing multi-source data collection on the road environment to obtain an original data set, where the original data set includes: high-definition image stream, infrared thermal map, laser point cloud, millimeter-wave radar data, and environmental parameters; Performing dynamic time warping and multi-modal feature fusion processing on the original data set to obtain a spatio-temporal feature tensor; Performing multi-scale pyramid decomposition and depth feature extraction on the spatio-temporal feature tensor to obtain a road condition feature vector; Performing semantic parsing on the road condition feature vector through feature mapping and knowledge graph enhancement to obtain a road condition attribute set, where the road condition attribute set includes road surface type, damage status, attachment distribution, and traffic flow parameters; Performing multi-level risk quantification and spatio-temporal propagation simulation on the road condition attribute set to obtain a hierarchical warning index and a risk distribution map; Performing multi-objective trade-off and causal chain reasoning on the hierarchical warning index and the risk distribution map to obtain an adaptive traffic control instruction set.

2. The road condition monitoring and recognition method based on machine vision according to claim 1, wherein, The performing multi-source data collection on the road environment to obtain an original data set, where the original data set includes: high-definition image stream, infrared thermal map, laser point cloud, millimeter-wave radar data, and environmental parameters, includes: Performing high-resolution scanning on the visible light scene in the road environment to obtain a high-definition image stream, and performing infrared imaging on the thermal radiation scene in the road environment to obtain an infrared thermal map; Performing three-dimensional laser scanning on the road environment to obtain a laser point cloud, and detecting moving targets in the road environment through a millimeter-wave radar to obtain millimeter-wave radar data; Performing multi-parameter collection on the meteorological conditions of the road environment to obtain environmental parameters; Performing time synchronization processing on the high-definition image stream, infrared thermal map, laser point cloud, millimeter-wave radar data, and environmental parameters through a timestamp alignment algorithm to obtain time-aligned data, and performing coordinate system conversion and spatial correspondence processing on the time-aligned data through a spatial registration algorithm to obtain spatially consistent data; Performing data cleaning and outlier detection on the spatially consistent data to obtain preprocessed data, and performing feature extraction on the preprocessed data through a multi-modal feature extraction algorithm to obtain the original data set.

3. The road condition monitoring and recognition method based on machine vision according to claim 1, characterized in that The performing dynamic time warping and multi-modal feature fusion processing on the original data set to obtain a spatio-temporal feature tensor, includes: Performing time series analysis on each data stream in the original data set to obtain a timestamp sequence, and performing alignment processing on the timestamp sequence to obtain a regularized time series; Performing interpolation processing on the regularized time series to obtain a data sequence with a unified sampling rate, and segmenting the data sequence with the unified sampling rate to obtain a time window data set; Performing feature extraction on each data modality in the time window data set to obtain a multi-modal feature set, and performing scale adjustment on the multi-modal feature set to obtain a normalized feature set; Performing cross-modal feature fusion on the normalized feature set through a multi-modal feature fusion algorithm to obtain a fusion feature vector, and performing temporal correlation analysis on the fusion feature vector to obtain a temporal feature matrix; Performing spatial relationship analysis on the temporal feature matrix through a spatial correlation analysis algorithm to obtain the spatio-temporal feature tensor.

4. The method for monitoring and recognizing road conditions based on machine vision according to claim 1, characterized in that Performing multi-scale pyramid decomposition and deep feature extraction on the spatio-temporal feature tensor to obtain a road condition feature vector, including: Performing multi-scale decomposition on the spatio-temporal feature tensor through a Gaussian pyramid algorithm to obtain a multi-scale feature pyramid, and performing differential processing on the multi-scale feature pyramid through a Laplacian pyramid algorithm to obtain multi-scale differential features; Enhancing the features of the multi-scale differential features through convolution operations to obtain an enhanced feature map, and reducing the dimension of the enhanced feature map through a pooling operation to obtain a compressed feature representation; Performing feature transformation on the compressed feature representation through a non-linear activation function to obtain an activated feature map, and performing multi-layer feature fusion on the activated feature map to obtain a fused feature representation; Performing key information extraction on the fused feature representation through an attention mechanism to obtain an attention-weighted feature, and reducing the dimension of the attention-weighted feature through a fully connected layer to obtain the road condition feature vector.

5. The method for road condition monitoring and recognition based on machine vision according to claim 1, wherein Performing semantic parsing on the road condition feature vector through feature mapping and knowledge graph enhancement to obtain a road condition attribute set, where the road condition attribute set includes road surface type, damage status, attachment distribution, and traffic flow parameters, including: Performing semantic space projection on the road condition feature vector through a feature mapping algorithm to obtain a semantic feature representation, and performing feature grouping on the semantic feature representation to obtain a preliminary semantic category; Performing concept alignment on the preliminary semantic category through a knowledge graph matching algorithm to obtain a semantic concept set, and constructing a semantic relationship network for the semantic concept set; Identifying the road surface type for the semantic relationship network to obtain the road surface type, and analyzing the road surface damage for the semantic relationship network to obtain the damage status; Identifying the road surface attachments for the semantic relationship network to obtain the attachment distribution, and calculating the traffic flow parameters for the semantic relationship network to obtain the traffic flow parameters; Integrating the road surface type, damage status, attachment distribution, and traffic flow parameters through an attribute aggregation algorithm to obtain the road condition attribute set.

6. The method for road condition monitoring and recognition based on machine vision according to claim 1, wherein Performing multi-level risk quantification and spatio-temporal propagation simulation on the road condition attribute set to obtain a hierarchical warning index and a risk distribution map, including: Identifying risk factors for the road condition attribute set to obtain a risk factor set, and sorting the importance of the risk factor set to obtain weighted risk factors; Performing hierarchical processing on the weighted risk factors through a multi-level analysis method to obtain a risk hierarchy structure, and dividing the risk levels of the risk hierarchy structure to obtain hierarchical risk indicators; Estimating the spatial distribution of the hierarchical risk indicators through a spatio-temporal interpolation algorithm to obtain a risk spatial distribution map, and performing dynamic evolution simulation on the risk spatial distribution map through a cellular automaton algorithm to obtain a risk propagation model; Performing multi-scenario simulation on the risk propagation model to obtain a risk evolution sequence, and calculating the probability distribution of the risk evolution sequence to obtain a risk probability map; Divide the risk probability map into risk regions to obtain the hierarchical warning indicators, and graphically represent the hierarchical warning indicators to obtain the risk distribution map.

7. The method for monitoring and recognizing road conditions based on machine vision according to claim 1, characterized in that Perform multi-objective trade-off and causal chain reasoning on the hierarchical warning indicators and the risk distribution map to obtain an adaptive traffic control instruction set, including: Construct a multi-objective function for the hierarchical warning indicators to obtain a multi-objective optimization problem, and set constraints for the multi-objective optimization problem to obtain an optimization problem with constraints; Perform multi-objective solution for the optimization problem with constraints through the Pareto optimization algorithm to obtain a candidate solution set, and evaluate the candidate solution set to obtain a weighted candidate solution; Identify the causal relationship of the risk distribution map to obtain a causal map, and perform probability reasoning on the causal map to obtain an event probability distribution; Evaluate the intervention effect of the event probability distribution through a causal intervention analysis algorithm to obtain an intervention plan set, and match the intervention plan set with the weighted candidate solution to obtain a control strategy set; Perform adaptive adjustment on the control strategy set to obtain the adaptive traffic control instruction set.

8. A road condition monitoring and recognition device based on machine vision, which is used to implement the road condition monitoring and recognition method based on machine vision as described in any one of claims 1-7, characterized in that, The road condition monitoring and recognition device based on machine vision includes: An acquisition module for collecting multi-source data of the road environment to obtain an original data set, where the original data set includes: high-definition image stream, infrared thermal map, laser point cloud, millimeter wave radar data, and environmental parameters; A fusion module for performing dynamic time warping and multi-modal feature fusion processing on the original data set to obtain a spatio-temporal feature tensor; An extraction module for performing multi-scale pyramid decomposition and deep feature extraction on the spatio-temporal feature tensor to obtain a road condition feature vector; An analysis module for semantically analyzing the road condition feature vector through feature mapping and knowledge graph enhancement to obtain a road condition attribute set, where the road condition attribute set includes road surface type, damage status, attachment distribution, and traffic flow parameters; A simulation module for performing multi-level risk quantification and spatio-temporal propagation simulation on the road condition attribute set to obtain hierarchical warning indicators and a risk distribution map; An inference module for performing multi-objective trade-off and causal chain reasoning on the hierarchical warning indicators and the risk distribution map to obtain an adaptive traffic control instruction set.

9. A computer device, characterized in that, Including: A processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the road condition monitoring and recognition method based on machine vision according to any one of claims 1 to 7 are executed.

10. A computer-readable storage medium, on which instructions are stored, characterized in that, When the instructions are executed by the processor, the road condition monitoring and recognition method based on machine vision according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Method for managing sky-ground multi-source heterogeneous data

    CN115203241A

  • Multi-source sensor data spatio-temporal feature fusion elevator guide system anomaly detection method

    CN118495280A