Multi-mode road abnormity real-time monitoring system and method thereof

By multimodally fusing lidar point cloud and video image data, combined with 5G edge computing and dynamic weight adjustment, high-precision road anomaly detection is achieved around the clock, solving the problems of poor lighting conditions and insufficient environmental adaptability in existing technologies, and improving detection accuracy and response speed.

CN120635841APending Publication Date: 2025-09-12HARBIN ENG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510728288.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing road anomaly detection methods lack accuracy under poor lighting conditions and find it difficult to achieve high-precision real-time monitoring in all weather conditions. Traditional methods are inefficient and have poor environmental adaptability.

Method used

By adopting multimodal fusion technology, combining lidar point cloud data and video image data, and using 5G edge computing technology, the YOLOv7-pointnet network model is used for data processing, and combined with the spatiotemporal attention mechanism and dynamic weight adjustment, high-precision road anomaly detection can be achieved around the clock.

Benefits of technology

It improves detection accuracy by 15%-20%, enhances environmental adaptability, reduces false detection rate by 40%, and increases response speed to within 150ms, making it suitable for a variety of scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635841A_ABST
    Figure CN120635841A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent traffic, in particular to a multi-mode road abnormity real-time monitoring system and a method thereof.The system comprises a laser radar module and a camera module which collect road point cloud and video data respectively, a 5G edge computing unit receives and processes the data, a YOLOv7-pointnet network model is adopted, and a real-time monitoring result is obtained; according to the model, an EfficentNet backbone feature extraction network and a point cloud feature extraction module CFFM are connected in series in a jumping manner to form a 3D feature extraction network, so that effective fusion of 2D and 3D features is realized; the data post-processing module optimizes the processing result and sends the result to the server through the data transmission module; the server judges road abnormal events, positions and classifies the road abnormal events, and plans alarm and maintenance paths, and the control module dynamically adjusts the confidence coefficient weight of the detection model according to illumination conditions; the system makes full use of image and point cloud features through a multi-modal fusion technology, and effectively improves the precision and real-time performance of road anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation, and in particular to a multimodal road anomaly real-time monitoring system and method thereof, which can monitor, identify and issue early warnings on abnormal conditions on the road surface in real time. Background Art

[0002] As urbanization continues to deepen, the scale of road infrastructure construction continues to expand, posing significant challenges to road maintenance. Road surface anomalies such as potholes, cracks, bulges, and depressions not only impact driving comfort but can also lead to traffic accidents, posing a threat to people's lives and property. Traditional road maintenance relies primarily on manual inspections, which are inefficient and susceptible to subjective factors.

[0003] Existing automated road anomaly detection methods fall into two main categories: image-based detection methods and lidar point cloud-based detection methods. Image-based detection methods are limited by lighting conditions, with detection performance significantly declining in inclement weather or low light conditions. While lidar point cloud-based detection methods are unaffected by lighting, they lack sensitivity for detecting low-level anomalies such as tiny cracks. Single-modality detection methods struggle to meet the demands for all-weather, high-precision road anomaly detection.

[0004] Furthermore, most existing detection systems utilize centralized processing architectures, which lack real-time performance and struggle to cope with complex and changing road environments. Effectively integrating data from multiple sensors to achieve high-precision, all-weather, real-time monitoring of road anomalies is a pressing challenge in the intelligent transportation sector. Summary of the Invention

[0005] The purpose of the present invention is to provide a multimodal road anomaly real-time monitoring system and method, aiming to overcome the problems of insufficient single-modal detection accuracy and poor environmental adaptability in the existing technology. By fusing lidar point cloud data and video image data and combining 5G edge computing technology, all-weather, high-precision, and low-latency road anomaly detection can be achieved.

[0006] The present invention proposes a multi-modal road anomaly real-time monitoring system, comprising:

[0007] LiDAR module, used to collect LiDAR point cloud data of roads;

[0008] A camera module for collecting road video data;

[0009] A 5G edge computing unit is communicatively connected to the lidar module and the camera module, and the 5G edge computing unit includes:

[0010] A data access module, configured to receive the laser radar point cloud data and the road video data;

[0011] A data processing module is used to process the lidar point cloud data and the road video data. The data processing module includes a YOLOv7-pointnet network model. The YOLOv7-point net network model includes a backbone feature extraction module network, a Neck, and an output module. The backbone feature extraction module network adopts an EfficientNet structure. The Neck part includes a 2D feature extraction network and a 3D feature extraction network. The 3D feature extraction network is composed of a point cloud feature extraction module CFFM connected in series with jumps. The number of channels of the CFFM includes M+1 groups of input channels and M groups of output channels, wherein one group of input channels is a global average pooling channel and the M groups of output channels are composed of grouped convolution GCs.

[0012] A data post-processing module, used for post-processing the results processed by the data processing module;

[0013] A data transmission module, used for transmitting the results processed by the data post-processing module to a server;

[0014] A control module is used to dynamically adjust the confidence of the detection model according to the lighting conditions;

[0015] The server is communicatively connected to the 5G edge computing unit, and is used to receive the results transmitted by the data transmission module and determine whether there is a road abnormality event. When there is a road abnormality event, it locates the abnormal event, determines the type of abnormal event, and plans corresponding alarm paths and maintenance paths.

[0016] Preferably, the 5G edge computing unit further includes: a power supply module for providing power to the 5G edge computing unit.

[0017] Preferably, the data processing module includes:

[0018] An image detection module, configured to perform real-time detection on the road video data;

[0019] A point cloud detection module, used to detect the laser radar point cloud data;

[0020] The abnormality judgment module is used to determine whether an abnormal event occurs by combining the laser radar point cloud data and the road video data.

[0021] Preferably, the point cloud detection module includes:

[0022] A point cloud preprocessing module, configured to perform grayscale conversion, threshold filtering, and downsampling processing on the laser radar point cloud data;

[0023] A point cloud feature extraction module is used to perform semantic segmentation and key point extraction on the pre-processed lidar point cloud data to obtain point cloud features;

[0024] A point cloud feature fusion module is used to combine the point cloud features and image features to provide global context information for the detection unit;

[0025] The point cloud detection model is a YOLOv7 model, whose network structure includes 5 downsampling modules, 5 upsampling modules, and 8 residual modules.

[0026] Preferably, the detection unit in the point cloud detection model includes:

[0027] The initial information receiving module is used to receive the image features transmitted from the downsampling module;

[0028] a fusion feature acquisition module, configured to fuse the image features received from the initial information receiving module with the global context obtained from the point cloud feature fusion module to obtain a fusion feature;

[0029] A detection task sharing module, used to fuse the fused features through a 1×1 convolution layer and an activation layer;

[0030] The image detection result output module is used to input the fused information into the neural network branches corresponding to each detection target to predict the target category.

[0031] Preferably, the abnormality judgment module further includes: a large model, and the large model includes a visual large model and a semantic segmentation model.

[0032] Preferably, the abnormality judgment module is further configured to perform spatiotemporal attention processing on the detection results to eliminate redundant detection frames in continuous video frames; the abnormality judgment module further comprises:

[0033] The spatiotemporal attention mechanism is used to calculate the spatial similarity of each pixel between the target image and the K frames before and after the target image and the video frames of the K frames before and after the target image, and define the pixels with spatial similarity greater than a threshold as a saliency detection frame; the spatiotemporal attention mechanism calculates the attention weight of each time point between the target image and the K frames through the saliency detection frame, and obtains the attention weight of the target image frame and the K frames before and after the target image frame; using the attention weight of the target image frame, the anomaly results predicted by the semantic segmentation model detection and the anomaly results predicted by the visual model detection are weighted averaged to obtain the final anomaly detection result.

[0034] Preferably, the weighted average formula is:

[0035] A(m,n)=A vis (m,n)×(1-P(am,n ))+A sem (m,n)×P(a m,n )×ω m,n ,

[0036] Among them, A is the final anomaly detection result, A vis is the abnormal result predicted by the visual model detection; A sem is the abnormal result predicted by the semantic segmentation model detection; m, n correspond to the coordinate position of each pixel in the final result; m∈[0,6], n∈[0,6]; ω m,n is the attention weight calculated by the spatiotemporal attention mechanism; P(a m,n ) is whether pixels m and n belong to the same target.

[0037] Preferably, the attention weight formula is:

[0038] ω i =(1-sI)(i∈[m,n])σ1+(1-sI)(i∈[n,m])σ2,

[0039] Among them, ω i is the attention weight of the current pixel; σ1 and σ2 are attention coefficients; sI is the pixel space similarity.

[0040] The operating method of the multimodal road anomaly real-time monitoring system includes the following steps:

[0041] S1: The LiDAR module and camera module collect LiDAR point cloud data and road video data in real time, and transmit the collected LiDAR point cloud data, road video data or abnormal results to the server through the data transmission module;

[0042] S2: The server determines whether there is a road abnormality event based on the abnormal result, and locates the abnormal event if there is a road abnormality event;

[0043] S3: Determine the type of abnormal event and plan the corresponding warning path and maintenance path. Based on the different types of road abnormalities identified, send an alert and repair to the corresponding maintenance department, and display the abnormal event location, type, and warning path on the client.

[0044] S4: Dynamically assign weights to the detection model and anomaly judgment module.

[0045] The present invention has the following beneficial effects:

[0046] 1. Improved detection accuracy: Multimodal fusion technology makes full use of the texture features of image data and the geometric features of point cloud data. Compared with single-modality detection methods, the detection accuracy is improved by 15% to 20%.

[0047] 2. Enhanced environmental adaptability: Through the light condition adaptive mechanism, the system can automatically adjust the weight of each sensor data according to different lighting conditions during the day / night, improving night detection performance by more than 35%.

[0048] 3. Reduced false detection rate: The spatiotemporal attention mechanism effectively filters out noise interference by analyzing the spatiotemporal correlation of the target in multiple frames, reducing the false detection rate by 40%.

[0049] 4. Improved response speed: The 5G edge computing architecture achieves detection response within 150ms, meeting real-time monitoring needs.

[0050] 5. Enhanced application flexibility: The system supports deployment at fixed monitoring points and mobile vehicle-mounted deployment, and is suitable for various scenarios such as urban roads and highways. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 Schematic diagram of the overall architecture of the multi-modal road anomaly real-time monitoring system of the present invention;

[0052] Figure 2 This is a schematic diagram of the structure of the YOLOv7-pointnet network model of the present invention;

[0053] Figure 3 This is a schematic diagram of the point cloud data preprocessing process of the present invention;

[0054] Figure 4 Schematic diagram of the workflow of the spatiotemporal attention mechanism of the present invention;

[0055] Figure 5 Schematic diagram of multimodal data fusion of the present invention;

[0056] Figure 6 This is a comparison chart of the detection effects of the system of the present invention under different lighting conditions;

[0057] Figure 7 Schematic diagram of the overall process of the method of the present invention. DETAILED DESCRIPTION

[0058] Please refer to the attached Figure 1-7 The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited thereto.

[0059] Reference Figure 1 The multimodal road anomaly real-time monitoring system provided by the present invention includes a lidar module 10, a camera module 20, a 5G edge computing unit 30, a server 40 and a client 50.

[0060] The laser radar module 10 is used to collect laser radar point cloud data of the road. In a preferred embodiment of the present invention, the laser radar module 10 uses a single-line laser radar with a detection frequency of 20Hz, which can provide high-precision three-dimensional point cloud data and effectively capture the geometric features of the road surface.

[0061] The camera module 20 is used to collect road video data. Preferably, the camera module 20 uses a high-definition camera with a resolution of 480×1280 pixels, which can provide clear road surface texture information and help identify minor road cracks and other anomalies.

[0062] The 5G edge computing unit 30 is in communication with the lidar module 10 and the camera module 20 to receive and process sensor data. The 5G edge computing unit 30 includes a data access module 31, a data processing module 32, a data post-processing module 33, a data transmission module 34, and a control module 35. Furthermore, the 5G edge computing unit 30 may also include a power supply module 36 for providing power to the 5G edge computing unit 30.

[0063] The data access module 31 is used to receive LiDAR point cloud data and road video data. In practical applications, the data access module 31 can receive sensor data via USB 3.0, Ethernet, or wirelessly, and supports parsing and processing of multiple data formats.

[0064] The data processing module 32 is used to process LiDAR point cloud data and road video data. The core of the data processing module 32 is the YOLOv7-PointNet network model, which will be described in detail in subsequent embodiments. The data processing module 32 also includes an image detection module 321, a point cloud detection module 322, and an anomaly detection module 323, responsible for processing video data and point cloud data, and detecting anomalies, respectively.

[0065] The data post-processing module 33 is used to perform post-processing on the results processed by the data processing module 32, including algorithms such as time sequence detection, spatiotemporal joint detection, and joint detection, so as to improve the reliability of the detection results.

[0066] The data transmission module 34 is used to transmit the processing results to the server 40. The data transmission module 34 includes a real-time transmission module and an offline transmission module, and can select different transmission modes according to network conditions and data importance.

[0067] The control module 35 is used to dynamically adjust the confidence of the detection model according to the lighting conditions to improve the adaptability of the system in different environments.

[0068] Server 40 is in communication with 5G edge computing unit 30 and is used to receive the results transmitted by data transmission module 34 and perform further analysis and processing. When server 40 determines that a road anomaly has occurred, it locates the anomaly, determines the type of anomaly, and plans the corresponding warning and maintenance routes.

[0069] Client 50 is used to display abnormal road conditions to users and supports PC and mobile access, making it convenient for maintenance departments to understand road conditions in a timely manner and take corresponding measures.

[0070] Reference Figure 2 One of the core technologies of the present invention is the YOLOv7-pointnet network model, which includes three parts: the backbone feature extraction module network, Neck and output module.

[0071] The backbone feature extraction module network uses the Efficient Net architecture, comprised of standard convolutions. Its channel count is 3 × scale × groups, where scale represents the scaling factor, controlling the network width; groups represents the number of groupings, controlling the refinement of feature extraction. In a preferred embodiment, scale is set to 1.2 and groups to 32. This parameter setting ensures that feature extraction capabilities are maintained while limiting computational resource consumption.

[0072] The Neck section consists of a 2D feature extraction network and a 3D feature extraction network. The 2D feature extraction network consists of standard convolutions in series with skip connections. The standard convolution channel count is 3×scale1×groups1 and 3×scale2×groups2, where scale1 and scale2 represent the scaling ratios of different paths, and groups1 and groups2 represent the number of groups in the corresponding path. In practice, scale1 is set to 1.0 and groups1 is set to 16; scale2 is set to 1.5 and groups2 is set to 24. This configuration effectively extracts multi-scale image features.

[0073] The 3D feature extraction network consists of a series of point cloud feature extraction modules (CFFMs) connected in series with skip connections. The CFFM (Cloud Feature Fusion Module) has M+1 input channels and M output channels. One input channel is a global average pooling channel used to capture global features, and the M output channels are grouped convolutions (GCs). M depends on the scale. Experiments show that an M value of 8 achieves optimal results when the scale is 1.2.

[0074] The specific structure of the CFFM module is as follows: First, global average pooling is performed on the input point cloud features to obtain a global feature vector; then, the original point cloud features are concatenated with the global features and feature extraction is performed using grouped convolution (GC); finally, a nonlinear activation function is applied in conjunction with skip connections to obtain enhanced point cloud features. This design simultaneously considers local geometric structure and global semantic information, improving feature expression capabilities.

[0075] The output module includes 2D and 3D prediction outputs. The 2D prediction module outputs 2D bbox coordinates and 2D mean Intersection over Union (mIoU) for anomaly localization and assessment in 2D images. The 3D prediction output includes 3D bbox coordinates, point features, and instance segmentation. 3D bbox coordinates are used to locate the target point cloud, while point features and instance segmentation are used to assist training and improve model performance.

[0076] Reference Figure 3 The preprocessing of laser radar point cloud data in the present invention includes the following steps:

[0077] First, we downsample the single-frame point cloud to 1,000 points using Kd-tree downsampling to reduce computational complexity. Kd-tree is a spatial partitioning data structure that ensures uniform spatial distribution of sampling points and avoids information loss.

[0078] Then, the mean of the three-dimensional coordinates of the lidar before downsampling is calculated and recorded as Subtract the mean from the coordinates of each point to make the center of gravity of the point cloud coordinates zero. This step can improve the robustness of the model to changes in the position of the point cloud. The specific calculation formula is:

[0079]

[0080] Among them, x i ,y i ,z i are the horizontal, vertical and vertical coordinates of each point before downsampling; is the mean of the corresponding coordinates; x′ i ,y′ i ,z′ i are the coordinates after the center of gravity returns to zero.

[0081] Next, we traverse the downsampled point cloud and calculate the local density of each point. Local density is an important indicator to characterize the geometric characteristics of the point cloud. The calculation formula is:

[0082]

[0083] Among them, ρ i is the local density of point i, Ni is the total number of points within a spherical neighborhood with radius r centered at point i, where r is the radius of the spherical neighborhood. In practice, r is typically set to 0.5 meters. This parameter setting ensures that the local feature capture is captured while limiting the amount of computation required.

[0084] Finally, the local density of each point is used as channel data and superimposed on the point cloud to form a 5-channel local feature [x, y, z, intensity, ρ], where intensity is the reflection intensity of the point cloud. This processed point cloud data contains richer feature information, which is beneficial for subsequent anomaly detection.

[0085] Reference Figure 4 The point cloud detection module 322 includes a point cloud preprocessing module 3221, a point cloud feature extraction module 3222, a point cloud feature fusion module 3223 and a point cloud detection model 3224.

[0086] Point cloud preprocessing module 3221 performs grayscale conversion, threshold filtering, and downsampling on the LiDAR point cloud data. Grayscale conversion maps the intensity values ​​of the point cloud to the range [0, 1]. Threshold filtering removes noise points, typically setting the intensity threshold to 0.1 to remove points with low reflection intensity. Downsampling uses the aforementioned Kd-tree method to limit the number of point clouds to 1000 points.

[0087] The point cloud feature extraction module 3222 is used to perform semantic segmentation and keypoint extraction on the preprocessed point cloud data to obtain point cloud features. This module uses the PointNet++ architecture, effectively capturing local and global features of the point cloud through hierarchical sampling and grouping operations. In one embodiment of the present invention, point cloud feature extraction uses a three-layer hierarchical structure, with each layer sampling 512, 256, and 128 points, respectively, resulting in feature dimensions of 64, 128, and 256, respectively.

[0088] The point cloud feature fusion module 3223 is used to combine point cloud features and image features to provide global context information for the detection unit. Feature fusion uses an attention mechanism to perform adaptive weighting based on the importance of different features. Specifically, for the point cloud feature F point and image features F image , fusion feature F fusion The calculation formula is:

[0089]

[0090] Among them, α is the weight coefficient, which ranges from [0, 1] and is used to balance the importance of point cloud features and image features; β is the interaction coefficient, which is used to control the intensity of the interaction between the two features; Indicates feature interaction operations, which are usually implemented using an attention mechanism. In a preferred embodiment of the present invention, the initial value of α is set to 0.5 and β is set to 0.3, and can be dynamically adjusted according to the actual detection effect.

[0091] Point cloud detection model 3224 is a YOLOv7 model. Its network architecture consists of five downsampling modules, five upsampling modules, and eight residual modules. The downsampling modules halve the feature map size through convolutions with a stride of 2; the upsampling modules increase the feature map size through transposed convolutions; and the residual modules contain multiple convolutional layers and skip connections to extract deep features. This design ensures a sufficiently large receptive field while maintaining detection accuracy, making it particularly suitable for detecting small objects.

[0092] Reference Figure 4 The detection unit in the point cloud detection model 3224 includes an initial information receiving module 32241, a fusion feature acquisition module 32242, a detection task sharing module 32243 and an image detection result output module 32244.

[0093] The initial information receiving module 32241 is used to receive image features passed from the downsampling module. These features are typically multi-scale and contain semantic information at different levels. In the YOLOv7 model, three feature maps of different scales are typically used, corresponding to shallow, mid-level, and deep features.

[0094] The fusion feature acquisition module 32242 is used to fuse the image features received from the initial information receiving module 32241 with the global context obtained from the point cloud feature fusion module 3223 to obtain the fusion feature. The fusion process adopts the attention mechanism, and the calculation formula is as follows:

[0095] F fused =F image +Attention(F image ,F context ),

[0096] Among them, F fused is the fused feature, F image is the image feature, F context is the global context feature, Attention is the attention operation, and its calculation process is:

[0097]

[0098] Among them, Softmax is a soft maximization function used to normalize the attention weight; · represents matrix multiplication; T represents the transpose operation.

[0099] The detection task sharing module 32243 is used to fuse the fused features through a 1×1 convolution layer and an activation layer. The 1×1 convolution is used to adjust the number of feature channels and fuse information from different channels. The activation layer usually uses the SiLU activation function, which is calculated as follows:

[0100] SiLU(x)=x·Sigmoid(x),

[0101] Among them, Sigmoid is an S-type function, and its calculation formula is:

[0102]

[0103] Image detection result output module 32244 is used to input the fused information into the neural network branches corresponding to each detection target to predict the target category. In the road anomaly detection task, there are four main types of anomalies: potholes, cracks, bulges, and depressions, each of which corresponds to a detection branch. Each branch consists of a classification head and a regression head. The classification head predicts the target category probability, while the regression head predicts the bounding box coordinates and size.

[0104] Reference Figure 5 The abnormality judgment module 323 also includes a large model 3231, and the large model 3231 includes a visual large model 32311 and a semantic segmentation model 32312.

[0105] The Vision Model 32311 is based on a Transformer architecture and is capable of capturing the global context of an image. In a preferred embodiment of the present invention, the Vision Model 32311 employs a ViT (Vision Transformer) architecture, which segments the input image into 16×16 blocks and then extracts features using a self-attention mechanism. The ViT model contains 12 Transformer blocks, each with a hidden dimension of 768 and 12 attention heads. This design effectively captures long-range dependencies in images and improves the accuracy of anomaly detection.

[0106] Semantic segmentation model 32312 is used to perform fine segmentation on point cloud data and identify different types of anomalies. In one embodiment of the present invention, semantic segmentation model 32312 employs a U-Net architecture, comprising a five-layer encoder and a five-layer decoder, with features transferred between each layer via skip connections. The encoder gradually reduces the feature map size and increases the number of channels through downsampling, while the decoder restores the feature map size through upsampling and outputs the segmentation results. This encoder-decoder architecture can simultaneously preserve global semantic information and local details, making it suitable for fine segmentation of road anomalies.

[0107] The results of the large visual model 32311 and the semantic segmentation model 32312 are combined through weighted fusion to fully leverage the strengths of both models. The weight coefficients during fusion are dynamically adjusted based on lighting conditions and anomaly types to ensure optimal results in different environments.

[0108] Reference Figure 6 ,The abnormality judgment module 323 also includes a spatiotemporal attention mechanism 3232 for performing spatiotemporal attention processing on the ,detection results and eliminating redundant detection frames in ,continuous video frames.

[0109] The spatiotemporal attention mechanism 3232 calculates the spatial similarity of each pixel between the target image and the K frames before and after it through the target image and its K frames. The calculation of spatial similarity is based on pixel value and position information, and the calculation formula is:

[0110]

[0111] Among them, sI(p i ,p j ) represents pixel p i and p j The spatial similarity between i ) represents pixel p i Pixel value; ||I(p i )-I(p j )|| 2 represents the Euclidean distance of pixel values; ||p i -p j || 2 represents the Euclidean distance of pixel position; σ c and σ s are the standard deviations of pixel values ​​and positions, respectively, which are used to control the sensitivity of the similarity measure. In practice, σ c Usually the value is 0.1,σ s The value is 10.

[0112] The spatiotemporal attention mechanism 3232 defines pixels with spatial similarity greater than a threshold as salient detection boxes. The threshold setting depends on the specific application scenario. In road anomaly detection, the threshold is usually set to 0.7, which effectively balances detection sensitivity and stability.

[0113] The spatiotemporal attention mechanism 3232 calculates the attention weight of each time point between the target image and K frames through the saliency detection box, and obtains the attention weight of the target image frame and its K adjacent frames before and after. The calculation formula of the attention weight is:

[0114] ω i=(1-sI)(i∈[m,n])σ1+(1-sI)(i∈[n,m])σ2,

[0115] Among them, ω i is the attention weight of the current pixel; σ1 and σ2 are attention coefficients, which control the attention of forward and backward time points, respectively; (i∈[m,n]) and (i∈[n,m]) are indicator functions, indicating whether time point i is within the interval [m,n] or [n,m]. In a preferred embodiment of the present invention, σ1 is 0.7, σ2 is 0.3, and K is 5. This set of parameters ensures temporal continuity while focusing on the latest detection results.

[0116] Using the attention weight of the image frame where the target is located, the spatiotemporal attention mechanism 3232 performs a weighted average of the anomaly results predicted by the semantic segmentation model 32312 and the anomaly results predicted by the visual model 32311 to obtain the final anomaly detection result. The calculation formula for the weighted average is:

[0117] A(m,n)=A vis (m,n)×(1-P(a m,n ))+A sem (m,n)×P(a m,n )×ω m,n ,

[0118] Among them, A(m,n) is the final anomaly detection result, corresponding to the degree of anomaly at the coordinate (m,n); A vis (m,n) is the abnormal result predicted by the visual model detection; A sem (m,n) is the abnormal result predicted by the semantic segmentation model detection; m, n correspond to the coordinate position of each pixel in the final result, usually in the range of m∈[0,6], n∈[0,6], indicating that the image is divided into a 7×7 grid; ω m,n The attention weight calculated for the spatiotemporal attention mechanism; P(a m,n ) is the probability that the pixel (m,n) and its adjacent pixels belong to the same target, which is used to smooth the detection results.

[0119] This spatiotemporal attention mechanism can effectively utilize the spatiotemporal information in video sequences, improve the stability and accuracy of detection, and is particularly suitable for continuous monitoring of road anomalies.

[0120] Reference Figure 7 The operating method of the multi-modal road anomaly real-time monitoring system of the present invention comprises the following steps:

[0121] S1: The laser radar module 10 and the camera module 20 collect laser radar point cloud data and road video data in real time, and transmit the collected laser radar point cloud data, road video data or abnormal results to the server 40 through the data transmission module 34.

[0122] In this step, the LiDAR module 10 collects point cloud data at a 20Hz frequency, with each frame containing tens of thousands of points. Each point includes spatial coordinates (x, y, z) and reflection intensity information. The camera module 20 collects video data with a resolution of 480×1280 pixels and a frame rate of 30fps. The data transmission module 34 transmits this data to the server 40 in real time using the 5G network, with transmission latency controlled to less than 20ms.

[0123] S2: The server 40 determines whether there is a road abnormality event based on the abnormality result, and locates the abnormality event if there is a road abnormality event.

[0124] After receiving the abnormality results, server 40 first determines the abnormal event. This determination is based on an abnormality score. When the score exceeds a preset threshold (typically 0.75), an abnormal event is determined. Server 40 then uses GPS information and road map data to accurately locate the abnormal event, with a positioning accuracy of up to 1 meter.

[0125] S3: Determine the type of abnormal event and plan the corresponding warning path and maintenance path; based on the different types of road abnormalities determined, send an alarm and repair to the corresponding maintenance department, and display the abnormal event location, type and warning path on the client 50.

[0126] In this step, the server 40 determines the type of abnormal event based on the detection results, which mainly includes four types: potholes, cracks, bulges and depressions. For different types of abnormalities, the system adopts different processing strategies:

[0127] Road cracks are repaired by the maintenance department's road maintenance vehicles. These vehicles automatically receive repair requests and navigate to the location of the abnormality for repair. Based on the severity of the crack, the system recommends repair methods such as grouting, sealing, or milling and resurfacing.

[0128] Maintenance personnel promptly clean up any bumps or depressions on the road surface. The system pushes task information, including the location, type, and severity of the anomaly, to the maintenance personnel's mobile devices, and provides guidance on the optimal route.

[0129] For damaged roads, maintenance personnel clean them based on reported repair tasks and vehicle GPS locations. The system updates the status of damaged areas in real time, helping maintenance personnel complete their tasks efficiently.

[0130] Ruts and potholes are inspected by the maintenance department's road inspection vehicles. These vehicles automatically receive repair requests, navigate to the location of the anomaly, conduct inspections, and report the results to the maintenance department.

[0131] S4: Dynamically assign weights to the detection model and anomaly judgment module.

[0132] In this step, the system dynamically adjusts the weights of each module based on environmental conditions and detection results to improve system performance. This includes calculating the confidence level of point cloud data, image data, and fusion data under daytime lighting conditions, as well as the corresponding confidence levels under nighttime lighting conditions.

[0133] The confidence calculation formula for point cloud data under daylight conditions is:

[0134]

[0135] Among them, P point is the confidence score of the point cloud data of the camera under daylight conditions; p k is the matching probability between the pixel of the point cloud data and the point cloud data category k under daytime conditions; C point-k is the confidence of point cloud data category k under daytime conditions. Usually, the number of categories K=4, corresponding to four types of road anomalies.

[0136] The confidence calculation formula for image data under daylight conditions is:

[0137]

[0138] Among them, P image is the confidence score of the image data under daylight conditions; q k is the matching probability between the pixel of the image data under daytime conditions and the image data category k; C image-k is the confidence of image data category k under daytime conditions.

[0139] The confidence calculation formula for fusion data under daylight conditions is:

[0140]

[0141] Among them, P fusion is the confidence score of the fused point cloud data and image data; α ,k The confidence weight assigned to category k for point cloud data, β ,k The confidence weight assigned to class k for the image data.

[0142] The confidence weight calculation formula assigned to category k by point cloud data is:

[0143]

[0144] Where A is the focal length of the point cloud image acquisition device, usually 4.5 mm; A″ is the focal length of the depth information acquisition device of the point cloud image acquisition device, usually 3.6 mm; B′ is the field of view height from the camera to the shooting ground, usually 1.8 meters; F is the focal length of the camera, usually 4.5 mm; H is the depth data acquisition distance of the point cloud image acquisition device, usually 20 meters.

[0145] A similar calculation formula applies to nighttime lighting conditions, so I won't go into detail here. This dynamic weight allocation mechanism allows the system to adaptively adjust the weight of each sensor data based on varying environmental conditions, improving detection accuracy and stability.

[0146] In another embodiment of the present invention, the system can automatically adjust the weights of different sensor data according to lighting conditions. Figure 6 ,There are obvious differences in the detection effects under daytime and nighttime lighting conditions.

[0147] During daylight hours, image data typically provides more texture information, facilitating the detection of anomalies like fine cracks. Point cloud data, on the other hand, provides precise geometric information, making it suitable for detecting potholes and uneven surfaces. At night, the quality of image data significantly degrades, making point cloud data even more important.

[0148] The system implements dynamic weight distribution in the following ways:

[0149] First, the system determines the current lighting conditions through a light sensor or image analysis. Lighting judgment is usually based on the average brightness and contrast of the image. When the average brightness falls below a threshold (usually 50, ranging from 0-255), it is determined to be night lighting conditions.

[0150] The system then calculates the confidence level of the point cloud and image data based on the lighting conditions and adjusts the weights accordingly. At night, the weight of the point cloud data is typically increased to 0.7-0.8, while the weight of the image data is reduced to 0.2-0.3. During daylight hours, the weights of the two data types are more balanced, typically around 0.4-0.6.

[0151] Finally, the system uses the adjusted weights to fuse the data and obtain the final detection results. The fusion process adopts a weighted average method to ensure that the advantages of both data are fully utilized and the shortcomings of each are compensated.

[0152] Experimental results show that this dynamic weight allocation mechanism can significantly improve the system's detection performance under different lighting conditions, especially under night conditions, where the detection accuracy is increased by more than 35%.

[0153] The multimodal road anomaly real-time monitoring system of the present invention has been applied in multiple practical scenarios and has demonstrated good performance and stability.

[0154] In urban road monitoring, the system is deployed on main arterial roads with high traffic volume, monitoring road conditions in real time through fixed monitoring points. The system promptly detects anomalies such as potholes and cracks and automatically sends alerts to maintenance departments. During a three-month test, the system successfully detected 98.3% of road anomalies with an average response time of 150ms, significantly improving maintenance efficiency and road safety.

[0155] For highway monitoring, the system is deployed on-board, using patrol vehicles to dynamically monitor the highway. The system can detect road anomalies in real time while vehicles are in motion, recording their location and type. This approach is particularly suitable for rapid inspections over long distances, improving both efficiency and coverage.

[0156] In tests conducted under adverse weather conditions, the system demonstrated excellent environmental adaptability. Even in low-light conditions such as rain, fog, and at night, the system maintained high detection accuracy, improving performance by 20% to 35% over traditional single-modality detection methods.

[0157] The system also boasts excellent scalability and compatibility. Through simple configuration adjustments, it can adapt to road conditions and detection requirements in different regions. The system also provides a standard API interface, enabling easy integration with existing road management and maintenance systems to form a complete intelligent road maintenance solution.

[0158] Through these practical application cases, the effectiveness and practicality of the multimodal road anomaly real-time monitoring system of the present invention in actual environments have been demonstrated, providing innovative solutions for the fields of intelligent transportation and road maintenance.

[0159] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. Multimodal road anomaly real-time monitoring system, characterized by: include: LiDAR module, used to collect LiDAR point cloud data of roads; A camera module for collecting road video data; A 5G edge computing unit is communicatively connected to the lidar module and the camera module, and the 5G edge computing unit includes: A data access module, configured to receive the laser radar point cloud data and the road video data; A data processing module is used to process the lidar point cloud data and the road video data. The data processing module includes a YOLOv7-pointnet network model. The YOLOv7-point net network model includes a backbone feature extraction module network, a Neck, and an output module. The backbone feature extraction module network adopts an EfficientNet structure. The Neck part includes a 2D feature extraction network and a 3D feature extraction network. The 3D feature extraction network is composed of a point cloud feature extraction module CFFM connected in series with jumps. The number of channels of the CFFM includes M+1 groups of input channels and M groups of output channels, wherein one group of input channels is a global average pooling channel and the M groups of output channels are composed of grouped convolution GCs. A data post-processing module, used for post-processing the results processed by the data processing module; A data transmission module, used for transmitting the results processed by the data post-processing module to a server; A control module is used to dynamically adjust the confidence of the detection model according to the lighting conditions; The server is communicatively connected to the 5G edge computing unit, and is used to receive the results transmitted by the data transmission module and determine whether there is a road abnormality event. When there is a road abnormality event, it locates the abnormal event, determines the type of abnormal event, and plans corresponding alarm paths and maintenance paths.

2. The multimodal road anomaly real-time monitoring system according to claim 1, characterized in that: The 5G edge computing unit also includes: a power supply module, used to provide power to the 5G edge computing unit.

3. The multimodal road anomaly real-time monitoring system according to claim 1, characterized in that: The data processing module includes: An image detection module, configured to perform real-time detection on the road video data; A point cloud detection module, used to detect the laser radar point cloud data; The abnormality judgment module is used to determine whether an abnormal event occurs by combining the laser radar point cloud data and the road video data.

4. The multi-modal road anomaly real-time monitoring system according to claim 3, characterized in that: The point cloud detection module includes: A point cloud preprocessing module, configured to perform grayscale conversion, threshold filtering, and downsampling processing on the laser radar point cloud data; A point cloud feature extraction module is used to perform semantic segmentation and key point extraction on the pre-processed lidar point cloud data to obtain point cloud features; A point cloud feature fusion module is used to combine the point cloud features and image features to provide global context information for the detection unit; The point cloud detection model is a YOLOv7 model, whose network structure includes 5 downsampling modules, 5 upsampling modules, and 8 residual modules.

5. The multi-modal road anomaly real-time monitoring system according to claim 4, characterized in that: The detection unit in the point cloud detection model includes: The initial information receiving module is used to receive the image features transmitted from the downsampling module; a fusion feature acquisition module, configured to fuse the image features received from the initial information receiving module with the global context obtained from the point cloud feature fusion module to obtain a fusion feature; A detection task sharing module, used to fuse the fused features through a 1×1 convolution layer and an activation layer; The image detection result output module is used to input the fused information into the neural network branches corresponding to each detection target to predict the target category.

6. The multi-modal road anomaly real-time monitoring system according to claim 3, characterized in that: The abnormality judgment module also includes: a large model, and the large model includes a visual large model and a semantic segmentation model.

7. The multi-modal road anomaly real-time monitoring system according to claim 1, characterized in that: The abnormality judgment module is further configured to perform spatiotemporal attention processing on the detection results to eliminate redundant detection frames in continuous video frames; the abnormality judgment module further comprises: The spatiotemporal attention mechanism is used to calculate the spatial similarity of each pixel between the target image and the K frames before and after the target image and the video frames of the K frames before and after the target image, and define the pixels with spatial similarity greater than a threshold as a saliency detection frame; the spatiotemporal attention mechanism calculates the attention weight of each time point between the target image and the K frames through the saliency detection frame, and obtains the attention weight of the target image frame and the K frames before and after the target image frame; using the attention weight of the target image frame, the anomaly results predicted by the semantic segmentation model detection and the anomaly results predicted by the visual model detection are weighted averaged to obtain the final anomaly detection result.

8. The multi-modal road anomaly real-time monitoring system according to claim 7, characterized in that: The weighted average formula is: A(m,n)=A vis (m,n)×(1-P(a m,n ))+A sem (m,n)×P(a m,n )×ω m,n , Among them, A is the final anomaly detection result, A vis is the abnormal result predicted by the visual model detection; A sem is the abnormal result predicted by the semantic segmentation model detection; m, n correspond to the coordinate position of each pixel in the final result; m∈[0,6], n∈[0,6]; ω m,n is the attention weight calculated by the spatiotemporal attention mechanism; P(a m,n ) is whether pixels m and n belong to the same target.

9. The multi-modal road anomaly real-time monitoring system according to claim 8, characterized in that: The attention weight formula is: oh i =(1-sI)(i∈[m,n])σ1+(1-sI)(i∈[n,m])σ2, Among them, ω i is the attention weight of the current pixel; σ1 and σ2 are attention coefficients; sI is the pixel space similarity.

10. An operating method of a multimodal road anomaly real-time monitoring system, applied to the system according to any one of claims 1 to 9, characterized in that: The following steps are involved: S1: The LiDAR module and camera module collect LiDAR point cloud data and road video data in real time, and transmit the collected LiDAR point cloud data, road video data or abnormal results to the server through the data transmission module; S2: The server determines whether there is a road abnormality event based on the abnormal result, and locates the abnormal event if there is a road abnormality event; S3: Determine the type of abnormal event and plan the corresponding warning path and maintenance path. Based on the different types of road abnormalities identified, send an alert and repair to the corresponding maintenance department, and display the abnormal event location, type, and warning path on the client. S4: Dynamically assign weights to the detection model and anomaly judgment module.

Citation Information

Cited By

  • Intelligent traffic early warning system and method based on linkage of video inspection and flash warning

    CN120954241A

  • Vehicle control method and vehicle

    CN121425199A