Tunnel pavement throwing object early warning method based on multi-mode perception
Through multimodal perception and deep learning technology, combined with spatio-temporal graph neural network and Bayesian network, high-precision detection and risk prediction of tunnel spills are achieved, solving the problem of insufficient data accuracy and prediction capabilities in the existing technology, and improving the intelligence level of tunnel safety management.
Patent Information
- Application Number
- CN202510601232.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing tunnel sprinkler detection and risk assessment technology has complex environmental factors such as limited perspective, uneven light and smoke interference, resulting in insufficient data accuracy and comprehensiveness, and the lack of prediction capabilities of spatiotemporal evolution of sprinklers, making it difficult to accurately predict the spreading trend and potential risks of sprinklers.
Using a multimodal perception method, by acquiring visible light images, thermal radiation images and three-dimensional point cloud data, combined with deep learning models (such as YOLOv11, CNN, ViT, LSTM, ST-GNN, Transformer) for feature extraction and spatiotemporal prediction, a spatiotemporal graph neural network is built, the spatiotemporal feature vector is enhanced, and risk assessment is performed through Bayesian network.
It has achieved ultra-high-precision capture of the type, location and evolutionary trend of thrown objects, with the detection accuracy exceeding 95%, and the prediction error is compressed to the centimeter level, providing timely and effective risk warnings, overcoming the limitations of traditional methods, and improving the intelligence level of tunnel safety management.
Smart Images

Figure CN120236199A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tunnel safety management, and particularly to a spatio-temporal evolution prediction method for tunnel pavement spillage based on multi-modal perception. Background Art
[0002] With the rapid development of modern transportation infrastructure, tunnels, as an important part of highway and railway networks, their safety and stability have attracted increasing attention. However, the presence of spillage in tunnels (such as oil stains, gravel, accumulated water, etc.) not only poses a direct threat to driving safety but may also trigger secondary disasters such as fires, structural damage, or traffic congestion. In recent years, a number of technologies have been applied to the detection and management of tunnel spillage, including image recognition based on traditional cameras, temperature monitoring by infrared thermal imaging, and spatial scanning by lidar. These technologies have achieved certain results in specific scenarios. For example, a single camera system can initially identify the presence of spillage, infrared sensors can detect high-temperature objects, and LiDAR provides basic spatial distribution information. In addition, in terms of spillage prediction, existing technologies have also tried some preliminary methods, such as predicting the area change of spillage based on simple threshold analysis or estimating the position movement trend of spillage using traditional time series models (such as ARIMA). These prediction methods can provide certain references in static scenarios but are generally still relatively primitive.
[0003] However, the existing tunnel spillage detection and risk assessment technologies have significant technical problems. First, the detection means mainly rely on single sensors or traditional image processing methods, which are limited by complex environmental factors such as perspective limitations, uneven lighting, and smoke interference, resulting in a significant reduction in the accuracy and comprehensiveness of data. For example, in low-light conditions, it is difficult for cameras to capture the details of spillage; in high-temperature or smoky environments, the single data of infrared and LiDAR are difficult to comprehensively reflect the characteristics of spillage. Second, in the spatio-temporal evolution prediction of spillage, the defects of existing technologies are particularly prominent. The prediction method based on threshold analysis is too simplistic, ignoring the interaction between spillage and the dynamic influence of the environment; traditional time series models (such as ARIMA) can only handle the position change of a single dimension and lack the ability to model multi-modal features (such as temperature, type) and their spatial relationships. This lack of prediction ability makes it difficult for existing technologies to accurately predict the diffusion trend and potential risks of spillage, and thus unable to provide timely and effective warning information for tunnel managers. Summary of the Invention
[0004] The purpose of the present invention is to provide a warning and prediction method for tunnel pavement spillage based on multi-modal perception, which realizes the dynamic prediction of the spatio-temporal evolution of spillage and the warning of potential risks.
[0005] To achieve the above object, the present invention provides the following solutions:
[0006] A warning method for tunnel pavement spillage based on multi-modal perception, including:
[0007] Obtain multi-modal data of the target tunnel pavement, where the multi-modal data includes visible light images, thermal radiation images, and three-dimensional point cloud data;
[0008] Extract features from the visible light image to obtain the feature information of the spillage;
[0009] Input the three-dimensional point cloud data into the LSTM network to obtain the short-term spatial displacement of the spillage;
[0010] Construct an ST-GNN network based on the feature information of the spillage and the multi-modal data, input the short-term spatial displacement of the spillage into the ST-GNN network, and obtain an enhanced spatio-temporal feature vector;
[0011] Input the visible light image and the thermal radiation image into the Transformer network to obtain the long-term distribution heat map feature;
[0012] Combine the enhanced spatio-temporal feature vector and the long-term distribution heat map feature to obtain the spatio-temporal distribution prediction result of the spillage;
[0013] Obtain the tunnel traffic risk warning result according to the multi-modal data and the spatio-temporal distribution prediction result of the spillage.
[0014] Optionally, obtaining the multi-modal data of the target tunnel pavement includes:
[0015] Collect an initial visible light image through a high-resolution camera, perform non-local mean denoising and Retinex illumination compensation on the initial visible light image to obtain the visible light image;
[0016] Collect the thermal radiation image through an infrared thermal imaging sensor;
[0017] Collect initial three-dimensional point cloud data through a lidar, register and align the initial three-dimensional point cloud data with the visible light image to obtain the three-dimensional point cloud data.
[0018] Optionally, extracting features from the visible light image to obtain the feature information of the spillage includes:
[0019] Input the visible light image into the YOLOv11 network, CNN network, and ViT network respectively, and output the position and category of the spillage, local feature map, and global feature map;
[0020] Fuse the intermediate feature map of the YOLOv11 network, the local feature map of the CNN network, and the global feature map of the ViT network to obtain the feature information of the spillage.
[0021] Optionally, the method for fusing the intermediate feature map of the YOLOv11 network, the local feature map of the CNN network, and the global feature map of the ViT network is as follows:
[0022] ;
[0023] where Fusion is the comprehensive feature representation, and α, β, γ are weight coefficients. , , are the intermediate feature map of YOLOv11, the local feature map of CNN, and the global feature map of ViT, respectively.
[0024] Optionally, constructing a spatio-temporal graph neural network includes: constructing the spatio-temporal graph neural network with nodes representing the spatial positions of the spillages and edges representing the spatial relationships between adjacent spillages or regions, where the node feature representation includes the coverage area, the type of spillage, the temperature, and the feature information of the spillage.
[0025] Optionally, calculating the short-term spatial displacement of the spillage includes:
[0026] ;
[0027] where ΔA is the short-term spatial displacement of the spillage, and are the coverage areas at adjacent time points, and is the time interval.
[0028] Optionally, obtaining the tunnel traffic risk warning result according to the multi-modal data and the spatio-temporal distribution prediction result of the spillage includes: calculating the current traffic risk index and the heat risk index according to the multi-modal data, and calculating the future traffic risk index and the future heat risk index according to the spatio-temporal distribution prediction result of the spillage;
[0029] The current traffic risk index is:
[0030] ;
[0031] ;
[0032] where is the adjusted current traffic risk index, k is the heat sensitivity coefficient, is the temperature difference, and is the adjustment coefficient for the short-term spatial displacement of the spillage, A is the current coverage area, P is the current position weight, and R is the current traffic risk index. is the weight coefficient for the coverage area. is the weight coefficient for the position weight.
[0033] The heat risk index is:
[0034] ;
[0035] Among them, is the short-term temperature change value predicted by ST-GNN, is the adjustment coefficient for the short-term temperature change value, and H is the heat risk index.
[0036] The future traffic risk index is:
[0037] ;
[0038] The future hot wind risk index is:
[0039] ;
[0040] Among them, is the future traffic risk index, is the future hot wind risk index, , , are the predicted coverage area, position, and temperature change value at the future time step.
[0041] Optionally, obtaining the tunnel traffic risk warning result further includes: inputting the spillage type, current position, and short-term spatial displacement of the spillage into a Bayesian network to output the blockage and collision probabilities, where the Bayesian network is trained by historical spillage types, positions, short-term spatial displacements of the spillage, and the corresponding blockage and collision probabilities.
[0042] The beneficial effects of the present invention are: (1) Compared with the prior art, the present invention can deeply analyze the characteristics of spillage from three dimensions of vision, heat sensitivity, and space, comprehensively reveal its potential threats to traffic, fire, and chain accidents, overcome the limitations of the traditional single perspective, and endow this method with better risk insight.
[0043] (2) Compared with the prior art, through the integration of advanced image processing, multi-model collaboration, and spatio-temporal prediction technologies, the present invention realizes ultra-high-precision capture of the spillage type, position, and evolution trend, the detection accuracy breaks through 95%, and the prediction error is compressed to the centimeter level, far leading the rough judgment of traditional methods.
[0044] (3) Compared with the prior art, by virtue of the instantaneous optimization of sensor parameters and the reasoning ability of the detection model, the present invention compresses the total time consumption from data acquisition to risk warning to the millisecond level, overcomes the bottleneck of low efficiency in traditional manual inspections or single methods, and demonstrates the advantages of intelligent response.
[0045] (4) Compared with the prior art, through the dynamic optimization driven by real-time environmental feedback and the distributed model evolution, the present invention endows this method with excellent adaptability under complex conditions such as light and smoke, while continuously improving the generalization performance across scenarios, surpassing the rigid design of traditional fixed parameters, and becoming an exemplary intelligent evolution in the field of tunnel safety management.
[0046] In summary, the present invention first collaboratively collects data through multi-modal sensors (high-resolution cameras, infrared thermal imaging, LiDAR), and uses image processing technologies such as denoising (such as non-local means denoising), light compensation (such as Retinex algorithm), and geometric distortion correction to improve the quality of multi-source data and enhance the recognizability of spillage features. At the same time, this method efficiently classifies and accurately identifies spillage through deep learning models (such as YOLOv11, convolutional neural network CNN, vision transformer ViT), and can distinguish different types of spillage (such as garbage, stones, oil stains, etc.) in real time in complex environments, overcoming the limitations of traditional methods in detecting object shapes and dynamic changes. Brief Description of the Drawings
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0048] Figure 1 It is a flowchart of the multi-modal perception-based tunnel pavement spillage warning method according to the embodiment of the present invention. Detailed Embodiments
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0050] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0051] Such asFigure 1 As shown in Figure 1 , this embodiment provides a warning method for tunnel pavement spillage based on multi-modal perception, including:
[0052] Obtain multi-modal data of the target tunnel pavement, where the multi-modal data includes visible light images, thermal radiation images, and three-dimensional point cloud data;
[0053] Extract features from the visible light image to obtain the feature information of the spillage;
[0054] Input the three-dimensional point cloud data into the LSTM network to obtain the short-term spatial displacement of the spillage;
[0055] Construct an ST-GNN network based on the feature information of the spillage and the multi-modal data, and input the short-term spatial displacement of the spillage into the ST-GNN network to obtain an enhanced spatio-temporal feature vector;
[0056] Input the visible light image and the thermal radiation image into the Transformer network to obtain the long-term distribution heat map feature;
[0057] Combine the enhanced spatio-temporal feature vector and the long-term distribution heat map feature to obtain the spatio-temporal distribution prediction result of the spillage;
[0058] Obtain the tunnel traffic risk warning result according to the multi-modal data and the spatio-temporal distribution prediction result of the spillage.
[0059] Further, obtaining the multi-modal data of the target tunnel pavement includes:
[0060] Collect the initial visible light image through a high-resolution camera, perform non-local mean denoising and Retinex illumination compensation on the initial visible light image to obtain the visible light image;
[0061] Collect the thermal radiation image through an infrared thermal imaging sensor;
[0062] Collect the initial three-dimensional point cloud data through a lidar, and register and align the initial three-dimensional point cloud data with the visible light image to obtain the three-dimensional point cloud data.
[0063] Specifically, deploy various sensors at key positions in the tunnel (such as busy traffic areas, tunnel entrances and exits, and high-incidence areas of falling objects). Cameras (HRC) are generally installed on the top or side walls on both sides of the tunnel to ensure the widest field of view; infrared sensors are deployed at different heights to more accurately capture temperature differences; LiDAR is placed at key nodes, and multi-line scanning is used to improve the accuracy of obtaining three-dimensional point clouds.
[0064] Further, extracting features from the visible light image to obtain the feature information of the spillage includes:
[0065] Input the visible light image into the YOLOv11 network, CNN network, and ViT network respectively, and output the position and category of the spillage, local feature map, and global feature map respectively.
[0066] Fuse the intermediate feature map of the YOLOv11 network, the local feature map of the CNN network, and the global feature map of the ViT network to obtain the feature information of the spillage.
[0067] Specifically, in this embodiment, by combining local texture, global structure features, and a deep learning detection model, the spillage in the tunnel is accurately identified.
[0068] Further, the method for fusing the intermediate feature map of the YOLOv11 network, the local feature map of the CNN network, and the global feature map of the ViT network is as follows:
[0069] ;
[0070] where Fusion is the comprehensive feature representation, and α, β, γ are weight coefficients. 、 、 are the intermediate feature map of YOLOv11, the local feature map of CNN, and the global feature map of ViT respectively.
[0071] Further, constructing a spatio-temporal graph neural network includes: constructing a spatio-temporal graph neural network with nodes representing the spatial positions of the spillage and edges representing the spatial relationships between adjacent spillage or regions. Among them, the node features include the coverage area, spillage type, temperature, and the feature information of the spillage.
[0072] Specifically, compared with the temporal modeling of LSTM and the global trend analysis of Transformer, the unique advantage of ST-GNN lies in its sensitivity to the spatial structure, which can reveal the mutual influence between spillage (such as the chain reaction of diffusion), which is particularly crucial for complex scenarios with multiple objects coexisting in the tunnel environment. For example, when oil spills spread, ST-GNN can predict its indirect impact on the distribution of nearby accumulated water or gravel, thereby improving the comprehensiveness and accuracy of the prediction.
[0073] Further, calculating the short-term spatial displacement of the spillage includes:
[0074] ;
[0075] where ΔA is the short-term spatial displacement of the spillage. and are the coverage areas at adjacent time points. is the time interval.
[0076] Further, based on the multi-modal data and the spatio-temporal distribution prediction results of the spillage, obtaining the tunnel traffic risk warning results includes: calculating the current traffic risk index and the heat risk index according to the multi-modal data, and calculating the future traffic risk index and the future heat risk index according to the spatio-temporal distribution prediction results of the spillage;
[0077] The current traffic risk index is:
[0078] ;
[0079] ;
[0080] Among them, is the adjusted current traffic risk index, k is the heat sensitivity coefficient, is the temperature difference, is the adjustment coefficient of the spatial displacement of the spillage in the short term, A is the current coverage area, P is the current position weight, R is the current traffic risk index, is the weight coefficient of the coverage area, is the weight coefficient of the position weight;
[0081] The heat risk index is:
[0082] ;
[0083] Among them, is the short-term temperature change value predicted by ST-GNN, is the adjustment coefficient of the short-term temperature change value, H is the heat risk index;
[0084] The future traffic risk index is:
[0085] ;
[0086] The future heat risk index is:
[0087] ;
[0088] Among them, is the future traffic risk index, is the future heat risk index, 、 、 are the predicted coverage area, position, and temperature change value at the future time step.
[0089] Further, obtaining the tunnel traffic risk warning result further includes: inputting the type of the spillage, the current location, and the short-term spatial displacement of the spillage into a Bayesian network to output the probabilities of blockage and collision, where the Bayesian network is obtained by training with historical spillage types, locations, short-term spatial displacements of the spillage, and the corresponding probabilities of blockage and collision.
[0090] The method of this embodiment will be further described below:
[0091] The warning method for tunnel pavement spillage based on multi-modal perception specifically includes the following steps:
[0092] Step 1: Multi-modal image acquisition:
[0093] Deploy various sensors at key positions in the tunnel (such as busy traffic areas, tunnel entrances and exits, and high-incidence areas of falling objects). Cameras (HRC) are generally installed on the top or side walls on both sides of the tunnel to ensure the widest field of view coverage; infrared sensors are deployed at different heights to more accurately capture temperature differences; LiDAR is placed at key nodes, and multi-line scanning is used to improve the accuracy of obtaining three-dimensional point clouds. By jointly collecting multi-dimensional data through multiple sensors, the blind spot problem of a single camera in a complex environment is solved, and the comprehensiveness of the data is improved.
[0094] Multi-modal sensors (HRC cameras, infrared thermal imaging sensors, LiDAR) can effectively collect various environmental data through precise spatio-temporal synchronization and avoid information loss caused by a single sensor due to the viewing angle or local interference.
[0095] Through computer control for real-time calibration, ensure that the coordinate systems between the sensors are aligned, and automatically detect and correct lens distortion and time synchronization deviation. Use a special calibration scene (such as a standard calibration board) to verify the error range of each sensor, and perform precise correction through a data fusion algorithm.
[0096] According to the environmental light intensity , adjust the exposure time, and increase the exposure time in low-light environments .
[0097] According to the influence of environmental smoke concentration or obstacles, automatically adjust the scanning frequency of LiDAR to reduce data redundancy and optimize data accuracy.
[0098] HRC acquires visible light images ( ), for detecting spillage in the tunnel (such as garbage, fallen objects);
[0099] The infrared sensor obtains a thermal radiation image ( ), detecting changes in heat sources (such as hot spillage or high-temperature objects left by vehicles, temperature differences , is a preset threshold);
[0100] LiDAR generates three-dimensional point cloud data ( ) for scanning obstacles on the tunnel floor and providing spatial information for subsequent analysis.
[0101] Step 2: Multi-scale image preprocessing:
[0102] In this step, the RGB image and LiDAR point cloud data are systematically preprocessed to extract the visual, thermal, and spatial characteristics of the spillage and perform preliminary quantification. The following is the specific implementation process:
[0103] 1. Non-Local Means Denoising:
[0104] This method is aimed at the random noise in the RGB image and performs denoising processing through the similarity between image patches. Specifically, the weighted average value of the area around each pixel point is calculated, and the weights are dynamically adaptively calculated based on local texture and global structure to ensure that the details of the spillage (such as edges and textures) are retained while suppressing noise.
[0105] 2. Retinex illumination compensation:
[0106] Under the complex illumination conditions in the tunnel (such as strong light or shadows), the characteristics of the spillage in the RGB image are easily masked by uneven illumination. The Retinex algorithm effectively reduces illumination interference by decomposing the image into a reflection component and an illumination component.
[0107] 3. 3D point cloud registration and distortion correction:
[0108] The point cloud data generated by LiDAR is registered and aligned with the RGB image to eliminate geometric errors caused by different perspectives. The specific method is to optimize the transformation matrix T (including rotation and displacement) so that the three-dimensional coordinates in the point cloud correspond to the pixel positions in the image.
[0109] The registered point cloud data not only provides the accurate spatial distribution of the spillage but also lays a foundation for subsequent quantitative analysis.
[0110] 4. Multi-modal data processing and feature extraction:
[0111] RGB image feature enhancement: The RGB image processed by denoising and illumination compensation is further used to extract the visual features of the spillage (such as edges, colors, textures) to provide high-quality input for type recognition.
[0112] Infrared image temperature difference analysis: The infrared image detects the temperature difference ( ) using the preset threshold ( Identify abnormally high-temperature areas (such as high-temperature liquids or combustibles), mark potentially heat-sensitive substances, and provide a basis for fire risk assessment.
[0113] LiDAR point cloud spatial quantization: The registered point cloud data extracts the spatial area of the spill through segmentation and clustering algorithms, calculates its geometric parameters (such as coverage area, volume, thickness), and combines the position and depth information of the RGB image to generate spatial distribution characteristics.
[0114] Step 3: Intelligent detection and identification of spills:
[0115] Through the image preprocessing completed in Step 2, a high-quality RGB image is received, and parallel YOLOv11, CNN, and ViT networks are used to extract multi-level features. Finally, precise identification and classification of spills are achieved through dynamic fusion. This process includes differentiating different types of spills (such as garbage, obstacles, leakage substances, etc.) and judging the size, shape, etc. of objects based on morphological complexity.
[0116] 1. Multi-feature extraction:
[0117] In this embodiment, three parallel network branches are used to process the input RGB image, and each network focuses on extracting different levels of features. The following is a detailed description:
[0118] YOLOv11: The input of the YOLOv11 model is the preprocessed RGB image, and the output is the position and category of the spill (such as garbage, stones, dropped tools, etc.). The YOLOv11 model has strong real-time performance and multi-category recognition ability, and is suitable for object detection tasks that are complex and rapidly changing in tunnel environments. It directly extracts spatial features from the image, detects and classifies various objects, and is especially suitable for quickly locating spills.
[0119] In this embodiment, the YOLOv11 model is used for fast and high-precision object detection, especially suitable for detecting different types of spills (such as garbage, stones, dropped tools, etc.). The YOLOv11 model directly extracts spatial features from the image and detects and classifies various objects.
[0120] CNN branch: The input of the CNN branch is the preprocessed RGB image, and the output is a local feature map or feature vector, capturing the detailed information of the spill (such as edges, textures, etc.).
[0121] Convolutional neural network (CNN) focuses on extracting local texture and edge features and is good at identifying small or clearly edged spills. By inputting the calibrated image, the convolutional layer of CNN extracts local features such as edges and textures, thus capturing the details of the spill.
[0122] ViT Branch: The input of the ViT branch is the preprocessed RGB image, and the output is the global feature representation for identifying large or irregular spills.
[0123] The Vision Transformer (ViT) model captures the global structural features in the image through the self-attention mechanism, which is particularly suitable for identifying large or irregular spills. ViT can help identify objects with long-range associations or complex topologies, such as long strip-shaped garbage or large dropped objects. The three networks (YOLOv11, CNN branch, and ViT branch) each independently receive the same preprocessed RGB image as input and extract different types of features (spatial features, local features, global features) respectively.
[0124] 2. Dynamic Feature Fusion:
[0125] Through the integration of the dynamic feature fusion mechanism, the intermediate feature map of YOLOv11 (containing spatial and semantic information) is combined with the local feature map / vector of the CNN branch and the global feature representation of the ViT branch to form a comprehensive feature representation.
[0126] The specific formula is as follows:
[0127] ;
[0128] Among them, is the intermediate feature map of YOLOv11, is the local feature map of CNN, is the global feature map of ViT, and α, β, γ are weight coefficients dynamically optimized through reinforcement learning. If used for a supervised task, the fused features are optimized through the mean squared error loss function L:
[0129] ;
[0130] Among them, is the fused feature vector, is the true feature label corresponding to the i-th sample, and N is the number of samples.
[0131] The fused features integrate the spatial and semantic information of YOLOv11, the local features of CNN (texture, edges), and the global features of ViT (structure, context), providing a more comprehensive spill characterization for subsequent tasks, which is used to improve the prediction accuracy or the accuracy of risk assessment. The detected spill location and category information will be used as the direct input for risk assessment in step 4. The fused feature representation can assist the subsequent model in more accurately analyzing the potential threats of spills.
[0132] 3. Semi-Supervised Spill Classification:
[0133] Adopt a semi-supervised learning strategy, combine labeled data and unlabeled data for training, and perform consistency regularization on the processing of unlabeled data to reduce the dependence on manually labeled data, enabling the model to learn potential patterns from unlabeled data and improve generalization ability. For example, the model can learn more stable feature representations in an unsupervised manner on unlabeled data and adjust the weights of YOLOv11 according to the results of semi-supervised learning to enhance the detection ability for complex or small-scale spillages (such as debris). Among them:
[0134] Labeled data ( ) calculate the cross-entropy loss ( );
[0135] Unlabeled data ( ) enhance generalization through consistency regularization.
[0136] Step 4: Predict the impact of spillages on tunnel safety and evaluation:
[0137] 1. Prediction of the spatio-temporal evolution of spillages:
[0138] Method: Based on spatio-temporal graph neural network (ST-GNN), LSTM, and Transformer models.
[0139] Input data:
[0140] RGB image: Visual features of the spillage (such as type);
[0141] Thermal radiation image: Temperature information of the spillage (such as temperature difference ΔT);
[0142] LiDAR point cloud: Spatial position and geometric parameters of the spillage (such as coverage area, volume, thickness).
[0143] Modeling process:
[0144] (1) Analysis of position and distribution changes (short-term evolution):
[0145] Based on the time series of LiDAR point cloud data, predict the potential diffusion trend of the spillage in the tunnel. For example, accumulated water may spread to low-lying areas due to gravity, and gravel may be displaced due to passing vehicles.
[0146] Use LSTM to capture the instantaneous changes in the position of the spillage, and calculate the short-term spatial displacement of the spillage by comparing the point cloud data in consecutive time periods:
[0147] ;
[0148] where ΔA is the short-term spatial displacement of the spillage, and is the coverage area at adjacent time points, is the time interval.
[0149] (2)State evolution estimation (long-term diffusion):
[0150] Combining the temperature change of the infrared image and the feature evolution of the RGB image, predict the change of the spill state over time. For example, the accumulated water may gradually decrease (the temperature tends to be stable), and the oil stain may become more difficult to identify due to diffusion (the color feature fades). Input the visible light image and the thermal radiation image into the Transformer network, and model the long-term state evolution trend of the spill through the self-attention mechanism to generate a long-term distribution heat map . This heat map reflects the change of the coverage range and temperature distribution of the spill in the future time period, such as the probability distribution of oil stain diffusion or accumulated water reduction.
[0151] (3)Spatio-temporal graph network enhancement:
[0152] Construct the LiDAR point cloud data into a spatio-temporal graph structure, where the nodes represent the spatial positions of the spills (based on the coordinates and geometric parameters in step 1, such as the coverage area), and the edges represent the spatial relationships between adjacent spills or regions (determined by the Euclidean distance or adjacency, and the distance threshold is a preset value). The node features include the coverage area obtained from the three-dimensional point cloud data, the type obtained from the visible light image (provided by YOLOv11), the temperature obtained from the thermal radiation image, and the feature vector obtained by fusing the YOLOv11, CNN, and ViT networks.
[0153] The ST-GNN updates the state of each node by combining the time dimension through graph convolution operations, capturing the dynamic interactions and spatial dependencies between the spills. The fused feature vector enhances the node representation, enabling the model to use more comprehensive spill information (such as local details and global structure) for spatio-temporal distribution prediction. For example, a large-area spill (such as an oil stain) may drive the distribution change of nearby small debris (such as gravel), or the aggregation effect of multiple spills in a curved area.
[0154] Compared with the time series modeling of LSTM and the global trend analysis of Transformer, the unique advantage of ST-GNN lies in its sensitivity to the spatial structure, which can reveal the mutual influence between the spills (such as the chain reaction of diffusion), which is particularly crucial for complex scenarios with multiple objects coexisting in the tunnel environment. For example, when the oil stain diffuses, ST-GNN can predict its indirect impact on the distribution of nearby accumulated water or gravel, thus improving the comprehensiveness and accuracy of the prediction.
[0155] The short-term spatial displacement (ΔA) of the LSTM is used as the temporal input of the ST-GNN to drive the temporal evolution of node features, and the output of the ST-GNN enhances the spatio-temporal feature representation. , including the predicted feature vectors of each spill at each time step (including updated position, coverage area, and distribution probability).
[0156] Dynamic interaction relationship: The spatial influence weight between spills (such as the updated value of the adjacency matrix) is used for generating the heatmap for subsequent risk assessment.
[0157] The long-term distribution heatmap of the Transformer Provides global constraints for the ST-GNN. Specifically, by embedding the heatmap features into the graph convolution weights of the ST-GNN, it restricts the spatial range of node updates and ensures the consistency between local predictions and global trends.
[0158] Output integration:
[0159] The enhanced features of the ST-GNN And the heatmap of the Transformer Form a complete spatio-temporal distribution prediction result through weighted fusion :
[0160] ;
[0161] Among them, , is the weight optimized by reinforcement learning.
[0162] 2. Traffic safety risk assessment:
[0163] Assessment framework: Based on multi-modal data fusion and dynamic risk quantification model.
[0164] Assessment process:
[0165] (1) Traffic safety impact assessment (current risk):
[0166] Based on the current coverage area (from LiDAR), position weight (from RGB and LiDAR), and short-term dynamic trend (from LSTM) at the current moment, calculate the current traffic risk index R:
[0167] ;
[0168] Among them: A is the current coverage area (unit: square meters), calculated from the LiDAR point cloud data, reflecting the spatial occupancy degree of the spill; P is the current position weight (value range [0, 1]), determined by the RGB image and LiDAR data. For example, the weight in the center of the lane is 1, and the weight at the edge is 0. is the short-term distribution change rate (unit: square meters per second), calculated by LSTM based on the LiDAR point cloud time series, reflecting the immediate diffusion or movement trend of the spillage; is the weight coefficient of the coverage area, indicating the contribution degree of the coverage area to the current traffic risk, with a value range of [0.5, 2.0], determined by regression analysis of historical traffic data, and used to quantify the direct impact of the spillage area on driving safety; is the weight coefficient of the position weight, indicating the contribution degree of the position to the current traffic risk, with a value range of [1.0, 3.0], calibrated according to the tunnel traffic flow and regional importance, highlighting the risks at key positions (such as the center of the lane); is the adjustment coefficient of the spatial displacement of the spillage in the short term, indicating the gain effect of the dynamic change of the spillage on the current risk, with a value range of [0.1, 1.0], optimized through experimental verification to ensure a reasonable weight for the dynamic trend.
[0169] Combined with infrared data, for high-temperature spillages (such as combustibles), an additional heat risk factor is added to adjust the risk index:
[0170] ;
[0171] Among them, is the adjusted risk index, k is the heat sensitivity coefficient, is the temperature difference.
[0172] (2) Heat risk analysis (current risk):
[0173] Based on the current coverage area A, temperature difference and the short-term temperature change value (local temperature update extracted from ), calculate the current heat risk index H:
[0174] ;
[0175] Among them, is the short-term temperature change value predicted by ST-GNN, is the adjustment coefficient of the short-term temperature change value.
[0176] (3) Secondary risk inference (current risk):
[0177] Using the Bayesian network, based on historical multimodal data (the historical database stored in step 6, such as the correlation between spillage types and accidents), input the current type, location, and into the Bayesian network, with blockage and collision as the output nodes, to predict the probability of a chain reaction caused by the spillage, such as traffic congestion or secondary collision.
[0178] (4)Future risk prediction:
[0179] Based on the spatio-temporal distribution prediction result F final , evaluate the possible risk changes caused by the spillage in the future time period. Specifically, use F final The predicted coverage area change value A pred and the position change value P pred to calculate the future traffic risk index:
[0180] ;
[0181] Where: A pred is the coverage area at the predicted time step (unit: square meter), derived from the enhanced features of ST-GNN in F final , reflecting the possible spread or shrinkage of the spillage; P pred is the position weight at the predicted time step, with a value range of [0,1], determined by the position change trend in F final . For example, the weight in the center of the lane is 1 and the edge is 0.
[0182] Future heat risk: Combine the predicted temperature change value final in F to calculate the future heat risk index:
[0183] ;
[0184] Future secondary risks: Input the spillage type and the predicted position change value in F final into the Bayesian network to infer the future blockage and collision probabilities.
[0185] Risk grading criteria:
[0186] Low risk: Coverage area < 0.1m², temperature < 50°C, position in the edge area;
[0187] Medium risk: Coverage area 0.1 - 2m², or temperature 50 - 100°C, position close to the lane;
[0188] High risk: Coverage area > 2m² and located in the center of the lane, or temperature > 100°C.
[0189] Step 5: Adaptive optimization and collaborative learning:
[0190] Execution entities: Federated learning server and edge computing nodes.
[0191] Function: Dynamically optimize performance and achieve multi-node knowledge sharing.
[0192] 1. Reinforcement learning for dynamic adjustment (automatically adjust parameters according to environmental feedback to balance detection accuracy and efficiency).
[0193] During the process of spill monitoring, reinforcement learning algorithms are used to optimize the working efficiency and performance of multiple components. By interacting with the environment (such as light, temperature, humidity, etc.), this method can automatically adjust parameters to balance detection accuracy and efficiency. The specific parameters include:
[0194] (1) Sensor layer: Camera exposure time. LiDAR scan frequency ( , where is the attenuation coefficient, f base is the reference scan frequency, representing the default frequency without smoke interference, usually set to the maximum supported frequency of the LiDAR hardware, such as 10 Hz or 20 Hz, depending on the specific device, and C 烟雾 is the smoke concentration, which is measured in real time by an infrared sensor or an environmental monitoring module, in units of ppm (parts per million) or percentage, and the value range is usually [0, 1000 ppm]).
[0195] (2) Computing layer: Edge node inference frequency (balancing computing latency and detection accuracy ). Data upload threshold (only transmitting high-risk spill data to reduce bandwidth occupancy).
[0196] Objective function of edge node inference frequency:
[0197] ;
[0198] where is the trade-off coefficient (known quantity).
[0199] Reinforcement learning continuously interacts with the environment to optimize the parameter adjustment strategy, thereby improving the speed and accuracy of spill detection. Through the dynamic adjustment of reinforcement learning, this method can adapt to environmental changes in real time (such as light, weather, temperature, humidity, etc.), and improve work efficiency while maintaining high accuracy.
[0200] 2. Federated model update (to improve the generalization ability of the model while protecting data privacy, and optimize the detection and prediction models in steps 3 and 4).
[0201] Federated Learning integrates the model information of each node through distributed training while ensuring that the data does not leave the local device. Each edge node (such as sensors, devices, etc.) trains a model with local data (such as spill images of a certain tunnel section) to generate gradients , then upload the model gradients instead of the raw data, which can ensure data privacy and security. By aggregating the model updates of all edge nodes, the central server will output a global model that has better generalization ability than the local models of individual nodes, thus improving the adaptability and accuracy of the spill detection process in different environments.
[0202] Specific content of the global model:
[0203] Each node uploads the model gradients ( ), and the server aggregates the global model.
[0204] This method effectively avoids the risk of uploading sensitive data to the cloud, and at the same time improves the intelligence level and performance of the overall process through joint training.
[0205] Step 6: Visualization report and decision support:
[0206] The executing entity: database and visualization engine.
[0207] Function: Convert the detection results into actionable maintenance strategies.
[0208] 1. Multi-dimensional data storage (structurally store historical data, support long-term trend analysis).
[0209] Stored fields: Spill location (three-dimensional coordinates), geometric parameters, risk level.
[0210] All detected spill information (such as location, geometric parameters, temperature data, point cloud data, risk level, etc.) will be stored in the database in real time for subsequent query and analysis.
[0211] The database adopts an efficient index structure, which can support fast retrieval and multi-dimensional analysis, and support filtering and query based on time, location and risk level. Historical data can be used for trend analysis, which helps to evaluate the long-term stability of the tunnel.
[0212] Through data storage and analysis, this method can provide location-based spill prediction, which is convenient for regular inspection and maintenance of specific areas.
[0213] 2. Intelligent report generation (provide intuitive decision-making basis):
[0214] Through visualization reports such as heat maps and 3D models, this method provides detailed status and risk assessment of spills. The heat map highlights high-risk areas, helping maintenance personnel quickly locate key areas; the 3D model visualizes the spill distribution, helping engineers analyze the impact of spills on the structure from multiple angles.
[0215] Based on the organic combination of multi-modal sensor data fusion, deep learning technology, and spatio-temporal evolution prediction, the present invention has successfully solved the problems of insufficient detection accuracy and lack of dynamic risk prediction ability in the prior art under complex environments. It not only achieves high-precision detection and classification of spillage, but also fills the technical gap in spatio-temporal evolution analysis through an advanced prediction model, with significant technical advantages and broad application prospects, providing a solution for modern tunnel safety management.
[0216] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A tunnel pavement spillage warning method based on multimodal perception, characterized in that: include: Acquire multimodal data of the target tunnel pavement, wherein the multimodal data includes visible light images, thermal radiation images, and three-dimensional point cloud data; Extracting features from the visible light image to obtain feature information of the scattered objects; The three-dimensional point cloud data is input into the LSTM network to obtain the spatial displacement of the scattered objects in a short period of time; Constructing an ST-GNN network according to the characteristic information of the scattered object and the multimodal data, inputting the spatial displacement of the scattered object in a short period of time into the ST-GNN network to obtain an enhanced spatiotemporal feature vector; Inputting the visible light image and the thermal radiation image into a Transformer network to obtain long-term distribution heat map features; Combining the enhanced spatiotemporal feature vector and the long-term distribution heat map feature, obtaining a spatiotemporal distribution prediction result of the scattered objects; A tunnel traffic risk warning result is obtained based on the multimodal data and the spatiotemporal distribution prediction result of the scattered objects.
2. The tunnel pavement spillage warning method based on multimodal perception according to claim 1 is characterized in that: Acquiring multimodal data of the target tunnel pavement includes: An initial visible light image is collected by a high-resolution camera, and non-local mean denoising and Retinex illumination compensation are performed on the initial visible light image to obtain the visible light image; Collecting the thermal radiation image by means of an infrared thermal imaging sensor; Initial three-dimensional point cloud data is collected by laser radar, and the initial three-dimensional point cloud data is aligned with the visible light image to obtain the three-dimensional point cloud data.
3. The tunnel pavement spillage warning method based on multimodal perception according to claim 1 is characterized in that: Extracting features from the visible light image to obtain feature information of the scattered objects includes: The visible light image is input into the YOLOv11 network, the CNN network and the ViT network respectively, and the location and category of the spilled object, the local feature map and the global feature map are output respectively; The intermediate feature map of the YOLOv11 network, the local feature map of the CNN network, and the global feature map of the ViT network are fused to obtain feature information of the scattered objects.
4. The tunnel pavement spillage warning method based on multimodal perception according to claim 3 is characterized in that: The method for fusing the intermediate feature map of the YOLOv11 network, the local feature map of the CNN network, and the global feature map of the ViT network is: ; Among them, Fusion is a comprehensive feature representation, α, β, γ are weight coefficients, , , They are the intermediate feature map of YOLOv11, the local feature map of CNN, and the global feature map of ViT.
5. The tunnel pavement spillage warning method based on multimodal perception according to claim 4 is characterized in that: Constructing a spatiotemporal graph neural network includes: constructing the spatiotemporal graph neural network by using nodes to represent the spatial position of spilled objects and edges to represent the spatial relationship between adjacent spilled objects or regions, wherein the node feature representation includes coverage area, type of spilled objects, temperature and feature information of the spilled objects.
6. The tunnel pavement spillage warning method based on multimodal perception according to claim 1 is characterized in that: Calculation of short-term spatial displacement of spilled objects includes: ; Among them, ΔA is the spatial displacement of the scattered objects in a short period of time, and is the coverage area of adjacent time points, is the time interval.
7. The tunnel pavement spillage warning method based on multimodal perception according to claim 6 is characterized in that: According to the multimodal data and the spatiotemporal distribution prediction result of the scattered objects, obtaining the tunnel traffic risk warning result includes: calculating the current traffic risk index and the heat risk index according to the multimodal data, and calculating the future traffic risk index and the future heat risk index according to the spatiotemporal distribution prediction result of the scattered objects; The current traffic risk index is: ; ; in, is the adjusted current traffic risk index, k is the thermal sensitivity coefficient, is the temperature difference, is the adjustment coefficient of the spatial displacement of the scattered objects in the short term, A is the current coverage area, P is the current position weight, R is the current traffic risk index, is the weight coefficient of coverage area, is the weight coefficient of the position weight; The heat risk index is: ; in, is the short-term temperature change value predicted by ST-GNN, is the adjustment coefficient of short-term temperature change value, H is the heat risk index; The future traffic risk index is: ; The future heat wave risk index is: ; in, is the future traffic risk index, For the future heat wave risk index, , , The coverage area, position, and temperature change values for the predicted future time steps.
8. The tunnel pavement spillage warning method based on multimodal perception according to claim 1 is characterized in that: Obtaining tunnel traffic risk warning results also includes: inputting the type of spilled objects, current location and short-term spatial displacement of spilled objects into the Bayesian network, and outputting the probability of congestion and collision, wherein the Bayesian network is obtained through historical spilled object types, locations and short-term spatial displacement of spilled objects and corresponding congestion and collision probabilities training.
Citation Information
Patent Citations
Differential video sequence and convolutional neural network fused spilled object detection method
CN115223106A
World coordinate positioning method for road spilled object in Gaussian mixture domain
CN115760898A
Pavement throwing event detection method based on artificial intelligence
CN117789141A
Dynamic target behavior prediction system based on 3D convolution and recurrent neural network
CN118506252A
Industrial thrown object monitoring method based on deep learning and graph neural network
CN118609069A
Cited By
Method for generating image data of thrown objects on expressway under aerial photography view angle
CN121527653A