Non-contact multi-mode fusion real-time sensing and early warning method for roof fall risk of roadway driving face
By using high-frequency acoustic wave sensors and high-definition video acquisition equipment on the coal mine tunnel excavation face, combined with deep reinforcement learning technology, real-time perception and early warning of multimodal data fusion are achieved, solving the problems of lag and high misjudgment rate of roof fall risk warning in existing technologies, and improving the safety and reliability of the tunnel excavation face.
Patent Information
- Application Number
- CN202510762755.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-12
AI Technical Summary
In coal mine tunnel excavation, single sensors or manual inspection methods are unable to accurately quantify the dynamic evolution of surrounding rock crack density, and multimodal data fusion is insufficient, resulting in delayed warning of roof fall risks and high misjudgment rate.
A non-contact multimodal fusion real-time perception method is adopted to obtain multimodal heterogeneous data through high-frequency acoustic wave sensors and high-definition video acquisition equipment, and feature extraction, dimensionality reduction optimization and deep reinforcement learning are performed to dynamically assess the risk of roof collapse and trigger early warning.
It has achieved real-time and accurate perception and dynamic early warning of the tunnel excavation face, significantly reduced the rate of roof collapse accidents, and improved the adaptability of roof stability assessment and the timeliness of early warning.
Smart Images

Figure CN120632784A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mining engineering safety technology, and in particular relates to a non-contact, multi-modal fusion, real-time perception and early warning method for roof collapse risks in tunnel excavation faces. Specifically, it involves technologies for monitoring surrounding rock stability in coal mine tunnels, analyzing rock mass fracture evolution, inverting acoustic wave attenuation characteristics, capturing dynamic deformations in video images, and integrating intelligent risk assessment systems. Background Art
[0002] At present, the monitoring of roof fall risks in coal mine tunneling working faces mainly relies on single physical sensors (such as displacement meters and pressure sensors) or manual inspection methods, supplemented by acoustic wave detection and local analysis of video images. Acoustic wave detection technology can preliminarily identify the distribution of cracks within the surrounding rock, but due to the limitations of the wave velocity inversion algorithm, it is difficult to accurately quantify the dynamic evolution of crack density; although video image monitoring can capture roof surface deformation or rock fall phenomena, it relies on manual experience judgment and has problems with timeliness and accuracy. In addition, existing technologies mostly use isolated data source analysis and lack the ability to integrate and process multi-modal heterogeneous data such as acoustic waves and videos. It is difficult to fully reflect the multi-factor coupling mechanism of roof instability, resulting in a high lag and misjudgment rate in risk warning. Summary of the Invention
[0003] In order to solve the above technical problems, the present invention proposes a non-contact multimodal fusion real-time perception and early warning method for the risk of roof collapse on the tunnel excavation face to solve the problems existing in the above-mentioned existing technologies.
[0004] To achieve the above objectives, the present invention provides a non-contact multi-modal fusion real-time perception and early warning method for roof collapse risk in a tunnel excavation face, comprising:
[0005] Acquiring multimodal heterogeneous data, wherein the multimodal heterogeneous data is data of different formats from different data sources; the multimodal heterogeneous data includes physical characteristics reflecting different aspects of the tunnel excavation surface;
[0006] Extract features from multimodal heterogeneous data to obtain multimodal features, perform key feature screening and dimensionality reduction optimization on the multimodal features, and calculate the roof fall risk latent variables based on the key features after dimensionality reduction optimization;
[0007] The temporal variation characteristics of the roof fall risk latent variable are extracted, and decisions are made based on the temporal variation characteristics to obtain the roof risk stability level and support strategy.
[0008] Optionally, the multimodal heterogeneous data includes sound wave data and image data.
[0009] Optionally, feature extraction also includes:
[0010] Preprocess multimodal heterogeneous data, including standardization, denoising and data enhancement.
[0011] Optionally, the feature extraction process includes:
[0012] Performing energy entropy analysis on the sound wave data to obtain sound wave video energy entropy;
[0013] Extracting non-stationary attenuation spectrum features according to the acoustic wave data, and processing the non-stationary attenuation spectrum features through identification and quantification methods to obtain anomaly coefficients;
[0014] Surface fractal dimension and texture features are extracted from the image data, and a deformation gradient tensor is constructed according to the graphic data. Crack topology length and bifurcation angle are extracted according to the image data.
[0015] Optionally, the process of key feature screening and dimensionality reduction optimization for multimodal features includes:
[0016] The multimodal features are screened according to the relevance of the multimodal features, and the screened features are subjected to dimensionality reduction, and the dimensionality reduction-optimized features are further subjected to dimensionality reduction optimization according to information redundancy to obtain key features after dimensionality reduction optimization.
[0017] Optionally, the process of obtaining the roof fall risk latent variable includes:
[0018] The reduced-dimensional features are fused and enhanced through a dual-stream cross-attention mechanism, wherein the modal weights are obtained through the cross-attention mechanism based on the residuals between the multimodal features and the key features after dimensionality reduction optimization, and the key features after dimensionality reduction optimization are enhanced. Multimodal fusion is performed based on the modal weights and feature enhancement results to calculate the roof collapse risk latent variable.
[0019] Optionally, the process of making decisions based on temporal variation characteristics includes:
[0020] The time series variation characteristics of the roof fall risk latent variable are extracted by the method used to extract time series characteristics. The time series variation characteristics are processed and decided by the deep enhancement model to obtain the roof risk stability level and support strategy.
[0021] Optionally, an LSTM model is used as a method for extracting time series features.
[0022] Optionally, t-SNE is used to visualize feature distribution for further dimensionality reduction optimization.
[0023] On the other hand, the present invention also provides a non-contact multimodal fusion real-time perception and early warning system for the risk of roof collapse on a tunnel excavation face, which is used to execute the above method.
[0024] Compared with the prior art, the present invention has the following advantages and technical effects:
[0025] Through multimodal data fusion and deep reinforcement learning technology, the present invention can realize real-time and accurate perception and dynamic early warning of the risk of roof fall on the tunnel excavation face; through the coordinated collection of high-frequency acoustic wave sensing and visual monitoring, it can effectively identify the coupled evolution laws of hidden crack expansion, surface deformation and stress release; based on the deep reinforcement learning model of gated fusion mechanism and online learning, it can dynamically optimize the acoustic-visual feature weights and significantly improve the adaptive ability of roof stability assessment; relying on the intelligent early warning system coordinated by edge computing and cloud computing, it can realize millisecond-level risk classification early warning and support strategy optimization, significantly reduce the roof fall accident rate, and provide efficient and reliable intelligent technical means for the safety management and control of coal mine tunnel excavation. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0027] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention. DETAILED DESCRIPTION
[0028] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0029] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0030] The present invention proposes a non-contact multimodal fusion real-time perception and early warning method for the risk of roof fall on the tunnel excavation face, aiming to solve the problems that the existing technology is difficult to accurately identify the surrounding rock structure of the tunnel roof and the precursors of the risk of roof fall, and lacks a systematic early warning solution for the risk of roof fall. Specifically, the method uses non-contact detection technology to arrange acoustic wave sensors and video image acquisition equipment on the tunnel excavation face, collect acoustic wave data and video image data in real time, pre-process, extract features and analyze the multimodal data, and fuse the feature information of the acoustic wave data and video image data to establish a dynamic correlation model between the features of multimodal heterogeneous data and the stability of the roof. Based on the fusion results, the risk of roof fall on the tunnel excavation face is evaluated in real time, different risk level thresholds are set, and a warning signal is issued in a timely manner when the risk assessment result exceeds the corresponding threshold. At the same time, we develop a software and hardware integration system that integrates multimodal data collection, processing, analysis, evaluation and early warning functions to achieve real-time and accurate perception and early warning of roof fall risks on tunnel excavation faces, and provide intelligent decision-making support for early identification of roof falls, risk classification and support optimization, thereby effectively reducing the probability of roof fall accidents in coal mine tunnels and ensuring the safety of coal mine workers and the normal production of coal mines.
[0031] The present invention provides a non-contact multi-modal fusion real-time perception and early warning method for roof collapse risk on a tunnel excavation face, comprising the following steps:
[0032] The first step is to non-contactly collect multi-modal heterogeneous data of the exposed surrounding rock of the tunnel excavation working face, build a non-contact monitoring system consisting of a high-frequency acoustic wave sensor array and a high-definition industrial camera, and collect the acoustic wave reflection waveform, surface texture deformation, crack development characteristics and water distribution signal of the tunnel surrounding rock in real time to accurately identify stress release characteristics.
[0033] The second step is to extract and analyze the multimodal data feature values of the exposed surrounding rock of the tunnel. The time-frequency energy entropy and non-stationary attenuation spectrum characteristics of the acoustic data are extracted through wavelet packet decomposition and empirical mode decomposition, and the crack density and acoustic waveguide anti-anomaly coefficient are quantified by combining deep convolutional network. At the same time, the fractal dimension, texture features and deformation gradient tensor of the video image are extracted using multi-scale Gabor filter and improved optical flow method, and high correlation features are screened through mutual information and random forest algorithm. Principal component analysis and t-SNE dimensionality reduction are used to construct a low-dimensional feature space. Finally, a two-stream cross-attention mechanism is designed, and the acoustic and visual features are fused with GRU and adversarial generative network to generate a 256-dimensional fusion feature vector, and the latent variable of roof collapse risk is calculated to realize a multimodal comprehensive evaluation of the stability of the exposed surrounding rock of the tunnel.
[0034] The third step is to dynamically correlate the characteristic values of the exposed surrounding rock data with the roof stability by constructing a multimodal feature-stability label dataset and combining it with the acoustic characteristic flow (D c , Z a 、Ewp ) and deformation characteristics (▽u, F d ) to generate a multi-dimensional time series feature matrix, and use the SMOTE algorithm to balance the category distribution; design a dual-branch time series feature extraction and deep Q network (DQN) hybrid architecture, reduce overfitting bias through prioritized experience replay and dual Q network, and realize dynamic decision-making of roof stability classification and support strategy; use the pre-trained ResNet-LSTM model for transfer learning, and combine the elastic weight solidification algorithm to realize online incremental training; quantify the dynamic correlation between multimodal features and roof stability through the multi-head attention mechanism, combine the SHAP value to analyze the contribution of key features, and establish quantitative mapping rules; use 5-fold time series cross-validation and Bayesian optimization to improve model performance, and deploy TensorRT to optimize the lightweight model, trigger response actions based on the hierarchical early warning mechanism, and realize real-time prediction and closed-loop feedback optimization.
[0035] The present invention proposes a non-contact multi-modal fusion real-time perception and early warning method for roof collapse risk in a tunnel excavation face, which specifically includes the following contents:
[0036] Step 1: Non-contact collection of multi-modal heterogeneous data of exposed surrounding rock at the tunnel excavation working face
[0037] The specific processing steps are as follows:
[0038] (1) Acoustic wave detection module
[0039] High-frequency acoustic wave sensors are deployed to actively emit acoustic wave pulse detection signals and record reflected waveforms, and the time difference method is used to invert the internal crack density and acoustic wave attenuation characteristics of the surrounding rock.
[0040] (2) Video monitoring module
[0041] Industrial-grade high-definition cameras are deployed to clearly capture the texture and crack details of the surrounding rock surface. Auxiliary LED light sources are also used to ensure clear imaging even in low-light conditions. The cameras are installed in protective covers with good sealing and impact resistance to adapt to the harsh mining environment.
[0042] (3) Multimodal data preprocessing
[0043] Wavelet threshold denoising is used to eliminate environmental noise in acoustic wave data, and Z-score normalization is used to eliminate dimensional differences. The Retinex algorithm is used to compensate for uneven illumination in video images, combined with morphological closing operations to repair crack and fracture areas, and the gray-level co-occurrence matrix is used to standardize texture features to ensure the consistency of multimodal data input.
[0044] Step 2: Extraction and analysis of multimodal data features of exposed surrounding rock in tunnels
[0045] (1) Extraction of time-frequency features of acoustic wave data
[0046] Extracting the time-frequency energy entropy E of sound waves based on wavelet packet decomposition wp :
[0047] E wp =-∑p i log p i #(1)
[0048] Among them, pi represents the proportion of the energy of the i-th frequency band to the total energy after wavelet packet decomposition.
[0049] The non-stationary attenuation spectrum features are extracted by combining empirical mode decomposition; the phase distortion and harmonic components of the reflected waveform are identified through a pre-trained deep convolutional network, and the crack density (D c , unit: bar / m 3 ) and the acoustic waveguide anti-anomaly coefficient Z a :
[0050] Z a =Z 实测 / Z 理论 #(2)
[0051] Among them, Z 实测 Indicates the measured acoustic impedance. This is the acoustic impedance value obtained through actual measurement, reflecting the true acoustic characteristics of the tunnel surrounding rock. 理论 Represents the theoretical acoustic impedance. This is the acoustic impedance value calculated based on an ideal model or prior knowledge, usually assuming that the surrounding rock is uniform and continuous.
[0052] (2) Video image texture and deformation feature extraction
[0053] The multi-scale Gabor filter bank is used to extract the surface fractal dimension (F d ) and local binary pattern (LBP) texture features;
[0054] The multi-scale Gabor filter bank consists of multi-band (scale) and multi-directional Gabor filters. Each filter captures the characteristics of different spatial frequencies and texture directions by adjusting the wavelength (such as λ = 4, 8, 16 pixels) and direction angle (0°, 45°, 90°, 135°) of the Gaussian kernel.
[0055] Combined with the improved Lucas-Kanade optical flow method to track the displacement vector field (accuracy ±0.1mm), the deformation gradient tensor is constructed
[0056] (3) Feature selection and cross-modal dimensionality reduction
[0057] Screening high correlation features (D c 、 F d ), and the sound-video features are reduced to 32 dimensions through principal component analysis.
[0058] t-SNE is used to visualize feature distribution. By optimizing the KL divergence of high-dimensional feature similarity probability distribution and low-dimensional mapping space, high-dimensional redundant features (such as linearly related sound wave frequency bands or video textures) are significantly overlapped or clustered in the low-dimensional space during visualization, thereby indirectly revealing the information redundancy between features. Redundant features with a discrimination value lower than a certain threshold (such as 5%) for classification tasks are quantitatively identified and eliminated, and a low-dimensional high-information feature space is constructed to obtain feature data after dimensionality reduction optimization (D' c 、 F' d ).
[0059] (4) Feature fusion and enhanced modeling
[0060] The roof collapse risk latent variable is calculated based on the above-mentioned dimension reduction and optimization data through the dual-stream cross-attention mechanism. The implementation process of the dual-stream cross-attention mechanism is as follows:
[0061] Design of a two-stream cross attention mechanism: Based on the multimodal feature vector (D' c 、 F' d ), reconstruct the roof fall risk latent variable R l :
[0062]
[0063] The α, β, and γ weights are dynamically generated by a two-stream cross-attention network. The specific steps are as follows:
[0064] (a) Input feature splitting and interaction
[0065] Physical feature flow (main branch): input the multimodal feature vector after dimensionality reduction (D' c 、 F' d ), a core physical quantity that characterizes crack density, deformation gradient, and surface fractal dimension.
[0066] Residual feature flow (auxiliary branch): the residual of the original high-dimensional data and the dimensionality reduction feature (D c -D' c 、 F d -F' d ), capturing nonlinear noise and potential weak risk signals.
[0067] (b) Cross-attention calculation
[0068] Physical features as query vectors Where W q is the learnable parameter matrix.
[0069] Residual features as key vector K and value vector V: Generate key vector With value vector
[0070] The interaction weights of the attention scores are calculated by scaling the dot product attention:
[0071]
[0072] where d k is the dimension of the key vector, used to stabilize the gradient.
[0073] (c) Dynamic generation of weights
[0074] Attention output mapping, mapping the attention output to the initial weight through the fully connected layer:
[0075] [α0,β0,γ0]=Linear(Attention(Q,K,V)) (5)
[0076] And normalize the weights:
[0077]
[0078] Step 3: Dynamic correlation modeling of exposed surrounding rock data characteristic values and roof stability
[0079] (1) Construction of multimodal feature-stability label dataset
[0080] Based on historical roof fall accident data and real-time monitoring results, the expert system labels the roof stability level (Level I: stable, Level II: warning, Level III: high risk), and the acoustic wave characteristics (D c , Z a 、E wp ) and deformation characteristics ( F d ) to form a multi-dimensional time series feature matrix, generate a training sample set through a sliding window, and use the SMOTE algorithm to balance the category distribution.
[0081] (2) Deep reinforcement learning model architecture design
[0082] A hybrid architecture of dual-branch temporal feature extraction + deep Q network dynamic decision-making is adopted. The dual-branch temporal feature extraction is divided into a perception branch and a decision branch:
[0083] The input of the perception branch is the multimodal feature vector (D' c , ▽u′, F'd ) obtained as the roof fall risk latent variable R l , combined with LSTM (long short-term memory network) to capture dynamic changes in time series.
[0084] The decision branch is based on a deep Q network and is used for dynamic decision making. The input is the roof collapse risk latent variable R extracted by the perception branch. l The temporal variation characteristics of the roof stability are output as roof stability classification (such as level I, level II, level III) and support strategy actions (such as "enhanced monitoring" and "emergency support").
[0085] (3) Transfer learning and online incremental training
[0086] Domain-adaptive transfer learning is performed based on the pre-trained ResNet-LSTM hybrid model. New data is collected in real time through a sliding window mechanism, and an elastic weight solidification algorithm is used to implement online incremental training, ensuring continuous optimization of the model in a dynamic environment.
[0087] Based on a pretrained ResNet-LSTM hybrid model (trained on a historical roof fall dataset), the model employs a domain adaptation method, aligning the feature distributions of new and old data using the Maximum Mean Difference (MMD) loss. A sliding window mechanism is then used to cache the latest monitoring data in real time (window length T = 60 seconds). During online training, the Elastic Weight Consolidation (EWC) algorithm is used to calculate the Fisher information matrix and impose constraints on key parameters, preventing new data from overwriting old knowledge. This allows for incremental updates of model parameters and ensures high-precision predictions despite changing rock mass conditions.
[0088] (4) Dynamic correlation modeling and quantitative relationship analysis
[0089] The multi-head attention mechanism is used to quantify the dynamic correlation between multimodal features and roof stability. The contribution of key features is analyzed through SHAP value (confidence threshold p=0.95), and a quantitative mapping rule between stability level and feature threshold is established.
[0090] Introducing the Multi-Head Attention mechanism at the top level of the model to calculate the sound wave characteristics (D c , Z a 、E wp ) and deformation characteristics (▽u, F d ) features, and quantify the dynamic impact of different modalities on stability (e.g., when the attention weight of displacement mutation is > 0.7, a Level II warning is triggered). SHAP value analysis (feature contribution based on game theory) is further used to screen non-significant features with p < 0.05, and to establish quantitative mapping rules between stability levels and key features (e.g., Z a >2.5 and Fd >1.2mm / s is determined as Level III), and the threshold logic is output through the decision tree interpretability module to support expert review.
[0091] (5) Cross-validation and model performance optimization
[0092] 5-fold time series cross-validation is used to evaluate the generalization of the model. Bayesian optimization is combined to search for hyperparameters. Dropout regularization and early stopping are used to prevent overfitting. The model performance is optimized by comprehensively considering accuracy, F1-score, and response delay.
[0093] (6) Real-time prediction and early warning decision generation
[0094] A lightweight model optimized with TensorRT is deployed, triggering response actions based on a graded early warning mechanism (Level I: Maintain monitoring; Level II: Encrypted sampling + audible and visual alarms; Level III: Emergency shutdown + automatic support), and optimizing long-term decision-making strategies through a closed-loop feedback mechanism. TensorRT optimization significantly improves real-time performance through model compression and inference acceleration: First, the trained ResNet-LSTM hybrid model for roof stability grading is optimized through graph structure optimization, including layer fusion (such as Conv-BN-ReLU merging), kernel automatic tuning (selecting the optimal GPU-compatible computing kernel), and FP16 / INT8 quantization calibration. While ensuring controllable accuracy loss (such as INT8 quantization error <1%), the model size is compressed by 60%-80%, and video memory usage is reduced by more than 50%. Second, TensorRT's dynamic batching and multi-stream parallel inference mechanisms are utilized to support concurrent processing of data from multiple monitoring points, meeting the real-time requirements of high-frequency monitoring scenarios.
[0095] 1. To address the problem of surrounding rock disturbance caused by the destructive installation of traditional monitoring technologies, this paper proposes a non-contact, multimodal data acquisition technology for efficient efficiency. By integrating acoustic wave detection and video image acquisition equipment, this technology enables simultaneous, non-contact acquisition of the acoustic characteristics of surrounding rock fractures and roof crack morphology. This technology eliminates the problem of interference with surrounding rock stability during data acquisition, ensures the authenticity and reliability of monitoring data, and provides a high-quality data foundation for surrounding rock stability assessment.
[0096] 2. This invention proposes a novel multimodal data acquisition solution that integrates data from multiple sensors, including acoustic and optical sensors, to comprehensively monitor exposed surrounding rock at the tunneling face. Acoustic detection captures the acoustic characteristics of rock fractures in real time, combined with video image acquisition technology to analyze the propagation morphology of roof cracks. This approach constructs a dynamic multimodal data coupling model, overcoming the lag in responding to hidden damage often associated with traditional monitoring methods. This multimodal data fusion acquisition method is a core innovation of this proposal and possesses significant practical and patent value.
[0097] 3. This paper proposes a multimodal data-driven intelligent decision-making model, innovatively constructing it. By integrating multimodal data such as acoustic waves and video images, combined with a deep reinforcement learning algorithm, it enables dynamic assessment of surrounding rock stability and risk prediction. This model addresses the challenge of insufficient analysis of latent damage characteristics, significantly improving the intelligent level of assessment and providing a highly efficient technical solution for mine safety production.
[0098] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A non-contact multimodal fusion real-time perception and early warning method for roof collapse risk in tunnel excavation face, characterized by: include: Acquiring multimodal heterogeneous data, wherein the multimodal heterogeneous data is data of different formats from different data sources, and the multimodal heterogeneous data includes physical characteristics reflecting different aspects of the tunnel excavation surface; Extract features from multimodal heterogeneous data to obtain multimodal features, perform key feature screening and dimensionality reduction optimization on the multimodal features, and calculate the roof fall risk latent variables based on the key features after dimensionality reduction optimization; The temporal variation characteristics of the roof fall risk latent variable are extracted, and decisions are made based on the temporal variation characteristics to obtain the roof risk stability level and support strategy.
2. The method according to claim 1, characterized in that The multimodal heterogeneous data includes acoustic wave data and image data.
3. The method according to claim 1, characterized in that Feature extraction also includes: Preprocess multimodal heterogeneous data, including standardization, denoising and data enhancement.
4. The method according to claim 2, characterized in that The feature extraction process includes: Energy entropy analysis is performed on the acoustic wave data to obtain acoustic wave video energy entropy; non-stationary attenuation spectrum features are extracted based on the acoustic wave data, and the non-stationary attenuation spectrum features are processed through identification and quantification methods to obtain anomaly coefficients; surface fractal dimension and texture features are extracted from the image data, and a deformation gradient tensor is constructed based on the graphic data; crack topology length and bifurcation angle are extracted based on the image data.
5. The method according to claim 2, characterized in that The process of key feature screening and dimensionality reduction optimization of multimodal features includes: The multimodal features are screened according to the relevance of the multimodal features, and the screened features are subjected to dimensionality reduction, and the dimensionality reduction-optimized features are further subjected to dimensionality reduction optimization according to information redundancy to obtain key features after dimensionality reduction optimization.
6. The method according to claim 1, characterized in that The process of obtaining the roof fall risk latent variable includes: The reduced-dimensional features are fused and enhanced through a dual-stream cross-attention mechanism, wherein the modal weights are obtained through the cross-attention mechanism based on the residuals between the multimodal features and the key features after dimensionality reduction optimization, and the key features after dimensionality reduction optimization are enhanced. Multimodal fusion is performed based on the modal weights and feature enhancement results to calculate the roof collapse risk latent variable.
7. The method according to claim 1, characterized in that The process of making decisions based on temporal variation characteristics includes: The time series variation characteristics of the roof fall risk latent variable are extracted by the method used to extract time series characteristics. The time series variation characteristics are processed and decided by the deep enhancement model to obtain the roof risk stability level and support strategy.
8. A non-contact multi-modal fusion real-time perception and early warning system for roof collapse risk in tunnel excavation faces, characterized by: Used to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Method and system for monitoring working face of thin coal seam based on computer vision
CN121095236A