A method and device for real-time prediction of geological anomalies in coal mines

By integrating multi-source sensors and deep learning models, real-time and accurate identification of geological anomalies ahead of coal mine tunneling has been achieved, solving the problem of insufficient real-time performance and accuracy of geological exploration in existing technologies, and improving the safety and efficiency of coal mining.

CN121211280BActive Publication Date: 2026-05-26CHINA COAL TECH & ENG GRP CHONGQING RES INST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA COAL TECH & ENG GRP CHONGQING RES INST CO LTD
Filing Date
2025-11-25
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies in coal mine geological exploration suffer from insufficient real-time continuity, detection resolution, information fusion depth, and intelligent interpretation, leading to the inability to accurately identify geological anomalies ahead of the working face.

Method used

By integrating stress sensors, gamma-ray sensors, ground-penetrating radar sensors, and pose sensors, multi-source data is collected and spatiotemporally aligned. Data fusion is performed using Gaussian kernel functions and deep learning models to generate a high spatiotemporal resolution three-dimensional geological voxel model, enabling dynamic real-time prediction of lithology and geological anomalies.

Benefits of technology

It enables dynamic real-time updates and accurate identification of geological conditions ahead of tunneling, significantly improving the timeliness and accuracy of geological support for coal mines and supporting unmanned and intelligent mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121211280B_ABST
    Figure CN121211280B_ABST
Patent Text Reader

Abstract

This invention discloses a method and device for real-time prediction of geological anomalies in coal mines, belonging to the fields of geological exploration, geophysical logging, and artificial intelligence. Data is synchronously collected and spatiotemporally aligned by multi-source sensors integrated into the tunneling equipment. The data sequence is normalized and rasterized into three-dimensional voxels, and a voxel grid containing lithology, stress anomalies, and reflection features is formed through Gaussian kernel function fusion. This grid is input into a deep learning model based on an encoding-cross-modal attention fusion-decoding architecture. Gamma, stress, and radar features are gradually fused through a cross-modal attention mechanism to generate unified semantic features. Based on these features, a high-resolution three-dimensional geological voxel model is constructed, and the lithology category of each voxel and the probability of the existence of geological anomalies such as fractures, water-rich areas, and collapse columns are predicted simultaneously. This achieves dynamic real-time updating and visualization of the geological conditions ahead of the tunnel, significantly improving the accuracy and timeliness of geological prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of geological exploration, geophysical logging and artificial intelligence, and in particular to a method and apparatus for real-time prediction of geological anomalies in coal mines. Background Technology

[0002] As coal mining shifts to deeper and more complex geological conditions, and "safety, efficiency, and intelligence" become the core goals of the industry's development, traditional production models are facing severe challenges. Among these challenges, the uncertainty of geological conditions is the biggest technical bottleneck restricting the improvement of the inherent safety level and production efficiency of coal mines. Thus, the concept of "transparent geology" has emerged.

[0003] "Transparent geology" does not refer to physical transparency, but rather to the precise, high-precision, and visualized perception and prediction of mine geological bodies (including coal and rock strata occurrence, geological structure, hydrogeology, stress environment, etc.) through advanced detection technologies and information processing methods, all in all times and spaces. Its strategic significance is mainly reflected in:

[0004] (1) Ensuring inherent safety: Major disasters in coal mine production, such as water inrush, gas outburst, roof collapse, and rock burst, are all closely related to unexplored, hidden geological factors that cause disasters (such as faults, fracture zones, water-rich areas, collapse columns, and high-stress areas). "Transparent geology" aims to reveal these "invisible killers" in advance and accurately, providing a basis for decision-making in disaster prevention and early warning, and transforming safety management from post-event emergency response to pre-event proactive prevention, which is the fundamental prerequisite for achieving inherent safety in coal mines.

[0005] (2) Improve production efficiency: Accurate geological models can guide the tunneling and mining faces to achieve precise geological orientation, ensuring that the equipment always operates efficiently in the coal seam, maximizing resource recovery rate, while reducing ineffective rock tunneling, reducing equipment wear and energy consumption. In addition, by avoiding or pre-treating unfavorable geological bodies, production interruptions can be effectively reduced, ensuring the continuity and stability of mining operations.

[0006] (3) Empowering Intelligent Mining: Unmanned and intelligent mining is the future direction of the coal industry. Whether it is the autonomous coal cutting of coal mining machines, the automatic following of hydraulic supports, or the autonomous navigation of tunneling robots, all their intelligent decision-making relies on a real-time, accurate, and high-resolution three-dimensional geological environment model. Without the "digital twin map" provided by "transparent geology", intelligent equipment is like "blind men touching an elephant", unable to operate safely and efficiently.

[0007] To achieve the above goals, the coal mining industry has applied various geological exploration technologies, but all have significant limitations and cannot fully meet the requirements of "transparent geology." The main challenges are as follows:

[0008] (1) Limitations of traditional macroscopic detection techniques:

[0009] Borehole exploration: As the most direct means of obtaining geological information, borehole exploration is essentially a discrete detection method that "uses points to represent areas." In structurally complex areas, a limited number of boreholes makes it difficult to control changes in geological bodies between two boreholes, and small to medium-sized anomalies such as faults, collapse columns, and karst fissures with scales smaller than the borehole spacing are easily missed. This inherent defect of insufficient spatial sampling results in limited accuracy and reliability of the geological models constructed using this method.

[0010] 3D seismic exploration: This technology can provide a large-scale macroscopic structural framework for mines and is effective in identifying large faults and folds. However, its core problem lies in its resolution bottleneck. Limited by the wavelength of seismic waves and acquisition conditions, its vertical and lateral resolution is usually on the order of meters or even ten meters, making it difficult to effectively identify small- and medium-scale fault structures (elevation less than 5 meters), fracture zones, and thin layers of water-rich sandstone that pose a direct threat to mining safety. Its results are static and cannot reflect the dynamic changes in the geological environment during mining operations.

[0011] (2) Shortcomings of advanced detection technology for working faces:

[0012] Mine geophysical methods (such as direct current methods and transient electromagnetic methods): These methods are widely used to detect anomalies such as water-bearing anomalies ahead of the working face. Their main problems are high ambiguity and limited accuracy. The interpretation of geophysical anomalies is not unique; for example, a low-resistivity anomaly may be a water-bearing area or caused by mudstone or conductive minerals, leading to low reliability of the identification results. Furthermore, their detection accuracy and range are greatly affected by the roadway environment, often only providing a general anomaly range and making it difficult to accurately locate the boundaries and morphology of the anomaly.

[0013] Advanced drilling: Drilling ahead of the working face can obtain accurate information, but the process interrupts the normal mining cycle, affecting production efficiency. At the same time, it is still discrete point detection and cannot achieve continuous and full-coverage detection of the geological body ahead.

[0014] (3) The multi-source information fusion layer is shallow and the level of intelligence is low:

[0015] Currently, coal mine geologists have attempted to combine data from boreholes, geophysical exploration, and seismic surveys for comprehensive geological interpretation. However, this integration largely remains at the level of data comparison and manual interpretation, representing a superficial "physical superposition." The complex and nonlinear intrinsic relationships between various data points have not been fully explored. For example, the coupling relationship between stress field changes, abrupt lithological changes, and water content is difficult to quantify accurately using only manual experience. There is a lack of an intelligent model capable of automatically learning these multidimensional characteristics and achieving deep information fusion.

[0016] (4) Existing measurement while drilling (MWD) technology has limited applications:

[0017] In current coal mine geological drilling, measurement-while-drilling (MWD) technology based on parameters such as natural gamma and resistivity is used for geological guidance, guiding the drill bit along coal seams or specific strata. However, existing applications have relatively singular purposes, mainly focusing on stratigraphic correlation and trajectory control. They typically utilize data from single sensors and lack a comprehensive, integrated solution for predicting various geological hazards (fractures, water, stress anomalies, etc.) by integrating sensors based on different physical principles (such as mechanics, nuclear physics, and electromagnetics).

[0018] In summary, existing technologies have significant shortcomings in real-time continuity, detection resolution, information fusion depth, and the level of intelligent interpretation, resulting in persistent problems of "unclear visibility and inaccurate detection" ahead of the working face during tunneling and mining. Therefore, the industry urgently needs a new technology that combines real-time, continuous, multi-physics data acquisition with powerful intelligent analysis models to fundamentally solve these problems and truly achieve dynamic, real-time, and transparent geological assurance. Summary of the Invention

[0019] The main objective of this invention is to provide a method and device for real-time prediction of geological anomalies in coal mines. By integrating multi-source sensors to synchronously collect data and constructing a three-dimensional geological voxel model, it aims to achieve high-precision dynamic prediction and visualization of lithology and various types of geological anomalies ahead of tunneling, thereby improving the geological support level and safe production capacity of coal mining.

[0020] To achieve the above objectives, the present invention provides a real-time prediction method for geological anomalies in coal mines. This method includes: S1, simultaneously acquiring stress data, multi-channel gamma-ray data, two-dimensional radar profile data, and equipment pose data along the tunneling path using stress sensors, gamma-ray sensors, ground-penetrating radar sensors, and pose sensors integrated at the front end of the tunneling equipment to obtain multi-source data; and performing spatiotemporal alignment of the multi-source data based on the transformation relationship between the equipment coordinate system and the geographic coordinate system to generate a multi-dimensional data sequence with spatiotemporal labels; S2, normalizing and rasterizing the multi-dimensional data sequence using three-dimensional voxel rasterization, and weighting and fusing the data within the same voxel using a Gaussian kernel function to form a multi-dimensional data sequence containing lithology, stress, gamma-ray, ground-penetrating radar, and pose sensors. S3, input the voxel grid of force anomalies and reflection features into a deep learning model based on an encoding-cross-modal attention fusion-decoding architecture, perform cross-modal attention query on stress features through gamma features, and then perform a second cross-modal attention query on radar features using the fused gamma-stress features, finally generating a unified fusion feature containing the semantic association of the multi-source data; S4, based on the unified fusion feature, generate a high spatiotemporal resolution three-dimensional geological voxel model through a three-dimensional upsampling network, and use a multi-task output layer to synchronously predict the lithology category of each voxel and the existence probability of three geological anomalies: fractures, water-rich areas, and collapse columns, so as to dynamically update and visualize the geological conditions ahead of the tunneling.

[0021] Furthermore, the step of simultaneously acquiring stress data, multi-channel gamma-ray data, two-dimensional radar profile data, and equipment pose data along the tunneling path by integrating stress sensors, gamma-ray sensors, ground-penetrating radar sensors, and pose sensors at the front end of the tunneling equipment to obtain multi-source data, and performing spatiotemporal alignment of the multi-source data based on the transformation relationship between the equipment coordinate system and the geographic coordinate system to generate a multidimensional data sequence with spatiotemporal labels includes: S11, calculating the spatial coordinates of the sensors in the geographic coordinate system using a formula: ;in, The coordinates of the reference point for the tunneling equipment. The rotation matrix is ​​calculated from the attitude angles. S12 is a fixed offset vector of the sensor relative to the reference point; S13 is a Kalman filter data fusion of a military-grade fiber optic inertial navigation unit and a high-precision odometer, which outputs the three-dimensional coordinates and attitude angle of the probe center point at a frequency of 100Hz to control the positioning accuracy to be better than 0.1 meters.

[0022] Furthermore, the step of normalizing and three-dimensional voxel rasterizing the multidimensional data sequence, and using a Gaussian kernel function to weight and fuse the data within the same voxel to form a voxel grid containing lithology, stress anomalies, and reflection characteristics includes: S21, applying a Z-score normalization formula to the stress data and gamma data. ,in, and S22 represents the mean and standard deviation over a sliding window; S23 uses a Gaussian kernel function to perform a weighted average on the radar image data, with the weights calculated based on the spatial distribution density of data points within the voxel.

[0023] Further, the step of inputting the voxel mesh into a deep learning model based on an encoding-cross-modal attention fusion-decoding architecture, performing cross-modal attention queries on stress features using gamma features, and then performing a second cross-modal attention query on radar features using the fused gamma-stress features, ultimately generating a unified fused feature containing the semantic association of the multi-source data, includes: S31, employing a multi-head self-attention mechanism, where each sub-module contains 8 attention heads, and calculating the attention score between different modal features using a formula: ;in, For querying the matrix, The key matrix, For value matrices, Let be the dimension of the key vector. To transform the similarity score into a function of probability distribution; S32, the decoder contains 4 layers of cross-modal attention modules, each followed by a residual connection and a layer normalization operation.

[0024] Furthermore, the step of generating a high spatiotemporal resolution 3D geological voxel model based on the unified fusion features through a 3D upsampling network, and simultaneously predicting the lithology category and the probability of existence of three geological anomalies (fractures, water-rich areas, and collapse columns) for each voxel using a multi-task output layer to dynamically update and visualize the geological conditions ahead of the tunneling includes: S41, using a 3D transposed convolutional layer for progressive upsampling, reducing the number of channels to half of the previous layer after each upsampling layer; S42, weighting and optimizing the joint loss function of lithology classification, fracture detection, water-rich area identification, and collapse column detection using a formula: ;in, These are hyperparameters used to balance different tasks. , , , The importance of.

[0025] Furthermore, the real-time prediction method for coal mine geological anomalies of the present invention also includes: S5, deploying the model on an underground edge computing server, using an NVIDIA Jetson AGX Orin module for real-time inference, and controlling the inference latency to within 500ms; S6, transmitting tunneling data via the mine 5G industrial ring network using the MQTT protocol to ensure synchronous updates of tunneling data with the ground dispatch center.

[0026] Another aspect of the present invention provides a real-time prediction device for geological anomalies in coal mines. This device includes: a multi-source data acquisition and spatiotemporal alignment module, used to simultaneously acquire stress data, multi-channel gamma-ray data, two-dimensional radar profile data, and equipment pose data along the tunneling path using stress sensors, gamma-ray sensors, ground-penetrating radar sensors, and pose sensors integrated at the front end of the tunneling equipment to obtain multi-source data; and to perform spatiotemporal alignment of the multi-source data based on the transformation relationship between the equipment coordinate system and the geographic coordinate system to generate a multi-dimensional data sequence with spatiotemporal labels; and a multi-dimensional data normalization and voxel fusion module, used to normalize and rasterize the multi-dimensional data sequence into three-dimensional voxels, and to perform weighted fusion of the data within the same voxel using a Gaussian kernel function to form a data sequence containing lithology and stress anomalies. The system includes a voxel mesh for reflection features; a cross-modal attention feature fusion module, which inputs the voxel mesh into a deep learning model based on an encoding-cross-modal attention fusion-decoding architecture, performs cross-modal attention queries on stress features using gamma features, and then performs secondary cross-modal attention queries on radar features using the fused gamma-stress features, ultimately generating a unified fusion feature containing the semantic associations of the multi-source data; and a 3D geological modeling and multi-task prediction module, which generates a high spatiotemporal resolution 3D geological voxel model based on the unified fusion feature through a 3D upsampling network, and uses a multi-task output layer to synchronously predict the lithology of each voxel and the probability of the existence of three geological anomalies: fractures, water-rich areas, and collapse columns, so as to dynamically update and visualize the geological conditions ahead of the tunneling.

[0027] Furthermore, the multi-source data acquisition and spatiotemporal alignment module is also used to: employ formulas Calculate the spatial coordinates of the sensor in the geographic coordinate system, where, The coordinates of the reference point for the tunneling equipment. The rotation matrix is ​​calculated from the attitude angles. The sensor is a fixed offset vector relative to a reference point; by fusing military-grade fiber optic inertial navigation unit with Kalman filter data from a high-precision odometer, the three-dimensional coordinates and attitude angle of the probe center point are output at a frequency of 100Hz to control the positioning accuracy to be better than 0.1 meters.

[0028] Furthermore, the multidimensional data normalization and voxel fusion module is also used to: apply the Z-score normalization formula to stress data and gamma data. ,in, and The mean and standard deviation are given on the sliding window; a Gaussian kernel function is used to perform a weighted average on the radar image data, and the weights are calculated based on the spatial distribution density of data points within the voxel.

[0029] Furthermore, the cross-modal attention feature fusion module is also used to: employ a multi-head self-attention mechanism, wherein each sub-module contains 8 attention heads, through a formula... Calculate attention scores between different modal features; the decoder contains 4 layers of cross-modal attention modules, each followed by residual connections and layer normalization operations.

[0030] Furthermore, the 3D geological modeling and multi-task prediction module is also used for: progressive upsampling using 3D transposed convolutional layers, reducing the number of channels to half of the previous layer after each upsampling layer; and using the formula... We perform weighted optimization on the joint loss function for lithological classification, fracture detection, water-rich area identification, and collapse column detection.

[0031] Furthermore, the real-time prediction device for coal mine geological anomalies of the present invention also includes: an edge computing deployment module, used to fuse Kalman filter data from a military-grade fiber optic inertial navigation unit and a high-precision odometer to output the three-dimensional coordinates and attitude angle of the probe center point at a frequency of 100Hz, so as to control the positioning accuracy to be better than 0.1 meters; and a data transmission module, used to transmit tunneling data via a mining 5G industrial ring network using the MQTT protocol to ensure that the tunneling data is updated synchronously with the ground dispatch center.

[0032] Compared with the prior art, the technical solution provided by the present invention has at least the following beneficial effects:

[0033] 1. Multi-source data collaborative perception: By integrating multiple sensors at the front end of the tunneling equipment, stress, gamma ray, ground penetrating radar and pose data are collected simultaneously, and spatiotemporal alignment is achieved based on coordinate system transformation to construct a multi-dimensional data sequence with spatiotemporal labels, forming a comprehensive three-dimensional perception capability of the geological environment ahead of the tunneling.

[0034] 2. Deep fusion of multimodal features: A deep learning model based on the "encoding-cross-modal attention fusion-decoding" architecture is adopted. By performing cross-modal attention query on stress features through gamma features, and then using the fused features to perform secondary attention query on radar features, the effective semantic association and deep fusion of multi-source geological data are realized.

[0035] 3. Real-time and accurate identification of geological anomalies: Based on unified fusion features, a three-dimensional geological voxel model with high spatiotemporal resolution is generated through a three-dimensional upsampling network. The multi-task output layer is used to synchronously predict the lithology of each voxel and the probability of the existence of various geological anomalies, realizing dynamic real-time updates and accurate identification of geological conditions ahead of tunneling, which significantly improves the timeliness and accuracy of geological support for coal mines.

[0036] This invention discloses a method and apparatus for real-time prediction of geological anomalies in coal mines. Through multi-source sensor data acquisition and deep fusion, it achieves accurate identification and three-dimensional visualization of geological anomalies ahead of tunneling. It also constructs a multi-dimensional voxel mesh incorporating lithological characteristics, stress anomalies, and reflection features, and achieves semantic-level fusion of multi-source geological data based on a cross-modal attention mechanism. This effectively solves the technical problems of data isolation and insufficient identification accuracy in traditional geological exploration methods. It significantly improves the real-time performance and accuracy of geological anomaly prediction, providing a reliable technical guarantee for safe and efficient coal mining. Attached Figure Description

[0037] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0038] Figure 1 A flowchart of a real-time prediction method for geological anomalies in coal mines provided in an embodiment of the present invention;

[0039] Figure 2 A detailed flowchart of a real-time prediction method for geological anomalies in coal mines provided in an embodiment of the present invention;

[0040] Figure 3 A detailed architecture diagram of a multimodal Transformer geological large model (GLM) for a real-time prediction method of coal mine geological anomalies provided in an embodiment of the present invention;

[0041] Figure 4 This is a schematic diagram of the structure of a real-time prediction device for geological anomalies in coal mines provided in an embodiment of the present invention. Detailed Implementation

[0042] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0043] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0044] The core idea of ​​this invention is to construct a dynamic perception and intelligent interpretation system that integrates multi-source geological sensing data. This system maps stress distribution, gamma-ray response, and radar reflection characteristics to a unified three-dimensional voxel space, enabling precise characterization of geological structures and anomalies ahead of the tunnel. Based on a cross-modal attention fusion mechanism, the system effectively establishes deep semantic relationships between different geological parameters, transforming traditional single-source geological exploration methods into an intelligent analysis paradigm characterized by multi-source data collaboration, deep feature fusion, and spatiotemporal continuity. Ultimately, through high-resolution three-dimensional geological modeling and multi-task collaborative prediction, real-time dynamic identification and visualization of lithological distribution and various geological anomalies are achieved, significantly improving the geological support capabilities and safety production level of coal mine tunneling faces.

[0045] The following describes, with reference to the accompanying drawings, a method and system for selecting collection scripts based on multimodal delinquency risk profiling according to an embodiment of the present invention.

[0046] Example 1

[0047] This embodiment provides a method for real-time prediction of geological anomalies in coal mines. For example... Figure 1 As shown, the method includes the following steps:

[0048] S1, by integrating stress sensors, gamma-ray sensors, ground-penetrating radar sensors and pose sensors at the front end of the tunneling equipment, simultaneously collects stress data, multi-channel gamma-ray data, two-dimensional radar profile data and equipment pose data along the tunneling path to obtain multi-source data, and performs spatiotemporal alignment of the multi-source data based on the transformation relationship between the equipment coordinate system and the geographic coordinate system to generate a multi-dimensional data sequence with spatiotemporal labels.

[0049] Specifically, in some implementations, the steps of this invention utilize multiple sensors integrated into the front end of the tunneling equipment, including stress sensors, gamma-ray sensors, ground-penetrating radar sensors, and pose sensors, to achieve synchronous acquisition and spatiotemporal alignment of geological information along the tunneling path. The core of this step lies in unifying sensor data from different physical principles, sampling frequencies, and spatial distributions into a consistent spatiotemporal coordinate framework, thereby providing structured and alignable input for subsequent multimodal data fusion and artificial intelligence modeling.

[0050] Specifically, the data from each sensor is timed according to its own timestamp during the tunneling process. Data was collected. The stress sensor continuously recorded formation stress values ​​with a spatial resolution of 0.25 meters. The gamma-ray sensor collects count values ​​from 16 sectors at a frequency of 2 Hz. Ground-penetrating radar collects one image every 0.1 meters. The device displays a two-dimensional cross-sectional view, while the pose sensor outputs the three-dimensional coordinates of the device's center point at a frequency of 100 Hz. and attitude angle By integrating an inertial navigation system (IMU) with an odometry system, the positioning accuracy of the device's pose data can reach within 0.1 meters.

[0051] Furthermore, based on the transformation relationship between the device coordinate system and the geographic coordinate system, all sensor data are spatiotemporally aligned. Specifically, for any given time... Collected sensor data Its spatial location Calculated using the following formula:

[0052] ;

[0053] in, It is a fixed offset vector of the sensor relative to the device reference point. This is a rotation matrix composed of attitude angles, used to transform the offset vector in the device coordinate system to the geographic coordinate system. Using this formula, data from various sensors can be mapped to a unified three-dimensional spatial coordinate system, thereby generating a multidimensional data sequence with spatiotemporal labels. ,in, It represents the three-dimensional spatial coordinates of the i-th data point in a unified geographic coordinate system, and serves as the spatial reference for the fusion and association of all data. It is data from the main sensing sensors and is the core data for subsequent tasks such as scene understanding and target detection. It is data from the Global Navigation Satellite System, serving as the spatiotemporal reference benchmark for the entire system. The data comes from the inertial measurement unit and is crucial for motion estimation during brief GNSS signal outages and for accurately calculating the rotation matrix R.

[0054] Specifically, the spatiotemporal alignment process must ensure the synchronization accuracy of the timestamps and spatial coordinates of the data from each sensor. For example, the timestamp error between gamma-ray and stress data should be controlled within 50 milliseconds to guarantee data alignment in the time dimension. Furthermore, the number of depth sampling points in the radar image... Typically 1024, number of scan channels The value is 64, and its spatial resolution is approximately 0.1 meters per pixel.

[0055] Specifically, in practical applications, this step is typically deployed in underground coal mine tunneling operations, suitable for continuous, real-time geological modeling under complex geological conditions. By unifying multi-source data into a geographic coordinate system, it provides a structured and interpretable data foundation for subsequent voxelization processing and deep learning model input.

[0056] Specifically, this step achieves spatiotemporal consistency of multi-source heterogeneous data, providing the model with high-precision, high-resolution input data, thereby significantly improving the accuracy and real-time performance of geological modeling. Its technical value lies in solving the fusion difficulties caused by asynchronous data and inconsistent coordinates in traditional methods, providing reliable data support for "transparent geology."

[0057] Furthermore, S1 also includes:

[0058] S11, using formula Calculate the spatial coordinates of the sensor in the geographic coordinate system, where, The coordinates of the reference point for the tunneling equipment. The rotation matrix is ​​calculated from the attitude angles. This is the fixed offset vector of the sensor relative to the reference point.

[0059] Specifically, this decoder employs a cross-modal attention structure. Its core idea is to use features from one modality as a query to focus on the keys and values ​​of another modality, thereby achieving information interaction and enhancement. In this invention, gamma-ray features are first used... As a basic query, stress features are queried through a cross-attention mechanism. The calculation method is as follows:

[0060]

[0061] Furthermore, through this mechanism, the model can learn the nonlinear correlation between gamma features and stress features. For example, in regions with high gamma values ​​(typically corresponding to mudstone), stress features may exhibit specific distribution patterns. The fused features... It is then used for cross-modal attention in the next layer, at which point... As a query, focus on radar image features The calculation method is as follows:

[0062]

[0063] Furthermore, through this layer-by-layer fusion structure, the model can gradually integrate local and global information from different modalities, ultimately outputting a fused feature. Its dimensions are It contains semantic information from all sensor data.

[0064] Specifically, cross-attention modules typically contain 8 attention heads, with an embedding dimension of [missing information]. Furthermore, residual connections and layer normalization (LayerNorm) are used to enhance the training stability of the model. In practical applications, this step is run on a downhole edge computing server (such as NVIDIA Jetson AGX Orin), with processing latency controlled within 500 milliseconds to meet real-time modeling requirements.

[0065] Specifically, a cross-modal attention mechanism was used to achieve deep semantic fusion of multi-source heterogeneous data, overcoming the limitations of traditional methods that rely on parallel data processing and human experience. The fused features output provide high-quality input for subsequent 3D geological decoding and multi-task prediction, significantly improving the model's accuracy and robustness in identifying key geological bodies such as coal-rock interfaces, fractures, and water-rich areas.

[0066] The S12 integrates military-grade fiber optic inertial navigation unit with Kalman filter data from a high-precision odometer to output the three-dimensional coordinates and attitude angle of the probe center point at a frequency of 100Hz, thereby controlling the positioning accuracy to be better than 0.1 meters.

[0067] Specifically, in some implementations, high-frequency, high-precision output of the three-dimensional coordinates and attitude angles of the probe's center point is achieved by fusing military-grade fiber optic inertial navigation units (IMUs) with Kalman-filtered data from high-precision odometers. This step is a crucial foundation for the entire system to achieve "transparent geology" modeling and real-time prediction, and its technical implementation relies on the fusion processing of multi-source sensor data and dynamic positioning algorithms.

[0068] Specifically, the IMU provides angular velocity and acceleration information, while the high-precision odometry provides kinematic displacement estimation. Because the IMU suffers from drift error, and the odometry may accumulate errors in complex terrain due to slippage or skidding, a Kalman filter algorithm is used to fuse the data from both to improve the robustness and accuracy of the positioning. Specifically, the system runs the Kalman filter at a frequency of 100Hz, integrating the angular velocity and acceleration of the IMU to obtain the attitude angle (roll angle). Pitch angle Azimuth Based on velocity estimation and odometer displacement information, the system employs a two-stage iterative optimization process of state prediction and observation update to output the real-time three-dimensional coordinates of the probe's center point. With attitude angle.

[0069] Furthermore, the system requires the IMU to have military-grade accuracy, with its attitude angle output error being less than [missing information]. The 3D coordinate positioning error is better than 0.1 meters. Odometry resolution is typically in the millimeter range, with a sampling frequency of 100Hz. The output displacement data needs to be filtered to eliminate slip noise. The covariance matrix of the Kalman filter... Process noise and observation noise Offline calibration is required based on the sensor characteristics to ensure that the filter has good convergence and stability during dynamic tunneling.

[0070] Specifically, this step is widely used in underground measurement-while-drilling (MWD) systems in coal mines, especially at the front end of tunneling machines or geological drilling equipment, to track the spatial position and attitude of the probe in real time. Through high-frequency positioning output, the system can accurately map multi-source data such as stress, gamma, and radar collected during drilling into three-dimensional geological space, providing a reliable spatial reference for subsequent voxel modeling and AI prediction.

[0071] Specifically, it achieves high dynamic and high-precision probe pose estimation, providing fundamental support for the spatiotemporal alignment of multi-source data. Its output is directly used in formulas. The pose parameter calculation in the model ensures that all sensor data can be fused and modeled in a unified geographic coordinate system, significantly improving the spatial consistency and predictive reliability of the geological model.

[0072] S2, normalize and rasterize the multidimensional data sequence into three-dimensional voxels, and use a Gaussian kernel function to weight and fuse the data within the same voxel to form a voxel grid that includes lithology, stress anomalies and reflection characteristics.

[0073] Specifically, in some implementations, normalizing and rasterizing the multidimensional data sequence is a key preprocessing step in constructing a transparent geological model of a coal mine. This step aims to eliminate differences in dimensions and dynamic range between data from different sensors, while mapping discrete spatiotemporal data to a unified three-dimensional spatial structure, providing structured input for subsequent deep learning models.

[0074] Furthermore, the normalization process employs the Z-score standardization method for the stress data. and Gamma Data The processing is performed using the following formula:

[0075]

[0076] in, and These represent the mean and standard deviation of the stress data within the sliding window, respectively. This method enhances the ability to identify outliers, making the model more sensitive to geological abrupt changes. For ground-penetrating radar data... Then, pixel value normalization is used to map it to the [0, 1] interval to adapt to the input requirements of the image encoder.

[0077] Furthermore, the three-dimensional voxel rasterization process constructs a three-dimensional mesh centered on the current position of the tunneling equipment. Its spatial resolution is typically set to This is consistent with the spatial sampling resolution of the fiber optic stress sensor. Each voxel... Each sensor data point is assigned a spatial coordinate system, and all sensor data points are mapped to their corresponding voxel locations based on their spatiotemporal labels. Since different sensors have varying sampling frequencies and spatial distributions, the same location may contain multiple data points. Therefore, a Gaussian kernel function is used to weight and fuse these multi-source data. The Gaussian kernel function assigns different weights based on the distance between the data point and the voxel center, thus achieving smooth spatial fusion and improving the model's ability to perceive local geological features.

[0078] Specifically, this step is widely used in real-time geological modeling during coal mine tunneling. For example, every time the tunneling machine advances 0.5 meters, the system normalizes and rasterizes sensor data from the past 10 meters, forming a voxel grid that includes lithology, stress anomalies, and reflection characteristics. This grid serves as input to a deep learning model to predict the geological structure within a 50-meter radius ahead.

[0079] Specifically, this step effectively improves the fusionability of multi-source data and the stability of model input through standardization and spatial alignment. Meanwhile, Gaussian weighted fusion enhances the representativeness of data within voxels, laying a solid foundation for subsequent multimodal feature encoding and cross-attention fusion. This is a key preprocessing step for achieving dynamic modeling and disaster early warning of "transparent geology."

[0080] Furthermore, S2 also includes:

[0081] S21, Z-score normalization formula is used for stress data and gamma data.

[0082]

[0083] in, and These are the mean and standard deviation over the sliding window.

[0084] Specifically, in the data preprocessing stage of this invention, stress data and gamma-ray data are normalized using the Z-score standardization formula. The core purpose is to eliminate the differences in dimensions and distributions of data from different sensors, thereby enhancing the model's sensitivity to outliers and improving the accuracy and robustness of multimodal data fusion.

[0085] Furthermore, Z-score normalization is a data preprocessing method based on statistical properties, and its formula is:

[0086]

[0087] in, Represents the first in the original stress or gamma data One measurement value, This represents the mean of the data within the sliding window. This represents the standard deviation within the window. This formula transforms the raw data into a standard normal distribution with a mean of 0 and a standard deviation of 1, allowing for comparison and fusion of data from different sensors within a unified numerical range. In practice, the length of the sliding window is typically set based on the data acquisition frequency and tunneling speed. For example, for gamma data, if the sampling frequency is... The sliding window can then be set to include the most recent... sampling points (i.e.) (Data within seconds) to balance real-time performance and statistical stability.

[0088] Specifically, the key parameters for Z-score standardization include the sliding window size and the mean. and standard deviation The calculation method is as follows. In this invention, the sliding window uses time alignment to ensure consistent time resolution across different sensor data. Furthermore, the standardized data range is theoretically [missing information]. However, in practical applications, due to the distribution characteristics of geological data, most data points will fall within... This effectively highlights outliers (such as stress mutations or abnormal increases in gamma values).

[0089] Furthermore, this step is primarily used to process the stress sequences continuously acquired along the tunneling path. and gamma ray vector sequence ,in The number of sectors of the gamma detector (usually 1000). or By standardizing the data using Z-scores, these data are unified to the same numerical scale, providing a consistent feature representation for subsequent Transformer encoder inputs, thereby improving the model's ability to fuse multi-source data.

[0090] Specifically, standardization effectively eliminates offset and scaling differences in sensor data, enabling the model to more accurately identify geological anomalies during training and inference. For example, high-stress or high-gamma regions (such as mudstone layers) are more easily captured by the model after standardization, revealing their potential association with geological hazards (such as fissures and water-rich areas), thus providing high-quality input features for subsequent cross-modal attention fusion. Therefore, Z-score standardization is one of the key preprocessing steps in this invention for achieving high-precision, real-time geological modeling and hazard prediction.

[0091] S22 uses a Gaussian kernel function to perform a weighted average on the radar image data, with the weights calculated based on the spatial distribution density of data points within the voxel.

[0092] Specifically, in the data preprocessing stage, the radar image data is weighted using a Gaussian kernel function, with the weights calculated based on the spatial distribution density of data points within voxels. The core objective of this step is to assign differentiated weights to data points at different locations in the radar image through a spatial density-aware mechanism, thereby preserving key geological information, suppressing noise interference, and improving the accuracy and robustness of 3D geological modeling during voxelization.

[0093] Specifically, this step first involves rasterizing the radar image data (B-scan profile), that is, mapping continuous detection data onto a three-dimensional voxel grid. In the middle. Each voxel represents a spatial resolution unit, the size of which is determined by the resolution parameter of the voxel grid. Decision, usually set To match the spatial sampling accuracy of stress and gamma data, when multiple radar data points fall into the same voxel, a Gaussian kernel function is used to weight the reflection amplitudes of these points, with the weight proportional to the spatial distribution density of the data points within the voxel.

[0094] Furthermore, let voxels be... Contains The spatial coordinates of the radar data points are: The reflection amplitude at each point is Then the comprehensive attribute value of the voxel. It can be represented as:

[0095]

[0096] Among them, weight Calculated using the Gaussian kernel function:

[0097]

[0098] Specifically, Indicates the first The distance from each data point to the center of the voxel The standard deviation of the Gaussian kernel is typically set based on the spatial resolution and noise level of the radar image, for example... This ensures that neighboring points have higher weights, while the weights of points far from the center decay rapidly.

[0099] Furthermore, due to multipath reflection, medium attenuation, and noise interference in radar images, simply using averages or maximum values ​​can easily lead to distortion of geological features. By introducing a Gaussian weighted average based on spatial density, the local reflection characteristics of geological anomalies (such as fissures and water-rich areas) can be effectively preserved, while smoothing noise in non-critical areas and improving the model's ability to identify geological structures.

[0100] Specifically, this method significantly improves the representation quality of radar data in a three-dimensional voxel grid, providing a more reliable data foundation for subsequent multimodal feature fusion. In the "encoding-fusion-decoding" architecture of this invention, this step is one of the key links in realizing multi-source data alignment and fusion, ensuring the consistency and complementarity of data from different sensors in the spatial dimension, thereby enhancing the model's perception accuracy and prediction capability of the coal mine geological environment.

[0101] S3, the voxel grid is input into a deep learning model based on an encoding-cross-modal attention fusion-decoding architecture. The stress features are subjected to cross-modal attention query through gamma features, and the radar features are subjected to a second cross-modal attention query using the fused gamma-stress features. Finally, a unified fusion feature containing the semantic association of the multi-source data is generated.

[0102] Specifically, the voxel grid is input into a deep learning model based on an "encoding-cross-modal attention fusion-decoding" architecture. Gamma features are used to perform cross-modal attention queries on stress features, and then the fused gamma-stress features are used to perform a second cross-modal attention query on radar features, ultimately generating a unified fused feature containing semantic associations of multi-source data. This step is the core component of this invention for achieving deep fusion of multi-source sensor data and geological semantic modeling.

[0103] Specifically, the model employs a multimodal Transformer encoder to extract features from gamma, stress, and radar data. The gamma and stress data, as one-dimensional sequences, are mapped to a unified embedding space through patching and linear projection. ,in For sequence length, The embedding dimension is set to 512 in this invention. The radar image is then processed into blocks using a Vision Transformer (ViT), with each image block being [size missing]. Pixels, after being flattened, are mapped to... The sequence is transformed into a dimensional vector, and a [CLS] marker is added to the front of the sequence to obtain global features. Location information is encoded through learnable location embeddings to preserve the spatial order of the data.

[0104] Furthermore, in the cross-modal fusion process, gamma features are first used... as query vector For the encoded stress features Perform cross-attention computation, i.e. , , Through the formula ,in, For querying the matrix, The key matrix, For value matrices, Let be the dimension of the key vector. To transform similarity scores into a function of probability distributions, this model can extract contextual information related to gamma features from stress characteristics, thereby enhancing the understanding of the coupling relationship between lithology and stress. Subsequently, the fused... As a new query vector, for radar features Perform a second cross-modal attention query, i.e. , , Furthermore, by integrating reflection structure information from radar images, the ability to identify geological anomalies (such as fissures and water-rich areas) can be improved.

[0105] Specifically, the model employs a 12-layer Transformer encoder, with each layer containing 8 attention heads, and the fusion module is a 4-layer Transformer decoder structure. (Embedding dimension) Input sequence length Number of radar image blocks By using residual connections and layer normalization, the stability and convergence of model training are ensured.

[0106] Specifically, this step is mainly used in practical applications for 3D geological modeling within a 50m × 20m × 20m area in front of the tunneling equipment. On a downhole edge computing server (such as NVIDIA Jetson AGX Orin), the model inference latency is controlled within 500 milliseconds, meeting real-time requirements. The fused features are used to drive the multi-task output head, realizing joint modeling of lithological classification and anomaly probability prediction.

[0107] Specifically, this step utilizes a cross-modal attention mechanism to achieve deep interaction between three heterogeneous data types—gamma, stress, and radar—in the feature space, effectively enhancing the model's ability to perceive complex geological structures. Experiments show that this fusion strategy achieves a lithological classification accuracy exceeding 96% and an intersection-over-union (IoU) ratio greater than 0.85 for anomaly identification, significantly outperforming traditional data parallel fusion methods.

[0108] Furthermore, S3 also includes:

[0109] S31 employs a multi-head self-attention mechanism, where each sub-module contains 8 attention heads, and the attention scores between different modal features are calculated using a formula:

[0110] ;

[0111] in, For querying the matrix, The key matrix, For value matrices, Let be the dimension of the key vector. This is a function that transforms similarity scores into a probability distribution.

[0112] Specifically, in the multimodal feature encoder of this invention, a multi-head self-attention (MHSA) mechanism is employed to extract features from one-dimensional sequential data (such as stress data and gamma data). This mechanism enables the model to capture local and global dependencies of data from different representation subspaces by computing multiple attention heads in parallel, thereby enhancing the modeling ability for complex geological features.

[0113] Furthermore, each submodule contains eight attention heads, each independently performing a linear transformation on the input sequence to generate a query, key, and value matrix. Specifically, the input sequence is first mapped to the embedding dimension through a linear projection layer. The feature space is used to form word embedding vectors. Then, each attention head projects the word embeddings onto the feature space. A dimensional query, key, and value vector space, where This ensures that the dimension of each head matches the total embedding dimension. The attention score is calculated using the formula:

[0114]

[0115] The calculation yielded, where These are query, key, and value matrices, respectively. Used to normalize attention weights It is a scaling factor used to alleviate the gradient vanishing problem caused by excessively large inner product results.

[0116] Furthermore, the multi-head mechanism concatenates the outputs of the eight attention heads and integrates them through a learnable linear transformation layer to output the final attention features. This process not only enhances the model's ability to model long-range dependencies in sequences but also improves its sensitivity to local anomalies (such as stress mutations and gamma value anomalies). In this invention, this mechanism is embedded in a module consisting of a 12-layer Transformer encoder, with each layer followed by residual connections and layer normalization to stabilize the training process and accelerate convergence.

[0117] Specifically, this step plays a crucial role in this invention, enabling the model to extract high-dimensional, structured geological feature representations from one-dimensional sensor data acquired during drilling, providing high-quality input for subsequent cross-modal fusion and three-dimensional geological modeling. Through parallel computing with eight heads, the model achieves richer feature interactions within a limited embedding dimension, significantly improving its ability to identify coal and rock strata boundaries, stress anomalies, and lithological abrupt changes.

[0118] The S32 decoder contains four layers of cross-modal attention modules, each followed by a residual connection and a layer normalization operation.

[0119] Specifically, in this invention, the decoder comprises four layers of cross-modal attention modules, each followed by residual connections and layer normalization operations, and is a key component for realizing deep fusion of multi-source sensor data and 3D geological modeling. This module, based on the Transformer architecture, achieves information interaction between different modal features through a cross-modal attention mechanism, thereby enhancing the model's ability to perceive complex geological structures.

[0120] Furthermore, the input to the cross-modal attention module is a high-dimensional feature sequence of three modalities. , and These correspond to the encoding results of gamma, stress, and radar images, respectively. For embedded dimensions, For sequence length, This refers to the number of image patches. The fusion process uses gamma features. As the initial query, it is sequentially linked to stress features. and radar characteristics Cross-modal attention computation is performed. Specifically, in the first-layer cross-modal attention module, , , Through the attention formula:

[0121]

[0122] Calculate the fused features This feature preserves the lithological information of the gamma data and introduces the mechanical response characteristics of the stress data. In the second layer, , , Further integrate the reflection pattern information of radar images to form This process is repeated in four cross-modal attention modules, each employing a multi-head attention mechanism with eight heads to enhance the model's ability to capture features from different subspaces.

[0123] Specifically, to ensure the stability and convergence of model training, each cross-modal attention module is followed by a residual connection and layer normalization. The residual connection adds the original input to the attention output to avoid gradient vanishing; layer normalization standardizes the feature vectors to have a mean of 0 and a variance of 1, accelerating the training process and improving the model's generalization ability.

[0124] Specifically, in practical applications, this step is deployed on downhole edge computing servers (such as NVIDIA Jetson AGXOrin) to support real-time prediction of a 50m × 20m × 20m three-dimensional space ahead during tunneling. Through a cross-modal attention mechanism, the model can automatically learn the coupling relationship between lithology, stress, and radar reflection. For example, it can identify stress anomaly patterns associated with high gamma-value areas (mudstone) or use reflection arc features in radar images to assist in determining fault boundaries. The final output is a fused feature. It provides high-quality input for subsequent 3D upsampling and multi-task prediction, significantly improving the accuracy and reliability of geological modeling.

[0125] S4. Based on the unified fusion features, a three-dimensional geological voxel model with high spatiotemporal resolution is generated through a three-dimensional upsampling network. A multi-task output layer is used to synchronously predict the lithology category of each voxel and the probability of the existence of three geological anomalies: fractures, water-rich areas, and collapse columns, so as to dynamically update and visualize the geological conditions ahead of the tunneling.

[0126] Specifically, this step, based on unified fusion features, generates a high spatiotemporal resolution 3D geological voxel model through a 3D upsampling network. It then employs a multi-task output layer to simultaneously predict the lithology of each voxel and the probability of the existence of three geological anomalies: fractures, water-rich areas, and collapse columns. This enables dynamic, real-time updates and visualization of the geological conditions ahead of the tunnel. This step is the core output of the "transparent geology" modeling system of this invention, exhibiting high real-time performance and prediction accuracy.

[0127] Specifically, the decoder first fuses the features Feature reshaping is performed through linear layers, transforming it from a two-dimensional feature sequence. Convert into a low-resolution 3D feature volume Subsequently, an upsampling operation is performed using a multi-layered 3D transposed convolution structure to progressively increase the spatial resolution of the voxel mesh. Each transposed convolution layer is followed by batch normalization and ReLU activation functions to enhance the model's non-linear expressiveness and accelerate training convergence. The final output 3D voxel model has a resolution of [resolution missing]. It covers a space of approximately 50 meters × 20 meters × 20 meters in front of the tunnel, meeting the requirements for high-precision geological modeling.

[0128] Furthermore, in the multi-task output layer, the model outputs two predictions in parallel for each voxel: the lithology classification probability and the probability of the presence of three geological anomalies. The lithology classification head uses the Softmax activation function and outputs... The probability distribution of each lithology category is calculated using the following formula:

[0129]

[0130] The anomaly probability head uses the Sigmoid activation function to independently predict the existence probability of cracks, water-rich areas, and collapse columns. The formula is as follows:

[0131]

[0132] The loss functions for each task are cross-entropy loss and binary cross-entropy loss, respectively, and the joint loss function is defined as:

[0133]

[0134] in, These are the weight coefficients for each task, used to balance the training priority of different tasks.

[0135] Specifically, this step runs on an edge computing server (such as NVIDIA Jetson AGXOrin) downhole in the tunneling machine, with model inference latency controlled within 500 milliseconds to ensure real-time updates of the geological model. The prediction results are rendered using 3D visualization software, with coal seams represented as black solids, rock strata as yellow, water-rich areas overlaid with blue semi-transparent cloud maps, and fractures and collapse columns presented as red semi-transparent structures, facilitating intuitive identification of potential risk areas by operators.

[0136] Specifically, through the decoding and multi-task output mechanism of the deep learning model, high-precision, multi-attribute joint prediction of complex geological structures is achieved, providing dynamic and visualized geological environment perception capabilities for intelligent coal mine tunneling, and significantly improving the accuracy and response efficiency of disaster early warning.

[0137] Furthermore, S4 also includes:

[0138] S41 uses a three-dimensional transposed convolutional layer for progressive upsampling, reducing the number of channels to half of the previous layer after each upsampling.

[0139] Specifically, in the decoder module, progressive upsampling using 3D transposed convolutional layers is a key step in gradually restoring the low-resolution fused features to the target 3D geological grid resolution. This process achieves a mapping from the compressed feature space to the high-resolution voxel space through a series of transposed convolutional operations. At the same time, the number of channels is reduced to half of the previous layer after each upsampling, thus gradually approximating the final geological attribute output.

[0140] Specifically, a 3D transposed convolutional layer is essentially a deconvolution operation. Its core principle is to "decompress" and spatially reconstruct the features by expanding the spatial dimension of the feature map while reducing the number of channels. Specifically, the size of the input feature volume is... ,in The depth, height, and width of the current feature body. Where is the number of channels. The output size of each transpose convolution operation is determined by the following formula:

[0141]

[0142]

[0143]

[0144] in, The stride is the step size. The kernel size is the convolution kernel size. For padding. In this invention, it is typically set to... , , This achieves a 2x upsampling rate. The number of channels is halved after each layer of operation, for example, from... Gradually down to This continues until the number of output channels is reached, which is the number of geological attribute categories (such as lithological classification and anomaly probability).

[0145] Furthermore, each transposed convolutional layer is followed by batch normalization and the ReLU activation function to enhance the model's stability and nonlinear expressive power. Batch normalization standardizes the features of each channel to a mean of 0 and a variance of 1, while the ReLU function preserves positive features and suppresses negative responses, helping the model focus on meaningful geological features.

[0146] Specifically, in practical applications, this step runs on an underground edge computing server (such as NVIDIA Jetson AGXOrin), with inference latency controlled within 500 milliseconds to meet real-time modeling requirements. Through progressive upsampling, the model can gradually recover the spatial details of the geological body, thereby achieving high-precision prediction of key geological elements such as coal seams, faults, and water-rich areas in the final output 3D voxel mesh. This provides tunneling equipment with a dynamic and visualized "transparent geological" environment model, significantly improving the safety and intelligence level of coal mining.

[0147] S42 uses a formula to perform weighted optimization of the joint loss function for lithology classification, fracture detection, water-rich area identification, and collapse column detection:

[0148] ;

[0149] in, These are hyperparameters used to balance different tasks. , , , The importance of.

[0150] Specifically, in this invention, the multi-layer Transformer decoder structure with cross-attention mechanism is a key module for realizing deep fusion of multi-source sensor data and 3D geological modeling. This step fuses high-dimensional features from different physical fields (stress, gamma rays, ground-penetrating radar) layer by layer by constructing the interaction relationship between multimodal features, thereby generating 3D voxel prediction results with geological semantics.

[0151] Furthermore, this decoder employs a cross-modal attention mechanism. The core idea is to use features from one modality as a query vector to focus on the key and value vectors of another modality, thereby achieving information complementarity and enhancement. Specifically, it first uses gamma-ray features... As a basic query, it is used across attention modules and stress features. The interaction is performed, and the calculation formula is as follows:

[0152]

[0153] Through this mechanism, the model can learn the nonlinear correlation between gamma features and stress features. For example, in regions with high gamma values ​​(potentially corresponding to mudstone), stress features may exhibit specific distribution patterns. Furthermore, the fused features... As a new query vector, it continues to be used with ground-penetrating radar features. Perform cross-attention interactions:

[0154]

[0155] Specifically, this process can be designed as a multi-layered, multi-directional structure. For example, stress-based reverse lookups of radar can be introduced in subsequent layers to achieve bidirectional information flow between multiple modalities. Each layer is configured with residual connections and layer normalization after the attention module to improve the training stability and generalization ability of the model.

[0156] Specifically, through a multi-layer cross-attention mechanism, the model outputs a unified feature representation that integrates information from all sensors. This feature is fed into the decoder module, where it is progressively upsampled through linear mapping and 3D transposed convolution to restore the resolution to the target 3D geological mesh. The number of channels is Each transposed convolutional layer is followed by batch normalization (BatchNorm) and ReLU activation function to enhance the model's nonlinear expressive power.

[0157] Specifically, through a cross-modal attention mechanism, deep semantic fusion of multi-source heterogeneous data is achieved, overcoming the limitations of traditional methods that rely on parallel data comparison and manual interpretation based on experience. In practical applications, this architecture is deployed on underground edge computing platforms to support real-time prediction of geological bodies within a 50-meter radius ahead during tunneling, providing a high-precision, dynamically updated "transparent geological" model for intelligent coal mining.

[0158] S5 deploys the model on a downhole edge computing server, uses the NVIDIA Jetson AGX Orin module for real-time inference, and keeps the inference latency within 500ms.

[0159] Specifically, in some implementations, deploying the trained deep learning model on a downhole edge computing server and using the NVIDIA Jetson AGX Orin module for real-time inference is a key step in achieving "transparent geology" dynamic modeling and advanced prediction in this invention. The core objective of this step is to ensure that the model can still perform the fusion and interpretation of multi-source drilling-while-drilling sensor data with high real-time performance in downhole environments characterized by high noise, high humidity, limited power supply, and constrained computing resources, thereby providing immediate geological decision support for tunneling operations.

[0160] Specifically, the Jetson AGX Orin module, as an edge computing platform, is equipped with an NVIDIA Orin system-on-a-chip (SoC) boasting up to 275 TOPS of AI computing power. It supports the TensorRT acceleration engine, enabling efficient execution of deep learning models based on the Transformer architecture. Before model deployment, quantization processing (such as FP32 to INT8 conversion) is required to reduce computational load and improve inference speed. Simultaneously, the model input data undergoes preprocessing, including spatiotemporal alignment, Z-score normalization, and voxelization, ultimately forming a three-dimensional tensor with dimensions of [missing information]. ,in For low-resolution three-dimensional space dimensions, This represents the number of feature channels after fusion.

[0161] Furthermore, inference latency is strictly controlled to within 500ms to meet the real-time requirement of performing geological predictions every 0.5 meters of tunneling. This latency metric is based on the inference performance test results of Jetson AGX Orin optimized with TensorRT, ensuring that the end-to-end response time of data acquisition, transmission, and model inference does not exceed the system's set threshold under the support of an underground 5G industrial ring network. In addition, the input data window length for model inference is 10 meters, and the output prediction range is a three-dimensional space of 50 meters × 20 meters × 20 meters ahead, with a voxel resolution of 0.25 meters, consistent with the ground truth model during the training phase.

[0162] Specifically, this step applies to real-time geological modeling systems for underground coal mine tunneling operations. Every time the tunneling machine advances a certain distance (e.g., 0.5 meters), the edge server triggers a model inference, inputting the latest collected stress, gamma, and radar data into the model and outputting the lithological classification and anomaly probability prediction of the geological voxels ahead. The prediction results are rendered in real-time using 3D visualization software, representing key geological elements such as coal seams, strata, water-rich areas, and fracture zones with different colors and transparency, providing intuitive geological risk warnings for the driver and dispatch center.

[0163] Specifically, by deploying high-performance AI modules at the edge, local real-time processing and inference of multi-source sensor data are achieved, significantly reducing reliance on ground-based central computing and improving the system's response speed and robustness. Meanwhile, a latency of less than 500ms ensures the continuity and safety of tunneling operations, enabling the "transparent geology" model to play a role in real-time early warning and decision support in actual production, thereby effectively improving the level of intelligent coal mining.

[0164] S6 transmits tunneling data via the MQTT protocol through a mining 5G industrial ring network, ensuring that the tunneling data is updated synchronously with the ground dispatch center.

[0165] Specifically, in some implementations, transmitting data via the MQTT protocol through a mining 5G industrial ring network is a key communication link for achieving real-time synchronization of tunneling data with the ground dispatch center. This step is technically based on an Industrial Internet of Things (IIoT) architecture, combining a highly reliable 5G wireless communication network with the lightweight MQTT (Message Queuing Telemetry Transport) protocol to ensure low-latency, high-fidelity data transmission to the ground system even in complex underground electromagnetic environments and under conditions of high dust and humidity.

[0166] Furthermore, the multi-source sensors integrated at the front end of the tunneling equipment (including stress, gamma-ray, ground-penetrating radar, and pose sensors) undergo data aggregation and preliminary processing through a mine-use explosion-proof edge computing server. This server encapsulates the sensor data into structured message bodies and publishes them to the 5G industrial ring network via the MQTT protocol. The MQTT protocol uses a publish-subscribe model, with the ground dispatch center acting as a subscriber, receiving the tunneling data stream in real time through preset topics. This protocol features low bandwidth consumption, low power consumption, and high real-time performance, making it suitable for underground communication environments.

[0167] Specifically, the transmission rate of the 5G industrial ring network should be no less than 100 Mbps, and the latency should be controlled within 50 ms to meet the real-time requirements of tunneling data. MQTT protocol message transmission adopts QoS (Quality of Service) level 1 to ensure that messages are delivered at least once, while a heartbeat mechanism (Keep Alive) is used to maintain connection stability; the heartbeat interval is recommended to be set to 30 seconds. The data packet format must comply with the ISO / IEC 14882 standard and use JSON or a binary protocol (such as Protobuf) for serialization to improve transmission efficiency and parsing speed.

[0168] Specifically, this step is widely used in intelligent tunneling systems for coal mines, particularly in real-time geological modeling and disaster early warning scenarios in deep mining and complex geological areas. During tunneling, the edge computing server uploads processed sensor data to the ground data center via a 5G ring network. Upon receiving the data, the ground system can immediately use it for model training, data fusion, and updating of the 3D geological model. This communication link also supports the issuance of remote control commands, enabling dynamic adjustment of the tunneling path.

[0169] Specifically, this step ensures the real-time nature and integrity of the tunneling data, providing high-quality data input for subsequent multimodal data fusion and AI model inference. Through the lightweight design of the MQTT protocol and the high bandwidth of the 5G network, the system can achieve stable data transmission in the harsh underground environment, thereby supporting the ground dispatch center's real-time perception and decision-making response to underground geological conditions, significantly improving the safety and intelligence level of coal mining.

[0170] This invention discloses a real-time prediction method for geological anomalies in coal mines, which enables high-precision, real-time 3D modeling and anomaly prediction of geological bodies ahead of the coal mine working face. It achieves 500ms-level real-time inference through an NVIDIA Jetson AGX Orin module deployed on an underground edge computing server, and combines a mining 5G industrial ring network and MQTT protocol to ensure low-latency synchronization between tunneling data and the ground dispatch center. This significantly improves the efficiency of underground geological information processing and system response capabilities, and further enhances the safety and intelligence of mining operations.

[0171] Example 2

[0172] To achieve the above-mentioned invention, embodiments of the present invention also provide a detailed flowchart of a real-time prediction method for geological anomalies in coal mines, such as... Figure 2 As shown, it includes:

[0173] By acquiring, fusing, and intelligently interpreting drilling-while-drilling sensor data with various physical properties in real time, a dynamically updated three-dimensional "transparent" geological model is constructed, enabling high-precision, real-time, and advanced detection and early warning of geological conditions ahead of the tunneling face (or borehole). The specific steps are as follows:

[0174] S101 integrates a sensor array at the front end of the tunneling equipment to simultaneously collect the following data:

[0175] S1011, stress data ( ): A continuous sequence of stress values ​​along the tunneling path. ,in It is the stress measurement value at a certain location (unit: MPa).

[0176] S1012, Gamma-ray data ( ): At each measurement point, collect Each sector (usually) The gamma-ray count (unit: API) of 10 ... .

[0177] S1013, Directional Ground Penetrating Radar Data ( The acquired B-scan two-dimensional profile image is in the form of a single data structure. The matrix, where It is the number of sampling points in the time window (depth). It represents the number of scan channels along a certain direction. The values ​​in the matrix are the amplitudes of the reflected waves.

[0178] S1014, Pose and Odometry Data ( ): including timestamps 3D coordinates and roll angle Pitch angle azimuth .

[0179] S1015, Spatiotemporal Alignment and Calibration: This is the foundation of all data fusion. It unifies all sensor data into a unified geographic coordinate system. For time... Data collected from any sensor Its spatial coordinates The calculation method is as follows:

[0180]

[0181] in, The master device reference point is provided by the inertial navigation system (IMU) and the odometer. The three-dimensional coordinates at any given time. It is the fixed offset vector of the sensor relative to the main device reference point (in the device coordinate system). It is based on The rotation matrix calculated from the attitude angle at a given time is used to transform the offset vector in the device coordinate system to the geographic coordinate system.

[0182] This step yields a series of data points with spatiotemporal labels. Then, a series of preprocessing steps are performed on these data.

[0183] S102, Data Preprocessing. This includes the following sub-steps:

[0184] S1021, Data Normalization: Eliminate the influence of different sensor data dimensions and normalize the data.

[0185] Furthermore, Z-score normalization is used for stress and gamma data to highlight outliers:

[0186]

[0187] in, and These represent the mean and standard deviation of stress and gamma data over a sliding window, respectively.

[0188] Furthermore, for radar image data, its pixels are normalized to the [0,1] interval.

[0189] S1022, Data Voxelization.

[0190] Construct a three-dimensional voxel mesh centered on the current tunneling location. Its resolution is Data points that have undergone spatiotemporal alignment are assigned to the corresponding voxels. For data points falling into the same voxel... Multiple data points are weighted and averaged using a Gaussian kernel function to obtain the comprehensive attribute value of the voxel.

[0191] S103, Model Design. Specifically, the core of this invention is a deep learning model based on an "encoding-fusion-decoding" architecture.

[0192] S1031, Multi-modal Feature Encoder. Specifically, it includes:

[0193] S10311, a one-dimensional sequence encoder (for stress and gamma data).

[0194] Specifically, for the stress sequences and multi-channel gamma sequences acquired along the tunneling path, a standard Transformer encoder was used for feature extraction. This included:

[0195] S103111, Input Embedding. Specifically, it includes:

[0196] S1031111, Patching / Tokenization: This involves segmenting a continuous one-dimensional data sequence (e.g., a sequence of length...) into segments and tokenizing data. ) divided into A length of Each data patch is considered a single word.

[0197] S1031112, Linear Projection: Maps each data segment to a learnable linear projection layer (fully connected layer) The feature space of 3D is used to form word embeddings.

[0198] S1031113, Position Encoding: The Transformer itself lacks sequence order awareness, therefore positional information must be introduced. A positional encoding vector is added to each token embedding. A fixed sine-cosine positional encoding is used:

[0199]

[0200]

[0201] in, It is the position of the word in the sequence. It is the coded dimension index. This is the embedding dimension. The final input vector is the sum of the word embedding and the positional encoding.

[0202] S103112, Transformer encoder module. Specifically includes:

[0203] The input vector sequence is fed into a... A Transformer encoder consisting of stacked identical layers. Each layer contains two core submodules:

[0204] S1031121, Multi-head Self-Attention Mechanism: Allows the model to compute attention scores between different positions in a sequence, thereby capturing dependencies within the sequence (such as the correlation between a stress peak and another stress valley several meters ahead). Its core calculation formula is:

[0205]

[0206] in All are obtained by linear transformation of the input sequence. The multi-head mechanism, on the other hand, divides and projects the input sequence multiple times, and computes attention in parallel to capture information from different representation subspaces.

[0207] S1031122, Feedforward Neural Network: A fully connected network consisting of two linear layers and an activation function (GELU) used to perform nonlinear transformations on the output of the self-attention module.

[0208] Specifically, each submodule is followed by a residual connection and layer normalization. After... After layer encoding, the stress sequence and gamma sequence are encoded into high-dimensional feature sequences. and .

[0209] S1032, a two-dimensional image encoder (ViT), is used to detect radar data.

[0210] Specifically, for two-dimensional ground-penetrating radar B-scan profiles, Vision Transformer (ViT) is used for feature extraction to capture the global geological patterns of the image. This includes:

[0211] S10321, Image Segmentation and Embedding:

[0212] S103211, Image Patching: Patching the input radar image Divide into a series of fixed-size two-dimensional image patches, for example, a size of... Pixels. The number of image patches is

[0213]

[0214] S103212, Flattening and Linear Projection: Flatten each 2D image patch into a 1D vector, and then map it to a linear projection layer. The embedding space of the dimension forms the image patch embedding.

[0215] S103213, [CLS] Label Addition: Add a learnable special classification label [CLS] embedding to the beginning of the embedding sequence. The corresponding vector of this label at the Transformer output will be used as the aggregate representation of the entire image.

[0216] S103214, Location Embedding: Add a learnable one-dimensional location embedding to each image patch embedding (including the [CLS] tag) to preserve the spatial location information of the image patch.

[0217] S10322, Transformer encoder module:

[0218] Specifically, the processed vector sequence is fed into a standard Transformer encoder (composed of MHSA and FFN) with the same structure as the one-dimensional sequence encoder described above. The self-attention mechanism enables the model to associate any two image patches in the image; for example, directly associating the strong reflection arc in the upper left corner with the signal attenuation zone in the lower right corner. This is highly effective for identifying the global features of large geological anomalies (such as the boundary of a collapse column). Finally, the ViT encoder outputs a high-dimensional feature sequence of the radar image. .

[0219] S1033, a multi-layer Transformer decoder architecture that crosses attention mechanisms.

[0220] Specifically, in order to achieve three different modal features Deep fusion. Design a multi-layer Transformer decoder structure based on a cross-attention mechanism (such as...). Figure 3 (As shown).

[0221] S10331, Working principle: Similar to self-attention ( (All from the same source) Different, cross-attention It comes from a mode, and and It comes from another modality. This allows one modality to "query" or "follow" information from another modality.

[0222] S10332, the fusion process. Specifically, it includes:

[0223] S103321, characterized by gamma features (The most representative lithology) is used as the base query.

[0224] S103322, firstly, using Go to query stress characteristics ,Right now

[0225]

[0226]

[0227]

[0228] Calculated attention output This represents "gamma features enhanced by stress characteristics," and the model can learn, for example, "specific stress patterns associated with high gamma value regions (mudstone)."

[0229] S103323, then, the result of the previous step As a new query, query radar features. .Right now

[0230]

[0231]

[0232]

[0233] Output This further integrates radar information.

[0234] S103324, this process can be designed as a multi-layered, multi-directional structure (e.g., using radar to query stress simultaneously), and through residual connections and layer normalization, a unified fusion feature containing all sensor information and processed through deep interaction is finally obtained. .

[0235] S10333, Decoder: Obtain the fused result After obtaining a two-dimensional feature sequence, it needs to be decoded into a three-dimensional geological voxel model. This specifically includes:

[0236] S103331, Feature Reshaping: ... It passes through a linear layer and is reshaped into a low-resolution 3D feature volume. For example, from... Remodeling to .

[0237] S103332, 3D Upsampling: A series of 3D transposed convolutional layers are used to progressively upsample low-resolution features while reducing the number of channels, ultimately restoring the target's 3D geological grid resolution. Each transposed convolutional layer is followed by batch normalization and an activation function (ReLU).

[0238] S10334, Multi-task Output Layer: At the end of the decoder, parallel output heads are set up to predict each voxel. This part is designed as follows:

[0239] S103341, Lithology Classification Header: Using the Softmax activation function, the output is... The probability of each lithology category is calculated, and the probability of belonging to each lithology category is calculated using the Softmax activation function. .

[0240]

[0241] The loss function used is the classification cross-entropy loss.

[0242] S103342, Anomaly Probability Header: Performs probability predictions for three types of geological anomalies (fractures, water-rich areas, and collapse zones). It outputs three independent probability values ​​for each voxel, each obtained through a Sigmoid activation function, ranging from [0,1].

[0243]

[0244] The loss function employs a binary cross-entropy loss for each anomaly.

[0245] S103343, Joint Loss Function: The model is optimized by weighting the loss of each task.

[0246]

[0247] in, It is a hyperparameter used to balance the importance of different tasks.

[0248] Specifically, the outputs of the multi-task prediction head are integrated. In the 3D visualization software, each voxel is assigned a corresponding color and texture based on its lithological classification probability, and a semi-transparent risk cloud map is rendered based on the anomaly probability value (the higher the probability, the darker or less transparent the color). As the equipment advances, new data is continuously input into the model, and the model continuously makes "rolling" predictions for the unknown areas ahead, realizing dynamic and real-time updates to the geological model.

[0249] This invention discloses a real-time prediction method for geological anomalies in coal mines. By constructing a deep fusion and intelligent interpretation system for multi-source sensor data, it effectively solves the core defects of traditional geological exploration methods, such as data isolation, delayed identification, and insufficient analytical accuracy. It achieves intelligent processing throughout the entire process, from data acquisition and feature fusion to 3D geological modeling, significantly improving the accuracy of geological anomaly identification and the system's real-time perception capabilities. This optimizes tunneling efficiency while ensuring safe coal mine production, and enhances the decision-making reliability and engineering applicability of the geological prediction system in complex mining environments.

[0250] Example 3

[0251] To achieve the above invention, this embodiment of the invention also provides a specific implementation process for a real-time prediction method for geological anomalies in coal mines, which specifically includes the following steps:

[0252] S111, Selection of Multifunctional Sensor Probes: The application of this algorithm requires the presence of stress sensing units, gamma-ray sensing units, ground-penetrating radar sensing units, and attitude sensing units during drilling. A simple example is as follows:

[0253] Specifically, the stress sensing unit is a fiber optic sensor that employs distributed Brillouin optical time-domain analyzer technology. The specific operation involves winding an armored sensing optical cable around the probe housing both circumferentially and axially, with a spatial sampling resolution set to 0.25 meters.

[0254] Furthermore, a ring imaging array consisting of 16 high-sensitivity NaI scintillation crystal detectors is incorporated into the gamma-ray sensing unit. Each crystal corresponds to a 22.5° sector, enabling 360° detection of the circumferential strata. The data sampling frequency is set to 2Hz.

[0255] Furthermore, a shielded dual-polarization directional ground-penetrating radar antenna with a center frequency of 100MHz was integrated into the ground-penetrating radar sensing unit. The requirement for this frequency was to balance detection depth (approximately 30-50 meters) and resolution. Then, the radar system was set to trigger 50 times per second, collecting and averaging one B-scan profile image for every 0.1 meters of tunneling; the time window length was set to 1024 sampling points.

[0256] Furthermore, the attitude sensing unit is an advanced sensor that combines a military-grade fiber optic inertial navigation unit (IMU) with a high-precision odometer built into the tunneling machine. It can perform data fusion through a Kalman filter algorithm. Specifically, it outputs the three-dimensional coordinates (X, Y, Z) and attitude angles (roll, pitch, azimuth) of the probe center point at a frequency of 100 Hz, with a positioning accuracy better than 0.1 meters.

[0257] S112, Data Transmission and Computing Platform Setup. Specifically, this includes: Downhole System: All sensor data from within the probe is aggregated via high-bandwidth cables to a mining explosion-proof edge computing server on the tunneling machine. This server is equipped with an NVIDIA Jetson AGX Orin industrial module (or equivalent GPU) responsible for initial data alignment, preprocessing, and model inference. Data is uploaded with low latency via the mining 5G industrial ring network using the MQTT protocol; Surface System: A data center is established at the ground dispatch center, configured with high-performance servers (such as an NVIDIA A100 GPU array) for storing massive amounts of raw data and for model training and iteration.

[0258] S113, System Calibration. This specifically includes: rigorous coordinate system calibration on the ground, and precise measurement of the three-dimensional offset vectors of each sensor relative to the IMU reference point. The relationship between amplitude and rotation is established. The radar system is calibrated in a model with a known dielectric constant to establish an accurate correspondence between amplitude, reflection coefficient, travel time, and depth.

[0259] S114, Training Set Construction. Specifically, this includes: First, selecting an area within the mining area whose geological conditions have been meticulously determined through high-density drilling (10-meter spacing) and 3D seismic exploration as a "digital test field"; further, based on detailed borehole cores, logging curves, and geological reports, manually constructing a high-precision 3D "true-value geological model" with a voxel resolution of 0.25m x 0.25m x 0.25m; within this test field, driving a tunneling machine equipped with this system for no-load test runs, or collecting a large amount of multi-source sensor data while drilling through simulated drilling; then, data annotation. Matching each set of sensor data sequences with spatiotemporal labels to the true-value geological model. Assigning a precise label to each voxel in front of the model: Lithology: 0 (rock), 1 (coal); Fractures: 0 (no), 1 (yes); Water Abundance: 0 (no), 1 (yes); Collapse: 0 (no), 1 (yes). Ultimately, tens of thousands of training sample pairs (input sensor data sequences, output three-dimensional geological labels) are formed.

[0260] S115, Model Training. Specifically, this step employs the full Transformer architecture geological large model (GLM-T1) described in this invention; further, the specific hyperparameters are set as follows: input data segment / image patch size: one-dimensional sequence data segment length P=16; radar image patch size is 16x16 pixels; embedding dimension... Encoder layer number 12 layers; 8 multi-head self-attention heads; 4 layers for cross-modal fusion modules; Optimizer: AdamW optimizer with a learning rate of 1e-4, coupled with a cosine annealing learning rate scheduling strategy; Batch Size: 16; Epochs: 200 epochs of training on an A100 GPU array. Furthermore, minimizing the joint loss function is employed. The training objective was to set the model's lithological classification accuracy to over 96% on an independent validation set and the Intersection over Union (IoU) ratio for identifying anomalous bodies to be greater than 0.85.

[0261] S116, Model Inference and Application. Specifically, this includes: first, system initialization; before the tunneling operation begins, loading the trained GLM-T1 model weight file onto the downhole edge computing server. The operator inputs the starting coordinates in the "Transparent Geology" module of the tunneling machine control interface; furthermore, a real-time "tunneling-detection-modeling" loop is performed, with the following steps:

[0262] Specifically, data acquisition and streaming are performed first. For every 0.5 meters the tunneling machine advances, the system automatically captures sensor data streams from a past 10 meters. Preprocessing is then applied to the edge data, enabling the underground server to complete spatiotemporal alignment, normalization, and rasterization of the corresponding data within 100 milliseconds. This data is then packaged into a 3D tensor conforming to the GLM-T1 model input format. The preprocessed tensor is then fed into the GLM-T1 model for a forward propagation. On Jetson AGX Orin, the latency of a complete inference operation is controlled to within 500 milliseconds. Furthermore, the model outputs a 3D prediction voxel grid for a 50m x 20m x 20m space ahead. Each voxel contains the lithology classification probability and the probability of the presence of three anomalies. Finally, the analyzed results are rendered into a 3D geological model and pushed in real-time to the explosion-proof display screen in the tunneling machine operator's cab and the large screen in the ground control center. The visualization scheme is as follows: coal seams are displayed as bright black entities; rocks are displayed in yellow. Water-rich areas are displayed as overlaid blue semi-transparent cloud maps; the higher the probability, the darker and less transparent the blue. Fault and fracture zones are displayed as red semi-transparent sheet-like or network-like structures.

[0263] This invention presents an application example of a real-time prediction method for geological anomalies in coal mines. By constructing a deep fusion and intelligent interpretation system for multi-source sensor data, it effectively addresses the core shortcomings of traditional geological exploration methods, such as data isolation, delayed identification, and insufficient analytical accuracy. This method achieves intelligent processing throughout the entire process, from data acquisition and feature fusion to 3D geological modeling. It significantly improves the accuracy of geological anomaly identification and the system's real-time perception capabilities, optimizing tunneling efficiency while ensuring safe coal mine production, and enhancing the decision-making reliability and engineering applicability of the geological prediction system in complex mining environments.

[0264] Example 4

[0265] This invention also provides a real-time prediction device 10 for geological anomalies in coal mines, such as... Figure 4 As shown, the device includes:

[0266] The multi-source data acquisition and spatiotemporal alignment module 100 is used to simultaneously acquire stress data, multi-channel gamma-ray data, two-dimensional radar profile data and equipment pose data along the tunneling path by integrating stress sensors, gamma-ray sensors, ground-penetrating radar sensors and pose sensors at the front end of the tunneling equipment to obtain multi-source data. Based on the transformation relationship between the equipment coordinate system and the geographic coordinate system, the multi-source data is spatiotemporally aligned to generate a multi-dimensional data sequence with spatiotemporal labels.

[0267] Specifically, using the formula Calculate the spatial coordinates of the sensor in the geographic coordinate system, where, The coordinates of the reference point for the tunneling equipment. The rotation matrix is ​​calculated from the attitude angles. The sensor is a fixed offset vector relative to a reference point; by fusing military-grade fiber optic inertial navigation unit with Kalman filter data from a high-precision odometer, the three-dimensional coordinates and attitude angle of the probe center point are output at a frequency of 100Hz to control the positioning accuracy to be better than 0.1 meters.

[0268] The multidimensional data normalization and voxel fusion module 200 is used to normalize the multidimensional data sequence and perform three-dimensional voxel rasterization. It uses a Gaussian kernel function to perform weighted fusion of the data within the same voxel to form a voxel grid that includes lithology, stress anomalies and reflection characteristics.

[0269] Specifically, the Z-score normalization formula is applied to the stress data and gamma data. ,in, and The mean and standard deviation are given on the sliding window; a Gaussian kernel function is used to perform a weighted average on the radar image data, and the weights are calculated based on the spatial distribution density of data points within the voxel.

[0270] The cross-modal attention feature fusion module 300 is used to input the voxel grid into a deep learning model based on an encoding-cross-modal attention fusion-decoding architecture, perform cross-modal attention query on stress features through gamma features, and then perform a second cross-modal attention query on radar features using the fused gamma-stress features, and finally generate a unified fusion feature containing the semantic association of the multi-source data.

[0271] Specifically, a multi-head self-attention mechanism is adopted, where each sub-module contains 8 attention heads, which are determined by the formula... Calculate attention scores between different modal features; the decoder contains 4 layers of cross-modal attention modules, each followed by residual connections and layer normalization operations.

[0272] The 3D geological modeling and multi-task prediction module 400 is used to generate a high spatiotemporal resolution 3D geological voxel model based on the unified fusion features through a 3D upsampling network, and to use a multi-task output layer to synchronously predict the lithology category of each voxel and the probability of the existence of three geological anomalies: fractures, water-rich areas, and collapse columns, so as to dynamically update and visualize the geological conditions ahead of the tunneling.

[0273] Specifically, progressive upsampling is performed using a 3D transposed convolutional layer, with the number of channels reduced to half of the previous layer after each upsampling layer; this is achieved through the formula... We perform weighted optimization on the joint loss function for lithological classification, fracture detection, water-rich area identification, and collapse column detection.

[0274] Furthermore, it also includes: an edge computing deployment module, used to fuse military-grade fiber optic inertial navigation unit with Kalman filter data from high-precision odometer, and output the three-dimensional coordinates and attitude angle of the probe center point at a frequency of 100Hz to control the positioning accuracy to be better than 0.1 meters; and a data transmission module, used to transmit tunneling data via the mining 5G industrial ring network using the MQTT protocol to ensure the synchronous update of tunneling data with the ground dispatch center.

[0275] This invention discloses a real-time prediction device for geological anomalies in coal mines. By constructing a collaborative acquisition and intelligent fusion system for multi-source sensor data, it effectively overcomes the technical limitations of isolated data and delayed identification in traditional geological exploration. The device achieves fully automated processing from data acquisition and feature fusion to 3D geological modeling, significantly improving the accuracy of geological anomaly identification and system response speed. While ensuring safe coal mine production, it optimizes tunneling efficiency and enhances the engineering applicability and decision-making reliability of the geological prediction system in complex mining environments.

[0276] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0277] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0278] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

Claims

1. A real-time prediction method of coal mine geological anomaly bodies, characterized in that, include: S1. By integrating stress sensors, gamma-ray sensors, ground-penetrating radar sensors, and pose sensors at the front end of the tunneling equipment, stress data, multi-channel gamma-ray data, two-dimensional radar profile data, and equipment pose data along the tunneling path are collected simultaneously to obtain multi-source data. Based on the transformation relationship between the equipment coordinate system and the geographic coordinate system, the multi-source data is spatiotemporally aligned to generate a multi-dimensional data sequence with spatiotemporal labels. S2, normalize and rasterize the multidimensional data sequence, and use a Gaussian kernel function to weight and fuse the data within the same voxel to form a voxel grid that includes lithology, stress anomaly and reflection characteristics. S2 specifically includes: S21, Z-score normalization formula is used for stress data and gamma data: wherein and are the mean and standard deviation over the sliding window; S22, the radar image data is weighted by a Gaussian kernel function, and the weight is calculated based on the spatial distribution density of data points within the voxel; Three-dimensional voxel rasterization includes: A three-dimensional mesh centered on the current position of the tunneling equipment is constructed, with a spatial resolution consistent with the spatial sampling resolution of the stress sensor. Each voxel is assigned a spatial coordinate, and the sensor data points are mapped to the corresponding voxel positions according to their spatiotemporal labels. S3, inputting the voxel mesh into a deep learning model based on an encoding-cross-modal attention fusion-decoding architecture to generate a high spatiotemporal resolution three-dimensional geological voxel model, including: During the encoding stage, a one-dimensional sequence encoder is used to extract features from stress sequences and multi-channel gamma sequences, and a two-dimensional image encoder is used to extract features from ground-penetrating radar B-scan profiles to capture the global geological patterns of the images. In the feature fusion stage, the gamma feature F g As the query vector Q, the encoded stress feature is calculated F s Cross attention calculation is performed, and the fused gamma-stress feature is used again As the new query vector, the radar feature F r Second cross-modal attention query is performed to finally generate a unified fusion feature containing multi-source data semantic association ; In the decoding stage, the unified fused features are reshaped through a linear layer, transforming them from a two-dimensional feature sequence into a low-resolution three-dimensional feature volume. A multi-layer three-dimensional transposed convolutional structure is then used for progressive upsampling. Each transposed convolutional layer is followed by batch normalization and ReLU activation functions, ultimately outputting a high spatiotemporal resolution three-dimensional geological voxel model. S4 employs a multi-task output layer to synchronously predict the lithology category of each voxel and the probability of the existence of three geological anomalies: fractures, water-rich areas, and collapse columns, so as to dynamically update and visualize the geological conditions ahead of the tunneling.

2. The method as described in claim 1, characterized in that, The S1 further includes: S11, The spatial coordinates of the sensor in the geographic coordinate system are calculated using the following formula: ; in, The coordinates of the reference point for the tunneling equipment. The rotation matrix is ​​calculated from the attitude angles. This is a fixed offset vector of the sensor relative to the reference point; The S12, through the fusion of military-grade fiber optic inertial navigation unit and Kalman filter data from a high-precision odometer, outputs the three-dimensional coordinates and attitude angle of the probe center point at a frequency of 100Hz.

3. The method as described in claim 1, characterized in that, The S3 further includes: S31 employs a multi-head self-attention mechanism, where each sub-module contains 8 attention heads, and the attention scores between different modal features are calculated using a formula: ; in, For querying the matrix, The key matrix, For value matrices, Let be the dimension of the key vector. This is a function that transforms similarity scores into a probability distribution. The S32 decoder contains four layers of cross-modal attention modules, each followed by a residual connection and a layer normalization operation.

4. The method as described in claim 1, characterized in that, The S4 further includes: S41, the joint loss function for lithology classification, fracture detection, water-rich area identification, and collapse column detection is weighted and optimized using the following formula: ; in, These are hyperparameters used to balance different tasks. , , , The importance of.

5. The method as described in claim 1, characterized in that, Also includes: S5 deploys the model on an downhole edge computing server, uses the NVIDIA Jetson AGX Orin module for real-time inference, and keeps the inference latency within 500ms; S6 transmits tunneling data via the MQTT protocol through a mining 5G industrial ring network, ensuring that the tunneling data is updated synchronously with the ground dispatch center.

6. A real-time prediction device for geological anomalies in coal mines, characterized in that, include: The multi-source data acquisition and spatiotemporal alignment module is used to simultaneously acquire stress data, multi-channel gamma-ray data, two-dimensional radar profile data and equipment pose data along the tunneling path by integrating stress sensors, gamma-ray sensors, ground-penetrating radar sensors and pose sensors at the front end of the tunneling equipment to obtain multi-source data. Based on the transformation relationship between the equipment coordinate system and the geographic coordinate system, the module performs spatiotemporal alignment on the multi-source data to generate a multi-dimensional data sequence with spatiotemporal labels. The multidimensional data normalization and voxel fusion module is used to normalize the multidimensional data sequence and perform three-dimensional voxel rasterization. It uses a Gaussian kernel function to perform weighted fusion of the data within the same voxel to form a voxel grid that includes lithology, stress anomaly and reflection characteristics. The multidimensional data normalization and voxel fusion module is specifically used for: The stress and gamma data were normalized using the Z-score formula: in, and The mean and standard deviation over the sliding window; The radar image data is weighted by a Gaussian kernel function, and the weights are calculated based on the spatial distribution density of data points within the voxel. Three-dimensional voxel rasterization includes: A three-dimensional mesh centered on the current position of the tunneling equipment is constructed, with a spatial resolution consistent with the spatial sampling resolution of the stress sensor. Each voxel is assigned a spatial coordinate, and the sensor data points are mapped to the corresponding voxel positions according to their spatiotemporal labels. A cross-modal attention feature fusion module is used to input the voxel mesh into a deep learning model based on an encoding-cross-modal attention fusion-decoding architecture to generate a high spatiotemporal resolution three-dimensional geological voxel model, including: During the encoding stage, a one-dimensional sequence encoder is used to extract features from stress sequences and multi-channel gamma sequences, and a two-dimensional image encoder is used to extract features from ground-penetrating radar B-scan profiles to capture the global geological patterns of the images. In the feature fusion stage, using gamma features F g As the query vector Q, F is calculated across attention for the encoded stress features. s Perform cross-attention calculations and then utilize the fused gamma-stress features. As a new query vector, for radar feature F r A second cross-modal attention query is performed to ultimately generate a unified fusion feature that includes semantic associations between multiple data sources. ; In the decoding stage, the unified fused features are reshaped through a linear layer, transforming them from a two-dimensional feature sequence into a low-resolution three-dimensional feature volume. A multi-layer three-dimensional transposed convolutional structure is then used for progressive upsampling. Each transposed convolutional layer is followed by batch normalization and ReLU activation functions, ultimately outputting a high spatiotemporal resolution three-dimensional geological voxel model. The 3D geological modeling and multi-task prediction module is used to synchronously predict the lithology of each voxel and the probability of the existence of three geological anomalies: fractures, water-rich areas, and collapse columns, using a multi-task output layer, so as to dynamically update and visualize the geological conditions ahead of the tunneling.

7. The apparatus as claimed in claim 6, characterized in that, The multi-source data acquisition and spatiotemporal alignment module is also used for: The spatial coordinates of the sensor in the geographic coordinate system are calculated using the following formula: ; in, The coordinates of the reference point for the tunneling equipment. The rotation matrix is ​​calculated from the attitude angles. This is a fixed offset vector of the sensor relative to the reference point; By fusing military-grade fiber optic inertial navigation unit with Kalman filter data from high-precision odometer, the three-dimensional coordinates and attitude angle of the probe center point are output at a frequency of 100Hz.

8. The apparatus as claimed in claim 6, characterized in that, The cross-modal attention feature fusion module is also used for: A multi-head self-attention mechanism is adopted, in which each sub-module contains 8 attention heads, and the attention scores between different modal features are calculated using a formula: ; in, For querying the matrix, The key matrix, For value matrices, Let be the dimension of the key vector. This is a function that transforms similarity scores into a probability distribution. The decoder contains four layers of cross-modal attention modules, each followed by a residual connection and a layer normalization operation.

9. The apparatus as claimed in claim 6, characterized in that, The three-dimensional geological modeling and multi-task prediction module is also used for: The joint loss function for lithology classification, fracture detection, water-rich area identification, and collapse column detection is weighted and optimized using a formula: ; in, These are hyperparameters used to balance different tasks. , , , The importance of.

10. The apparatus as claimed in claim 6, characterized in that, Also includes: The edge computing deployment module is used to fuse Kalman filter data from a military-grade fiber optic inertial navigation unit and a high-precision odometer to output the three-dimensional coordinates and attitude angle of the probe center point at a frequency of 100Hz, so as to control the positioning accuracy to be better than 0.1 meters. The data transmission module is used to transmit tunneling data via the MQTT protocol through the mining 5G industrial ring network, ensuring that the tunneling data is updated synchronously with the ground dispatch center.

Citation Information

Patent Citations

  • Underground geologic structure imaging system based on azimuth gamma rays

    CN119960069A