A Deep Learning-Based Smart Coal Mine Safety Early Warning Method and System

By using a deep learning-based multi-source data processing and early warning system, the problems of false alarms and missed alarms in traditional coal mine safety monitoring methods have been solved. This system enables efficient fusion of multi-source data and dynamic risk assessment, thereby improving the accuracy and timeliness of coal mine safety early warning.

CN122135529APending Publication Date: 2026-06-02INNER MONGOLIA YIWANG INFORMATION TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INNER MONGOLIA YIWANG INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-03-03
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Traditional coal mine safety monitoring methods rely on single parameter thresholds for judgment, which can easily lead to false alarms or missed alarms. Furthermore, existing machine learning methods are insufficient in terms of multi-source data fusion and dynamic risk assessment, making it difficult to effectively improve the accuracy and timeliness of coal mine safety early warnings.

Method used

A deep learning-based intelligent coal mine safety early warning system is adopted. Multimodal input tensors are generated by temporal alignment and spatial mapping of multi-source heterogeneous sensor data. Risk prediction is performed using a spatiotemporal graph convolution-attention fusion network. Preliminary feature extraction is performed at the edge computing gateway and complex model inference is performed on the ground server cluster to generate safety early warning instructions.

Benefits of technology

It significantly improves the accuracy and foresight of coal mine disaster early warning, enabling warnings to be issued 10 to 30 minutes in advance, avoiding the lag and false alarm/missed reporting problems of traditional methods, and forming a closed-loop response mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135529A_ABST
    Figure CN122135529A_ABST
Patent Text Reader

Abstract

This invention discloses a deep learning-based intelligent coal mine safety early warning method and system. It includes collecting multi-source heterogeneous sensor data, performing spatiotemporal alignment and mapping to generate multimodal input tensors, inputting these tensors into a spatiotemporal graph convolutional-attention fusion network for risk evolution prediction, and generating multi-level early warning instructions based on risk levels. This application can achieve deep fusion of multi-physics coupling features, providing accurate early warnings 10 to 30 minutes in advance, significantly improving the accuracy, foresight, and robustness of coal mine disaster early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent mining and safe production technology, and in particular to a method and system for intelligent coal mine safety early warning based on deep learning. Background Technology

[0002] Coal mine safety has always been a key focus of the energy industry. With the development of intelligent technologies, more and more monitoring and early warning methods are being introduced into mine safety management. Traditional coal mine safety monitoring mainly relies on sensor networks to collect environmental parameters such as gas concentration, temperature, humidity, and wind speed in real time, and then trigger alarms based on threshold judgment mechanisms. This method can reflect abnormal changes in the mine environment to a certain extent, but in the complex and ever-changing underground environment, the judgment of a single parameter threshold is easily interfered with, leading to false alarms or missed alarms. In addition, mine disasters are often caused by the coupling of multiple factors, and early warning models that rely solely on rules are insufficient to fully explore the potential correlations between multi-source heterogeneous data.

[0003] In recent years, some studies have attempted to apply machine learning methods to coal mine safety risk assessment, improving early warning accuracy by constructing classification or regression models. However, these methods typically rely on manual feature extraction, requiring sophisticated data preprocessing and feature engineering. Furthermore, their generalization ability is limited when faced with high-dimensional, nonlinear, and time-series mine monitoring data. Deep learning technology, with its powerful automatic feature extraction capabilities and ability to model complex patterns, has achieved significant results in fields such as image recognition and natural language processing. However, its application in coal mine safety early warning is still in the exploratory stage, particularly lacking systematic solutions in areas such as multi-source data fusion, dynamic risk evolution modeling, and real-time early warning response. Therefore, there is an urgent need for a smart coal mine safety early warning method and system that integrates deep learning technology to improve the accuracy and timeliness of mine disaster prediction. Summary of the Invention

[0004] The purpose of this invention is to provide a smart coal mine safety early warning method and system based on deep learning, which solves the problems mentioned in the background technology.

[0005] This invention is implemented as follows: a deep learning-based intelligent coal mine safety early warning method and system, comprising: collecting multi-source heterogeneous sensor data from multiple monitoring areas underground in a coal mine, wherein the multi-source heterogeneous sensor data includes gas concentration, carbon monoxide concentration, temperature, humidity, wind speed, roof displacement, microseismic signals, and video image sequences, and the monitoring area is a physical area in the mine where sensor nodes are deployed; performing time alignment and spatial mapping processing on the multi-source heterogeneous sensor data to generate a multimodal input tensor with unified spatiotemporal coordinates, wherein the time alignment is completed through interpolation and a sliding window mechanism, and the spatial mapping establishes a correlation between each sensor data and its corresponding roadway coordinates or mining face position; and inputting the multimodal input tensor into a pre-trained spatiotemporal graph convolutional-attention fusion network, wherein the spatiotemporal graph convolutional-attention fusion network includes a graph... The system comprises a graph construction module, a spatiotemporal convolutional encoder, a cross-modal attention fusion layer, and a risk evolution prediction head. The graph construction module constructs a dynamic adjacency matrix based on the tunnel topology and the physical connections between monitoring areas. The spatiotemporal convolutional encoder extracts local temporal features from the time-series data of each monitoring area. The cross-modal attention fusion layer calculates the correlation weights between different modalities of data and performs weighted fusion. The risk evolution prediction head outputs a disaster risk level sequence within a preset future time window. Based on the comparison between the disaster risk level sequence and preset multi-level threshold intervals, corresponding safety warning instructions are generated. The safety warning instructions include four categories: normal operation, enhanced inspection, local production restriction, and emergency evacuation. The safety warning instructions are then sent to the underground personnel positioning terminal and the ground dispatch center via an intrinsically safe communication link.

[0006] Secondly, this invention provides a deep learning-based intelligent coal mine safety early warning system, comprising a multi-source sensor node array, an edge computing gateway, an underground fiber optic ring network, a ground server cluster, and an early warning execution terminal. The multi-source sensor node array is deployed at mine roadway intersections, coal faces, return airways, and key support points. Each sensor node integrates a gas detection module, a temperature and humidity sensor, a MEMS microseismometer, a laser displacement meter, and an explosion-proof high-definition camera. Each module is connected to a local embedded controller via a CAN bus. The embedded controller is equipped with a timestamp synchronization unit to receive 1PPS pulse signals from the Beidou underground time synchronization module to achieve microsecond-level time alignment. The edge computing gateway connects to multiple sensor nodes via an RS485 interface and is installed in an explosion-proof cabinet in a substation or central pump room. The edge computing gateway incorporates a lightweight spatiotemporal feature extraction model to process raw sensor data. Preliminary noise reduction, normalization, and feature compression are performed, and the signal is connected to an underground fiber optic ring network via a gigabit Ethernet interface. The underground fiber optic ring network adopts a dual-ring redundant topology, is laid along the main haulage roadway, and is connected to a ground server cluster via a mine-use explosion-proof optical transceiver. The ground server cluster includes a data preprocessing server, a model inference server, and an early warning decision server. The data preprocessing server receives data streams uploaded from the edge computing gateway and performs spatial coordinate mapping and missing value imputation. The model inference server loads the spatiotemporal graph convolutional-attention fusion network and performs forward propagation calculations on the preprocessed multimodal input tensors. The early warning decision server parses the risk level sequence output by the model, matches it with the corresponding early warning strategy rule base, and generates structured safety early warning instructions. The early warning execution terminal includes an intrinsically safe mine-use handheld terminal, an audible and visual alarm column, and a dispatch screen. The handheld terminal is connected via Wi-Fi 6. The Mesh network receives early warning commands and vibrates to alert the user. The sound and light alarm column is installed at the entrance of the tunnel. It switches between red, yellow and green LED lights and buzzer frequency according to the warning level. The dispatch screen displays a mine-wide risk heat map and highlighted affected areas in real time at the ground control center.

[0007] Thirdly, the present invention provides a coal mine safety early warning device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the deep learning-based intelligent coal mine safety early warning method described in the first aspect.

[0008] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the deep learning-based intelligent coal mine safety early warning method described in the first aspect.

[0009] Fifthly, the present invention provides a computer program product that, when run on a coal mine safety early warning device, causes the coal mine safety early warning device to execute the deep learning-based intelligent coal mine safety early warning method described in the first aspect.

[0010] The advantages of this invention compared to existing technologies are as follows: In this invention, multi-source heterogeneous sensor data are mapped to spatial coordinates through timestamp synchronization units to form a multimodal input tensor with a unified spatiotemporal reference, avoiding feature misalignment problems caused by inconsistent sampling frequencies or missing location information in traditional methods; the spatiotemporal graph convolution-attention fusion network utilizes the physical connection relationship of the tunnel to construct a dynamic graph structure, retaining spatial proximity constraints during graph convolution, and automatically learning the nonlinear coupling relationship between gas inrush and roof displacement and microseismic activity through a cross-modal attention mechanism, without the need for manually setting rules or feature combinations; the risk evolution prediction head adopts a sequence-to-sequence architecture, outputting risk levels for multiple future time steps, supporting early warnings 10 to 30 minutes in advance, overcoming the lag defect of existing threshold triggering mechanisms that can only respond to anomalies that have already occurred; the edge computing gateway completes preliminary feature extraction underground, reducing uplink data bandwidth usage, while the ground server cluster centrally executes complex model inference, balancing real-time performance and computational accuracy; the early warning execution terminal ensures effective transmission of instructions through multi-channel feedback (handheld terminal vibration, audible and visual alarms, and large-screen visualization of the dispatching system), forming a closed-loop response mechanism. This method overcomes the limitations of relying on single-parameter threshold judgments or shallow machine learning models. By deeply fusing spatiotemporal context and multi-physics coupling features, it significantly improves the accuracy, foresight, and robustness of coal mine disaster early warning. (See attached figures.) Figure 1 This is a schematic diagram of the overall architecture of a deep learning-based intelligent coal mine safety early warning system.

[0011] Reference numerals in the attached diagram: 1. Multi-source sensor node array; 2. Edge computing gateway; 3. Downhole fiber optic ring network; 4. Ground server cluster; 5. Early warning execution terminal. Detailed Implementation

[0012] In practical deployment, the deep learning-based intelligent coal mine safety early warning method and system of this invention first deploys a multi-source sensor node array at mine roadway intersections, coal mining faces, return airways, and key support points. Each sensor node integrates a gas detection module, a temperature and humidity sensor, a MEMS microseismometer, a laser displacement meter, and an explosion-proof high-definition camera. All the aforementioned sensor modules are connected to a local embedded controller via a CAN bus. The embedded controller is equipped with a timestamp synchronization unit, which is connected to the Beidou underground time synchronization module via a hardware interface, receiving its output 1PPS pulse signal to achieve alignment of all sensor data at the microsecond level of time accuracy. The gas detection module is used to collect real-time methane and carbon monoxide concentrations; the temperature and humidity sensor is used to acquire ambient temperature and relative humidity; the MEMS microseismometer is used to record microseismic signals generated by rock fractures or fault slippage; the laser displacement meter is used to monitor the displacement changes of the roof or roadway sides; and the explosion-proof high-definition camera continuously captures video image sequences at a frame rate of 25 frames per second. All sensor data is timestamped in the local embedded controller and packaged according to a preset data encapsulation format.

[0013] Multiple sensor nodes establish physical connections with the edge computing gateway via RS485 communication interfaces. The edge computing gateway is installed in an explosion-proof cabinet in the mine substation or central pump room. It is equipped with an ARM Cortex-A72 quad-core processor, 8GB LPDDR4 memory, and 128GB eMMC storage media, and runs a lightweight Linux operating system. The edge computing gateway has a built-in lightweight spatiotemporal feature extraction model, which uses MobileNetV3 as the backbone network and performs pruning and quantization processing for the characteristics of underground data. When the edge computing gateway receives the raw data packets uploaded by the sensor nodes, it first performs data integrity verification to remove abnormal frames caused by communication interference; then it performs sliding window mid-value filtering on numerical data such as gas concentration, temperature and humidity, and displacement to suppress high-frequency noise, and performs YUV color space conversion and histogram equalization on video image sequences; then it normalizes all modal data to the [0,1] interval and compresses high-dimensional features to a fixed dimension through principal component analysis (PCA); finally, it uploads the processed feature vectors to the underground fiber optic ring network through a gigabit Ethernet interface. The RS485 link between the edge computing gateway and the sensor node adopts the Modbus RTU protocol, with a baud rate of 115200bps, 8 data bits, no parity bit, and 1 stop bit, to ensure communication stability during long-distance transmission.

[0014] The underground fiber optic ring network is laid along the top of the main haulage roadway, employing a dual-ring redundant topology, consisting of two independent single-mode optical fibers forming the main and backup loops. Every 500 meters along the fiber optic ring network, a mine-use explosion-proof optical transceiver is installed. These transceivers connect to the gigabit Ethernet ports of the edge computing gateways via SC / APC interfaces, converting electrical signals to optical signals. All data streams from the edge computing gateways converge into the fiber optic ring network via the optical transceivers and are transmitted bidirectionally along the loop until they reach the surface exit at the main and auxiliary shafts. At the surface exit, another set of mine-use explosion-proof optical transceivers converts the optical signals back to electrical signals and connects them to the surface server cluster via a 10 Gigabit Ethernet switch. The surface server cluster is deployed in the mine dispatch center's computer room and includes three types of dedicated servers: a data preprocessing server, a model inference server, and an early warning decision server. All three types of servers utilize dual Intel Xeon Silver 4310 processors, 128GB DDR4 ECC memory, and 2TB NVMe solid-state drives, interconnected via an InfiniBand network with a bandwidth of 100Gbps.

[0015] After receiving the multimodal data stream from the underground fiber optic ring network, the data preprocessing server first parses the spatial metadata field in each data packet. This field contains the sensor node ID and its corresponding roadway coordinates (X, Y, Z). The coordinate system is based on the main shaft opening as the origin, with the X-axis pointing east, the Y-axis pointing north, and the Z-axis pointing vertically downwards. Based on the pre-stored 3D digital twin model of the mine, the server maps the spatial location of each sensor node to a specific node in the roadway topology map. For missing sensor data (e.g., no video frames for a certain period due to equipment failure), the server uses cubic spline interpolation to complete the data in the time dimension and performs spatial interpolation correction by combining similar data from neighboring nodes. After completing the spatiotemporal alignment, the server organizes the data from all monitored areas within the same time window (e.g., 5 minutes) into a four-dimensional tensor with dimensions [N, T, M, F], where N is the number of monitored areas, T is the number of time steps, M is the number of modal types (M=8, corresponding to gas, CO, temperature, humidity, wind speed, roof displacement, microseismic activity, and video), and F is the feature dimension of each mode. This tensor is fed into the model inference server as a multimodal input tensor.

[0016] The model inference server loads a pre-trained spatiotemporal graph convolutional-attention fusion network. The network's graph construction module first reads a mine roadway topology database, which stores the physical connections between monitoring areas in the form of an adjacency list. For example, there is a direct connection between the coal face and the return airway, while there is no direct connection between two parallel tunneling roadways. Based on this, the graph construction module dynamically generates an adjacency matrix A, where an element A_ij=1 indicates that region i and region j are structurally adjacent, otherwise it is 0. Simultaneously, this module adjusts the adjacency matrix with weights based on the current airflow direction and wind speed data, making the information propagation more consistent with the actual gas diffusion path. The spatiotemporal convolutional encoder consists of three stacked ST-Conv blocks, each containing a graph convolutional layer and a temporal convolutional layer. The graph convolutional layer is implemented using a Chebyshev polynomial approximation with order K=3, used to aggregate spatial features of neighboring regions; the temporal convolutional layer uses causal dilated convolution with dilation rates of 1, 2, and 4, and a kernel size of 3, to capture dynamic patterns at different time scales. After passing through the spatiotemporal convolutional encoder, each monitoring region outputs a 128-dimensional spatiotemporal feature vector at each time step.

[0017] The cross-modal attention fusion layer receives spatiotemporal feature sequences from eight modalities. First, it maps each modality's features to a unified 64-dimensional query, key, and value space using linear projection. For any two modalities m and n, at time step t and region i, the attention score S_{m,n}(i,t) = softmax(Q_m(i,t)·K_n(i,t)^T / √d), where d=64, is calculated. This score reflects the degree of attention modality m pays to modality n at a specific spatiotemporal location. Subsequently, the value vectors of all modalities are weighted and summed according to the attention weights to obtain the fused feature vector. This process is executed in parallel across the region, time, and modality dimensions, ultimately outputting a feature tensor that integrates multimodal contextual information. The risk evolution prediction head consists of two LSTM layers and a fully connected classification layer. The LSTM has 256 hidden units, and the fully connected layer outputs the risk level probability distribution for the next 12 time steps (corresponding to the next 30 minutes, with each step being 2.5 minutes). The risk levels are divided into four levels: Level 0 (normal operation), Level 1 (enhanced inspection), Level 2 (partial production restriction), and Level 3 (emergency evacuation).

[0018] After receiving the risk level sequence output by the model inference server, the early warning decision server compares it step-by-step with preset multi-level threshold intervals. The threshold intervals are defined as follows: if the predicted level at any future time step is ≥3, a Level 3 early warning is immediately triggered; if the predicted level at three consecutive time steps is ≥2, a Level 2 early warning is triggered; if the predicted level at a single time step is 1 and the previous time period was Level 0, a Level 1 early warning is triggered. The early warning strategy rule base is stored in a PostgreSQL database. Each rule includes the early warning level, a list of affected areas, a response action code, and the target of the instruction. The server generates structured security early warning instructions based on the matching rules. The instruction format is a JSON object, containing an instruction ID, timestamp, early warning level, area coordinate range, suggested measures text, and a list of target terminals.

[0019] Safety warning commands are distributed to various warning execution terminals via an intrinsically safe communication link used in mining. The intrinsically safe handheld terminal is an Android device with an Ex ib 1 Mb explosion-proof rating, equipped with a Wi-Fi 6 Mesh communication module, operating on the 2.4GHz frequency band, and supporting OFDMA and MU-MIMO technologies. The handheld terminal periodically scans surrounding Mesh nodes and automatically joins the self-organizing network composed of APs within the tunnel. Upon receiving a warning command, the terminal CPU parses the JSON content. If the warning level is ≥2, a vibration motor is activated to vibrate continuously for 5 seconds, and a red warning box pops up on the lock screen; if it is a level 1 warning, only a yellow warning is displayed in the notification bar. Audible and visual alarm posts are installed at the entrances of each main tunnel, integrating three-color LED lights (red, yellow, and green) and a piezoelectric buzzer. The alarm posts are connected to the nearest edge computing gateway via an RS485 bus, receiving commands using a custom protocol. When a Level 3 warning is received, the red LED flashes at a frequency of 2Hz, and the buzzer emits a continuous 120dB, 1kHz sound; for a Level 2 warning, the yellow LED remains constantly lit, and the buzzer sounds intermittently at 1Hz; for a Level 1 warning, the green LED flashes slowly, with no sound alert. The dispatch screen is a 65-inch 4K LCD display deployed on the ground dispatch center console, running a customized web application. This application maintains a long-term connection with the warning decision server via WebSocket, receiving real-time risk heat map data from the entire mine. The heat map uses a two-dimensional plan view of the mine as its base map, employing a red-orange-yellow-green gradient to represent the current and future risk levels of each area. Affected areas are displayed as overlaid semi-transparent, high-brightness blocks, with the predicted peak time and level value marked.

[0020] Throughout the system's operation, a multi-source sensor node array continuously collects raw data, an edge computing gateway performs local feature extraction and compression, an underground fiber optic ring network ensures highly reliable data transmission, a ground server cluster completes complex model inference and decision generation, and an early warning execution terminal transmits multi-channel commands. All components are tightly coupled through standardized interfaces and protocols, forming a complete closed loop from sensing, transmission, computation to response. The system completes a full early warning cycle every 5 minutes, including data acquisition, preprocessing, model inference, decision generation, and command issuance. All equipment complies with the "Coal Mine Safety Regulations" and GB 3836 series explosion-proof standards, and the communication link meets the MT / T 1117-2011 mining Ethernet technology requirements, ensuring long-term stable operation in the high-humidity, high-dust, and strong electromagnetic interference underground environment. To better enable those skilled in the art to fully understand and implement this invention, the specific implementation principle of this invention is further supplemented below with a specific application scenario.

[0021] In the actual deployment of a high-gas outburst mine, the first sensing node was deployed at the intersection of the return airway 30 meters behind the coal face. This node integrates an infrared gas sensor, an electrochemical carbon monoxide detection module, an SHT35 temperature and humidity chip, a MEMS triaxial microseismometer, a laser triangulation sensor, and a 1080P explosion-proof camera. Each sensing unit is connected to a local embedded controller via a CAN2.0B bus. The controller has an embedded STM32H743 chip, whose TIM1 advanced timer channel is connected to the 1PPS signal pin output by the Beidou underground time synchronization module. The rising edge of each second triggers an interrupt service routine, aligning the system's local time with the Beidou standard time to within ±1μs, thereby ensuring strict synchronization between the gas concentration sampling time, the microseismic event occurrence time, and the video frame capture time. When the coal mining machine cuts the coal face, causing local rock mass fracturing, the MEMS microseismometer detects a vibration signal with a main frequency of 800Hz at time t0. At the same time, the laser displacement meter records the roof subsidence rate as 2.3mm / min at the same time stamp, while the gas sensor detects that the concentration suddenly rises from 0.3% to 0.8% at t0+1.2 seconds. All of the above multi-source data are stamped with the same timestamp and packaged into a unified data package.

[0022] The data packet was uploaded via an RS485 link using the Modbus RTU protocol to an edge computing gateway located in the explosion-proof cabinet of the central pump room. The gateway's built-in lightweight MobileNetV3 model performed YUV422 to YUV420 conversion on the video stream and then histogram equalization to enhance the texture features of coal dust dispersion under low illumination. At the same time, it applied a sliding median filter with a window length of 9 to the numerical sequences of gas, CO, and displacement to remove instantaneous jump values ​​caused by electromagnetic interference. Subsequently, all modal data were normalized to the [0,1] interval, and the original 128-dimensional video features and 8-dimensional numerical features were compressed into a 32-dimensional joint vector using PCA, which was then sent to the underground fiber optic ring network via gigabit Ethernet. Because the fiber optic ring network adopts a dual-ring redundancy structure, even if the main ring is interrupted due to roof collapse in the middle of the transport roadway, the data can still be transmitted in reverse through the backup ring, ensuring uninterrupted communication.

[0023] After receiving the data from this area, the ground data preprocessing server parses its spatial metadata field "NodeID=WF-07, X=1250, Y=-320, Z=480" and calls the mine's 3D digital twin model to confirm that the node is located at the 7th monitoring point on the return air side of the coal face, with the adjacent area including the mining face (WF-06) and the return air incline (RF-01). When it is found that the node is missing video frames in the t0+5 minute period (because the camera lens is covered by coal slurry), the server first uses cubic spline interpolation to complete the video feature trend from t0+4 to t0+6 minutes in the time dimension, and then corrects the missing values ​​by distance-weighted average based on the spatial correlation of the video features of the neighboring node WF-06 during the same period. Finally, a 5-minute window (T=120 steps, one step every 2.5 seconds) four-dimensional tensor input model inference server is constructed.

[0024] The spatiotemporal graph convolutional-attention fusion network loaded by the model inference server first reads the alleyway topology database to generate an initial adjacency matrix A, where A[WF-07][WF-06]=1, A[WF-07][RF-01]=1, and the rest are 0. Subsequently, based on real-time anemometer data (airflow from WF-06 to WF-07 and then to RF-01), the weight of A[WF-07][RF-01] is increased to 1.2, and A[WF-07][WF-06] is decreased to 0.8 to reflect the gas diffusion direction. In the spatiotemporal convolutional encoder, the Chebyshev graph convolutional layer (K=3) aggregates the gas and microseismic features of the three regions WF-06, WF-07, and RF-01, and the causal dilatational convolutional layers (dilatation rates 1 / 2 / 4) capture the concentration increase trends at 2.5-second, 5-second, and 10-second scales, respectively. Cross-modal attention layer calculations show that at t0+2 minutes and in region WF-07, the attention score of the microseismic mode to the gas mode reaches 0.78, indicating that the system automatically identifies a strong correlation between rock mass fracturing and gas release. The fused features are then used by the LSTM prediction head to output the risk probability for the next 12 steps, with P(level=3)=0.92 for the 8th step (i.e., the next 20 minutes).

[0025] The early warning decision server compares the threshold rules and immediately triggers a Level 3 early warning if the predicted level is ≥3 at any time step. The system retrieves the policy ID=ALM-301 from the PostgreSQL rule base and generates a JSON instruction containing "regional coordinate range: X∈[1200,1300], Y∈[-350,-300], Z∈[470,490]" and "target terminals: handheld terminal group G-WF, ALM-07 sound and light column, and dispatch screen". This instruction is sent to the handheld terminal of the sampling inspector via an intrinsically safe Wi-Fi 6 Mesh network. After the terminal CPU parses the instruction, it starts the vibration motor to vibrate continuously for 5 seconds and pops up a red box on the lock screen. At the same time, the instruction is transmitted to the ALM-07 sound and light column via RS485 bus, driving the red LED to flash at 2Hz and triggering a buzzer with a sound pressure level of 120dB. The dispatch screen web application receives the heat map data via WebSocket, overlays a semi-transparent red highlight block on the corresponding coordinate area of ​​the base map, and marks it "peak level 3, expected in 20 minutes".

[0026] Throughout the process, microsecond-level time synchronization ensures that the causal relationships of multiphysics events can be accurately modeled, the dynamic weighted adjacency matrix makes graph convolution propagation conform to the real airflow path, the cross-modal attention mechanism automatically focuses on the key coupling features of disaster precursors, and the edge-cloud collaborative architecture reduces downhole bandwidth pressure while ensuring the inference capability of complex models. Ultimately, it achieves accurate early warning of gas outburst risk 20 minutes in advance, avoiding the shortcomings of traditional threshold methods that miss complex disasters due to monitoring only a single parameter.

[0027] All contents not described in detail in the specification are existing technologies known to those skilled in the art, and the model parameters of each electrical appliance are not specifically limited; conventional equipment can be used. Electrical control components not mentioned in this technical solution are not shown in the figures because they are existing technologies, and will not be described here.

[0028] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method and system for intelligent coal mine safety early warning based on deep learning, characterized in that, include: Multi-source heterogeneous sensor data from multiple monitoring areas in a coal mine are collected. This data includes methane concentration, carbon monoxide concentration, temperature, humidity, wind speed, roof displacement, microseismic signals, and video image sequences. The monitoring areas are physical regions within the mine where sensor nodes are deployed. The multi-source heterogeneous sensor data undergoes time alignment and spatial mapping to generate a multimodal input tensor with unified spatiotemporal coordinates. Time alignment is achieved through interpolation and a sliding window mechanism, while spatial mapping establishes a correlation between each sensor data point and its corresponding roadway coordinates or mining face location. The multimodal input tensor is then input into a pre-defined system. The trained spatiotemporal graph convolutional-attention fusion network includes a graph construction module, a spatiotemporal convolutional encoder, a cross-modal attention fusion layer, and a risk evolution prediction head. The graph construction module constructs a dynamic adjacency matrix based on the roadway topology and the physical connection relationships between monitoring areas. The spatiotemporal convolutional encoder extracts local temporal features from the time-series data of each monitoring area. The cross-modal attention fusion layer calculates the correlation weights between different modalities of data and performs weighted fusion. The risk evolution prediction head outputs a disaster risk level sequence within a preset future time window. The disaster risk level sequence is compared with the preset multi-level threshold range to generate corresponding safety warning instructions. The safety warning instructions include four categories: normal operation, enhanced inspection, local production restriction, and emergency evacuation. The safety warning instructions are sent to the underground personnel positioning terminal and the ground dispatch center through the intrinsically safe communication link for mining.

2. The intelligent coal mine safety early warning method based on deep learning as described in claim 1, characterized in that, The time alignment is achieved by receiving a 1PPS pulse signal output by the Beidou downhole timing module to achieve microsecond-level time synchronization. The spatial mapping is based on a coordinate system with the main wellhead as the origin, the X-axis pointing east, the Y-axis pointing north, and the Z-axis pointing vertically downward, which maps the sensor node ID to the specific node position in the roadway topology map.

3. The intelligent coal mine safety early warning method based on deep learning as described in claim 1, characterized in that, In the spatiotemporal graph convolutional-attention fusion network, the graph construction module generates an adjacency matrix based on the mine roadway topology database and adjusts the adjacency matrix by weighting it in conjunction with the current airflow direction and wind speed data; the spatiotemporal convolutional encoder is composed of stacked graph convolutional layers with Chebyshev multinomial approximation and temporal convolutional layers with causal dilation; the cross-modal attention fusion layer projects each modality feature onto a query, key, and value space of a unified dimension and calculates the inter-modal attention score to weightedly fuse multimodal features; the risk evolution prediction head consists of two LSTM layers and a fully connected classification layer, outputting the risk level probability distribution for the next 12 time steps, with each time step corresponding to 2.5 minutes.

4. The intelligent coal mine safety early warning method based on deep learning as described in claim 1, characterized in that, The comparison rules for the multi-level threshold intervals include: if the predicted level is 3 at any future time step, an emergency evacuation order is triggered; if the predicted level is 2 for three consecutive time steps, a local production restriction order is triggered; if the predicted level is 1 at a single time step and the predicted level for the previous period is 0, an enhanced inspection order is triggered.

5. A deep learning-based intelligent coal mine safety early warning system, characterized in that, It includes a multi-source sensor node array (1), an edge computing gateway (2), an underground fiber optic ring network (3), a ground server cluster (4), and an early warning execution terminal (5); the multi-source sensor node array (1) is deployed at the intersection of mine roadways, coal mining faces, return airways, and key support points. Each sensor node integrates a gas detection module, a temperature and humidity sensor, a MEMS micro-vibration meter, a laser displacement meter, and an explosion-proof high-definition camera. Each module is connected to a local embedded controller via a CAN bus. The embedded controller is equipped with a timestamp synchronization unit to receive 1PPS pulse signals from the Beidou underground time synchronization module. The edge computing gateway (2) is connected to multiple sensor nodes via an RS485 interface and installed in an explosion-proof cabinet in a substation or central pump room. It has a built-in lightweight spatiotemporal feature extraction model to reduce noise, normalize, and compress features of the original sensor data. It is connected to the underground fiber optic ring network (3) via a gigabit Ethernet interface. The underground fiber optic ring network (3) adopts a dual-ring redundant topology structure and is laid along the main transport roadway. It is connected to the ground server cluster (4) via a mine explosion-proof optical transceiver. The ground server cluster (4) includes a data preprocessing server, a model inference server, and an early warning decision server. The data preprocessing server performs spatial coordinate mapping and missing value imputation. The model inference server loads a spatiotemporal graph convolution-attention fusion network for forward propagation calculation. The early warning decision server matches the early warning strategy rule base to generate structured safety early warning instructions. The early warning execution terminal (5) includes an intrinsically safe handheld terminal, an audible and visual alarm column, and a dispatch screen. The handheld terminal is connected via Wi-Fi 6. The Mesh network receives early warning commands and vibrates to alert the user. The sound and light alarm pillars are installed at the entrance of the tunnel and switch between red, yellow and green LED lights and buzzer frequency according to the warning level. The dispatch screen displays a mine-wide risk heat map and highlighted affected areas in real time at the ground control center.

6. The intelligent coal mine safety early warning system based on deep learning as described in claim 5, characterized in that, The edge computing gateway (2) runs a lightweight Linux operating system, equipped with an ARM Cortex-A72 quad-core processor, 8GB LPDDR4 memory and 128GB eMMC storage medium. Its built-in lightweight spatiotemporal feature extraction model is based on the MobileNetV3 backbone network and has been pruned and quantized. The RS485 interface adopts the Modbus RTU protocol with a baud rate of 115200 bps, 8 data bits, no parity bit, and 1 stop bit.

7. The intelligent coal mine safety early warning system based on deep learning as described in claim 5, characterized in that, The data preprocessing server, model inference server, and early warning decision server in the ground server cluster (4) all use dual Intel Xeon Silver 4310 processors, 128GB DDR4 ECC memory, and 2TB NVMe solid-state drives, and are interconnected through an InfiniBand network with a bandwidth of 100Gbps; the early warning strategy rule base is stored in a PostgreSQL database, and each rule includes the early warning level, a list of affected areas, a response action code, and the target of the instruction.