A coal and gas outburst risk early intelligent grading early warning method under sparse samples

By employing a text-data dual-channel adaptive multimodal learning model that combines large language models with task transfer learning, the problem of high-precision early warning of coal and gas outburst risks under sparse samples was solved. This model enables advanced intelligent hierarchical early warning and multi-indicator data fusion analysis of coal and gas outburst risks, thereby improving the accuracy and generalization ability of early warnings.

CN122114595APending Publication Date: 2026-05-29CHINA UNIV OF MINING & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNIV OF MINING & TECH
Filing Date
2025-12-30
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Under sparse sample conditions, existing technologies are unable to achieve high-precision early warning and strong generalization capabilities for coal and gas outburst risks. Especially when there are significant differences in geological conditions and gas occurrence conditions among different coal mines, existing methods suffer from missed and false alarms in coal and gas outburst risk monitoring, and their ability to integrate and analyze multiple indicators is insufficient.

Method used

A text-data dual-channel adaptive multimodal learning model combining a large language model and task transfer learning is adopted. By preprocessing gas concentration, electromagnetic radiation and acoustic emission signals, a TimeAFFN model is established for risk warning. Cross-modal attention and feature fusion networks are used for the fusion analysis and prediction of multi-indicator data.

Benefits of technology

It achieves high-precision early warning of coal and gas outburst risk under sparse sample conditions, improves the accuracy and generalization ability of early warning, and enables advanced graded early warning to ensure accurate identification of coal and gas outburst disaster risk.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114595A_ABST
    Figure CN122114595A_ABST
Patent Text Reader

Abstract

The application discloses a kind of coal and gas outburst risk early intelligent grading early warning method under sparse sample, collect time series data in coal mine working face and coal seam geological parameters as initial data set and carry out pre-processing, data-data large language model double channel is carried out data fusion, and multi-index collaborative prediction model TimeAFFN is established, coal and gas outburst multi-disaster risk fusion early warning model based on task transfer learning TimeAFFN model is established, the original data of each index and prediction data are input into fusion early warning model, and the risk probability of each disaster current time point and future time period is obtained synchronously, and comprehensive risk identification early warning is carried out by coal and gas outburst grading identification early warning method.The application can realize high-precision early warning and strong generalization ability of large language model under small sample condition, and the risk probability of current and future coal and gas outburst disaster risk is graded early warning by fusion early warning model, which better guarantees the accurate identification of coal and gas outburst disaster risk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for early warning of coal and gas outburst disasters in coal mines, and more particularly to an advanced intelligent classification and early warning method for coal and gas outburst risk under sparse samples. Background Technology

[0002] Coal remains a vital component of the global energy mix. With the gradual depletion of shallow coal reserves, deeper mining has become an inevitable trend. Increased operating depth leads to more complex geological conditions, and higher ground stress and gas pressure exacerbate the risk of coal and gas outbursts, potentially causing severe casualties and property damage. Reducing the impact of abnormal data, comprehensively analyzing the response characteristics of different monitoring indicators before the occurrence of gas dynamic disaster risks, and achieving intelligent prediction and early warning of multiple monitoring indicators, as well as the accurate generation of risk interpretability and decision-making solutions, are crucial prerequisites for the prevention and control of coal and gas outburst risks.

[0003] Coal and gas outbursts are influenced by geological structures, coal seam structures, gas conditions, and coal properties, and their evolution leads to dynamic changes in coal and rock stress, gas emission, and deformation and fracturing. Therefore, to improve the predictive ability of outburst risks, previous studies have proposed using static index methods (such as drill cuttings volume, gas desorption index, firmness coefficient, permeability coefficient, and initial gas emission velocity in boreholes) to predict outburst hazards. However, conventional index acquisition is affected by the measurement points and human factors, resulting in low accuracy and the inability to achieve real-time online monitoring. With the improvement of mining technology and the urgent need for dynamic monitoring, researchers have used gas emission and concentration combined with geophysical indicators (such as electromagnetic radiation, acoustic emission, and microseismic activity) to achieve real-time dynamic monitoring of outburst risks. However, microseismic activity is mainly used to monitor high-energy rupture events (such as rockbursts) and is less effective for monitoring gas-dominated outbursts. Acoustic emission and electromagnetic radiation signals combined with gas concentration can reflect coal and rock stress, deformation and fracturing, and gas seepage and emission in real time, thus having wide applications in coal and gas outburst risk monitoring and early warning. However, acoustic emission and electromagnetic radiation signals are easily affected by noise interference, and the monitoring systems currently established in mines still mainly use statistical methods such as trend method, threshold method and evidence chain fusion, which have problems such as insufficient multi-indicator fusion analysis capability and poor generalization ability, and are prone to missed reports and false reports.

[0004] With the increase in computing power, deep learning methods have begun to be applied to real-time early warning of coal and gas outburst risks. Researchers input data such as coal firmness coefficient, gas pressure, coal seam gas content or methane concentration, electromagnetic radiation, and acoustic emission into models like TabNet, HPO-BiLSTM, physical information neural networks, and ETO-TSMixer to quantify and learn the characteristics of coal and gas outburst risks. Furthermore, some researchers convert electromagnetic radiation and acoustic emission signals into frequency domain images, then input them into CNN models to identify abnormal frequency domain locations in the images, thereby identifying outburst risks. These methods have made valuable explorations in the field of coal and gas outburst risk early warning. However, these methods mostly utilize tabular models to fuse and analyze non-time-series indicators or use neural network models to fuse and predict coal and gas outburst risks from multiple real-time monitoring time-series indicators. Limited by the sparsity of non-time-series signals, it is difficult to achieve effective comprehensive analysis of non-time-series and time-series signals. Furthermore, the limited number of samples in actual engineering projects and the significant differences in geological and gas occurrence conditions among different coal mines mean that the effectiveness of the aforementioned methods in different coal mines still needs to be verified. In recent years, however, Large Language Models (LLMs) have demonstrated superior performance in data analysis, intelligent reasoning, and sparse sample cross-task transfer learning due to their powerful knowledge understanding and cross-modal representation capabilities. Therefore, how to achieve high-precision early warning and strong generalization ability of LLMs under small sample conditions is an important direction for current research on intelligent early warning of coal and gas outburst risks. Summary of the Invention

[0005] Purpose of the invention: To address the above-mentioned problems, the purpose of this invention is to provide an advanced intelligent classification and early warning method for coal and gas outburst risks under sparse samples. This method can combine task transfer learning and a text-data dual-channel big language model to construct an adaptive multimodal learning model, which can be used to realize advanced early warning of coal and gas outburst disaster risks and data fusion analysis and prediction of various monitoring indicators.

[0006] Technical solution: A method for advanced intelligent classification and early warning of coal and gas outburst risk under sparse samples, comprising the following steps:

[0007] Step 1: Collect time-series data from the coal mine working face, including gas concentration, electromagnetic radiation, and acoustic emission, and obtain coal seam geological parameters, including gas pressure, gas content, coal seam firmness coefficient, permeability, and initial venting velocity of the borehole, as the initial dataset;

[0008] Step 2: Continuous Fourier transform is used to reduce and smooth the acoustic emission and electromagnetic radiation noise signals. Then, box plot statistical criteria are used to identify abnormal observations in the gas concentration sequence and the identified abnormal points are set to null values. A multiple interpolation method based on chain equations is introduced to perform conditional modeling and interpolation for all missing terms.

[0009] Step 3: Establish a multi-indicator collaborative prediction model TimeAFFN, input the preprocessed data into the model, and simultaneously obtain the future development trends of gas concentration, electromagnetic radiation and acoustic emission;

[0010] Step 4: Establish a coal and gas outburst multi-hazard risk fusion early warning model based on the TimeAFFN model of task transfer learning. Input the original data and predicted data of each indicator into the fusion early warning model to obtain the risk probability of each hazard at the current time point and in the future time period, and perform risk classification, identification and early warning.

[0011] Furthermore, in step two, the determination of abnormal gas concentration points includes the following steps:

[0012] S211: Let the gas concentration sample sequence be X = {x1, x2, ..., x...} n}, where n is the total length of the sequence, and its first quartile Q1, second quartile Q2, and third quartile Q3 are defined as follows:

[0013]

[0014] Where Q2 is the median;

[0015] S212: Use interquartile ranges to construct outlier determination intervals. Interquartile ranges are denoted as:

[0016] IQR = Q3 - Q1;

[0017] Given a threshold coefficient k for the interquartile range, where k takes the value 1.5, the anomaly determination for any sample value x is as follows:

[0018]

[0019] Set all anomalies to empty after identification.

[0020] Ideally, in step two, the imputation of missing items includes the following steps:

[0021] S221: Suppose the monitoring data contains N variables, denoted as:

[0022] X i , i = 1, 2, ..., N;

[0023] Where i is the variable index, if variable X i If there are missing observations, then the set of location indices corresponding to the missing entries is denoted as M. i and M iMissing items are considered as objects to be imputed. During the algorithm initialization phase, all missing items are initially filled with random values ​​or simple statistics. Variables include acoustic emission, electromagnetic radiation, and gas concentration.

[0024] S222: Let the objective variable X... i The predictors are the set of the remaining variables:

[0025] X -i ={X1, ...,X i-1 X i+1 , ...,X N};

[0026] Where X -i Indicates the difference from X i All variables other than X are used as input to the conditional regression model. -i For X i The missing part (index set M) i Construct a conditional regression model; at the r-th iteration, the model's parameter vector... It is obtained by sampling from its fully conditional posterior distribution:

[0027]

[0028] Where, θ i For the fitted variable X i Required model parameters This represents all observations of other variables that have been updated in the r-th iteration; given sampling parameters Under the condition that variable X i Missing observations in the dataset are used to generate imputed values ​​through their posterior predicted distribution:

[0029]

[0030] Where f(·) is the probability distribution corresponding to the selected conditional model;

[0031] S223: Establish a set:

[0032]

[0033] It includes the division by X i The latest estimates of all variables except those in the current iteration;

[0034] S224: Each variable is updated sequentially through the above steps. After multiple iterations, the interpolation results gradually converge. By repeatedly performing independent interpolation multiple times, the uncertainty of the estimate is further characterized.

[0035] Furthermore, in step three, the establishment of the prediction model TimeAFFN includes the following steps:

[0036] S31: Construct a dual-channel data-large language model. In the data channel, the data is encoded using a batch normalization layer and a linear layer and then input into the Transformer encoder.

[0037] S32: In the large language model channel, multiple sequence data and prompts that form word vectors are combined using a pre-trained large language model and a Transformer encoder to extract language semantic features;

[0038] S33: The features obtained from the two channels are then fused using an adaptive feature fusion network, and finally the fused data is used to output the prediction result using a Transformer decoder.

[0039] Ideally, in step S31, the encoding of the data channel includes the following steps:

[0040] S311: Let the matrix of the input sequence be:

[0041]

[0042] Where T is the sequence length, i.e. the number of data steps, and M is the number of indicators;

[0043] S312: In the data extraction channel, a pre-processed batch normalization + linear layer is first used to achieve non-linear feature extraction and dimensionality increase to d for multiple indicators.

[0044]

[0045] Among them, W e For the linear mapping weights, b e Here, d is the bias term, d is the feature dimension after dimensionality increase, and E is the feature dimension after dimensionality increase. (data) This is a hidden representation of the data channel;

[0046] S313: Input the upgraded data into the Transformer encoder for multi-head attention encoding. Its computational principle conforms to the principle proposed in the Transformer paper.

[0047]

[0048] Enc(·) is the Transformer Encoder module, F (data) This is the vector representation of the time sequence encoded by the Transformer encoder.

[0049] Ideally, in step S32, feature extraction for the large language model channels includes the following steps:

[0050] S321: Use the Prompt template as the basis for a large language model;

[0051] S322: Obtain word vector encodings through a large language model, and input the vector encodings into a Transformer encoder for feature extraction.

[0052]

[0053] in, Let L be the text feature vector output after large language encoding, H be the embedding dimension projection that maintains consistency with the input metrics, and E be the text feature vector output after large language encoding. (prompt) F is the sequence of word vector representations generated by a large language model for input prompt text. (lang) These are high-order features of the language channel, which are then represented as vectors by multi-head attention from the Transformer encoder.

[0054] Ideally, in step S33, the encoding of the data channel includes the following steps:

[0055] S331: Use cross-modal attention on the two channels of data to align the language and data modality features:

[0056]

[0057]

[0058] in, This is a cross-modal attention matrix, where Softmax is used for row-wise normalization to generate attention weights. To align A to the data time scale via a cross-modal attention matrix.

[0059] S332: The aligned data from the two channels is concatenated and then input into the AFFN network model. The AFFN network model includes an optional shared module as a supplement for shared feature learning. This module can be enabled or disabled through settings, and its definition is:

[0060]

[0061]

[0062] Among them, F t The final fused output feature at time step t, where t = 1, 2, ..., T represents the t-th time step of the time series. This represents the feature vector of the data channel at time step t. The feature vectors of the language channel aligned to time step t. For the i-th shared expert network, a feature transformation module is uniformly shared across all tasks, where i = 1, 2, ..., n. s There are n s Number of shared experts For the j-th task-specific expert network, used to model the characteristics of the task itself, where j = 1, 2, ..., n h There are n h Number of experts specific to each task For the dynamic weights of the i-th shared expert, For the dynamic weights of experts specific to the j-th task, the following conditions must be met:

[0063]

[0064] Among them W g This is the learnable weight matrix for the gated network, used to generate the Softmax input;

[0065] S333: Features of fusion F t The residuals are updated to the backbone feature space through normalization operations, a process consistent with the residual update mechanism of the Transformer.

[0066]

[0067] Where F is the feature sequence fused by AFFN at all time steps, Dropout(·) is the random deactivation of the input features, and Norm(·) is Layer Normalization, used for feature normalization and numerical stabilization. This is the feature sequence after residual update, which is the intermediate representation output by the AFFN module;

[0068] S334: After multi-layer module learning, the features fused from the last layer are concatenated and input into the Transformer-decoder to perform conditional prediction of the time series target. The results are output using independent task towers for AE, EMR, and Gas. The mathematical model is as follows:

[0069]

[0070] in, This is the prediction result for the m-th indicator. There are M indicators in total, where m is one of the corresponding indicators. m The prediction task tower for indicator m is composed of a feedforward neural network, with Dec being a Transformer decoder. For the last layer of data channel features, This is the last layer of language channel features. This represents the concatenated output features of each channel of the model at the last layer.

[0071] Furthermore, in step four, a risk fusion early warning model is established. First, the parameters of the optimal multi-indicator prediction model are saved. Then, the prediction head in the network structure is converted into a classification head, and the saved optimal model parameters are loaded. At this point, the initial parameters outside the model's classification head already have a multi-indicator fusion foundation. After the conversion, the model's loss function is changed from the predicted MSE to the classification loss cross-entropy function. Specifically, the following steps are included:

[0072] S41: Obtain the fused representation through the TimeAFFN model after multi-index prediction. This means that the input of the average pooling aggregation to the linear classification layer is used to map the output to a binary risk probability via the sigmoid function, as expressed in the following expression:

[0073]

[0074] Where Pool(·) is the average pooling operator for the time dimension, z is the global feature representation after time dimension compression, w is the weight vector of the linear classification layer, b is the bias term of the classification layer, s is the unnormalized classification score, σ(·) is the Sigmoid function used to map real numbers to the probability interval of 0–1, and y∈{0,1} is the true label, where 1 indicates risk and 0 indicates no risk. The optimization loss for the classification task;

[0075] S42: For classification prompts, multiple statistical features such as kurtosis, peak factor, skewness, peak density, variance, fractal box dimension, and spectral entropy are further incorporated into the prediction prompts.

[0076] S43: The threshold determination method based on maximizing the Youden index is adopted. By simultaneously considering the model's sensitivity, i.e., the true positive rate, and its specificity, i.e., the true negative rate, the comprehensive diagnostic capability is calculated under different candidate thresholds, and the threshold that maximizes the Youden index is selected as the risk warning point.

[0077] Ideally, in step S43, the determination of the risk warning point includes the following steps:

[0078] S431: Define the true positive rate and the false positive rate as follows:

[0079]

[0080] Where TP is the number of samples that are actually warnings and the model predicts them as warnings; FN is the number of samples that are actually warnings but are predicted as not warnings; FP is the number of samples that are actually not warnings but are predicted as warnings; and TN is the number of samples that are actually not warnings and are predicted as not warnings.

[0081] S432: Calculate the Youden index at different thresholds:

[0082] J(δ) = TPR(δ) - FPR(δ);

[0083] Where δ is the classification threshold, and the Youden index reflects the model's overall discriminative ability at the current threshold, with its maximum point corresponding to the optimal risk assessment threshold:

[0084]

[0085] S433: Predicting probability With the optimal threshold δ * By comparing the results, the final classification of risk events is obtained:

[0086]

[0087] Where I(·) is an indicator function, which takes the value 1 when the condition is met, and 0 otherwise.

[0088] Furthermore, in step four, the graded identification and early warning is divided into levels based on the distribution of current risk probability and future risk probability. Using TimeAFFN's multi-indicator prediction and a risk early warning model based on transfer learning, the risk probability of coal and gas outburst disasters at the current time point and in the future time period can be calculated. If the disaster risk probability does not exceed the threshold in either time period, no warning is issued. If the disaster risk probability at the current time point exceeds the threshold, and the disaster risk warning probability in the future time period also exceeds the threshold, then a level one warning is issued for the corresponding disaster. In other cases, a level two warning is issued.

[0089] Beneficial effects: Compared with the prior art, the advantages of the present invention are:

[0090] This invention combines task transfer learning with a large language model's text-data dual-channel adaptive multi-modal TimeAFFN for coal and gas outburst disaster risk fusion early warning and data fusion analysis and prediction of various monitoring indicators. It comprehensively considers the mutual influence patterns among various indicators and the risk information of different indicators before the occurrence of outburst disasters to provide advanced early warning of coal and gas outburst disaster risks. First, the invention preprocesses the original data sequence, providing complete and high-quality sequence data for subsequent prediction and early warning. Then, it fuses and predicts gas concentration, acoustic emission, and electromagnetic radiation indicators. This prediction model effectively learns the correlation between different indicators and their statistical characteristics on the prediction of each indicator by inputting time-series indicator data and prompts embedded with data features into the model. It has high data prediction accuracy and provides important future data support for subsequent multi-hazard risk fusion early warning. Finally, task transfer learning further fine-tunes the model based on the knowledge learned for the prediction task to adapt to the requirements of the classification task. It's important to note that during task transfer learning, this invention modifies the TimeAFFN prediction head into a classification head (TimeAFFN-Classifier) ​​and redesigns the prompt that integrates non-time-series geological parameters and multiple statistical indicators. The model also accepts multimodal input data and outputs classification results. By simultaneously inputting raw and future data into the fusion early warning model, which comprehensively learns the correlation between time-series and non-time-series indicators and the multi-statistical precursor features of time-series indicators, the fusion analysis capability of the early warning task can be improved, resulting in higher accuracy in identifying coal and gas outburst disaster risks. Furthermore, this invention utilizes the risk probabilities of current and future coal and gas outburst disaster risks obtained from the fusion early warning model for graded early warning, further ensuring the accurate identification of coal and gas outburst disaster risks. Attached Figure Description

[0091] Figure 1 This is a flowchart of the method of the present invention;

[0092] Figure 2 The model diagram for the TimeAFFN fusion prediction model. Detailed Implementation

[0093] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0094] A method for advanced intelligent classification and early warning of coal and gas outburst risk under sparse samples, such as Figure 1 , 2 As shown, it includes the following steps:

[0095] Step 1: Data preprocessing;

[0096] This invention first employs continuous Fourier transform to denoise and smooth acoustic emission and electromagnetic radiation noise signals. Then, it uses box plot statistical criteria to identify anomalous observations in the gas concentration sequence, setting the identified outliers as null values ​​(i.e., missing data). Finally, it introduces a multiple imputation by chained equations (MICE) method to conditionally model and imputate all missing terms. Its mathematical principles can be summarized as follows.

[0097] Let the gas concentration sample sequence be X = {x1, x2, ..., x...} n}, where n is the total length of the sequence, and its first, second, and third quartiles are defined as follows:

[0098]

[0099] Using interquartile differences:

[0100] IQR = Q3 - Q1;

[0101] An outlier detection interval can be constructed. Given a threshold coefficient k (usually 1.5) for the interquartile range multiple, the outlier determination for any sample value x is as follows:

[0102]

[0103] All outliers are set to empty after identification to facilitate subsequent interpolation.

[0104] Multiple Imputation by Chained Equations (MICE) is an iterative algorithm that estimates missing data through conditional modeling of each variable. Its core idea is to use a generalized linear regression framework to infer the missing portion of the target variable from the conditional distributions of the remaining variables. Let the monitoring data contain N variables (such as acoustic emission, electromagnetic radiation, gas concentration, etc.), denoted as:

[0105] X i , i = 1, 2, ..., N;

[0106] Where i is the variable index, if variable X i If there are missing observations, then the set of location indices corresponding to the missing entries is denoted as M. i and M i Missing terms in the variable X are considered as objects to be imputed. During the algorithm initialization phase, all missing terms are initially filled with random values ​​or simple statistics; the variables include acoustic emission, electromagnetic radiation, and gas concentration. Let the target variable X... i The predictors are the set of the remaining variables:

[0107] X -i ={X1, ..., X} i-1 X i+1 , ..., X N};

[0108] Using this as a predictor, for X i A conditional regression model is constructed based on the missing observations. At the r-th iteration, the model's parameter vector... Sampling from its fully conditional posterior distribution

[0109]

[0110] Where θ i For the fitted variable X i The required model parameters (such as linear regression coefficients, bias, or variance, etc.), This represents all observations of other variables that have been updated in the r-th iteration. Given sampling parameters... Under the condition that variable X i Missing observations in the dataset are used to generate imputed values ​​through their posterior predicted distribution:

[0111]

[0112] Where f(·) is the probability distribution (such as normal distribution, logistic distribution, etc.) corresponding to the selected conditional model. Set:

[0113]

[0114] Includes division by X i The chain equation represents the latest estimates of all variables except those mentioned above in the current iteration. Since each variable is updated sequentially in the manner described, the chain equation ensures that the missing portion of each variable is estimated under the conditional structure constraints of the remaining variables, thus maintaining overall correlation. After multiple iterations, the imputation results gradually converge. Further characterizing the uncertainty of the estimation can be achieved by repeatedly performing independent imputation.

[0115] Step 2: Construct an adaptive fusion prediction model TimeAFFN based on a large language model and data dual-channel approach;

[0116] This invention addresses coal and gas outburst disaster monitoring indicators. Drawing on experience from multimodal fusion (such as image-text multimodal fusion), it processes data and prompts separately, then fuses them using attention or gating units. This approach decouples data processing from prompt processing using a large language model, fully leveraging the advantages of different processing models. By comprehensively analyzing the coupling evolution patterns between them, it can more accurately predict the future trends of each indicator. Therefore, this invention proposes a novel multi-indicator prediction method for coal and gas outbursts based on a large language model using a multimodal fusion approach. The key features of this method include designing a prompt embedding sequence information and time-frequency features, a feature extraction method based on a dual-channel data-large language model, and a data adaptive fusion architecture incorporating MOE (Multimodal Expression Encryption).

[0117] Specifically, in the data channel, batch normalization and linear layers are used to encode the data before it is input into the Transformer encoder; in the large language model channel, multi-sequence data and prompts forming word vectors are combined using a pre-trained large language model and a Transformer encoder to extract linguistic semantic features; then, the features obtained from the two channels are fused through an adaptive feature fusion network, and finally, the fused data is used to output the prediction result using a Transformer decoder. The core principle is as follows:

[0118] Suppose the matrix of the input sequence is Where T is the sequence length, i.e., the number of data steps, and M is the number of indicators. In the data extraction channel, a pre-processed batch normalization + linear layer is first used to achieve non-linear feature extraction and dimensionality increase to d for multiple indicators.

[0119]

[0120] Among them W e For the linear mapping weights, b e Here, d is the bias term, d is the feature dimension after dimensionality increase, and E is the feature dimension after dimensionality increase. (data) This is a hidden representation of the data channels. The upgraded data is then input into the Transformer encoder for multi-head attention encoding, the computation of which conforms to the principle proposed in the Transformer paper:

[0121]

[0122] Enc(·) is the Transformer Encoder module, F (data) This is the vector representation of the temporal sequence encoded by the Transformer encoder. For the large language model part, the Prompt template is used as the basis, and the template style is as follows:

[0123] From [t1] to o[t2], the value is the [sequence value] per [time interval]. The overall trend is the [trend].

[0124] Then, word vector encodings are obtained through a large language model, and these vector encodings are input into a Transformer encoder for feature extraction.

[0125]

[0126] in Let L be the text feature vector output after large language encoding, H be the embedding dimension projection that maintains consistency with the input metrics, and E be the text feature vector output after large language encoding. (prompt) F is the sequence of word vector representations generated by a Large Language Model (LLM) for input prompt text. (lang) The high-order features of the language channel are represented as vectors after multi-head attention by the Transformer encoder. The temporal-encoded language is input into the AFFN network, which adaptively learns and fuses multimodal features from different channels through an expert network and gating units. Before inputting into this module, cross-modal attention is applied to the two channels to align the language and data modal features.

[0127]

[0128]

[0129] This is a cross-modal attention matrix, where Softmax is used for row-wise normalization to generate attention weights. To align language features to the data time scale via a cross-modal attention matrix A, the aligned data is concatenated into dual-channel data and input into the AFFN. We provide an optional shared module in the AFFN as a supplement to shared feature learning, which can be enabled or disabled through settings. Its definition is as follows:

[0130]

[0131] Where F t The final fused output feature at time step t, where t = 1, 2, ..., T represents the t-th time step of the time series. This represents the feature vector of the data channel at time step t. The language channel is aligned to the feature vector at time step t. For the i-th "shared expert network", a feature transformation module is uniformly shared across all tasks, where i = 1, 2, ..., n. s (Number of shared experts) The j-th "task-specific expert network" is used to model the characteristics of the task itself (such as differences in AE / EMR / Gas), where j = 1, 2, ..., n h (Number of mission-specific experts). For the dynamic weights of the i-th shared expert, The dynamic weights of the expert for the j-th task satisfy:

[0132]

[0133] Among them W g The learnable weight matrix of the gated network is used to generate the softmax input, and the fused features F t The residual and normalization operations are used to update the backbone feature space, a process consistent with the residual update mechanism of the Transformer, which helps to preserve information and stabilize gradients.

[0134]

[0135] Where F is the feature sequence fused by AFFN at all time steps, Dropout(·) randomly deactivates the input features, and Norm(·) is Layer Normalization, used for feature normalization and numerical stabilization. The feature sequence after residual update is the intermediate representation output by the AFFN module. After learning through multiple modules, the features fused from the last layer are concatenated and input into the Transformer-decoder for conditional prediction of the time series target. The results are output using independent task towers for AE, EMR, and Gas. The calculation method is as follows:

[0136]

[0137] in For the prediction result of the m-th indicator (there are M indicators in total, and m is one of the corresponding indicators), τ m The prediction task tower for metric m is constructed from a feedforward neural network. Dec is the Transformer decoder. For the last layer of data channel features, This is the last layer of language channel features. This represents the concatenated output features of each channel of the model at the last layer.

[0138] Step 3: Establish a multi-hazard risk fusion early warning model for coal and gas outbursts based on the TimeAFFN model using task transfer learning.

[0139] This invention transforms the TimeAFFN model's task objective from prediction to classification, requiring changes to the prompt word rules and the conversion of the final linear layer into a classification head. Since classification tasks may involve relatively little data, while TimeAFFN has a large number of parameters, it's necessary to pre-train the model using prediction tasks and then fine-tune it in classification tasks. Specifically, the parameters of the optimal multi-metric prediction model are saved, and then the prediction head in the network structure is converted into a classification head. The saved optimal model parameters are then loaded. At this point, the initial parameters outside the classification head already have a multi-metric fusion foundation, allowing for high classification performance with fewer iterations. After the conversion, the model's loss function is changed from the predicted MSE to the classification loss cross-entropy function, as detailed below:

[0140] The TimeAFFN model, after multi-index prediction, can yield a fused representation. This means that the input of the average pooling aggregation to the linear classification layer is used to map the output to a binary risk probability via the sigmoid function, as expressed in the following expression:

[0141]

[0142] Where Pool(·) is the average pooling operator for the time dimension, z is the global feature representation after time dimension compression, w is the weight vector of the linear classification layer, b is the bias term of the classification layer, s is the unnormalized classification score (logit value), σ(·) is the Sigmoid function used to map real numbers to the probability interval of 0–1, and y∈{0,1} is the true label, where 1 indicates risk and 0 indicates no risk. The optimization loss for the classification task.

[0143] For the classification prompt, we further incorporate multi-dimensional statistical features such as kurtosis, peak factor, skewness, peak distribution density, variance, fractal box dimension, and spectral entropy into the prediction prompt, as detailed below:

[0144]

[0145]

[0146] To more rationally select the optimal model suitable for field applications, this invention employs a threshold determination method based on maximizing the Youden index. This method simultaneously considers the model's sensitivity (true positive rate) and specificity (true negative rate), calculates the comprehensive diagnostic capability under different candidate thresholds, and selects the threshold that maximizes the Youden index as the risk warning point. First, the True Positive Rate (TPR) and False Positive Rate (FPR) are defined as follows:

[0147]

[0148] Where TP is the number of samples that are actually warnings and the model predicts them as warnings; FN is the number of samples that are actually warnings but are predicted as not warnings; FP is the number of samples that are actually not warnings but are predicted as warnings; and TN is the number of samples that are actually not warnings and are predicted as not warnings.

[0149] Next, the Youden index is calculated at different thresholds:

[0150] J(δ) = TPR(δ) - FPR(δ);

[0151] Where δ is the classification threshold. The Youden index reflects the model's overall discriminative ability at the current threshold, with its maximum point corresponding to the optimal risk assessment threshold.

[0152]

[0153] Finally, predict the probability With the optimal threshold δ * By comparing the results, the final classification of risk events is obtained:

[0154]

[0155] Where I(·) is an indicator function, which takes the value 1 when the condition is met, and 0 otherwise.

[0156] Step 4: Conduct risk classification, identification, and early warning;

[0157] The tiered early warning system classifies risks based on the distribution of current and future risk probabilities. Specifically, it uses TimeAFFN's multi-indicator prediction and a risk warning model based on transfer learning to calculate the risk probability of coal and gas outburst disasters at the current time and in the future. If the risk probability does not exceed the threshold in either time period, no warning is issued. If the risk probability exceeds the threshold at the current time and also exceeds the threshold in the future time period, a Level 1 warning is issued for the corresponding disaster. All other cases are subject to a Level 2 warning.

Claims

1. A method for advanced intelligent classification and early warning of coal and gas outburst risk under sparse samples, characterized in that... Includes the following steps: Step 1: Collect time-series data from the coal mine working face, including gas concentration, electromagnetic radiation, and acoustic emission, and obtain coal seam geological parameters, including gas pressure, gas content, coal seam firmness coefficient, permeability, and initial venting velocity of the borehole, as the initial dataset; Step 2: Continuous Fourier transform is used to reduce and smooth the acoustic emission and electromagnetic radiation noise signals. Then, box plot statistical criteria are used to identify abnormal observations in the gas concentration sequence and the identified abnormal points are set to null values. A multiple interpolation method based on chain equations is introduced to perform conditional modeling and interpolation for all missing terms. Step 3: Establish a multi-indicator collaborative prediction model TimeAFFN, input the preprocessed data into the model, and simultaneously obtain the future development trends of gas concentration, electromagnetic radiation and acoustic emission; Step 4: Establish a coal and gas outburst multi-hazard risk fusion early warning model based on the TimeAFFN model of task transfer learning. Input the original data and predicted data of each indicator into the fusion early warning model to obtain the risk probability of each hazard at the current time point and in the future time period, and perform risk classification, identification and early warning.

2. The method for advanced intelligent classification and early warning of coal and gas outburst risk under sparse samples according to claim 1, characterized in that, In step two, the determination of abnormal gas concentration points includes the following steps: S211: Let the gas concentration sample sequence be X = {x1, x2, ..., x...} n }, where n is the total length of the sequence, and its first quartile Q1, second quartile Q2, and third quartile Q3 are defined as follows: Where Q2 is the median; S212: Use interquartile ranges to construct outlier determination intervals. Interquartile ranges are denoted as: I QR =Q3-Q1; Given a threshold coefficient k for the interquartile range, where k takes the value 1.5, the anomaly determination for any sample value x is as follows: Set all anomalies to empty after identification.

3. The method for advanced intelligent classification and early warning of coal and gas outburst risk under sparse samples according to claim 1 or 2, characterized in that, In step two, the imputation of missing items includes the following steps: S221: Suppose the monitoring data contains N variables, denoted as: X i ,i=1,2,...,N; Where i is the variable index, if variable X i If there are missing observations, then the set of location indices corresponding to the missing entries is denoted as M. i and M i Missing items are considered as objects to be imputed. During the algorithm initialization phase, all missing items are initially filled with random values ​​or simple statistics. Variables include acoustic emission, electromagnetic radiation, and gas concentration. S222: Let the objective variable X... i The predictors are the set of the remaining variables: X -i ={X1,...,X i-1 X i+1 ..., X N }; Where X -i Indicates the difference from X i All variables other than X are used as input to the conditional regression model, based on X. -i For X i The missing part is the index set M i Construct a conditional regression model; at the r-th iteration, the model's parameter vector... It is obtained by sampling from its fully conditional posterior distribution: Where, θ i For the fitted variable X i Required model parameters This represents all observations of other variables that have been updated in the r-th iteration; given sampling parameters Under the condition that variable X i Missing observations in the dataset are used to generate imputed values ​​through their posterior predicted distribution: Where f(·) is the probability distribution corresponding to the selected conditional model; S223: Establish a set: It includes the division by X i The latest estimates of all variables except those in the current iteration; S224: Each variable is updated sequentially through the above steps. After multiple iterations, the interpolation results gradually converge. By repeatedly performing independent interpolation multiple times, the uncertainty of the estimate is further characterized.

4. The method for advanced intelligent classification and early warning of coal and gas outburst risk under sparse samples according to claim 1, characterized in that, In step three, the establishment of the prediction model TimeAFFN includes the following steps: S31: Construct a dual-channel data-large language model. In the data channel, the data is encoded using a batch normalization layer and a linear layer and then input into the Transformer encoder. S32: In the large language model channel, multiple sequence data and prompts that form word vectors are combined using a pre-trained large language model and a Transformer encoder to extract language semantic features; S33: The features obtained from the two channels are then fused using an adaptive feature fusion network, and finally the fused data is used to output the prediction result using a Transformer decoder.

5. The method for advanced intelligent classification and early warning of coal and gas outburst risk under sparse samples according to claim 4, characterized in that, In step S31, the encoding of the data channel includes the following steps: S311: Let the matrix of the input sequence be: Where T is the sequence length, i.e. the number of data steps, and M is the number of indicators; S312: In the data extraction channel, a pre-processed batch normalization + linear layer is first used to achieve non-linear feature extraction and dimensionality increase to d for multiple indicators. Among them, W e For the linear mapping weights, b e Here, d is the bias term, d is the feature dimension after dimensionality increase, and E is the feature dimension after dimensionality increase. (data) This is a hidden representation of the data channel; S313: Input the upgraded data into the Transformer encoder for multi-head attention encoding. Its computational principle conforms to the principle proposed in the Transformer paper. Enc(·) is the Transformer Encoder module, F (data) This is the vector representation of the time sequence encoded by the Transformer encoder.

6. The method for advanced intelligent classification and early warning of coal and gas outburst risk under sparse samples according to claim 4, characterized in that, In step S32, feature extraction of the large language model channels includes the following steps: S321: Use the Prompt template as the basis for a large language model; S322: Obtain word vector encodings through a large language model, and input the vector encodings into a Transformer encoder for feature extraction. in, Let L be the text feature vector output after large language encoding, H be the embedding dimension projection that maintains consistency with the input metrics, and E be the text feature vector output after large language encoding. (prompt) F is the sequence of word vector representations generated by a large language model for input prompt text. (lang) These are high-order features of the language channel, which are then represented as vectors by multi-head attention from the Transformer encoder.

7. The method for advanced intelligent classification and early warning of coal and gas outburst risk under sparse samples according to claim 4, characterized in that, In step S33, the encoding of the data channel includes the following steps: S331: Use cross-modal attention on the two channels of data to align the language and data modality features: in, This is a cross-modal attention matrix, where Softmax is used for row-wise normalization to generate attention weights. To align A to the data time scale via a cross-modal attention matrix. S332: The aligned data from the two channels is concatenated and then input into the AFFN network model. The AFFN network model includes an optional shared module as a supplement for shared feature learning. This module can be enabled or disabled through settings, and its definition is: Among them, F t The final fused output feature at time step t, where t = 1, 2, ..., T represents the t-th time step of the time series. This represents the feature vector of the data channel at time step t. The feature vectors of the language channel aligned to time step t. For the i-th shared expert network, a feature transformation module is uniformly shared across all tasks, where i = 1, 2, ..., ns, and there are n such modules. s Number of shared experts For the j-th task-specific expert network, used to model the characteristics of the task itself, where j = 1, 2, ..., n h There are n h Number of experts specific to each task For the dynamic weights of the i-th shared expert, For the dynamic weights of experts specific to the j-th task, the following conditions must be met: Among them W g This is the learnable weight matrix for the gated network, used to generate the Softmax input; S333: Features of fusion F t The residuals are updated to the backbone feature space through normalization operations, a process consistent with the residual update mechanism of the Transformer. Where F is the feature sequence fused by AFFN at all time steps, Dropout(·) is the random deactivation of the input features, and Norm(·) is Layer Normalization, used for feature normalization and numerical stabilization. This is the feature sequence after residual update, which is the intermediate representation output by the AFFN module; S334: After multi-layer module learning, the features fused from the last layer are concatenated and input into the Transformer-decoder to perform conditional prediction of the time series target. The results are output using independent task towers for AE, EMR, and Gas. The mathematical model is as follows: in, This is the prediction result for the m-th indicator. There are M indicators in total, where m is one of the corresponding indicators. m The prediction task tower for indicator m is composed of a feedforward neural network, with Dec being a Transformer decoder. For the last layer of data channel features, This is the last layer of language channel features. This represents the concatenated output features of each channel of the model at the last layer.

8. The method for advanced intelligent classification and early warning of coal and gas outburst risk under sparse samples according to claim 1, characterized in that, In step four, a risk fusion early warning model is established. First, the parameters of the optimal multi-indicator prediction model are saved. Then, the prediction head in the network structure is converted into a classification head, and the saved optimal model parameters are loaded. At this point, the initial parameters outside the model's classification head already have a multi-indicator fusion foundation. After the conversion, the model's loss function is changed from the predicted MSE to the classification loss cross-entropy function. The specific steps include: S41: Obtain the fused representation through the TimeAFFN model after multi-index prediction. This means that the input of the average pooling aggregation to the linear classification layer is used to map the output to a binary risk probability via the sigmoid function, as expressed in the following expression: Where Pool(·) is the average pooling operator for the time dimension, z is the global feature representation after time dimension compression, w is the weight vector of the linear classification layer, b is the bias term of the classification layer, s is the unnormalized classification score, σ(·) is the Sigmoid function used to map real numbers to the probability interval of 0–1, and y∈{0,1} is the true label, where 1 indicates risk and 0 indicates no risk. The optimization loss for the classification task; S42: For classification prompts, multiple statistical features such as kurtosis, peak factor, skewness, peak density, variance, fractal box dimension, and spectral entropy are further incorporated into the prediction prompts. S43: The threshold determination method based on maximizing the Youden index is adopted. By simultaneously considering the model's sensitivity, i.e., the true positive rate, and its specificity, i.e., the true negative rate, the comprehensive diagnostic capability is calculated under different candidate thresholds, and the threshold that maximizes the Youden index is selected as the risk warning point.

9. The method for advanced intelligent classification and early warning of coal and gas outburst risk under sparse samples according to claim 8, characterized in that, In step S43, the determination of risk warning points includes the following steps: S431: Define the true positive rate and the false positive rate as follows: Where TP is the number of samples that are actually warnings and the model predicts them as warnings; FN is the number of samples that are actually warnings but are predicted as not warnings; FP is the number of samples that are actually not warnings but are predicted as warnings; and TN is the number of samples that are actually not warnings and are predicted as not warnings. S432: Calculate the Youden index at different thresholds: J(δ) = TPR(δ) - FPR(δ); Where δ is the classification threshold, and the Youden index reflects the model's overall discriminative ability at the current threshold, with its maximum point corresponding to the optimal risk assessment threshold: S433: Predicting probability With the optimal threshold δ * By comparing the results, the final classification of risk events is obtained: Where I(·) is an indicator function, which takes the value 1 when the condition is met, and 0 otherwise.

10. The method for advanced intelligent classification and early warning of coal and gas outburst risk under sparse samples according to claim 1, characterized in that, In step four, the graded identification and early warning is based on the distribution of current risk probability and future risk probability. The risk probability of coal and gas outburst disaster at the current time point and in the future time period can be calculated using TimeAFFN's multi-indicator prediction and a risk early warning model based on transfer learning. If the disaster risk probability does not exceed the threshold in either time period, no warning is issued. If the disaster risk probability at the current time point exceeds the threshold, and the disaster risk warning probability in the future time period also exceeds the threshold, a first-level warning is issued for the corresponding disaster. In other cases, a second-level warning is issued.