Multi-model fusion coal yield prediction method and system based on conflict resolution

Through a multi-model fusion method based on TRIZ theory, combined with CNN, BiGRU and Transformer models, the spatial and temporal characteristics of coal output data were extracted, and the problem of insufficient prediction accuracy and reliability of coal output in the existing technology was solved, and higher prediction accuracy and robustness were achieved.

CN120031192APending Publication Date: 2025-05-23SHANDONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510100688.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing coal output prediction methods have challenges in data quality and quantity limitation, model training efficiency and accuracy, and it is difficult to effectively improve the prediction accuracy.

Method used

The multi-model fusion method based on TRIZ theory is adopted to determine the basic structure of the coal production prediction model through the construction of contradiction analysis and technical contradiction matrix table. Combining CNN, BiGRU and Transformer models, the spatial and temporal characteristics of coal output data are extracted and model parameters are optimized to improve prediction accuracy.

Benefits of technology

Through multi-model fusion and parameter optimization, the accuracy and reliability of coal output prediction are significantly improved, long-term dependencies and timing information can be better captured, and the robustness of the model is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031192A_ABST
    Figure CN120031192A_ABST
Patent Text Reader

Abstract

The invention provides a multi-model fusion coal yield prediction method and system based on conflict resolution, and the method comprises the steps: carrying out the contradictory pair analysis of a coal yield prediction algorithm, determining technical parameters, constructing a technical contradictory matrix table, and determining a basic structure of a coal yield prediction model according to the technical contradictory matrix table; the method comprises the following steps: acquiring coal yield related data in a set time period, and preprocessing to obtain a training data set; building a prediction model, training the prediction model by using the training data set, and optimizing parameters of the prediction model by calculating an error value between an expected output value and a predicted output value of the prediction model until the prediction model converges to obtain an optimal prediction model; and processing the coal yield related data in the target time period by using the optimal prediction model to obtain a coal yield prediction result in the target time period. The prediction accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of coal production time series prediction, and specifically relates to a multi-model fusion coal production prediction method and system based on conflict resolution. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Time series forecasting is a method to predict future coal production based on historical data by analyzing the dependencies between time series data and the interference effects of random fluctuations. There are many existing coal production forecasting methods, which are based on different theoretical frameworks and technical means to improve the accuracy and practicality of forecasts. The following is a summary of several major coal production forecasting methods:

[0004] Artificial neural network (ANN) is a computational model that simulates the structure and function of the human brain neural network and can make predictions by learning nonlinear relationships in historical data.

[0005] Grey prediction method is a method for processing incomplete information or a small amount of data. It is suitable for situations where the amount of data is small or the data sequence is incomplete. This method predicts the system behavior by constructing a grey system model and using known information.

[0006] Regression analysis is a method of predicting the change of dependent variables by establishing a mathematical relationship model between independent variables and dependent variables. In coal production forecasting, multiple factors that affect coal production can be selected as independent variables, such as coal seam thickness, mining technical conditions, market demand, etc., and coal production can be predicted through regression analysis.

[0007] This is a relatively new method for predicting coal production based on the principle of electricity-to-coal conversion, which is based on the principle of electricity-to-coal conversion (i.e. there is a certain conversion relationship between electricity consumption and coal production) and is combined with an artificial intelligence model for prediction. This method obtains electricity consumption data of coal enterprises and uses artificial intelligence algorithms to mine the potential laws in the data to predict coal production.

[0008] The above methods have achieved certain results in coal production forecasting. Thanks to their flexibility and adaptability, they can handle various complex data sets and optimize the forecasting effect by adjusting parameters. However, the above models are also limited by data quality and quantity, as well as challenges in model training efficiency and accuracy. Summary of the invention

[0009] In order to solve the above problems, the present invention proposes a multi-model fusion coal production prediction method and system based on conflict resolution, which can improve the prediction accuracy.

[0010] According to some embodiments, the present invention adopts the following technical solutions:

[0011] A multi-model fusion coal production prediction method based on conflict resolution includes the following steps:

[0012] Using TRIZ theory, we analyze the contradictions of the coal production prediction algorithm, determine the technical parameters, build a technical contradiction matrix, and determine the basic structure of the coal production prediction model based on the technical contradiction matrix.

[0013] Obtain coal production related data within a set time period and perform preprocessing to obtain a training data set;

[0014] A prediction model is built, and the prediction model is trained using the training data set. The parameters of the prediction model are optimized by calculating the error value between the expected output value and the predicted output value of the prediction model until the prediction model converges to obtain an optimal prediction model; the prediction model comprises an input layer, a CNN layer, a BiGRU layer, a Transformer layer and an output layer, wherein the input layer is used to receive feature data of coal production-related data, the CNN layer is used to extract spatial features in the coal production data, the BiGRU layer is used to extract time features of the received data, the Transformer layer is used to generate prediction results in combination with the extracted time features, and the output layer is used to output the prediction results;

[0015] The optimal prediction model is used to process the coal production related data in the target time period to obtain the coal production prediction result in the target time period.

[0016] As an optional implementation method, the process of conducting contradiction analysis on the coal production prediction algorithm, determining technical parameters, and constructing a technical contradiction matrix table includes: determining technical parameters including accuracy, device complexity, reliability and time waste, constructing a contradiction matrix table of accuracy and device complexity, and a contradiction matrix table of reliability and time waste.

[0017] As an optional implementation, the preprocessing process includes deleting useless tag fields and feature data with serious missing values, using the KNN nearest neighbor method to fill in the values ​​of feature data with less missing data, and then using the XGBoost method to filter the feature data, and normalizing the filtered data.

[0018] As an optional implementation, the CNN layer includes a convolution layer and a pooling layer. The input layer receives coal production-related features, first passes through the convolution layer, learns the spatial structure information in the data through multiple convolution kernels, extracts local features in the data, and then reduces the number of parameters and calculations through the pooling layer.

[0019] As an optional implementation, the BiGRU layer consists of two independent GRU units, one that processes data forward in time series and the other that processes data in reverse time series. Through this bidirectional structure, the BiGRU layer can capture both forward and backward information of sequence data, thereby better understanding and predicting patterns in the sequence. The BiGRU model generates the final output by combining the hidden states in both the forward GRU and reverse GRU directions. The formula is as follows:

[0020]

[0021] in, represents the output result of the forward propagation hidden state, represents the output result of the backward propagation hidden state, GRU() is the gated recurrent unit, x t represents input, w t is the output weight of the hidden layer of the GRU unit where information is forward propagated at time t, v t is the hidden output weight of the GRU unit that propagates information backward at time t; b t Represents the bias corresponding to the BiGRU hidden state at time t.

[0022] As a further step, the internal structure formula of the GRU layer is as follows:

[0023] r t =σ(W r ·[h t-1 ,x t ]);

[0024] z t =σ(W z ·[h t-1 ,x t ]);

[0025]

[0026] Among them, σ is the sigmoid function, which transforms the data into a value in the range of 0-1, thus acting as a gating signal; t In, W r is the weight matrix, and the weight matrix is ​​used to input x t and hidden state h t-1 The concatenated matrices are multiplied and then r is obtained through the sigmoid function. t Value, r t The smaller the value, the more information needs to be forgotten in the previous moment, and the more information is discarded; update gate z t The closer the value is to 1, the more data is memorized.

[0027] As an optional implementation, the Transformer layer uses an encoder-decoder structure for parallel processing of input sequences to capture global timing information; the encoder structure consists of a multi-head attention layer and a feedforward neural network layer, and each sublayer contains a residual connection and a layer normalization layer; the decoder structure has one more layer of masked multi-head attention than the encoder structure, and a residual connection and a layer normalization layer are also added between each sublayer;

[0028] The encoder structure encodes the input sequence into a high-dimensional feature vector representation, and the decoder structure decodes the vector into the target sequence.

[0029] As a further step, the encoder structure is composed of 6 encoder layers stacked together, and the input data of each encoder layer is obtained after data preprocessing to obtain data X i , after transformation in the form of a matrix, the query vector, key vector and value vector are obtained, which enter the multi-head attention layer, calculate the attention values ​​respectively, and finally concatenate them;

[0030] Next, the data is passed into the residual connection and layer normalization layer, and then into the fully connected layer. Finally, it passes through the residual connection and layer normalization layer again and enters the next encoder layer. The encoder layers are superimposed.

[0031] As a further step, the calculation process of the residual connection and the layer normalization layer includes:

[0032] Add&Norm=LayerNorm(X+MultiHeadAttention(X));

[0033] Add&Norm=LayerNorm(X+FeedForward(X));

[0034] Among them, X represents the input of the multi-head attention or feedforward neural network, MultiHeadAttention(X) and FeedForward(X) represent the output; Add refers to X+MultiHeadAttention(X). As a residual connection technology, Norm is used for layer normalization. By standardizing the input of each layer of neurons, the mean and variance of the input data remain consistent.

[0035] As an optional implementation, the calculation of the fully connected layer includes:

[0036] FFN(X)=max(0,X·W 1 +b 1 )·W 2 +b 2 ;

[0037] The fully connected layer is a two-layer neural network, X represents the output of the Add&Norm layer, and W 1 and b 1 is the parameter of the first fully connected layer, W 2 and b 2 are the parameters of the second fully connected layer.

[0038] As an optional embodiment, when training the prediction model, the root mean square error, mean absolute error, mean absolute percentage error and determination coefficient are used to evaluate the prediction value of the prediction model;

[0039] The Adam optimizer is used to optimize the prediction model to minimize the loss value of the loss function of the prediction model.

[0040] A multi-model fusion coal production prediction system based on conflict resolution, comprising:

[0041] The prediction analysis module is configured to use TRIZ theory to perform contradiction analysis on the coal production prediction algorithm, determine technical parameters, construct a technical contradiction matrix table, and determine the basic structure of the coal production prediction model according to the technical contradiction matrix table;

[0042] The data preprocessing module is configured to obtain data related to coal production within a set time period and perform preprocessing to obtain a training data set;

[0043] The model building and training module is configured to build a prediction model, train the prediction model using the training data set, optimize the parameters of the prediction model by calculating the error value between the expected output value and the predicted output value of the prediction model, until the prediction model converges to obtain the optimal prediction model; the prediction model includes an input layer, a CNN layer, a BiGRU layer, a Transformer layer and an output layer, the input layer is used to receive feature data of coal production related data, the CNN layer is used to extract spatial features of received data, the BiGRU layer is used to extract temporal features of received data, the Transformer layer is used to generate prediction results in combination with the extracted temporal features, and the output layer is used to output the prediction results;

[0044] The prediction module is configured to use the optimal prediction model to process the coal production related data in the target time period to obtain the coal production prediction result in the target time period.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] The present invention first uses TRIZ theory to construct technical contradiction pairs, from which it is inspired to construct a coal production prediction method. According to the principle of the invention, improvements are made, multiple models are fused, and a multi-model fusion coal production prediction method based on TRIZ conflict resolution is constructed, which has the following advantages and helps to capture long-term dependencies: the Transformer model is based on its self-attention mechanism and can learn the long-term dependencies between elements in the input sequence; it can capture time series information, BiGRU is good at capturing time information in sequence data, and CNN is good at capturing spatial information in sequence data, which is crucial for analyzing the law of changes in coal production over time;

[0047] By combining the advantages of Transformer, BiGRU and CNN, the model can more comprehensively understand the feature representation and timing information in the input data, thereby improving the accuracy and reliability of the prediction; the deep learning model is adopted, which has strong robustness and can resist the influence of noise and outliers in the data on the prediction results to a certain extent.

[0048] The present invention can fully utilize the advantages of the three models, improve the accuracy and reliability of prediction, and provide more scientific and reasonable decision support for the planning and management of the coal industry.

[0049] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0051] Figure 1 It is a flow chart of a method for predicting coal production using multi-model fusion based on TRIZ conflict resolution in an embodiment;

[0052] Figure 2 A schematic diagram of the network structure of a GRU model in an embodiment;

[0053] Figure 3 A schematic diagram of the network structure of a BiGRU model in an embodiment;

[0054] Figure 4 A schematic diagram of the structure of a self-attention mechanism in an embodiment;

[0055] Figure 5 A schematic diagram of the structure of a multi-head attention mechanism in one embodiment;

[0056] Figure 6A schematic diagram of the encoder structure of a Transformer model in an embodiment;

[0057] Figure 7 Schematic diagram of the decoder structure of a Transformer model in an embodiment. DETAILED DESCRIPTION

[0058] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0059] It should be noted that the following detailed descriptions are all illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0060] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0061] In the absence of conflict, the embodiments in this application and the features in the embodiments may be combined with each other.

[0062] Embodiment 1

[0063] A multi-model fusion coal production prediction method based on TRIZ conflict resolution, such as Figure 1 As shown, the following steps are included:

[0064] Step 1: Use TRIZ theory to analyze the contradictions of the coal production prediction algorithm: For the coal production prediction method, in order to improve the prediction accuracy, it is usually necessary to increase the complexity of the network, including increasing the number of nodes, the number of layers or using more complex algorithms. This increase in complexity increases resource consumption, forming a contradiction between improved accuracy and increased complexity. Secondly, ensuring the reliability of the prediction results usually requires more data processing and verification steps, which will lead to a longer prediction time, forming a contradiction between reliability and wasted time. Determine the technical parameters: accuracy (42), complexity of the device (45), reliability (35), and waste of time (26). Construct two pairs of contradiction matrices based on the technical contradiction matrix table, as shown in Table 1 and Table 2.

[0065] Table 1 Conflict matrix between accuracy and device complexity

[0066]

[0067] Analyzing Table 1, the invention principle corresponding to "2" is the separation principle, which means separating the useful part (or attribute) or harmful part (or attribute) of the system from the whole system in a virtual or physical way, including extracting the "negative" part or attribute of the object or only extracting the necessary and useful part or attribute of the object; the invention principle corresponding to "16" is the principle of local action or overaction. If it is difficult to obtain 100% of the required effect, a slightly smaller or larger effect should be obtained; the invention principle corresponding to "18" is mechanical vibration, which makes the object vibrate and increases the vibration frequency, or replaces the mechanical vibrator with a piezoelectric vibrator; "26" corresponds to the replication principle, which replaces expensive and fragile objects with simple and cheap replicas; "3" corresponds to the invention principle of local characteristics, which means that different parts of an object should have different functions, and each part of an object should have the most suitable conditions for its work; "4" corresponds to the asymmetry principle, which converts symmetrical forms into asymmetric forms or strengthens their asymmetry; "28" corresponds to the mechanical system substitution principle, which replaces mechanical systems with light, sound, and olfactory systems, replaces mechanical fields with electric and magnetic fields, or replaces static fields with moving fields, replaces constant fields with time-varying fields, and replaces random fields with structured fields.

[0068] While ensuring that the accuracy does not drop, the two invention principles that are more in line with reducing the complexity of the model are 2 (separation principle) and 16 (local action or excess action principle). Principle 2 can be used to first identify and separate the unnecessary complex parts of the model, and then Principle 16 can be used to find opportunities to simplify the remaining parts without sacrificing too much accuracy. This comprehensive application helps to effectively reduce the complexity of the model while maintaining the accuracy of the model. The inspiration is: improve the original Transformer structure and remove the Decoder part that does not play a big role in time series prediction.

[0069] Table 2 Reliability and time waste contradiction matrix

[0070]

[0071] Analyzing Table 2, the invention principle corresponding to "10" is the pre-action principle, which means completing part or all of the work in advance, or placing the object in the most convenient position in advance so that it can be put into use immediately; "2" corresponds to the extraction principle, which means extracting the "interfering" part or characteristics from the system, or only extracting the required part or characteristics; "25" corresponds to the self-service principle, and the object can self-serve and complete auxiliary or maintenance work; "5" corresponds to the combination principle, which means establishing a connection between the functions, characteristics or parts of the system to produce a new and expected result. By combining existing functions, new functions can be generated. Including combining objects of the same type or objects that need to be operated continuously in space or combining parts of the same type or parts that need to be operated continuously in time; "4" corresponds to the asymmetric principle, which converts the symmetric form into an asymmetric form or strengthens its asymmetry; "30" corresponds to the flexible shell or film principle, which uses soft shells and films to replace common structures, or uses them to separate objects from the external environment; "3" corresponds to the local characteristic principle, and different parts of an object should have different functions, and each part of the object should have the conditions that are most suitable for its work.

[0072] Under the premise of ensuring the reliability of the prediction results, by combining the invention principles 5 (combination principle) and 3 (local characteristic principle), an efficient and accurate prediction model can be constructed. First, the functions of multiple algorithms are reasonably utilized using principle 3, and then, multiple algorithms are combined using principle 5, which can significantly reduce the time required for prediction while maintaining or improving the reliability of the prediction results. The inspiration of the two invention principles is that according to the characteristics of time series data, the network is divided into three modules and then combined. Each module implements different functions. For example, CNN is used to extract some deep spatial features, BiGRU is responsible for extracting the time series features of coal production data, and Transformer combines the extracted features with predictions.

[0073] In the time series prediction task, Transformer usually only needs to focus on predicting future values ​​or trends from past time series data. This prediction is usually one-way and does not require reverse prediction from the future to the past. Although the Decoder part is useful in generating sequence tasks (such as machine translation or text generation), this generation capability is not necessary in the time series prediction task. In addition, the Decoder part may introduce additional complexity and computational overhead in the time series prediction task, which may not bring substantial improvement to the prediction task itself. Therefore, according to Invention Principle 2 and Invention Principle 16, the original Transformer is improved, the Decoder part of the Transformer is removed, and it is replaced by a fully connected layer. According to the second pair of contradictions, Invention Principle 5 and Invention Principle 3 are derived. A one-dimensional convolutional neural network is selected to extract the spatial features in the coal production data. The one-dimensional convolutional neural network performs sliding convolution operations on the input data through the convolution kernel, which can capture the spatial features of the time series. This local perception feature makes the model more efficient when processing large-scale data. Through multi-layer convolution and pooling operations, effective features can be automatically extracted from the original time series data. However, CNN cannot fully capture the temporal dependency of coal production data. Therefore, the characteristics of the BiGRU network model are used to extract the temporal features of coal production data. Finally, Transformer combines the extracted features to generate predictions. The advantages of the three network models can complement each other to improve prediction accuracy.

[0074] Step 2: Preprocessing of coal production related data: Collect coal production related data for the past five years, delete useless label fields and features with serious missing values, and use the KNN nearest neighbor method to fill in the values ​​of features with less missing data to form a relatively complete data set. Since the coal production data has a large number of features, in order to avoid irrelevant features affecting the prediction accuracy, the XGBoost method is used for feature screening, and then the data is normalized to return the processed data set;

[0075] Step 3, construction of CNN-BiGRU-Transformer model: This embodiment proposes a hybrid model based on convolutional neural network (CNN)-bidirectional gated recurrent unit neural network (BiGRU) and Transformer network according to the principle of the invention. As a prediction model, the test data set is input into the trained CNN-BiGRU-Transformer model, and the predicted value of future coal production is obtained after feature extraction of CNN and BiGRU layers and feature integration and prediction of Transformer layer. This hybrid model can make more accurate predictions on coal production, and is compared with models such as SVR, XGBoost, LSTM, GRU and Transformer.

[0076] Step 4: Model performance evaluation: In order to scientifically evaluate the effect of the model on coal production prediction, multiple indicators will be selected to evaluate the effect, namely RMSE, MAE, MAPE and R 2 ;

[0077] like Figure 2 As shown, GRU is updated by gate z t and reset gate r t Composed of t Decide to forget and add some information, t It is used to determine the degree of forgetting previous information. Assume that there is an input x at the current time t t , the hidden state h passed down at time t-1 t-1 , which contains relevant information at previous moments. Combined with x t and h t-1 , GRU will get the output y of the hidden state at the current time t t and the hidden state h passed to the next moment t The internal structure formula of GRU is as follows:

[0078] r t =σ(W r ·[h t-1 ,x t ]) (1)

[0079] z t =σ(W z ·[h t-1 ,x t ]) (2)

[0080]

[0081] In formula (1) and formula (2), σ is the sigmoid function, which can transform the data into a value in the range of 0-1, thus acting as a gating signal. t In, W r is the weight matrix, and the weight matrix is ​​used to t and h t-1 The concatenated matrices are multiplied and then r is obtained through the sigmoid function. t , r t The value will be used in the next hidden state As shown in formula (3), r t The smaller the value, the more information needs to be forgotten in the previous moment, and the more information is discarded. The update gate formula is shown in formula (2), and the update memory expression is shown in formula (4). This stage uses the same gate z t Forgetting and remembering at the same time, tThe closer the value is to 1, the more data is memorized.

[0082] like Figure 3 The figure shows the BiGRU network structure diagram. BiGRU combines forward and backward propagation, taking into account past and future time series data at the same time, more effectively capturing dependencies in time series and significantly improving the ability to model time series data. The current hidden state is determined by the output of the forward and backward propagation hidden layers and the input at the current moment.

[0083] The Transformer network mainly relies on the attention mechanism to achieve parallel processing of input sequences and can directly capture global timing information.

[0084] like Figure 4 As shown in the figure, it is a structural diagram of the self-attention mechanism. The self-attention mechanism has a low degree of dependence on external information, and has a better effect in analyzing the characteristics and influencing factors related to coal production, and can better capture the intrinsic meaning of data or features. The self-attention mechanism calculates the correlation between each element in the sequence and all other elements, which is called weight. These weights reflect the relationship between elements.

[0085] In the attention mechanism, query vector Q (Query vector), key vector K (Key vector) and value vector V (Value vector) play an important role. The main task of the query vector is to calculate the attention weight; the key vector and value vector are used to represent different parts of the input sequence respectively. Specifically, by calculating the similarity between the query vector and the key vector, the importance of each key vector to the query vector can be determined, and then these weights are used to perform weighted summation on the value vector to obtain the final vector representation.

[0086] Assume that the input sequence is X = [x 1 ,x 2 ,...,x n ], the weight matrix is ​​W Q ,W k ,W V For each input element x i have:

[0087] Q i =X i ·W Q (5)

[0088] K i =X i ·W K (6)

[0089] V i =X i ·Wv (7)

[0090] In the above formula, Q i ,K i ,V i denote the query, key, and value vectors of the i-th element respectively.

[0091] For every pair of elements X i and X j , calculate an attention score, indicating that X i X j The attention score is calculated by taking the dot product of the query vector and the key vector, i.e., matrix multiplication, and dividing it by a scaling factor (usually the square root of the key vector dimension). The specific calculation is shown in formula (8):

[0092]

[0093] Then, the softmax function (i.e., activation function) is used to convert the attention score into a value between 0 and 1, and the sum is 1, thereby obtaining the attention weight:

[0094]

[0095] As shown in the following formula (10), the value vector of each element is multiplied by its corresponding attention weight, and then summed to obtain the final output Z i , these weights are used to combine the input word vectors to generate a new context-dependent word vector. This word vector contains not only the information of the current word, but also the information of its context, so that the model can better understand the meaning of each word in a specific context.

[0096]

[0097] One of the main drawbacks of the self-attention mechanism is that when encoding specific position information, it tends to over-focus on its own position, resulting in a relatively weak ability to capture effective information. Based on this problem, experts proposed a multi-head attention mechanism as a solution. By introducing multiple independent self-attentions, the model can focus on different positions of the input sequence at the same time, thereby capturing richer and more comprehensive information.

[0098] In some embodiments, a mask operation may be performed before the attention score is converted into a value between 0 and 1 and the sum is 1 through a softmax function (ie, an activation function).

[0099] like Figure 5As shown in Figure 1, it is a network structure diagram of the multi-head attention mechanism. In the multi-head attention mechanism, the input data is first divided into multiple heads. Each head has its own query, key and value vectors. After passing through the linear layer, each head performs self-attention calculation independently to obtain the self-attention output. Finally, all the outputs are spliced ​​together to form the final output, as shown in formula (11).

[0100] Z=W o Concat(Z 1 ,Z 2 ,...,Z h ) (11)

[0101] Among them, W o is the output weight matrix, Concat is the concatenation operation. Assuming there are h heads, the attention output of each head is a d-dimensional vector. After concatenation through the Concat function, an h×d-dimensional vector will be obtained. This vector contains the attention output information of all heads, which allows the model to take into account the information of all heads at the same time and obtain a more comprehensive representation.

[0102] After the concatenation function, it passes through a linear layer and then outputs.

[0103] like Figure 6 As shown in the figure, they correspond to the encoder and decoder structures respectively. The encoder consists of two sublayers, multi-head attention and feedforward neural network, each of which contains an Add&Norm layer; the decoder has one more layer of masked multi-head attention (MaskedMulti-Head Attention) than the encoder, and an Add&Norm layer is also added between each sublayer. The encoder encodes the input sequence into a high-dimensional feature vector representation, and the decoder decodes the vector into the target sequence.

[0104] The Transformer model uses an encoder-decoder structure with an attention mechanism as the core, which realizes parallel processing of input sequences and can directly capture global timing information. The Transformer encoder structure consists of a multi-head attention layer and a feedforward neural network layer, and each sub-layer contains an Add&Norm layer; the decoder structure has one more masked multi-head attention layer than the encoder structure, such as Figure 6 As shown in the figure, Add&Norm layers are also added between each sublayer. The encoder encodes the input sequence into a high-dimensional feature vector representation, and the decoder decodes the vector into the target sequence. In the Transformer network, in order to accelerate the convergence process of the model and improve performance, researchers introduced technologies such as residual connection and normalization.

[0105] The processing methods of the encoder input and decoder input of the Transformer model are the same. In the data preprocessing, a position encoding step is added to provide relative position information. The specific calculation formulas are shown in formulas (12) and (13):

[0106] PE (pos,2i) =sin(pos / 10000 2i / d ) (12)

[0107] PE (pos,2i+1) =cos(pos / 10000 2i / d ) (13)

[0108] In formula (12) and formula (13), pos represents the position of the word in the sentence, and d represents the dimension of PE, where the even dimension is represented by 2i and the cardinality dimension is represented by 2i+1. This allows PE to flexibly adapt to texts longer than all the sentences in the training set, and also helps the model calculate relative positions.

[0109] The encoder module of the Transformer model is composed of 6 encoder layers stacked together. The input data is obtained after data preprocessing. i , after transformation in the form of a matrix, we get Q, K, V, which enter the multi-head attention layer, calculate the attention values ​​respectively, and finally concatenate them.

[0110] Next, the data is passed to the Add&Norm layer, which consists of two parts, Add and Norm. The calculation is shown in formula (14) and formula (15):

[0111] Add&Norm=LayerNorm(X+MultiHeadAttention(X)) (14)

[0112] Add&Norm=LayerNorm(X+FeedForward(X)) (15)

[0113] In the above formula, X represents the input of the multi-head attention or feedforward neural network, and MultiHeadAttention(X) and FeedForward(X) represent the output. Add refers to X+MultiHeadAttention(X). As a residual connection technology, it plays an important role in multi-layer network training. By introducing residual connections, the model can focus on learning the difference between the current layer and the previous layer, which can effectively alleviate the gradient disappearance and representation bottleneck problems in deep network training. Norm is layer normalization, which standardizes the input of each layer of neurons so that the mean and variance of the input data remain consistent. This helps the model better capture the intrinsic characteristics of the data and speeds up the convergence of the model.

[0114] Then enter the Feed-Forward Networks fully connected layer. The formula of the fully connected layer is as follows:

[0115] FFN(X)=max(0,X·W 1 +b 1 )·W 2 +b 2 (16)

[0116] The fully connected layer of Transformer is a two-layer neural network. X represents the output of the Add&Norm layer, and W 1 and b 1 is the parameter of the first fully connected layer, W 2 and b 2 It is the parameter of the second fully connected layer. In the multi-head attention mechanism, matrix multiplication is mainly performed, which is all linear transformation. However, the learning ability of linear transformation is not as strong as that of nonlinear transformation. Since this layer adds the activation function ReLU, it can be used to increase the nonlinearity of the model and enhance the learning ability of the model.

[0117] Finally, after the Add&Norm layer, it enters the next encoder layer. By stacking multiple encoder layers, an encoder module can be formed.

[0118] Similarly, the decoder module is also composed of 6 decoder layers stacked together. The difference from the encoder module is that it has an additional masked multi-head attention. The principle of masked multi-head attention is similar to that of conventional multi-head attention, but a mask is added to mask specific values ​​to avoid the impact of parameter updates.

[0119] In this embodiment, in order to scientifically evaluate the effect of the model on coal production prediction, multiple indicators are selected to evaluate the effect, namely, root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE) and determination coefficient R-Square (R 2 ), the specific calculation formulas are shown in formula (17), formula (18), formula (19) and formula (20):

[0120]

[0121]

[0122] In the above four formulas, n represents the number of coal production samples. i Indicates the actual value of coal production; Represents the predicted value of coal production.

[0123] The root mean square error (RMSE) is the square root of the mean square error. The smaller the value, the closer the predicted value is to the true value. The mean absolute error (MAE) is used to measure the difference between the predicted value and the true value. The mean absolute percentage error (MAPE) can accurately calculate the deviation between the predicted value and the true value. The range of MAE and MAPE is [0, +∞), and the value is negatively correlated with the prediction effect. In addition, the determination coefficient R 2 Represents the correlation between the predicted value and the true value. The closer the value is to 1, the better the prediction effect.

[0124] In this embodiment, the model is optimized by an Adam optimizer with a learning rate of 0.001 to minimize the loss value of the loss function.

[0125] Different evaluation indicators in the data set have different dimensions. In this embodiment, data normalization is used to eliminate the errors caused by different dimensions on the prediction results. The deviation normalization method is used to normalize the data, and the calculation formula is as follows:

[0126]

[0127] Among them, x min Represents the minimum value of the sample data, x max It represents the maximum value. In this embodiment, the MinMaxScaler method provided by the Sklean library in Python can be used to map the data set to the interval [0,1].

[0128] Step 5: Use the trained hybrid model to process the coal production related data in the target time period to obtain the coal production forecast result in the target time period.

[0129] In summary, the multi-model fusion coal production forecasting method based on TRIZ conflict resolution makes full use of the time series characteristics of coal data and improves on the shortcomings of previous coal production forecasting methods, so good forecasting results can be obtained.

[0130] Embodiment 2

[0131] A multi-model fusion coal production prediction system based on conflict resolution, comprising:

[0132] The prediction analysis module is configured to use TRIZ theory to perform contradiction analysis on the coal production prediction algorithm, determine technical parameters, construct a technical contradiction matrix table, and determine the basic structure of the coal production prediction model according to the technical contradiction matrix table;

[0133] The data preprocessing module is configured to obtain data related to coal production within a set time period and perform preprocessing to obtain a training data set;

[0134] The model building and training module is configured to build a prediction model, train the prediction model using the training data set, optimize the parameters of the prediction model by calculating the error value between the expected output value of the prediction model and the predicted output value, until the prediction model converges to obtain the optimal prediction model; the prediction model includes an input layer, a CNN layer, a BiGRU layer, a Transformer layer and an output layer, the input layer is used to receive feature data of coal production related data, the CNN layer is used to extract spatial features of received data, the BiGRU layer is used to extract temporal features of received data, the Transformer layer is used to generate prediction results in combination with the extracted temporal features, and the output layer is used to output the prediction results;

[0135] The prediction module is configured to process the coal production related data in the target time period by using the optimal prediction model to obtain the coal production prediction result in the target time period.

[0136] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0137] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0138] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0139] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0140] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made by those skilled in the art within the spirit and principle of the present invention without creative labor shall be included in the protection scope of the present invention.

Claims

1. A multi-model fusion coal production prediction method based on conflict resolution, characterized in that: The following steps are involved: Using TRIZ theory, we analyze the contradictions of the coal production prediction algorithm, determine the technical parameters, build a technical contradiction matrix, and determine the basic structure of the coal production prediction model based on the technical contradiction matrix. Obtain coal production related data within a set time period and perform preprocessing to obtain a training data set; A prediction model is built, and the prediction model is trained using the training data set. The parameters of the prediction model are optimized by calculating the error value between the expected output value and the predicted output value of the prediction model until the prediction model converges to obtain an optimal prediction model; the prediction model comprises an input layer, a CNN layer, a BiGRU layer, a Transformer layer and an output layer, wherein the input layer is used to receive feature data of coal production-related data, the CNN layer is used to extract spatial features in the coal production data, the BiGRU layer is used to extract time features of the received data, the Transformer layer is used to generate prediction results in combination with the extracted time features, and the output layer is used to output the prediction results; The optimal prediction model is used to process the coal production related data in the target time period to obtain the coal production prediction result in the target time period.

2. The multi-model fusion coal production prediction method based on conflict resolution as claimed in claim 1 is characterized in that: The process of conducting contradiction analysis on the coal production prediction algorithm, determining technical parameters, and constructing a technical contradiction matrix table includes: determining technical parameters including accuracy, device complexity, reliability and time waste, constructing a contradiction matrix table of accuracy and device complexity, and a contradiction matrix table of reliability and time waste.

3. The method for predicting coal production based on multi-model fusion and conflict resolution as claimed in claim 1, characterized in that: The preprocessing process includes deleting useless tag fields and feature data with serious missing values, using the KNN nearest neighbor method to fill in the values ​​of feature data with less missing data, and then using the XGBoost method to screen the feature data, and normalizing the screened data.

4. The method for predicting coal production based on multi-model fusion and conflict resolution as claimed in claim 1, characterized in that: The CNN layer includes a convolution layer and a pooling layer. The input layer receives coal production-related features, first passes through the convolution layer, learns the spatial structure information in the data through multiple convolution kernels, extracts local features in the data, and then reduces the number of parameters and calculations through the pooling layer.

5. The method for predicting coal production based on multi-model fusion and conflict resolution as claimed in claim 1, characterized in that: The BiGRU layer consists of two independent GRU units, one processes data in the forward direction of the time series, and the other processes data in the reverse direction of the time series. Through the bidirectional structure, the BiGRU layer captures both the forward and backward information of the sequence data, so as to better understand and predict the patterns in the sequence. The BiGRU model generates the final output by combining the hidden states in the forward GRU and reverse GRU directions. Or further, the internal structure formula of the GRU layer is as follows: r t =σ(W r ·[h t-1 ,x t ]); z t =σ(W z ·[h t-1 ,x t ]); Among them, σ is the sigmoid function, which transforms the data into a value in the range of 0-1, thus acting as a gating signal; t In, W r is the weight matrix, and the weight matrix is ​​used to input x t and hidden state h t-1 The concatenated matrices are multiplied and then r is obtained through the sigmoid function. t Value, r t The smaller the value, the more information needs to be forgotten in the previous moment, and the more information is discarded; update gate z t The closer the value is to 1, the more data is memorized.

6. The method for predicting coal production based on multi-model fusion and conflict resolution as claimed in claim 1, characterized in that: The Transformer layer uses an encoder-decoder structure to process the input sequence in parallel and capture global temporal information. The encoder structure consists of a multi-head attention layer and a feedforward neural network layer, and each sublayer contains a residual connection and a layer normalization layer. The decoder structure has one more layer of masked multi-head attention than the encoder structure, and a residual connection and a layer normalization layer are also added between each sublayer. The encoder structure encodes the input sequence into a high-dimensional feature vector representation, and the decoder structure decodes the vector into the target sequence.

7. A multi-model fusion coal production prediction method based on conflict resolution as claimed in claim 6, characterized in that: The encoder structure is composed of 6 encoder layers stacked together. The input data of each encoder layer is obtained after data preprocessing. i , after transformation in the form of a matrix, the query vector, key vector and value vector are obtained, which enter the multi-head attention layer, calculate the attention values ​​respectively, and finally concatenate them; Next, the data is passed into the residual connection and layer normalization layers, and then into the fully connected layer. Finally, it passes through the residual connection and layer normalization layers again and enters the next encoder layer. Each encoder layer is superimposed. Or further, the calculation process of the residual connection and the layer normalization layer includes: Add&Norm=LayerNorm(X+MultiHeadAttention(X)); Add&Norm=LayerNorm(X+FeedForward(X)); Among them, X represents the input of the multi-head attention or feedforward neural network, MultiHeadAttention(X) and FeedForward(X) represent the output; Add refers to X+MultiHeadAttention(X). As a residual connection technology, Norm is used for layer normalization. By standardizing the input of each layer of neurons, the mean and variance of the input data remain consistent.

8. The method for predicting coal production based on multi-model fusion and conflict resolution as claimed in claim 1, characterized in that: The calculation of the fully connected layer includes: FFN(X)=max(0,X·W1+b1)·W2+b2; The fully connected layer is a two-layer neural network, X represents the output of the Add&Norm layer, W1 and b1 are the parameters of the first fully connected layer, and W2 and b2 are the parameters of the second fully connected layer.

9. The method for predicting coal production based on multi-model fusion and conflict resolution according to claim 1, characterized in that: When training the prediction model, the root mean square error, mean absolute error, mean absolute percentage error and coefficient of determination are used to evaluate the prediction value of the prediction model; The Adam optimizer is used to optimize the prediction model to minimize the loss value of the loss function of the prediction model.

10. A multi-model fusion coal production prediction system based on conflict resolution, characterized by comprising: The prediction analysis module is configured to use TRIZ theory to perform contradiction analysis on the coal production prediction algorithm, determine technical parameters, construct a technical contradiction matrix table, and determine the basic structure of the coal production prediction model according to the technical contradiction matrix table; The data preprocessing module is configured to obtain data related to coal production within a set time period and perform preprocessing to obtain a training data set; The model building and training module is configured to build a prediction model, train the prediction model using the training data set, optimize the parameters of the prediction model by calculating the error value between the expected output value and the predicted output value of the prediction model, until the prediction model converges to obtain the optimal prediction model; the prediction model includes an input layer, a CNN layer, a BiGRU layer, a Transformer layer and an output layer, the input layer is used to receive feature data of coal production related data, the CNN layer is used to extract spatial features of received data, the BiGRU layer is used to extract temporal features of received data, the Transformer layer is used to generate prediction results in combination with the extracted temporal features, and the output layer is used to output the prediction results; The prediction module is configured to use the optimal prediction model to process the coal production related data in the target time period to obtain the coal production prediction result in the target time period.

Citation Information

Cited By

  • Converter-based coke quality prediction model for samples containing missing coal parameters

    CN120297823A

  • Synthesis gas component content prediction method based on multi-algorithm fusion modeling

    CN122245487A