Converter-based coke quality prediction model for samples containing missing coal parameters
Through the missing mask, trainable placeholder embedding vector mechanism and transformer encoder, the problem that the coke mass prediction model cannot handle missing coal parameters is solved, and efficient prediction of missing coal parameters samples is achieved, which improves the applicability and prediction performance of the model.
Patent Information
- Application Number
- CN202510790119.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The existing coke quality prediction model cannot process samples containing missing coal parameters, resulting in poor applicability of the model in actual applications and hinders the development process of the global unified coke quality prediction model.
The missing mask and trainable placeholder embedding vector mechanism are used, and combined with the transformer encoder and multi-layer perceptron, a coke quality prediction model is constructed. The missing features are identified through the missing mask and the trainingable placeholder embedding vector is used to retain the missing structure information. The high-dimensional features are extracted using the transformer encoder, and finally prediction is performed through the multi-layer perceptron.
Effectively process samples containing missing coal parameters to improve the applicability and prediction performance of the model in practical applications, and significantly improve the accuracy and stability of coke quality prediction.
Smart Images

Figure CN120297823A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of coking, artificial intelligence, and data mining and processing, and particularly relates to a converter-based coke quality prediction model for samples with missing coal parameters. Background Art
[0002] In the past few decades, numerous coke quality prediction models have been successively proposed, ranging from traditional statistical regression algorithms to the currently popular artificial neural networks. These methods have made remarkable progress in improving the prediction performance of the models. However, existing models still face a key challenge in practical applications: when there are missing values in the parameters of newly input coal sample data and they cannot match the complete input parameters required by the training model, the model cannot make predictions. Due to differences in the coal sample property test procedures adopted by different coking plants and the complexity and time costs of various analysis methods, the integrity of coal sample parameters varies, and the existence of a large number of missing values makes it difficult for existing models to make predictions. This limitation not only severely restricts the application effect of the model in actual production but also hinders the development process of a globally unified coke quality prediction model.
[0003] Traditional artificial neural network-based coke quality prediction models can effectively extract the complex non-linear relationship between coal parameters and coke quality. However, due to the fixed number of neurons in the input layer, it is required that the number of features of all input coal sample data must be the same and cannot contain any missing values, otherwise matrix operations and gradient propagation cannot be performed. Therefore, how to retain the powerful feature extraction ability of the neural network while enabling it to process coal parameter samples with any number of missing values, thereby improving the applicability of the model and accurately and stably predicting coke quality is the core issue in the intelligentization of coal blending technology. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a converter-based coke quality prediction model for samples with missing coal parameters in view of the problem that the current coke quality prediction model cannot process samples with missing coal parameters due to fixed input features.
[0005] To achieve the above technical objectives and reach the above technical effects, the present invention is realized through the following technical solutions: A converter-based coke quality prediction model for samples with missing coal parameters, the model includes the following steps: Step S1: Construction of a sample data set containing missing values of coal parameters; Step S2: Embedding mapping of non-missing coal sample parameters, and simultaneously constructing a missing mask and a missing placeholder embedding to represent missing features; Step S3: Construction of a converter encoder for extracting high-dimensional features; Step S4: Perform mean pooling on the extracted features and output the predicted coke quality through a multi-layer perceptron. Further, in the step S1, the sample data set includes multi-dimensional coal parameters with missing values and corresponding coke quality values, where the coke quality values include the reactivity CRI and the post-reaction strength CSR.
[0006] Further, in the step S2, the missing mask length is equal to the feature dimension of each coal sample, used to identify whether each parameter is missing. Where a mask value of 0 indicates that the parameter is missing, and a value of 1 indicates that the parameter exists; for the non-missing coal sample parameters, by inputting the original numerical value into a set of trainable linear transformation layers, perform a projection operation from a scalar value to a fixed-dimensional embedding space. The linear transformation layer includes a fully connected network structure composed of a weight matrix and a bias vector, used to map a one-dimensional input scalar to a d-dimensional embedding vector, where d is a preset embedding dimension, and the weight and bias are learnable parameters, and are automatically optimized through the backpropagation algorithm during the model training process.
[0007] Further, in the step S2, the missing placeholder embedding is replaced by a trainable placeholder embedding vector E_miss with a unified dimension. E_miss is automatically learned through error backpropagation during the model training process, used to uniformly represent the embedding form of the missing features in the sample, so as to retain the missing structure information without introducing artificial filling values.
[0008] Further, in the step S3, the transformer encoder includes multiple stacked encoder layers, each layer including a multi-head self-attention mechanism, a feed-forward neural network module, a residual connection structure, and a layer normalization module, used to capture the high-order interaction relationships and context semantics between the coal sample features.
[0009] Further, the multi-head self-attention mechanism models the influence degree between different coal sample parameters in the sequence by mapping the input embedding representation into query Q, key K, and value V vectors respectively and calculating the attention distribution.
[0010] Further, the feed-forward neural network module is composed of multiple hidden layers and a single output layer. The hidden layer is followed by a non-linear activation function, used to perform per-position feature expansion and non-linear modeling on each embedding vector.
[0011] Further, the residual connection structure is used to perform an element-wise addition operation on the input and output of each sub-module containing the attention mechanism module and the feed-forward neural network module, thereby effectively alleviating the problems of gradient disappearance and feature degradation during the deep stacking process of the network, improving the stability and convergence performance of the model. On this basis, a layer normalization mechanism is introduced to perform normalization processing on each sample in its feature dimension, making its mean zero and variance one, which helps to maintain the stability of the feature distribution and enhance the expressive ability and training efficiency of the deep network.
[0012] Further, in step S4, the mean pooling process is to calculate the average of the variable-length embedded sequence features output by the transformer encoder in each dimension to obtain a fixed-length coal sample representation vector with a unified dimension.
[0013] Further, the multi-layer perceptron consists of an input layer, one or more hidden layers, and an output layer. The number of neurons in the output layer is two. One neuron outputs CRI, and the other neuron outputs CSR.
[0014] The beneficial effects of the present invention are as follows: 1) The method proposed in the present invention that adopts the missing mask and trainable placeholder embedding vector mechanism can retain the missing structure information in the data without explicit interpolation or sample rejection, and convert the coal sample parameters with different missing patterns into a vector sequence with consistent structure, which is input into the transformer encoder for medium and high-dimensional feature extraction, and then the coke quality prediction is realized through a multi-layer perceptron. This model can not only effectively deal with the problem of missing coal sample parameters, realize the coke quality prediction of incomplete samples, significantly improve the applicability of the model in practical applications, but also make full use of the effective information in the missing samples to further improve the overall prediction performance.
[0015] 2) The transformer encoder model proposed in the present invention is stacked by the multi-head self-attention mechanism and the feed-forward neural network module, which can capture the high-order interaction relationship and context-dependent structure between coal sample features. Compared with the traditional feature-independent processing method, the transformer structure can significantly improve the modeling ability for non-linear and non-stationary data, thereby enhancing the prediction performance and generalization ability of the coke quality prediction model. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a block diagram of the prediction model of the present invention; Figure 2 It is a flowchart of the prediction model of the present invention; Figure 3 It is a schematic diagram of the transformer encoder model in the prediction model of the present invention; Figure 4 It is a schematic diagram of the multi-layer perceptron model in the prediction model of the present invention; Figure 5 Schematic diagram of comparison between predicted values and true values of coke CRI and CSR indexes of the prediction model of the present invention. Specific implementation manners
[0017] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0018] As Figures 1 to 4As shown in the figure, a transformer-based coke quality prediction model for samples with missing coal parameters, the prediction model includes the following steps: constructing a sample data set containing missing values of coal parameters; performing embedding mapping on non-missing coal sample parameters, and at the same time constructing a missing mask and a missing placeholder embedding to represent missing features; constructing a transformer encoder to extract high-dimensional features; performing mean pooling processing on the extracted features and outputting the predicted coke quality result through a multi-layer perceptron. First, perform embedding mapping processing on the coal sample parameters with missing values and inconsistent parameter numbers in the input blended coal sample data. For this purpose, a missing mask vector of the same length as the feature dimension of each coal sample is constructed to identify whether each parameter is missing, where the mask value of 0 indicates that the parameter is missing, and the mask value of 1 indicates that the parameter exists. For non-missing coal sample parameters, by inputting the original numerical value into a group of trainable linear transformation layers, perform the projection operation from the scalar value to the fixed-dimensional embedding space. The linear transformation layer consists of a weight matrix and a bias vector, constituting a fully connected network structure, which can map a one-dimensional input scalar to a d-dimensional embedding vector, where d is the preset embedding dimension. The weight and bias are learnable parameters and can be automatically optimized through error backpropagation during model training. For missing coal sample parameters, a trainable placeholder embedding vector E_miss with a unified dimension is used for replacement. The E_miss is also one of the model parameters and can be obtained through backpropagation learning during training. It is used to uniformly represent the embedding form of missing features in the sample, so as to retain the missing structure information without introducing artificial filling values. Then, input the above embedding vector sequence into multiple stacked encoder layers that have been trained. The encoder layer is a transformer encoder structure, which consists of a multi-head self-attention mechanism, a feed-forward neural network module, a residual connection structure, and a layer normalization module, and is used to capture the high-order interaction relationships and context semantics between coal sample features. Among them, the multi-head self-attention mechanism models the influence degree between different coal sample parameters in the sequence by mapping the embedding representation into query Q, key K, and value V vectors respectively and calculating the attention distribution. The feed-forward neural network module consists of multiple hidden layers and a single output layer, and a non-linear activation function (such as ReLU or GELU) is connected after the hidden layer. The residual connection structure is used to perform an element-wise addition operation on the input and output of each sub-module (including the attention mechanism module and the feed-forward neural network module). And on this basis, a layer normalization mechanism is introduced to normalize each sample in its feature dimension so that its mean is zero and its variance is one. Perform mean pooling processing on the high-dimensional embedding feature sequence extracted by the transformer encoder, that is, calculate the average of the variable-length embedding sequence features output by the transformer encoder according to the dimension to obtain a fixed-length coal sample representation vector with a unified dimension. Finally, input the fixed-length representation vector into a trained multi-layer perceptron consisting of an input layer, one or more hidden layers, and an output layer, and output the predicted coke CRI and CSR indicators.
[0019] Example 1: Collect 195 industrial production data, and perform embedding mapping processing on the blended coal data parameters represented by numerical values with missing values and different coal parameters (moisture, ash, volatile matter, sulfur content, G value, X value, and Y value). To this end, construct a missing mask vector of the same length as the characteristic dimension of each coal sample to identify whether each parameter is missing. The mask value of 0 indicates that the parameter is missing, and the mask value of 1 indicates that the parameter exists. For non-missing coal sample parameters, by inputting the original numerical value into a group of trainable linear transformation layers, perform the projection operation from the scalar value to the fixed-dimensional embedding space. The linear transformation layer consists of a weight matrix and a bias vector, forming a fully connected network structure, and maps the one-dimensional input scalar to a 128-dimensional embedding vector. The weight and bias are learnable parameters and can be automatically optimized through error backpropagation during model training. For missing coal sample parameters, a trainable placeholder embedding vector E_miss of a unified dimension is used for substitution. E_miss is also one of the model parameters and can be learned through backpropagation during training. It is used to uniformly represent the embedding form of the missing features in the sample, so as to retain the missing structure information without introducing artificial filling values. Then, input the above embedding vector sequence into three trained stacked encoder layers. The encoder layer is a transformer encoder structure, which consists of a multi-head self-attention mechanism, a feed-forward neural network module, a residual connection structure, and a layer normalization module, and is used to capture the high-order interaction relationships and context semantics between coal sample features. Among them, the multi-head self-attention mechanism models the influence degree between different coal sample parameters in the sequence by mapping the embedding representation into query Q, key K, and value V vectors respectively and calculating the attention distribution. The feed-forward neural network module consists of three hidden layers and a single output layer, and the hidden layer is followed by a non-linear activation function (ReLU). The residual connection structure is used to perform an element-wise addition operation on the input and output of each sub-module (including the attention mechanism module and the feed-forward neural network module). And on this basis, a layer normalization mechanism is introduced to normalize each sample in its feature dimension so that its mean is zero and its variance is one. Perform mean pooling on the high-dimensional embedding feature sequence extracted by the transformer encoder, that is, average the variable-length embedding sequence features output by the transformer encoder by dimension to obtain a fixed-length coal sample representation vector of a unified dimension. Finally, input the fixed-length representation vector into a trained multi-layer perceptron consisting of an input layer, three hidden layers, and an output layer to output the predicted coke CRI and CSR indexes. Figure 5For the coke thermal quality parameters predicted by using the model of the present invention and the coke thermal quality parameters produced by actual industrial coke ovens, as can be seen from the figure, the prediction model established by using the present invention can successfully predict the sample data with missing values of coal parameters. The average absolute error values of the predicted CRI and CSR indexes of coke and the coke thermal quality parameters in industrial production are 3.88, and the prediction results are very close to the true values.
[0020] Comparative Example 1: Select 195 industrial production data used in Example 1, and input the blended coal data represented by numerical values with missing values with different coal parameter quantities (moisture, ash, volatile matter, sulfur content, G value, X value and Y value) into an artificial neural network model containing three fully connected layers. The last layer of the fully connected layer contains two neurons, corresponding to CRI and CSR respectively. Since the number of neurons in the input layer of the artificial neural network model is fixed, it is required that the number of feature quantities of all input coal sample data must be consistent and cannot contain any missing values. Therefore, it is impossible to predict the CRI and CSR indexes of coke for this industrial production data.
[0021] In summary, by using the method provided by the present invention for embedding and mapping the numerical parameters in the coal sample, and combining the missing information modeling mechanism and the transformer encoder structure for feature extraction, it is possible to realize the prediction of the coke thermal quality of the sample data with missing values of coal parameters, significantly improve the applicability of the model in actual production, and at the same time promote the development process of the global unified coke quality prediction model.
[0022] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A converter-based coke quality prediction model for samples with missing coal parameters, characterized in that The model includes the following steps: Step S1: Construct a sample data set containing missing values of coal parameters; Step S2: Perform embedding mapping on the non-missing coal sample parameters, and at the same time construct a missing mask and a missing placeholder embedding to represent the missing features; Step S3: Construct a Transformer encoder to extract high-dimensional features; Step S4: Perform mean pooling on the extracted features, and output the predicted coke quality through a multi-layer perceptron.
2. The coke quality prediction model based on a transducer for samples with missing coal parameters according to claim 1, characterized in that In the said Step S1, the sample data set contains multi-dimensional coal parameters with missing values and the corresponding coke quality values, where the coke quality values include the reactivity CRI and the post-reaction strength CSR.
3. The coke quality prediction model based on a transducer for samples with missing coal parameters according to claim 1, characterized in that In the said Step S2, the length of the missing mask is equal to the feature dimension of each coal sample, used to identify whether each parameter is missing. Among them, a mask value of 0 indicates that the parameter is missing, and a value of 1 indicates that the parameter exists; for the non-missing coal sample parameters, by inputting the original numerical value into a group of trainable linear transformation layers, a projection operation from a scalar value to a fixed-dimensional embedding space is performed. The linear transformation layer includes a fully connected network structure composed of a weight matrix and a bias vector, used to map a one-dimensional input scalar to a d-dimensional embedding vector, where d is a preset embedding dimension, and the weight and the bias are learnable parameters, and are automatically optimized through the backpropagation algorithm during the model training process.
4. The coke quality prediction model based on a transducer for samples with missing coal parameters according to claim 1, wherein In the said Step S2, the missing placeholder embedding is replaced by a trainable placeholder embedding vector E_miss with a unified dimension. E_miss is automatically learned through error backpropagation during the model training process, used to uniformly represent the embedding form of the missing features in the sample, so as to retain the missing structure information without introducing artificial filling values.
5. The coke quality prediction model based on a transducer for samples with missing coal parameters according to claim 1, wherein In the said Step S3, the Transformer encoder includes multiple stacked encoder layers, and each layer includes a multi-head self-attention mechanism, a feed-forward neural network module, a residual connection structure and a layer normalization module, used to capture the high-order interaction relationships and context semantics between the coal sample features.
6. The coke quality prediction model based on a transducer for samples with missing coal parameters according to claim 5, characterized in that The said multi-head self-attention mechanism models the influence degree between different coal sample parameters in the sequence by mapping the input embedding representation into query Q, key K and value V vectors respectively, and calculating the attention distribution.
7. The coke quality prediction model based on a transducer for samples with missing coal parameters according to claim 5, characterized in that The said feed-forward neural network module consists of multiple hidden layers and a single output layer, and a non-linear activation function is connected after the hidden layer, used to perform per-position feature expansion and non-linear modeling on each embedding vector.
8. The coke quality prediction model based on a transducer for samples with missing coal parameters according to claim 5, characterized in that, The said residual connection structure is used to perform an element-wise addition operation on the input and output of each sub-module, so as to effectively alleviate the problems of gradient disappearance and feature degradation during the deep stacking process of the network, improve the stability and convergence performance of the model, and introduce a layer normalization mechanism on this basis, used to normalize each sample in its feature dimension, so that its mean is zero and its variance is one, to maintain the stability of the feature distribution and enhance the expression ability and training efficiency of the deep network.
9. The coke quality prediction model based on a transducer for coal parameter samples with missing values according to claim 1, wherein In the said Step S4, the mean pooling process is to average the variable-length embedding sequence features output by the Transformer encoder in dimension to obtain a fixed-length coal sample representation vector with a unified dimension.
10. The coke quality prediction model based on a transducer for samples with missing coal parameters according to claim 9, characterized in that, The multi-layer perceptron consists of an input layer, one or more hidden layers, and an output layer. The number of neurons in the output layer is two. One neuron outputs CRI, and the other neuron outputs CSR.
Citation Information
Patent Citations
Coke quality prediction method and system based on machine learning
CN114841460A
Coke quality prediction method and system based on BP neural network
CN115186914A
Coke quality prediction method based on support vector machine and application
CN115619022A
Convolutional neural network coke thermal state quality prediction method based on coal data imaging
CN116504329A
Domain adaptive training method for improving applicability of coke thermal state quality prediction model
CN117669395A