Density Decision Method for Heavy Medium Separation Guided by the Fusion of Deep Learning and Physical Models
Through the method of integrating deep learning and physical model guidance, a deep neural network prediction model based on attention mechanism is constructed, which solves the real-time and stability problems of sorting density prediction in the existing technology, and achieves higher prediction accuracy and model generalization capabilities.
Patent Information
- Application Number
- CN202411497946.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-10-25
AI Technical Summary
The prior art is difficult to achieve real-time online prediction of sorting density during coal sorting and processing, and data-driven machine learning models lack stability and reliability in the face of unknown or extreme situations, and cannot ensure the physical rationality of the predicted results.
The method of fusion of deep learning and physical model guidance is adopted to improve the prediction accuracy of sorting density and the generalization ability of the model through data acquisition and preprocessing, construct theoretical mapping models, construct deep neural network prediction models based on attention mechanisms, and define composite loss functions.
It significantly improves the prediction accuracy of sorting density, enhances the generalization ability and interpretability of the model, can provide reliable prediction results under complex operating conditions, and improves the stability and efficiency of the coal preparation process.
Smart Images

Figure CN119049595B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and particularly relates to a heavy medium separation density decision-making method integrating deep learning and physical model guidance. Background Art
[0002] In the process of coal separation and processing, the separation density is a key index for evaluating coal separation. The separation density directly affects the quality of clean coal and the rational utilization of resources: too high a separation density may lead to an increase in the impurity content in clean coal, reducing the economic value of coal; while too low a separation density may cause waste of high-quality coal resources. Therefore, accurately predicting the separation density is of great significance for the efficient separation and processing of coal.
[0003] Traditionally, the prediction methods for separation density include empirical method, experimental method and numerical simulation method. Although these methods can provide high-precision measurement results, the process is complex, time-consuming, and cannot meet the requirements of real-time online prediction in modern coal processing systems, resulting in measurement lag and affecting rapid decision-making in actual production.
[0004] In recent years, the academic and industrial communities have begun to pay attention to using machine learning technology to improve the accuracy of separation density prediction. These methods analyze a large amount of historical data of coal preparation processes and use machine learning algorithms for prediction in order to improve the prediction accuracy. However, these data-driven machine learning models also have some problems. They rely too much on the quality of data and cannot ensure the stability and reliability of the models when facing unknown or extreme situations. In addition, these models may not be able to ensure the physical rationality of the prediction results. In fact, the coal preparation process is only one of the multiple factors affecting the ash content of clean coal products, and the quality of raw coal is the key factor determining the quality of clean coal products. According to coal preparation theory, there is a theoretical connection between the quality of raw coal and the separation density. The present invention is based on this theoretical connection to guide the data-driven model to achieve more stable and accurate separation density prediction. Summary of the Invention
[0005] In view of the above technical problems existing in the prior art, the present invention proposes a heavy medium separation density decision-making method integrating deep learning and physical model guidance, which is reasonably designed, overcomes the deficiencies of the prior art, and has good effects.
[0006] The technical solution of the present invention is as follows:
[0007] A heavy medium separation density decision-making method integrating deep learning and physical model guidance, comprising the following steps:
[0008] Step 1, data collection and preprocessing, and constructing a sample data set;
[0009] Step 2: Construct a theoretical mapping model and calculate the theoretical separation density;
[0010] Step 3: Construct a deep neural network prediction model based on the attention mechanism for predicting the separation density;
[0011] Step 4: Define a composite loss function;
[0012] Step 5: Divide the samples constructed in Step 1 into a training set and a test set according to a ratio, and perform model training based on the training set;
[0013] Step 6: Model evaluation and application.
[0014] Further, the specific process of Step 1 is as follows:
[0015] First, comprehensively collect the key production process parameters of the heavy medium separation process and the raw coal float-sink data through the production monitoring system of the coal preparation plant; the key production process parameters of the heavy medium separation process specifically include the sampling date, sampling time, magnetic substance content, slime content, separation density, clean coal ash content, cyclone inlet pressure, and combined medium tank liquid level; the raw coal float-sink data specifically includes the sampling date, sampling time, density grade yield, and density grade ash rate;
[0016] Subsequently, perform standardization and missing value processing on the collected data.
[0017] Further, in Step 2, according to the raw coal float-sink experiment data, construct a theoretical mapping relationship between the separation density and the clean coal ash content; first, determine the discrete mapping relationship between the two, and then through the interpolation method, smooth the off-line points to establish a continuous value theoretical mapping model between the separation density and the clean coal ash content; the specific discrete mapping relationship is as follows:
[0018] (1);
[0019] Wherein, is the density grade interval with the density grade serial number of , is an integer and ; the separation density is the upper bound of ; is the clean coal ash content corresponding to ; is the yield of the density grade area with the serial number of ; is the density grade ash content of the density grade area with the serial number of ; is the density grade area serial number index; is the maximum value function;
[0020] Using the continuous value theory mapping model, the theoretical separation density value under the set target clean coal ash content is calculated, and the theoretical separation density value is one of the inputs for the subsequent deep neural network prediction model.
[0021] Furthermore, in the said step 3, the deep neural network prediction model based on the attention mechanism includes a one-dimensional convolutional neural network layer, a bidirectional long short-term memory network layer, and an attention mechanism layer; the inputs of the model include the data after preprocessing in step 1 and the theoretical separation density value calculated in step 2, and the output of the model is the separation density.
[0022] Among them, the input of the one-dimensional convolutional neural network layer is the key coal preparation production process parameters at different times. By performing convolutional operations on the local areas of the key coal preparation production process parameter data at different times, the corresponding feature maps are output, and the effective short-term features on the key production process parameters and historical data are extracted; the calculation formula for one-dimensional convolution is:
[0023] (2);
[0024] Among them, represents the output value after convolution; are different elements in the input time series; , , are the weights of different convolutional kernels respectively;
[0025] The bidirectional long short-term memory network layer is composed of two long short-term memory networks with opposite directions, namely the forward layer and the backward layer, and is used to extract long-term dependence features in the time dimension; the bidirectional long short-term memory network layer includes an input layer, a forward layer, a backward layer, and an output layer; the bidirectional long short-term memory network layer processes the input data from the forward and backward directions of the sequence respectively, and generates a forward hidden state sequence and a backward hidden state sequence respectively; for each time step, the forward and backward hidden states are concatenated to generate the final hidden state, which is also used as the input of the attention mechanism layer.
[0026] The forward process is as follows:
[0027] (3);
[0028] (4);
[0029] (5);
[0030] (6);
[0031] (7);
[0032] (8);
[0033] Among them, is the state of the forward candidate cell at a certain moment; and are respectively the moment, the state of the forward storage cell at a certain moment; and and are respectively the forget gate, input gate, and output gate at a certain moment during the forward process; and and are respectively the biases corresponding to the forget gate, input gate, and output gate during the forward process; is the input state bias during the forward process; and and are respectively the weights of the forget gate, input gate, and output gate during the forward process; is the input state weight during the forward process; is the sigmoid function; is the tanh activation function; and are respectively the moment and the output values of the forward storage cell at a certain moment; is the input value at a certain moment;
[0034] The reverse process is:
[0035] (9);
[0036] (10);
[0037] (11);
[0038] (12);
[0039] (13);
[0040] (14);
[0041] Among them, is the state of the reverse candidate cell at a certain moment; and are respectively the moment, the state of the reverse storage cell at a certain moment; and , are respectively the forget gate, input gate, and output gate at the reverse process moment; , , are respectively the biases corresponding to the forget gate, input gate, and output gate of the reverse process; is the input state bias of the reverse process; , , are respectively the weights of the forget gate, input gate, and output gate of the reverse process; is the input state weight of the reverse process; , are respectively moment and the output values of the reverse storage unit at the moment;
[0042] Concatenate and to obtain the output vector of the bidirectional long short-term memory network layer at the moment;
[0043] During the calculation process of the attention mechanism layer, the hidden representation after dimensional transformation is :
[0044] (15);
[0045] Among them, is the weight parameter of the attention mechanism layer; is the bias parameter of the attention mechanism layer;
[0046] By calculating the similarity between and the context vector to measure the importance of , is randomly initialized and obtained through learning during the training process. Then, the following softmax function is used to normalize the importance of each input. The specific formula is:
[0047] (16);
[0048] Among them, is the normalized attention value of , is the output vector of the bidirectional long short-term memory network layer at the moment; is the hidden representation of after dimensional transformation; is the exponential function with base e; The summation variable index for the time instant;
[0049] Finally, the output of the attention mechanism layer is obtained by weighted averaging all the outputs of the bidirectional long short-term memory network layer:
[0050] (17);
[0051] The output of the attention mechanism layer undergoes dimensional transformation through a fully connected layer, and finally a one-dimensional sorting density is obtained.
[0052] Furthermore, in the step 4, the composite loss function fuses two parts: the empirical loss and the physical loss;
[0053] For the empirical loss part, the mean square error is selected as the measurement criterion to measure the difference between the predicted value and the actual value of the sorting density. The calculation formula is:
[0054] (18);
[0055] where is the mean square error value; is the total number of samples; represents the actual value of the sorting density of the th sample; is the predicted value of the sorting density of the th sample;
[0056] For the physical loss part, additional constraints based on the physical process of coal preparation are introduced; by pertinently disturbing the process parameters and raw coal float-sink data in the sample except for the sorting density, and comparing the prediction results before and after the disturbance, the physical loss is established to reflect the physical influence of the disturbed parameters on the sorting density; the physical loss part can evaluate the actual influence of the disturbance on the predicted value of the sorting density. The physical loss function is defined as:
[0057] (19);
[0058] where is the physical loss function; and are the sorting density values predicted by the deep neural network based on the attention mechanism for the input data before and after the disturbance respectively; is the ReLU function;
[0059] Finally, the composite loss function of the deep neural network prediction model based on the attention mechanism is set as the weighted sum of the empirical loss and the physical loss.
[0060] Further, in step 5, during the training process, the empirical loss and the physical loss are optimized simultaneously; the model hyperparameters are set as follows: the learning rate is 0.001; the optimizer is the Adam algorithm; the batch size is 32; the number of training epochs is 50; meanwhile, an early stopping mechanism is introduced to monitor the training process, and once the model performance no longer improves, the training is stopped to prevent overfitting of the model.
[0061] Further, in step 6, after the model training is completed, the trained model is evaluated using the test set. By comparing the differences between the predicted values and the actual values of the separation density, key indicators such as the mean square error, relative absolute error, and relative absolute error are calculated to measure the performance of the model; meanwhile, the differences between the predicted values and the actual values are visually displayed through charts; after the evaluation is completed, the optimal model during the evaluation process is selected and applied to actual production to achieve the prediction of the separation density and assist engineers in making decisions.
[0062] The beneficial technical effects brought by the present invention are as follows.
[0063] 1. By combining the machine learning model of the neural network with the theoretical model guided by physics, the present invention significantly improves the prediction accuracy of the separation density. This method can not only more accurately predict the separation density under complex working conditions, but also enhance the generalization ability of the model by introducing the prior knowledge of the physical model, enabling it to better adapt to different production conditions. Such optimization helps to improve the stability and efficiency of the coal preparation process. Especially in the case of scarce data or large changes in working conditions, the model can still provide reliable prediction results.
[0064] 2. By introducing the theoretical predicted values and the physical loss, the present invention constructs a composite loss function, making the prediction results of the model not only based on data but also follow physical laws. This method improves the interpretability of the model, making the prediction process more transparent and facilitating engineers to understand and trust the prediction basis of the model. In actual operation, this transparency and interpretability are crucial for engineers to make decisions and adjustments, and contribute to enhancing the intelligent level and operation accuracy of the entire coal preparation process. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 is the overall architecture diagram of the heavy medium separation density decision-making method integrating deep learning and physical model guidance of the present invention.
[0066] Figure 2 is the structural diagram of the deep neural network prediction model based on the attention mechanism of the present invention.
[0067] Figure 3 is the schematic diagram of the one-dimensional convolution process in the model of the present invention.
[0068] Figure 4 is the architecture diagram of the LSTM network in the model of the present invention.
[0069] Figure 5 It is the diagram of the bidirectional LSTM network architecture in the model of the present invention. Specific implementation manners
[0070] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners:
[0071] A density decision-making method for heavy medium separation integrating deep learning and physical model guidance, the process of which is as Figure 1 shown, and includes the following steps:
[0072] Step 1: Data collection and preprocessing to construct a sample data set; the specific process is as follows:
[0073] First, comprehensively collect the key production process parameters of the heavy medium separation process and the raw coal float-sink data through the production monitoring system of the coal preparation plant. The key production process parameters of the heavy medium separation process specifically include sampling date, sampling time, magnetic substance content, slime content, separation density, clean coal ash content, cyclone inlet pressure, and combined medium tank liquid level. The raw coal float-sink data specifically includes sampling date, sampling time, density grade yield, and density grade ash rate, etc.
[0074] Subsequently, conduct meticulous standardization and missing value processing on the collected data to ensure the integrity and consistency of the data, and provide high-quality input for the subsequent neural network model.
[0075] The data in the embodiments of the present invention is collected from a certain coal preparation plant under Shandong Energy Group, and includes the key production process parameters of the heavy medium separation process and the float-sink characteristics of the raw coal. Specifically, as shown in Table 1, the key production process parameters of the heavy medium separation process cover multiple dimensions such as sampling date, sampling time, magnetic substance content, slime content, separation density, cyclone inlet pressure, and combined medium tank liquid level, and these data are continuously sampled at 1-minute intervals.
[0076] Table 1 Example table of heavy medium separation process parameters
[0077] 。
[0078] Meanwhile, as shown in Table 2, the raw coal flotation data also includes information such as sampling date, sampling time, density grade, yield, and ash content of the density grade. In the daily production of coal preparation plants, operations are divided into three 8-hour shifts. According to production requirements, each shift usually conducts 2 to 3 raw coal flotation experiments. The time of the flotation experiment is uncertain, while the duration of the heavy medium separation process is about 4 minutes. In the present invention, samples are constructed based on the data of each raw coal flotation experiment, thus forming a sample data set. The output of the sample is the separation density corresponding to the clean coal collection moment, and the input is the raw coal flotation experiment data 4 minutes ago and the process parameters within the past 4 minutes. Through this method, a total of 5290 samples are generated, laying a reliable data foundation for subsequent multi-model training and testing.
[0079] Table 2 Example Table of Raw Coal Flotation Data
[0080] 。
[0081] Step 2: Calculation of theoretical separation density; Using the raw coal flotation experiment data, a theoretical mapping model is constructed to formalize the relationship between the separation density and the clean coal ash content. First, the present invention establishes a discrete mapping relationship between the separation density and the clean coal ash content, and further smooths the discrete data points through interpolation technology to establish a continuous-value theoretical mapping model. The specific discrete mapping relationship is as follows:
[0082] (1);
[0083] Among them, is the density grade interval with the density grade serial number of , is an integer and ; the separation density is the upper bound of ; is the clean coal ash content corresponding to ; is the yield of the density grade area with the density grade area serial number of ; is the density grade ash content of the density grade area with the density grade area serial number of ; is the density grade area serial number index; is the maximum value function;
[0084] Taking the raw coal flotation data at 6:20 on June 1st in Table 2 as an example, when the theoretical separation density is 1.4, the calculation formula for the clean coal ash content is (50.17 * 4.83 + 15.57 * 7.01) / (50.17 + 15.57). By calculating the clean coal ash content at different theoretical separation densities and interpolating the results, a continuous-value theoretical mapping model between the separation density and the clean coal ash content can be established. Using the continuous-value theoretical mapping model (referred to as the theoretical model), the theoretical separation density value under the set target clean coal ash content value can be calculated. This continuous-value theoretical separation density value (i.e., the theoretical prediction value) is not only used as one of the inputs for the subsequent deep neural network prediction model but also crucial for improving the prediction accuracy of the model.
[0085] Step 3: Construct a deep neural network prediction model based on the attention mechanism for predicting the final separation density; The present invention designs a deep neural network prediction model based on the attention mechanism (referred to as the neural network model), which is a deep learning model that integrates the convolutional neural network CNN and the bidirectional long short-term memory network BiLSTM. The model architecture diagram is as Figure 2 shown, including a one-dimensional convolutional neural network layer (1D CNN layer), a bidirectional long short-term memory network layer (BiLSTM layer), and an attention mechanism layer. The input of the model includes the data preprocessed in Step 1 and the theoretical separation density value calculated in Step 2, and the output of the model is the final separation density.
[0086] The input of the convolutional neural network layer is the key coal preparation production process parameters at different times;
[0087] Convolutional neural network layer:
[0088] Two-dimensional convolution is mainly used for image feature extraction, while one-dimensional convolution is mainly used for processing time series feature extraction. Figure 3 shows a schematic diagram of one-dimensional convolution, and its calculation formula is:
[0089] (2);
[0090] Among them, represents the output value after convolution, which participates in the subsequent calculation process as the hidden feature of the convolved area in the time series; is different elements in the input time series; 、 、 are different weights of the convolution kernel respectively. For other elements at different timestamps in the time series, define and as other elements at different timestamps in the time series, then is the corresponding output value after convolution of these elements, and the calculation method is the same as formula (2).
[0091] The convolutional neural network layer performs a convolution operation on local regions of key production process parameters at different times to extract short-term features in the feature dimension (production process parameters).
[0092] Bidirectional LSTM layer:
[0093] LSTM and bidirectional LSTM, as specific types of RNN networks, are used to model sequence data. The structure of LSTM is as Figure 4 shown, consisting of a forget gate, an input gate, and an output gate.
[0094] The equations for the forward propagation process are as follows:
[0095] (3);
[0096] (4);
[0097] (5);
[0098] (6);
[0099] (7);
[0100] (8);
[0101] Among them, is the forward candidate cell state at time , are respectively the state of the forward storage cell at time ; , , are respectively the forget gate, input gate, and output gate of the forward process at time ; , , are respectively the biases corresponding to the forget gate, input gate, and output gate of the forward process; is the bias of the input state of the forward process; , , are respectively the weights of the forget gate, input gate, and output gate of the forward process; is the weight of the input state of the forward process; is the sigmoid function; is the tanh activation function; , are respectively time and The output value of the time-forward storage unit is concatenated and used as the input to the attention mechanism layer; is the input value at time step
[0102] The bidirectional LSTM adopted in the present invention consists of two LSTMs with opposite directions, namely the forward layer and the backward layer, which are used to extract long-term dependence features in the time dimension. The structure of the bidirectional LSTM is as Figure 5 shown, including an input layer, a forward layer, a backward layer, and an output layer. The bidirectional LSTM extracts features of the input time series from both the forward and backward directions, generating a forward hidden state sequence and a backward hidden state sequence respectively. For each time step, the forward and backward hidden states are concatenated to form the final output of the bidirectional LSTM layer, which is also used as the input to the attention mechanism layer. Among them, represents the input sequence, represents the hidden state of the forward layer, represents the hidden state of the backward layer, represents the final output of the bidirectional LSTM layer.
[0103] The reverse process is as follows:
[0104] (9);
[0105] (10);
[0106] (11);
[0107] (12);
[0108] (13);
[0109] (14);
[0110] Among them, is the state of the reverse candidate cell at time step , are respectively the states of the reverse storage units at time steps ; , , are respectively the forget gate, input gate, and output gate at time step in the reverse process; , , are respectively the biases corresponding to the forget gate, input gate, and output gate in the reverse process; Input the state bias for the reverse process; , , are the weights of the forget gate, input gate, and output gate of the reverse process respectively; is the input state weight of the reverse process; , are respectively time and the output values of the reverse storage unit at time;
[0111] Concatenate and to obtain the output vector of the bidirectional long short-term memory network layer at time ;
[0112] Attention mechanism layer: To further enhance the feature representation and enable the model to dynamically focus on the key features in the input sequence, the present invention introduces an attention mechanism to assign different weights to the output of the bidirectional LSTM to amplify the key production process parameters related to the sorting density prediction.
[0113] During the calculation process of the attention mechanism layer, the hidden representation after dimensional transformation is :
[0114] (15);
[0115] where is the weight parameter of the attention mechanism layer; is the bias parameter of the attention mechanism layer;
[0116] By calculating the similarity between and the context vector to measure the importance of , is randomly initialized and learned during training, and then, the following softmax function is used to normalize the importance of each input. The specific formula is:
[0117] (16);
[0118] where is the normalized attention value of, is the output vector of the bidirectional long short-term memory network layer at time; is the hidden representation of after dimensional transformation; is the exponential function with base e; is the summation variable index at time;
[0119] Finally, the output of the attention mechanism layer is obtained by weighted averaging all the outputs of the bidirectional long short-term memory network layer:
[0120] (17);
[0121] The output of the attention mechanism layer undergoes dimensional transformation through a fully connected layer, and finally a one-dimensional sorting density is obtained.
[0122] Step 4, define a composite loss function; in order to improve the prediction accuracy of the model, the present invention designs a composite loss function, which ingeniously combines two key components: empirical loss and physical loss.
[0123] For the empirical loss part, the mean squared error (MSE) is selected as the measurement standard to measure the difference between the predicted sorting density value and the actual sorting density value. Its calculation formula is:
[0124] (18);
[0125] where, is the mean squared error value; is the total number of samples; represents the actual sorting density value of the th sample; is the predicted sorting density value of the
[0126] th sample. For the physical loss part, an additional constraint based on the coal preparation physical process is innovatively introduced. Specifically, by pertinently disturbing the process parameters and raw coal flotation data in the sample except for the sorting density, and comparing the prediction results before and after the disturbance, a physical loss is established to reflect the physical influence of the disturbed parameters on the sorting density; the physical loss part can evaluate the actual influence of these disturbances on the predicted sorting density value, and the physical loss function is defined as:
[0127] (19);
[0128] where, is the physical loss function; and are the sorting density values predicted by the input data before and after the disturbance through the deep neural network based on the attention mechanism respectively; is the ReLU function. The ReLU function ensures that the loss is calculated only when there is a deviation between the predicted value and the actual value, so that the model pays more attention to reducing the prediction error.
[0129] Finally, the composite loss function (also known as the final loss) of the deep neural network prediction model based on the attention mechanism is set as the weighted sum of the empirical loss and the physical loss, which can not only ensure that the model has a good fitting degree for historical data in a statistical sense, but also ensure that the model prediction results conform to the actual laws of the physical process. Through this innovative design of the composite loss function, the model of the present invention can achieve higher accuracy and robustness when predicting the separation density in the coal preparation process.
[0130] Step 5: Divide the samples constructed in Step 1 into a training set and a test set according to a ratio, and perform model training based on the training set. During the training process, the loss function optimizes both the empirical loss and the physical loss to ensure that the model's prediction is both accurate and conforms to the physical laws, thereby improving the generalization ability and prediction accuracy of the model.
[0131] In the embodiment of the present invention, the deep neural network prediction model based on the attention mechanism is trained using the large historical data of coal preparation production. Specifically, the samples constructed in Step 1 are divided into a training set and a test set at a ratio of 8:2. During the training process, a composite loss function is specially designed, which combines the empirical loss and the physical loss. The purpose of this is to make the model follow the basic laws of the physical process while making data-driven predictions, belonging to a physical model. In addition, the hyperparameters of the model are carefully adjusted to further optimize the performance: the learning rate is 0.001; the optimizer is the Adam algorithm; the batch size is 32; the number of training epochs is 50. At the same time, an early stopping mechanism is introduced to monitor the training process. Once the model performance no longer improves, the training is stopped to prevent overfitting of the model.
[0132] Step 6: Model evaluation and application; after the model training is completed, the trained model is evaluated using the test set. By comparing the differences between the predicted values and the actual values of the separation density, key indicators such as the mean square error (MSE), relative squared error (RSE), and relative absolute error (RAE) are calculated to measure the performance of the model. At the same time, the differences between the prediction and the actual values are visually displayed through charts to quickly evaluate the model effect. After the evaluation is completed, the optimal model during the evaluation process is selected and applied to actual production to achieve accurate prediction of the separation density, thereby improving the stability and efficiency of the coal preparation process, optimizing the production process, and assisting engineers in making decisions.
[0133] Certainly, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the essence of the present invention should also fall within the protection scope of the present invention.
Claims
1. A dense medium separation density decision method guided by deep learning and physical model, characterized in that: The steps include: Step 1: Data collection and preprocessing, building a sample data set; Step 2: construct a theoretical mapping model and calculate the theoretical sorting density; In step 2, a theoretical mapping relationship between sorting density and clean coal ash content is constructed based on the raw coal floating and sinking experimental data; first, a discrete mapping relationship between the two is determined, and then the offline points are smoothed by an interpolation method to establish a continuous value theoretical mapping model between sorting density and clean coal ash content; the specific discrete mapping relationship is as follows: (1); in, The density level is The density range of is an integer and ; Sorting density for The upper bound of for The corresponding clean coal ash content; The density level zone number is The yield of The density level zone number is Density grade ash; is the density level area serial number index; is the maximum value function; The continuous value theoretical mapping model is used to calculate the theoretical sorting density value under the set target clean coal ash value. The theoretical sorting density value is one of the inputs of the subsequent deep neural network prediction model. Step 3: Build a deep neural network prediction model based on the attention mechanism to predict the sorting density; In the step 3, the deep neural network prediction model based on the attention mechanism includes a one-dimensional convolutional neural network layer, a bidirectional long short-term memory network layer, and an attention mechanism layer; the input of the model includes the data preprocessed in step 1 and the theoretical sorting density value calculated in step 2, and the output of the model is the sorting density; Step 4: Define the composite loss function; In step 4, the composite loss function combines the empirical loss and the physical loss. In the physical loss part, additional constraints based on the physical process of coal preparation are introduced. By perturbing the process parameters and raw coal floating and sinking data in the sample except for the sorting density in a targeted manner and comparing the prediction results before and after the disturbance, the physical loss is established to reflect the physical influence of the disturbed parameters on the sorting density. The physical loss part can evaluate the actual influence of the disturbance on the predicted value of the sorting density. The physical loss function is defined as: (19); in, is the physical loss function; and They are the sorting density values predicted by the deep neural network based on the attention mechanism for the input data before and after the disturbance; is the ReLU function; Step 5: Divide the samples constructed in step 1 into training set and test set in proportion, and perform model training based on the training set; Step 6: Model evaluation and application.
2. The dense medium separation density decision method guided by deep learning and physical model according to claim 1 is characterized in that: The specific process of step 1 is as follows: First, the key production process parameters of the heavy medium separation process and the floating and sinking data of raw coal are comprehensively collected through the production monitoring system of the coal preparation plant; the key production process parameters of the heavy medium separation process specifically include sampling date, sampling time, magnetic content, coal slime content, separation density, clean coal ash content, cyclone inlet pressure, and combined medium barrel liquid level; the floating and sinking data of raw coal specifically include sampling date, sampling time, density grade yield and density grade ash rate; Subsequently, the collected data were standardized and missing values were processed.
3. The dense medium separation density decision method guided by deep learning and physical model according to claim 1 is characterized in that: In step 3, the input of the one-dimensional convolutional neural network layer is the key production process parameters of coal preparation at different times. By performing convolution operations on local areas of the key production process parameter data of coal preparation at different times, the corresponding feature maps are output to extract the effective short-term features of the key production process parameters and historical data; the calculation formula of the one-dimensional convolution is: (2); in, Represents the output value after convolution; are different elements in the input time series; , , are the weights of different convolution kernels respectively; The bidirectional long short-term memory network layer consists of two layers of long short-term memory networks in opposite directions, namely the forward layer and the backward layer, which are used to extract long-term dependent features in the time dimension; the bidirectional long short-term memory network layer includes an input layer, a forward layer, a backward layer and an output layer; the bidirectional long short-term memory network layer processes the input data from the forward and reverse directions of the sequence respectively, and generates a forward hidden state sequence and a reverse hidden state sequence respectively; for each time step, the forward and reverse hidden states are concatenated to generate the final hidden state, which is also used as the input of the attention mechanism layer; The forward process is as follows: (3); (4); (5); (6); (7); (8); in, for The state of the candidate unit is always positive; , They are time, The state of the forward storage unit at all times; , , Forward Process The forget gate, input gate, and output gate at each moment; , , They are the biases corresponding to the forget gate, input gate, and output gate in the forward process respectively; Enter the state bias for the forward process; , , They are the weights of the forget gate, input gate, and output gate in the forward process respectively; Enter the state weights for the forward process; is the sigmoid function; is the tanh activation function; , They are Moment and The output value of the positive storage unit at all times; for The input value at the moment; The reverse process is: (9); (10); (11); (12); (13); (14); in, for Reverse candidate unit status at all times; , They are time, The state of the storage unit is reversed at all times; , , Reverse process The forget gate, input gate, and output gate at each moment; , , They are the biases corresponding to the forget gate, input gate, and output gate in the reverse process respectively; Enter the state bias for the reverse process; , , They are the weights of the forget gate, input gate, and output gate in the reverse process respectively; Enter state weights for the reverse process; , They are Moment and The output value of the reverse storage unit at all times; Will and The bidirectional long short-term memory network layer is obtained by splicing Output vector at time ; During the calculation of the attention mechanism layer, The latent representation after dimension transformation is : (15); in, is the weight parameter of the attention mechanism layer; is the bias parameter of the attention mechanism layer; By calculation With context vector To measure the similarity between The importance of During the training process, it is randomly initialized and obtained through learning. Then, the importance of each input is normalized using the softmax function shown below. The specific formula is: (16); in, for The normalized attention value of For the bidirectional long short-term memory network layer Output vector at time instant; After dimension transformation The implicit representation of is an exponential function with base e; is the sum variable index of the moment; Finally, the output of the attention layer The weighted average of all outputs of the bidirectional long short-term memory network layer is obtained: (17); The output of the attention mechanism layer is transformed through the fully connected layer to finally obtain a one-dimensional sorting density.
4. The dense-medium separation density decision method guided by deep learning and physical model according to claim 1 is characterized in that: In step 4, the empirical loss part uses the mean square error as a measurement standard to measure the difference between the predicted value and the actual value of the sorting density. The calculation formula is: (18); in, is the mean square error value; is the total number of samples; Representative The actual value of the sorting density of each sample; It is The predicted value of sorting density for each sample; Finally, the composite loss function of the attention-based deep neural network prediction model is set to the weighted sum of empirical loss and physical loss.
5. The dense-medium separation density decision method guided by deep learning and physical model fusion according to claim 1 is characterized in that: In step 5, during the training process, the empirical loss and the physical loss are optimized simultaneously; the model hyperparameters are set: the learning rate is 0.001; the optimizer is the Adam algorithm; the batch size is 32; the number of training rounds is 50; and an early stopping mechanism is introduced to monitor the training process. Once the model performance no longer improves, the training is stopped to prevent the model from overfitting.
6. The dense medium separation density decision method guided by deep learning and physical model fusion according to claim 1 is characterized in that: In step 6, after the model training is completed, the trained model is evaluated using the test set, and the performance of the model is measured by comparing the difference between the predicted value and the actual value of the sorting density, calculating the mean square error, relative absolute error and relative absolute error key indicators; at the same time, the difference between the predicted value and the actual value is intuitively displayed through a chart; after the evaluation is completed, the optimal model in the evaluation process is selected and applied to actual production to realize the prediction of the sorting density and assist engineers in making decisions.
Citation Information
Patent Citations
Machining roughness prediction method and system based on physical information machine learning, medium and product
CN117909924A