Production process export quality prediction method based on performance-driven knowledge distillation

By adopting a performance-driven knowledge distillation method in the industrial production process and using twin neural networks and LSTM modules to train student models, the problem of model performance fluctuations and dimensional imbalance in quality indicator prediction is solved, and higher prediction stability and accuracy are achieved.

CN120013310APending Publication Date: 2025-05-16CENT SOUTH UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411863556.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the industrial production process, the real-time prediction of quality indicators faces the problems of model performance fluctuations and imbalance in the input and output dimensions, resulting in instability of the prediction results.

Method used

Using a performance-driven knowledge distillation method, through twin neural network structures and LSTM modules, student models are trained to learn the feature representation of high-performance teacher models and fine-tune under the guidance of soft measurement tasks.

Benefits of technology

The prediction stability and accuracy of industrial process production quality indicators are improved, and the stability and superiority of model performance can be guaranteed more than traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013310A_ABST
    Figure CN120013310A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of chemical production process control, and particularly discloses a production process export quality prediction method based on performance-driven knowledge distillation, which comprises the following steps: step 1, acquiring observable process data of an industrial production process and test data of export quality indexes, and forming a data set; 2, cleaning and standardizing the data set, and randomly dividing the data set into a training set, a verification set and a test set according to a certain proportion; 3, building a basic quality predictor, and embedding an LSTM module into the basic quality predictor through a twinning neural network framework; 4, training a plurality of basic twin LSTM quality predictors to obtain a teacher model; step 5, training student models with the same structure by taking performance as drive and combining knowledge distillation and supervised learning; and step 6, inputting process data of the test sample, and obtaining an output predicted value, so as to improve the accuracy of the soft measurement of the production quality index in the industrial process on the basis of ensuring the stability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of chemical production process control, and specifically discloses a method for predicting the outlet quality of a production process based on performance-driven knowledge distillation. Background Art

[0002] With the continuous improvement of the degree of automation in industrial processes, real-time measurement of quality indicators can provide timely guidance for the control and optimization of production processes, which is conducive to the stability and long-term operation of production processes. However, due to factors such as the environment and instrument performance, most quality variables can only be obtained through periodic laboratory offline testing. The detection delay caused by manual testing makes the control and optimization of the process lack the guidance of quality information. In order to solve this problem, it is necessary to realize real-time prediction of quality indicators. As the development of modern industry has brought massive amounts of data to quality prediction modeling, data-driven methods have developed more rapidly in industrial applications.

[0003] However, in the actual modeling process, the establishment of the model usually takes into account the complex temporal and spatial coupling relationships, and adds complex regularization terms. However, due to the extremely high dimensionality of the parameter space in the deep network, the parameter optimization problem becomes highly non-convex, and there is an extreme imbalance in the dimensions of the input space and the output space. Although models with different parameter distributions can achieve similar levels when fitting the training set and the validation set, their prediction results may be quite different when facing the test set where the working conditions change over time. Therefore, how to design the model structure and training method to avoid fluctuations in model performance is one of the challenging issues faced by the prediction of quality indicators in industrial production processes.

[0004] Therefore, in view of this, the inventor provides a method for predicting the export quality of a production process based on performance-driven knowledge distillation to solve the above problems. Summary of the invention

[0005] The purpose of the present invention is to provide a method for predicting the outlet quality of a production process based on performance-driven knowledge distillation, so as to improve the accuracy of soft measurement of industrial process production quality indicators while ensuring the stability of the model.

[0006] In order to achieve the above object, the basic scheme of the present invention provides a method for predicting the export quality of a production process based on performance-driven knowledge distillation, comprising the following steps:

[0007] Step 1: collect observable process data of industrial production process and laboratory data of export quality indicators, and form a data set;

[0008] Step 2: Clean and standardize the data set, and randomly divide it into training set, validation set, and test set according to a certain ratio;

[0009] Step 3: Build a basic quality predictor using a twin neural network framework and embed the LSTM module into it;

[0010] Step 4: train multiple basic twin LSTM quality predictors to obtain a teacher model;

[0011] Step 5: Driven by performance, combine knowledge distillation with supervised learning to train a student model with the same structure;

[0012] Step 6: Input the process data of the test sample and obtain the predicted value of the output.

[0013] Furthermore, in step 2, the cleaning process uses spline interpolation to fill in missing data, replace abnormal data, and standardize the data according to the z-score method after cleaning.

[0014] Furthermore, the expression for standardizing the data according to the z-score method is as follows:

[0015]

[0016] In the formula, X is the original data, Z is the standardized data, is the mean vector of all variables in the original data, and S is the standard deviation vector of all variables in the original data.

[0017] Furthermore, the proportions of the training set, validation set and test set are 70%, 15% and 15% respectively.

[0018] Furthermore, in step 3, two different original data are input into the sub-network of the twin neural network, and two different latent space feature maps are obtained, and then the loss function is used to ensure that the distance between two similar samples in the feature space is minimized.

[0019] Furthermore, the expression of the loss function is as follows:

[0020]

[0021] In the formula, x(p) and x(q) are the two original input data, z(p) and z(q) are the two latent space feature maps obtained, d{z(p),z(q)} is the distance between the feature pairs, δ is the minimum distance between dissimilar samples, and L sf is the twin network optimization loss function value.

[0022] Further, in step 4, the basic twin LSTM quality predictor is composed of adding a fully connected layer after the features output by the LSTM module of the twin neural network, and ensuring that the features are suitable for the quality prediction task through the following prediction loss:

[0023]

[0024] Where y(t) is the true value, is the predicted value, L pred is the predicted loss value. Then, all loss functions are reconciled using the following weight distribution method:

[0025] L=L sf +αL pred

[0026] Where L is the overall loss value of quality predictor training, and α is the harmonic parameter;

[0027] Train at least five base twin LSTM quality predictors as teacher models.

[0028] Furthermore, in step 5, the same training sample is input into the teacher model to obtain the feature vectors Quality prediction results The inverse of the deviation between the true value and the predicted value is calculated, and the weight of the teacher model is calculated based on the inverse. The weight of each teacher model is multiplied by its feature representation, so that the student model learns the feature representation of the excellent teacher model based on the KL divergence.

[0029] Furthermore, the expression for calculating the inverse of the deviation between the true value and the predicted value is as follows:

[0030]

[0031] where y is y O The corresponding true value, s is the inverse of the deviation.

[0032] The expression for calculating the weight of the teacher model based on the inverse is as follows:

[0033]

[0034] Where exp(·) represents the natural exponential function and a is the weight value of the teacher model.

[0035] The expression for multiplying the weight of each teacher model by its feature representation so that the student model can learn the feature representation of the excellent teacher model using the knowledge distillation loss built based on KL divergence is as follows:

[0036]

[0037] In the formula, z ada is the optimal feature, z S is the feature generated by the student model, L KL Represents the KL divergence value between two features.

[0038] Furthermore, in step 6, the loss function used in the training process is the minimum mean square error loss function:

[0039]

[0040] Where N is the number of samples, y is the true test value, is the predicted value, and L is the minimum mean square error loss between the true value and the predicted value.

[0041] The principle and effect of this scheme are:

[0042] 1. The present invention is based on a twin neural network structure, uses LSTM as a subnetwork, and the twin LSTM is guided by a soft measurement task. This helps to capture both the spatiotemporal features within the sample and the distribution differences between samples to ensure the effectiveness of the features constructed by the encoder.

[0043] 2. The present invention is based on a performance-driven knowledge distillation training method. First, the teacher model is trained, and then a higher weight is assigned to the high-performance teacher model, so that the student model can learn the feature representation of the high-performance teacher model and further improve the performance of the model under the guidance of the soft measurement task.

[0044] 3. In summary, the present invention effectively extracts the state information feature changes in the time and space dimensions reflected by the data in the industrial production process based on the twin LSTM network structure. After training multiple twin LSTMs, a performance-driven knowledge distillation strategy is used to select a model with superior performance from multiple teacher models for learning, and the feature construction process of the student model is guided and fine-tuned by soft measurement tasks, so that the student model finally has stable and superior performance. Compared with traditional methods, this method has better stability and higher accuracy in predicting industrial process production quality indicators. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0046] Figure 1 A schematic diagram of an alumina six-effect four-flash structure evaporation process equipment is shown;

[0047] Figure 2 A flowchart of a method for predicting the export quality of a production process based on performance-driven knowledge distillation proposed in an embodiment of the present application is shown;

[0048] Figure 3 A schematic diagram of the structure of a twin LSTM in a method for predicting the export quality of a production process based on performance-driven knowledge distillation proposed in an embodiment of the present application is shown;

[0049] Figure 4 A schematic flow chart of a performance-driven knowledge distillation training process in a production process export quality prediction method based on performance-driven knowledge distillation proposed in an embodiment of the present application is shown;

[0050] Figure 5 A schematic diagram of the prediction results of the quality index of the evaporation mother liquor at the outlet of the IV effect flash evaporator of a production process outlet quality prediction method based on performance-driven knowledge distillation proposed in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0051] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation mode, structure, characteristics and effects of the present invention are described in detail below in combination with the accompanying drawings and preferred embodiments.

[0052] A method for predicting the outlet quality of a production process based on performance-driven knowledge distillation, taking the alumina six-effect four-flash structure evaporation process as an example, the process equipment schematic diagram is as follows Figure 1 Shown: caustic soda stock solution The raw liquid enters the entire evaporation system from the flash evaporator, and then enters the VI-effect evaporator after flash evaporation in the raw liquid flash evaporator, and then flows to the I-effect evaporator. After the evaporation raw liquid is heat exchanged in the evaporation stage of the evaporator, it enters the I-effect flash evaporator, and then passes through multiple flash evaporators in sequence, and finally the mother liquid Flow out of the IV effect flash evaporator. In the evaporation stage of the evaporator, the new steam F s First, it enters the I-effect evaporator to provide a high-temperature basis for heat exchange of the caustic soda stock solution. After heat exchange, it flows into the II-effect evaporator and then flows to the VI-effect evaporator in turn.

[0053] The predicted quality index of this embodiment is the caustic soda concentration of the mother liquor at the outlet of the IV-effect flash evaporator. Due to the scaling of the pipeline and the expensive analyzer, this value cannot be accurately measured in real time by the sensor, and can only be obtained by manually collecting samples every 8 hours and sending them to the laboratory for analysis. There is a problem of detection delay, which makes the control and optimization of the process lack the guidance of quality information.

[0054] The prediction method in this embodiment is as follows Figure 2 As shown, the following steps are included:

[0055] Step 1: Collect observable process data of industrial production processes and laboratory data of export quality indicators and form a data set.

[0056] Step 2: Clean and standardize the data set, and randomly divide it into training set, validation set, and test set according to a certain ratio. Specifically, it includes:

[0057] Step 2.1, data cleaning uses spline interpolation to fill missing data and replace abnormal data;

[0058] Step 2.2, after cleaning, standardize the data according to the z-score method:

[0059]

[0060] Where X is the original data and Z is the standardized data. is the mean vector of all variables in the original data, and S is the standard deviation vector of all variables in the original data;

[0061] Step 2.3: randomly divide the processed data into the following proportions: 70% of the samples are training sets, 15% of the samples are validation sets, and 15% of the samples are test sets.

[0062] Step 3: Build a basic quality predictor, using the twin neural network framework and embedding the LSTM module into it. Specifically:

[0063] Step 3.1, the twin neural network framework consists of two sub-networks with shared structures and parameters. Two different original data x(p) and x(q) are input into two LSTM sub-networks respectively, and two different latent space feature maps z(p) and z(q) can be obtained accordingly;

[0064] Step 3.2, use the LSTM module as a sub-network of the twin neural network;

[0065] In step 3.3, the Siamese neural network uses the following loss function to minimize the distance between two similar samples in the feature space, while ensuring that the distance between dissimilar samples is at least δ:

[0066]

[0067] Where L sf is the twin network optimization loss function value, and d{z(p),z(q)} is the distance between feature pairs.

[0068] Step 4: Train multiple basic twin LSTM quality predictors to obtain the teacher model. Specifically, it includes:

[0069] Step 4.1, add a basic twin LSTM quality predictor consisting of a fully connected layer after the features output by the LSTM module of the twin neural network, and ensure that the features are suitable for the quality prediction task by minimizing the following prediction loss:

[0070]

[0071] Step 4.2, reconcile all loss functions of the base twin LSTM quality predictor with the following loss function:

[0072] L=L sf +αL pred

[0073] Where L is the overall loss value for quality predictor training and α is the tuning parameter.

[0074] In step 4.3, five basic twin LSTM quality predictors are trained in the above manner as teacher models.

[0075] Step 5: Driven by performance, combine knowledge distillation with supervised learning to train a student model with the same structure. Specifically, it includes:

[0076] Step 5.1: Input the same training sample into the five teacher models and obtain the feature vectors Quality prediction results Calculate the inverse of the deviation between the true value and the predicted value:

[0077]

[0078] where y is y O The corresponding true value, s is the inverse of the deviation.

[0079] Step 5.2, calculate the weight of the teacher model based on the inverse mentioned in step 5.1:

[0080]

[0081] Where exp(·) represents the natural exponential function and a is the weight value of the teacher model.

[0082] In step 5.3, the weight of each teacher model is multiplied by its feature representation, and the student model uses the knowledge distillation loss constructed based on KL divergence to learn the feature representation of the excellent teacher model:

[0083]

[0084] where z ada is the optimal feature, z S is the feature generated by the student model, L KL Represents the KL divergence value between two features.

[0085] Step 6: Take the process variables in the training set and the validation set as input and the corresponding quality indicators as labels, and use the training framework mentioned in step 5 for supervised training. Use the Adam optimizer, set the learning rate to 0.0005, and use L2 regularization to prevent overfitting to obtain the training model. The loss function used in the training process is the minimum mean square error loss function:

[0086]

[0087] Where N is the number of samples, y is the true test value, is the predicted value, and L is the minimum mean square error loss between the true value and the predicted value.

[0088] Step 7: Input the process data of the test sample and obtain the predicted value of the output.

[0089] The training method based on the twin network structure has two characteristics. Feature 1: The training process is not based on complex assumptions, and the sub-network structure is arbitrary. Feature 2: The goal of training can ensure that the distance between two similar samples in the feature space is close, while making the distance between dissimilar samples in the feature space large enough. Based on feature 1, using LSTM as a sub-network structure provides the possibility of capturing the historical time series change information of the data. Based on feature 2, the construction of feature distance provides the possibility of identifying inputs at different times and identifying the coupling relationship between different variables. Based on features 1 and 2, it is possible to capture the spatiotemporal features within the sample and the distribution differences between samples at the same time, and to construct features to improve the soft measurement performance. However, the existing soft measurement algorithms are difficult to face the scenario where the working conditions change over time, and cannot distinguish the relationship between the data distribution of the samples and the production time, which limits the performance of the soft measurement. The present invention is based on the twin neural network structure, uses LSTM as a sub-network, and the twin LSTM composed is guided by the soft measurement task. This helps to capture the spatiotemporal features within the sample and the distribution differences between samples at the same time to ensure the effectiveness of the characteristics constructed by the encoder.

[0090] The training method based on knowledge distillation has two characteristics. Feature 1: There can be multiple teacher models. Feature 2: When the structural gap between the performance of the teacher model and the performance of the student model is small, the student model has the opportunity to obtain higher performance than the teacher model. Based on feature 1, the method of learning from multiple teacher models provides the possibility of learning superior feature representation. Based on feature 2, the student model adopts the same structure and is guided based on the soft measurement task, which provides the possibility of obtaining better performance than the teacher model. Based on features 1 and 2, multi-teacher knowledge distillation provides the possibility of ensuring the stability and superiority of the student model performance at the same time. However, the existing methods focus more on reducing the number of parameters of the student model to achieve a faster response speed, and do not put forward higher requirements on the performance of the student model. The present invention is based on a performance-driven knowledge distillation training method, which first trains the teacher model, and then assigns a higher weight to the high-performance teacher model, so that the student model can learn the feature representation of the high-performance teacher model, and further improve the performance of the model under the guidance of the soft measurement task.

[0091] The structure of the twin LSTM used in the aluminum oxide evaporation process is shown below: Figure 3 As shown in the figure, the input data is first divided into samples, and then the encoder part is responsible for feature extraction of the samples. By minimizing the feature distribution error, the spatiotemporal features within the sample and the distribution differences between samples can be captured at the same time; the predictor part performs preliminary training on the soft measurement task to ensure that the features constructed by the encoder are suitable for the soft measurement task.

[0092] The performance-driven knowledge distillation training process applied to the alumina evaporation process is shown in the following figure. Figure 4 As shown in the figure, in this process, five twin LSTMs are first trained as teacher models, and then the performance of the teacher models is analyzed and higher weights are given to the features of the teacher models with better performance, prompting the student model to learn this feature representation; the features of the student model are further fine-tuned using soft measurement tasks, so that the student model can obtain a more stable performance and higher performance than the teacher model.

[0093] The prediction of the caustic soda concentration of the evaporation mother liquor at the outlet of the IV-effect flash evaporator in the alumina evaporation process is as follows: Figure 5 As shown in the figure, compared with the results of laboratory tests and some other traditional soft measurement methods for industrial process quality indicators, the production process outlet quality prediction method based on performance-driven knowledge distillation proposed in the present invention can accurately predict the caustic soda concentration based on general process data, and is suitable for extension to other types of industrial processes.

[0094] In summary, the present invention effectively extracts the state information feature changes in the time and space dimensions reflected by the data in the industrial production process based on the twin LSTM network structure. After training multiple twin LSTMs, a performance-driven knowledge distillation strategy is used to select models with superior performance from multiple teacher models for learning, and the feature construction process of the student model is guided and fine-tuned by soft measurement tasks, so that the student model finally has stable and superior performance. Compared with traditional methods, this method has better stability and higher accuracy in predicting industrial process production quality indicators.

[0095] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A method for predicting the export quality of a production process based on performance-driven knowledge distillation, characterized in that: The following steps are involved: Step 1: collect observable process data of industrial production process and laboratory data of export quality indicators, and form a data set; Step 2: Clean and standardize the data set, and randomly divide it into training set, validation set, and test set according to a certain ratio; Step 3: Build a basic quality predictor using a twin neural network framework and embed the LSTM module into it; Step 4: train multiple basic twin LSTM quality predictors to obtain a teacher model; Step 5: Driven by performance, combine knowledge distillation with supervised learning to train a student model with the same structure; Step 6: Input the process data of the test sample and obtain the predicted value of the output.

2. According to claim 1, a method for predicting the export quality of a production process based on performance-driven knowledge distillation is characterized in that: In step 2, the cleaning process uses spline interpolation to fill in missing data, replace abnormal data, and standardize the data according to the z-score method after cleaning.

3. The method for predicting the export quality of a production process based on performance-driven knowledge distillation according to claim 2, characterized in that: The expression for standardizing data based on the z-score method is as follows: In the formula, X is the original data, Z is the standardized data, is the mean vector of all variables in the original data, and S is the standard deviation vector of all variables in the original data.

4. A method for predicting the export quality of a production process based on performance-driven knowledge distillation according to any one of claims 1 to 3, characterized in that: The training set, validation set and test set account for 70%, 15% and 15% respectively.

5. The method for predicting the export quality of a production process based on performance-driven knowledge distillation according to claim 1, characterized in that: In step 3, two different original data are input into the sub-network of the twin neural network, and two different latent space feature maps are obtained, and then the loss function is used to ensure that the distance between two similar samples in the feature space is minimized.

6. The method for predicting the export quality of a production process based on performance-driven knowledge distillation according to claim 5, characterized in that: The expression of the loss function is as follows: In the formula, x(p) and x(q) are the two original input data, z(p) and z(q) are the two latent space feature maps obtained, d{z(p),z(q)} is the distance between the feature pairs, δ is the minimum distance between dissimilar samples, and L sf is the twin network optimization loss function value.

7. The method for predicting the export quality of a production process based on performance-driven knowledge distillation according to claim 1, characterized in that: In step 4, the basic twin LSTM quality predictor is composed of adding a fully connected layer after the features output by the LSTM module of the twin neural network, and ensuring that the features are suitable for the quality prediction task through the following prediction loss: Where y(t) is the true value, is the predicted value. Then, all loss functions are reconciled using the following weight distribution method: L=L sf +αL pred Where L is the overall loss value of quality predictor training, and α is the harmonic parameter; Train at least five base twin LSTM quality predictors as teacher models.

8. The method for predicting the export quality of a production process based on performance-driven knowledge distillation according to claim 7, characterized in that: In step 5, the same training sample is input into the teacher model to obtain the feature vectors Quality prediction results And calculate the inverse of the deviation between the true value and the predicted value, then calculate the weight of the teacher model based on the inverse, and multiply the weight of each teacher model by its feature representation, The student model is made to learn the feature representation of the excellent teacher model based on KL divergence.

9. The method for predicting the export quality of a production process based on performance-driven knowledge distillation according to claim 7, characterized in that: The expression for calculating the inverse of the deviation between the true value and the predicted value is as follows: where y is y O The corresponding true value, s is the inverse of the deviation. The expression for calculating the weight of the teacher model based on the inverse is as follows: Where exp(·) represents the natural exponential function, and a is the weight value of the teacher model; Multiply the weight of each teacher model by its feature representation, The expression that enables the student model to learn the feature representation of the excellent teacher model using the knowledge distillation loss built on KL divergence is as follows: In the formula, z ada is the optimal feature, z S is the feature generated by the student model, L KL Represents the KL divergence value between two features.

10. The method for predicting the export quality of a production process based on performance-driven knowledge distillation according to claim 9, characterized in that: In step 6, the loss function used in the training process is the minimum mean square error loss function: Where N is the number of samples, y is the true test value, is the predicted value, and L is the minimum mean square error loss between the true value and the predicted value.

Citation Information

Cited By

  • Product quality prediction model construction method based on industrial knowledge distillation

    CN121257605A

  • Complex multi-working-condition alcohol distillation soft measurement method and system based on neural network

    CN122412863A

  • Complex multi-working condition alcohol distillation soft measurement method and system based on neural network

    CN122412863B