A prediction method and device for a quality index
Through a model-independent meta-learning framework that adjusts adaptive parameters and variable scaling steps of the support set during the training and testing phases, the problem of quality indicator changes in deep learning soft measurements is solved, and more accurate prediction results are achieved.
Patent Information
- Application Number
- CN202110530591.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-15
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-05-15
AI Technical Summary
The existing soft measurement methods based on deep learning are inconsistent in the prediction effects on the training set and the test set, and cannot accurately capture the changing trends of quality indicators. Traditional methods cannot effectively solve the problem of the relationship between process variables and quality indicators.
A model-independent meta-learning framework with variable scaling step size is adopted. By adjusting the adaptive parameters and support sets respectively in the training stage and the testing stage, the gradient descent method is used to optimize the scaling value, and the adaptive modules related to the construction stage are used to adapt to data changes and realize the adaptive adjustment of the model.
It improves prediction accuracy on the test set, reduces the performance differences of the model on different data sets, and improves the prediction effect of quality indicators.
Smart Images

Figure QLYQS_3 
Figure QLYQS_16 
Figure QLYQS_19
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing, and specifically discloses a prediction method for quality indicators and a prediction device for quality indicators. Background Art
[0002] In the industrial production process, there are many quality indicators or key process variables that cannot be directly obtained due to various reasons. It is necessary to construct a regression model between the quality indicators and those secondary variables that are easy to measure, so as to predict the potential results of the quality indicators of interest. This process is usually also called soft sensing, which plays an important role in the processes of industrial production monitoring, control and optimization.
[0003] Traditional soft sensing methods include principal component regression (PCR), partial least squares regression (PLSR), support vector machine (SVR), etc. In the past few years, many successful applications of soft sensing have been proposed in the fields of chemical engineering, biochemical engineering, metallurgy and pharmaceutical industries. With the development of technologies such as big data, process data shows the characteristics of large data volume, multiple data sources and high data dimension, making traditional data-driven methods unable to make more accurate predictions due to limited characterization learning ability. At the same time, with the development of deep learning, the advantages of big data can be fully utilized through parallel computing to extract the characterization information of process data. Therefore, soft sensing based on deep learning algorithms has received more and more attention due to its non-linear extraction ability and advantages in the context of big data.
[0004] However, in the process of soft sensing based on deep learning, variables will change, which makes the model constructed on the training set perform poorly when applied to the test set, and it is unable to accurately predict the quality indicators of interest. Taking the continuous catalytic reforming of naphtha chemical process (CCR) as an example, Figure 1 is the prediction result graph of the neural network model in the training set and the test set in the continuous catalytic reforming of naphtha chemical process, Figure 2 is the scatter plot of the prediction effect of the neural network model in the training set and the test set in the continuous catalytic reforming of naphtha chemical process. Figure 1 、 2 show that the model trained on the training set can produce valuable predictions on the training set, but the prediction effect on the test set is very poor, and the prediction effect on the test set cannot capture the change trend of the quality indicators, which is significantly different from the reference label.
[0005] In the prior art, immediate learning is proposed to construct a local model of samples to reduce the adverse effects of irrelevant samples. However, it still cannot well achieve the original intention of obtaining a more accurate prediction model through existing data, and cannot solve the problem of the relationship change between process variables and quality indicators in the soft sensing process. Summary of the Invention
[0006] A brief overview of one or more aspects is given below to provide a basic understanding of these aspects. This overview is not an exhaustive survey of all contemplated aspects and is neither intended to identify key or decisive elements of all aspects nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to a more detailed description to follow.
[0007] The present invention provides a method for predicting quality indicators, characterized by including the following steps: inputting the data to be measured into a model-agnostic meta-learning framework with variable scaling step sizes; and in the model-agnostic meta-learning framework with variable scaling step sizes, determining the quality indicators corresponding to the data to be measured according to the first adaptive parameter a p (θ p , g p ) and the first support set S p corrected in the test phase, wherein the first adaptive parameter a p (θ p , g p ) is obtained by performing at least one correction iteration operation using multiple test samples in the test phase according to the first support set S t not corrected in the test phase, wherein θ p is the meta-parameter of the neural network model after the correction iteration operation, g p is the gradient parameter of the neural network model after the correction iteration operation, and the first support set S p corrected in the test phase is generated by selecting multiple sample data from multiple test samples according to the meta-parameter θ p corrected in the test phase, and the first support set S t not corrected in the test phase is generated by selecting multiple sample data from multiple test samples according to the second adaptive parameter a t (θ t , g t ) of the meta-parameter θ t of the neural network model after t training iteration operations, wherein θ t is the meta-parameter of the neural network model after t training iteration operations, g t is the gradient parameter of the neural network model after t training iteration operations, and the second adaptive parameter a t (θ t , g t ) is obtained by performing multiple training iteration operations using multiple training samples in the training phase according to the second support sets S0 to S t-1 corresponding to the current meta-parameters θ0 to θ t-1 , wherein θ0 to θ t-1is the meta-parameter of the neural network model after the corresponding training iteration operation, and each second support set S0~S t-1 are respectively based on the corresponding meta-parameters θ0~θ t-1 , multiple sample data are selected from multiple training samples to generate.
[0008] In one embodiment, preferably, the training phase includes the following steps: a neural network model f with meta-parameters θ θ As the basic model of the model-independent meta-learning framework with variable scaling steps; for the neural network model f θ Initialization is performed to determine the initial element parameter θ0 and the second support data set S0~S t-1 The window size N t ; Obtain multiple training samples to form a training sample set; Select N from the training sample set according to the initial meta-parameter θ0 t training samples to generate the initial second support set S0; multiple training samples are input into the neural network model f θ To calculate the loss function corresponding to the initial parameter θ0 According to the initial parameters θ0, the loss function The second support data set S0 and the scaling value α of the variable scaling step length are used to calculate the meta-parameter θ1 after the first training iteration operation; and the corresponding loss function is calculated according to the meta-parameter θ1. If the loss function Greater than or equal to the preset loss function threshold Then, according to the meta-parameter θ1 and its corresponding second support data set S1, the meta-parameter θ2 and its corresponding loss function after the next training iteration are calculated again. And so on until the loss function Less than the loss function threshold Then the meta-parameter θ after t training iterations is t Determined as the meta-parameters trained during the training phase.
[0009] In one embodiment, preferably, the step of calculating the meta-parameter θ1 after the first training iteration operation includes: using the loss function The scaling value α is derived, and the local minimum of the scaling value α is calculated using the gradient descent method as the optimal scaling value α1 corresponding to the meta-parameter θ1, wherein the optimal scaling value α1 indicates the optimal scaling step length for iterating the initial meta-parameter θ0 to the meta-parameter θ1; and according to the initial meta-parameter θ0, the loss function The initial second support set S0 and the optimal scaling value α1 are used to calculate the meta-parameter θ1.
[0010] In one embodiment, preferably, the training phase further includes the following steps: according to the calculated meta-parameter θi Reselect N t training samples from the training sample set to generate a corresponding second support set S i , where 1 ≤ i ≤ t - 1; calculate the loss function according to the meta-parameter θ i-1 、meta-parameter θ i and its corresponding optimal scaling value α i Calculate the loss function Gradient parameter g of the second support set S i ; and determine the adaptive parameter a after i training iteration operations according to the meta-parameter θ i ; and according to the meta-parameter θ i and the gradient parameter g i Determine the adaptive parameter a after i training iteration operations i (θ i , g i ).
[0011] In one embodiment, preferably, the steps of and so on include: in response to the loss function being greater than or equal to the loss function threshold According to the meta-parameter θ i and the loss function Gradient parameter g of the second support set S i , use the gradient descent method to calculate the optimal scaling value α corresponding to the meta-parameter θ after the next training iteration operation i ; and according to the meta-parameter θ i+1 、loss function Second support set S i+1 and the optimal scaling value α i i i+1 , calculate the meta-parameter θ i+1 M t .
[0012] In one embodiment, preferably, the training phase is also set with a maximum number of iterations M, and the training phase further includes the following steps: determining whether the current number of iterations has reached the maximum number of iterations M; and in response to the current number of iterations reaching the maximum number of iterations M, determining that the training phase is completed, and determining the meta-parameter θ after M training iteration operations M as the meta-parameter θ trained through the training phase t .
[0013] In one embodiment, preferably, the training phase further includes the following steps: dividing the training sample set according to the input task distribution p(T) to determine multiple batch tasks T b1 ~T bB , where each batch task T b1 ~T bB includes multiple training samples, and select N from the training sample set according to the initial meta-parameter θ0 tThe steps of generating the initial second support set S0 from N training samples include: selecting N training samples from the first batch of tasks T according to the initial meta-parameters θ0 to generate the initial second support set S0, inputting the multiple training samples into the neural network model f to calculate the loss function corresponding to the initial parameters θ0. b1 of the multiple training samples to select N t training samples to generate the initial second support set S0, and inputting the multiple training samples into the neural network model f θ to calculate the loss function corresponding to the initial parameter θ0. The steps of calculating the loss function corresponding to the initial parameter θ0 for the first batch of tasks T include: inputting the multiple training samples of the first batch of tasks T into the neural network model f b1 to calculate the loss function corresponding to the initial parameter θ0 for the first batch of tasks T. θ The steps of calculating the meta-parameters θ1 after the first training iteration operation include: calculating the meta-parameters θ after one iteration operation on the first batch of tasks T according to the initial parameters θ0, the loss function b1 the second support data set S0 and the corresponding scaling value α ; and calculating the meta-parameters θ after one iteration operation on each of the remaining batches of tasks T ~T Tb1 one by one according to the meta-parameters θ b1 , the loss function Tb1 , the second support data set S Tbi and the corresponding scaling value α , and determining the finally obtained meta-parameters θ Tbi as the meta-parameters θ1 after the first training iteration operation, where 1 ≤ i ≤ B - 1. Tb(i+1) In one embodiment, preferably, the steps of calculating the meta-parameters θ after iteration operations on each of the remaining batches of tasks T b2 ~T bB include: calculating the values of the loss functions Tb(i+1) corresponding to the meta-parameters θ TbB ~θ
[0014] ~θ b2 ~T bB after each iteration operation; determining that the current iteration is effective and recording the corresponding meta-parameters θ Tb(i+1) in response to the value of any loss function Tb2 ~θ TbB being less than the value of the previous loss function ; and determining that the scaling value a is too large and halving the scaling value a and performing the next iteration until the maximum number of iterations is reached or convergence to the local minimum α1 is achieved in response to the value of any loss function Tb(i+1) ~θ being greater than or equal to the value of the previous loss function . Tb(i+1) is too large, halving the scaling value a Tb(i+1) and performing the next iteration until the maximum number of iterations is reached or convergence to the local minimum α1.
[0015] In one embodiment, preferably, the test phase includes the following steps: determining the window size N of the first support data set S p ; obtaining a plurality of test samples to form a test sample set; according to the second adaptive parameter a p (θ t , g t ) of the meta-parameters θ t , selecting N t test samples from the test sample set to generate a first support set S p without being corrected by the test phase; inputting the plurality of test samples into the neural network model f t to calculate the loss function corresponding to the meta-parameters θ θ ; t calculating the meta-parameters θ according to the meta-parameters θ t , the loss function , the first support data set S t and the scaling value α of the corresponding variable scaling step size t after the first correction iteration operation; re-selecting N p test samples from the test sample set according to the meta-parameters θ p to generate a first support set S p corrected by the test phase; calculating the loss function p according to the meta-parameters θ t , the meta-parameters θ p and their corresponding optimal scaling value α p ; calculating the gradient parameter g p of the loss function under the first support set S p ; and determining the first adaptive parameter a p corrected by the test phase according to the meta-parameters θ p and the gradient parameter g p (θ p , g p ).
[0016] In one embodiment, preferably, the training phase and the test phase further include the following steps: obtaining a plurality of key variable data of the chemical process, wherein the chemical process includes a continuous catalytic naphtha reforming process, and the key variable data includes input variable data and output variable data; preprocessing the plurality of key variable data according to the 3σ criterion to remove outliers and extreme values therefrom; and dividing the preprocessed plurality of key variable data into a plurality of training samples and a plurality of test samples according to a preset ratio.
[0017] In one embodiment, preferably, the data to be measured is the input variable data of the continuous catalytic naphtha reforming process. The steps of determining the quality index corresponding to the data to be measured include: in the model-agnostic meta-learning framework with variable scaling step size, according to the first adaptive parameter a corrected in the test stage p (θ p , g p ) and the first support set S corrected in the test stage p , predict the output variable data corresponding to the input variable data.
[0018] The present invention also provides a prediction device for quality index, which is characterized by including: a memory; and a processor, the processor is connected to the memory and is configured to implement the prediction method of the quality index according to any one of the above.
[0019] The present invention also provides a computer-readable storage medium, on which computer instructions are stored. It is characterized in that when the computer instructions are executed by a processor, the prediction method of the quality index according to any one of the above is implemented. Description of the Drawings
[0020] After reading the detailed description of the embodiments of the present disclosure in conjunction with the following drawings, the above features and advantages of the present invention can be better understood. In the drawings, the components are not necessarily drawn to scale, and components with similar related characteristics or features may have the same or similar reference numerals.
[0021] Figure 1 is a prediction result diagram of the neural network in the training set and the test set in the continuous catalytic naphtha reforming chemical process;
[0022] Figure 2 is a scatter diagram of the prediction effect of the neural network in the training set and the test set in the continuous catalytic naphtha reforming chemical process;
[0023] Figure 3 is a flowchart of the prediction method in the training stage and the test stage shown according to an embodiment of the present invention;
[0024] Figure 4 is a schematic flowchart of the prediction method of the quality index shown according to an aspect of the present invention;
[0025] Figure 5 is a prediction result diagram of three quality index prediction methods;
[0026] Figure 6 is an error comparison diagram of three quality index prediction methods; and
[0027] Figure 7 is a schematic structural diagram of the quality index prediction device shown according to another aspect of the present invention. Detailed Implementation Manner
[0028] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Although the description of the present invention will be introduced in combination with preferred embodiments, this does not mean that the features of this invention are limited to this implementation manner. On the contrary, the purpose of introducing the invention in combination with the implementation manner is to cover other alternatives or modifications that may extend based on the claims of the present invention. In order to provide a deep understanding of the present invention, many specific details will be included in the following description. The present invention can also be implemented without using these details. In addition, in order to avoid confusing or obscuring the key points of the present invention, some specific details will be omitted in the description.
[0029] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0030] In addition, the "upper", "lower", "left", "right", "top", "bottom", "horizontal", and "vertical" used in the following description should be understood as the orientations shown in this paragraph and the related drawings. This relative term is only for the convenience of description, and it does not mean that the device described needs to be manufactured or operated in a specific orientation, so it should not be understood as a limitation to the present invention.
[0031] It can be understood that although terms such as "first", "second", and "third" can be used here to describe various components, regions, layers, and / or parts, these components, regions, layers, and / or parts should not be limited by these terms, and these terms are only used to distinguish different components, regions, layers, and / or parts. Therefore, the first component, region, layer, and / or part discussed below can be called the second component, region, layer, and / or part without departing from some embodiments of the present invention.
[0032] Aiming at the problem of the change in the relationship between quality indicators and process variables existing in the prior art, the present invention proposes a prediction method for quality indicators, and the adaptation module based on stages obtains adaptation values under gradients scaled by variable scaling values, solving the problem of the change in the relationship between process variables and quality indicators in soft measurement.
[0033] Figure 3 It is a flowchart of the prediction method in the training stage and the testing stage shown according to an embodiment of the present invention.
[0034] Please refer to Figure 3 , the quality index prediction method provided by the present invention adopts different method steps in the training stage and the prediction stage.
[0035] As Figure 3 shown, in the training stage, it includes:
[0036] Step 301: Obtain multiple training samples to form a training sample set.
[0037] The training samples here can be obtained by collecting the original data and performing corresponding preprocessing on it. In one embodiment, the original data is the data of multiple key variables of a chemical process, such as the continuous catalytic reforming process of naphtha (CCR). The following table gives the physical meanings of the key variable data in the embodiment of the continuous catalytic reforming process of naphtha.
[0038] Table 1 Key variable data of the continuous catalytic reforming process of naphtha
[0039]
[0040]
[0041]
[0042]
[0043] In this embodiment, 84 physical quantities are sampled above, and about 10,000 samples are selected as the training sample set. Since the original data contains measurement noise, corresponding preprocessing operations need to be performed before the subsequent steps. For example, non-positive values and other outliers of the prediction labels are removed by the 3-sigma criterion to obtain the preprocessed data, and then the preprocessed data of multiple key variables is divided into multiple training samples and multiple test samples according to a preset ratio to form the training sample set and the test sample set respectively.
[0044] As Figure 3 shown, after constructing the training sample set, the quality index prediction method further includes:
[0045] Step 302: Initialize the neural network model f θ , determine the initial meta-parameters θ0 and the window size N t of the second support set S t ;
[0046] Step 303: Select N t training samples from the training sample set according to the initial meta-parameters θ0 to generate an initial second support set S0; and
[0047] Step 304: Input multiple training samples into the neural network model f θ to calculate the loss function corresponding to the initial parameter θ0
[0048] First, the neural network model f with meta-parameters θ θ is used as the basic model of the model-agnostic meta-learning framework with variable scaling step sizes. We want to learn an initial θ = θ0 such that for the support set S b a small number of N gradient updates are performed on the data to obtain θ N After that, the network performs well on the target set T of this task b Here b is the index of a specific support set task in a batch of support set tasks. This set of N update steps is called the inner loop update process. After obtaining data from the support task S b the updated base network parameters can be expressed as:
[0049]
[0050] where α is the learning rate, is the adapted parameter of the neural network after adapting i times for task b, is the loss function of the support set of b after (i - 1) times (i.e., the previous step) of update, which is also called the inner loop process. The meta-objective can be expressed as:
[0051]
[0052] where B represents the batch size of the task. Therefore, the relationship between θ and θ0 is given by the above formula. The formula measures the quality of the initialization based on the total loss using this initialization θ0 across all tasks. The meta-objective function is minimized to optimize the initial parameter value θ0 such that this parameter contains cross-task knowledge.
[0053] As Figure 3 shown, the quality metric prediction method further includes:
[0054] Step 305: Calculate the meta-parameter θ1 after the first training iteration operation according to the initial parameter θ0, the loss function the second support data set S0, and the scaling value α of the variable scaling step size.
[0055] As Figure 1 and Figure 2As shown, we found in the experiment that if the same number of iterations used in the training process is used during the prediction process, there is generally a problem of inaccurate prediction results. To address this issue, the present invention proposes a unified adaptation method at different stages, that is, using a stage-related adaptation module to summarize the existing meta-learning adaptation methods, which is mainly manifested as an approximation strategy. Usually, at the test stage, the local optimal parameters for batch tasks are achieved through sufficient iterations, while at the training stage, we obtain an effective estimate through an estimation strategy.
[0056] Following the assumption of MAML, it is considered that the adaptation parameters can be estimated through one gradient descent during the training stage, and a variable scaling step size is set in the training adaptation method. Therefore, the sub-goal in the training stage is to find the optimal scaling value.
[0057] Therefore, we use the loss function to take the derivative of the scaling value α, and use the gradient descent method to calculate the local minimum of the scaling value α as the optimal scaling value α1 corresponding to the meta-parameter θ1. Among them, the optimal scaling value α1 indicates the optimal scaling step size for iterating the initial meta-parameter θ0 to the meta-parameter θ1; then, according to the initial meta-parameter θ0, the loss function the initial second support set S0, and the optimal scaling value α1, the meta-parameter θ1 is calculated.
[0058] The specific process is as follows: Since the optimal scaling value is not a scalar, a method of meta-learning using single-step adaptation by finding the best adaptation parameters related to the support set is proposed during the training stage. The adaptation parameters in the training mode are constrained by the gradient direction, and the adaptation size is a variable scaling value. The best scaling value is a free variable λ, which is estimated according to the support data in the inner loop of the training stage. In this way, our training stage adaptation method can be expressed by the following formula:
[0059]
[0060]
[0061]
[0062] a t (θ0,f,T b )=θ a ,g a
[0063] To estimate the best scaling value during the training stage, we regard the variable scaling value as an additional parameter λ. In the actual process, let θ s be the adaptation parameter obtained according to the scaling value Then our optimization goal is:
[0064]
[0065] We express the adapted parameters as a function of the scaling value:
[0066]
[0067] where α is the scaling value. Our sub-goal is then to find the optimal α. To estimate this optimal scaling value, an initial scaling value can be initialized, and then the local optimal solution can be found through gradient descent. The derivative of the loss function with respect to the adaptation step size is expressed in the following form:
[0068]
[0069] where θ0 is the initial parameter of the network, α is the adaptation step size scalar, and θ s (α) is the adapted parameter of the initial parameter in the support set under the adaptation step size.
[0070] Therefore, through gradient descent, a local minimum can be continuously updated, and this value can be estimated more accurately than a scalar independent of the data. The parameter update formula is as follows:
[0071]
[0072]
[0073] where α is the current scaling value, α lr is the learning rate of the scaling value, and α new is the scaling value after one gradient update according to α, and θ s (α new ) is the estimated value of the updated adapted parameter. To accelerate the search for the optimal scaling value, we record the value of the adaptation loss function in the inner loop. If the loss function with respect to the adapted parameter decreases, then the adaptation is effective, and we record the corresponding adapted parameter. Otherwise, it is considered that the scaling step size is too large, the scaling step size is halved, and the next iteration is performed. This continues until the maximum number of iterations is reached and convergence to the local minimum is achieved.
[0074] As Figure 3 shown, the quality index prediction method further includes:
[0075] Step 306: Calculate the corresponding loss function according to the meta-parameter θ1 If the loss function is greater than or equal to the preset loss function threshold then calculate the meta-parameter θ2 and its corresponding loss function after the next training iteration operation again according to the meta-parameter θ1 and its corresponding second support data set S1 and so on.
[0076] By repeating the above method, θ1, θ2, … can be calculated. Based on the calculated meta-parameters θ i Redo the selection of N t training samples from the training sample set to generate the corresponding second support set S i , where 1 ≤ i ≤ t - 1; Based on the meta-parameters θ i-1 , meta-parameters θ i and their corresponding optimal scaling value α i calculate the loss function under the second support set S i gradient parameter g i ; And based on the meta-parameters θ i and the gradient parameter g i determine the adaptive parameter α i (θ i , g i ) after i training iteration operations.
[0077] As Figure 3 shown, the quality index prediction method further includes:
[0078] Step 307: Determine whether the loss function is less than the loss function threshold
[0079] Step 308: Determine the meta-parameters θ t after t training iteration operations as the meta-parameters trained in the training stage.
[0080] In response to the loss function being greater than or equal to the loss function threshold Based on the meta-parameters θ i and the loss function under the second support set S i gradient parameter g i , use the gradient descent method to calculate the optimal scaling value α i+1 corresponding to the meta-parameters θ i+1 after the next training iteration operation; And based on the meta-parameters θ i , the loss function second support set S i and the optimal scaling value α i+1 , calculate the meta-parameters θ i+1 , that is, go back to step 306 and loop until the condition is satisfied.
[0081] In response to the loss function being less than the loss function threshold Then determine the meta-parameters θ t after t training iteration operations as the meta-parameters trained in the training stage.
[0082] In one embodiment, a maximum number of iterations M is also set in the training phase. The training phase further includes the following steps: determining whether the current number of iterations has reached the maximum number of iterations M; and in response to the current number of iterations reaching the maximum number of iterations M, determining that the training phase is completed and taking the meta-parameters θ after M training iteration operations M as the meta-parameters θ trained through the training phase t .
[0083] In one embodiment, the tasks can also be batched to perform batch processing. The training sample set is divided according to the input task distribution p(T) to determine multiple batch tasks T b1 ~T bB , where each batch task T b1 ~T bB includes multiple training samples;
[0084] The step of selecting N t training samples from the training sample set according to the initial meta-parameters θ0 to generate an initial second support set S0 includes: selecting N b1 training samples from the multiple training samples of the first batch task T t to generate an initial second support set S0;
[0085] The step of inputting multiple training samples into the neural network model f θ to calculate the loss function corresponding to the initial parameter θ0 includes: inputting the multiple training samples of the first batch task T into the neural network model f b1 to calculate the loss function corresponding to the initial parameter θ0 of the first batch task T θ b1
[0086] The step of calculating the meta-parameters θ1 after the first training iteration operation includes:
[0087] According to the initial parameter θ0, the loss function the second support data set S0 and the corresponding scaling value α Tb1 , calculate the meta-parameters θ b1 after one iteration operation on the first batch task T Tb1 ; and
[0088] According to the meta-parameters θ Tbi , the loss function the second support data set S Tbi and the corresponding scaling value α Tb(i+1) , calculate one by one the meta-parameters for the remaining batch tasks T b2 ~T bB The meta-parameter θ after one iteration operation Tb(i+1) , and the finally obtained meta-parameter θ TbB is determined as the meta-parameter θ1 after the first training iteration operation, where 1 ≤ i ≤ B - 1.
[0089] In one embodiment, for each of the remaining batches of tasks T b2 ~T bB the steps of calculating the meta-parameter θ after the iteration operation include: after each iteration operation, calculating the corresponding loss function Tb(i+1) of each meta-parameter θ Tb2 ~θ TbB respectively; value;
[0090] In response to the value of any loss function being less than the value of the previous loss function , it is determined that the current iteration is effective, and the corresponding meta-parameter θ Tb(i+1) is recorded; and
[0091] In response to the value of any loss function being greater than or equal to the value of the previous loss function , it is determined that the scaling value a Tb(i+1) of the current iteration is too large, and the scaling value a Tb(i+1) is halved and the next iteration is performed until the maximum number of iterations is reached or convergence to the local minimum α1 is achieved.
[0092] The present invention uses an optimal scalable step size to replace the fixed step size in the traditional model-agnostic meta-learning training stage, so that only one iteration can be used in the prediction stage to replace the infinite iterations in the traditional method, achieving the effect of simplifying the test process.
[0093] After obtaining the trained meta-parameter θ t in the training stage, enter the test stage, as Figure 3 shown, the test stage includes:
[0094] Step 309: Determine the window size N p of the first support dataset S p ; obtain multiple test samples to form a test sample set;
[0095] Step 310: According to the second adaptive parameter a t (θ t , g t ) of the meta-parameter θ t , select N p test samples from the test sample set to generate the first support set S t without correction in the test stage; and input the multiple test samples into the neural network model f θto calculate the loss function corresponding to the meta-parameter θ t and calculate the meta-parameter θ t after the first correction iteration operation according to the meta-parameter θ , the loss function t , the first support dataset S t and the corresponding scaling value α of the variable scaling step size p ;
[0096] Step 311: Re-select N p test samples from the test sample set to generate the first support set S p after being corrected in the test stage; calculate the gradient parameter g p of the loss function t under the first support set S p according to the meta-parameter θ p , the meta-parameter θ and its corresponding optimal scaling value α p ; and p
[0097] Step 312: Determine the first adaptive parameter a p after being corrected in the test stage according to the meta-parameter θ q and the gradient parameter g p (θ p , g p ).
[0098] The above method of the stage-related adaptation module can establish the connection between the adaptation parameter and the initial parameter through constraints, making it effective for us to perform gradient update on the initial parameter using the adapted parameter.
[0099] Figure 4 is a schematic flowchart of the prediction method of the quality index shown according to one aspect of the present invention.
[0100] As Figure 4 shown, the quality index prediction method provided by the present invention includes:
[0101] Step 401: Input the data to be measured into the model-agnostic meta-learning framework with variable scaling step size; and
[0102] Step 402: In the model-agnostic meta-learning framework with variable scaling step size, determine the quality index corresponding to the data to be measured according to the first adaptive parameter a p (θ p , g p ) after being corrected in the test stage and the first support set S p after being corrected in the test stage.
[0103] Among them, the first adaptive parameter a p (θ p , g p ) is obtained by performing at least one correction iteration operation using multiple test samples in the test stage, based on the first support set S without correction in the test stage t . Among them, θ p is the meta-parameter of the neural network model after the correction iteration operation, and g p is the gradient parameter of the neural network model after the correction iteration operation
[0104] The first support set S corrected in the test stage p is generated by selecting multiple sample data from multiple test samples according to the meta-parameter θ corrected in the test stage p
[0105] The first support set S without correction in the test stage t is generated by selecting multiple sample data from multiple test samples according to the second adaptive parameter a trained in the training stage t (θ t , g t ). Among them, θ t is the meta-parameter of the neural network model after t training iteration operations, and g t is the gradient parameter of the neural network model after t training iteration operations t
[0106] The second adaptive parameter a t (θ t , g t ) is obtained by performing multiple training iteration operations using multiple training samples in the training stage, based on the second support sets S0 to S corresponding to the current meta-parameters θ0 to θ t-1 . Among them, θ0 to θ t-1 are the meta-parameters of the neural network model after the corresponding number of training iteration operations, and each of the second support sets S0 to S t-1 is respectively generated by selecting multiple sample data from multiple training samples according to the corresponding meta-parameters θ0 to θ t-1 t-1
[0107] The steps to determine the quality index corresponding to the data to be measured include: in the model-agnostic meta-learning framework with variable scaling step size, according to the first adaptive parameter a corrected in the test stage p (θ p , g p ) and the first support set S corrected in the test stage p , predicting the output variable data corresponding to the input variable data
[0108] In the embodiments of the continuous catalytic reforming process (CCR) described above, the quality index or output variable data is RON Barrel (Research Octane Number Barrel).
[0109] Although the above methods are illustrated and described as a series of actions for simplicity of explanation, it should be understood and appreciated that these methods are not limited by the order of the actions, because according to one or more embodiments, some actions may occur in a different order and / or concurrently with other actions that are illustrated and described herein or that are not illustrated and described herein but are understood by those skilled in the art.
[0110] Figure 5 is the prediction result diagram of three quality index prediction methods; Figure 6 is the error comparison diagram of three quality index prediction methods.
[0111] As Figure 6 shown, RMSE and R 2 can be used to measure the accuracy of the method
[0112]
[0113]
[0114] More accurate calculation data are shown in Tables 2 and 3 below:
[0115] Table 2: Prediction Results on the Test Set
[0116] RMSE-T <![CDATA[R 2 -T]]> RMSE-P <![CDATA[R 2 -P]]> MAML 0.5740 -0.6701 0.1783 0.8774 Reptile 0.1913 0.8498 0.1769 0.8777 MAML-OASV 0.2727 0.8269 0.1388 0.9200
[0117] Table 3: Prediction Results on the Test Set by the Method of Adapting in the Training Phase and Prediction Phase with Meta-Learning
[0118] PLS NN MAML Reptile MAML-OASV RMSE 0.5855 0.7060 0.1787 0.1769 0.1388 <![CDATA[R 2 > -3.9751 -3.7765 0.8774 0.8777 0.9200
[0119] From Figure 5 , 6 and Tables 2 and 3, it can be seen that compared with other model-agnostic meta-learning methods, such as MAML and Reptile, the model-agnostic meta-learning method MAML-OASV based on the optimal scalable step size proposed by the present invention better solves the problem of the change of variable relationships in industrial processes and achieves the most effective prediction effect.
[0120] Figure 7 is a schematic structural diagram of a quality index prediction device according to another aspect of the present invention.
[0121] As Figure 7 shown, the present invention also provides a quality index prediction device 700, including a memory 701 and a controller 702 connected thereto, and the controller 702 is configured to implement the steps of any one of the above quality index prediction methods.
[0122] Although the controller 702 of the above embodiments can be implemented by a combination of software and hardware, it can be understood that the controller 702 can also be implemented in software or hardware. For hardware implementation, the controller 702 can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DAPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic devices for performing the above functions, or a selected combination of the above devices. For software implementation, the controller 702 can be implemented by independent software modules such as procedures and functions running on a general-purpose chip, where each module performs one or more of the functions and operations described herein.
[0123] The present invention also provides an embodiment of a computer-readable medium having computer instructions stored thereon, which when executed by a processor, implement the steps of any of the above quality index prediction methods.
[0124] Those skilled in the art will appreciate that information, signals, and data can be represented using any of a variety of different technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips described throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.
[0125] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and the design constraints imposed on the overall system. Skilled artisans may implement the described functionality in different ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.
[0126] The various illustrative logical modules and circuits described in connection with the embodiments disclosed herein can be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.
[0127] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read from, and write to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0128] In one or more exemplary embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code. The computer-readable medium includes both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. The storage media may be any available media that can be accessed by a computer. By way of example and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a web site, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. As used herein, the terms "disk" and "disc" include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0129] The foregoing description of the disclosure has been provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A prediction method for quality indicators, characterized in that, comprising the following steps, inputting the data to be measured into a model-agnostic meta-learning framework with variable scaling step sizes, wherein the data to be measured includes input variable data of a continuous catalytic naphtha reforming process; In the model-agnostic meta-learning framework with variable scaling step sizes, according to the first adaptive parameter corrected during the testing phase and the first support set corrected during the testing phase , determine the quality metric corresponding to the data to be tested, where the quality metric includes the output variable data of the continuous catalytic naphtha reforming process; and The first adaptive parameter corrected through the testing phase and the first support set corrected through the testing phase are determined as follows: Determine the first support set corrected through the testing phase with a window size ; obtaining a plurality of test samples to form a test sample set; According to the second adaptive parameter of the meta-parameter , select the test samples from the test sample set to generate a first support set without correction in the test phase , where the second adaptive parameter is obtained by using multiple training samples in the training phase and performing multiple training iteration operations according to the second support sets ~ corresponding to the current meta-parameters ~ . The ~ are the meta-parameters of the neural network model after the corresponding number of training iteration operations. Each of the second support sets ~ is respectively generated by selecting multiple sample data from the multiple training samples according to the corresponding meta-parameters ~ . The is the meta-parameter of the neural network model after t training iteration operations, and the is the gradient parameter of the neural network model after t training iteration operations; Input the multiple test samples into the neural network model to calculate the loss function corresponding to the meta-parameters ; ; According to the meta-parameters , the loss function , the first support set and the scaling value corresponding to the variable scaling step , calculate the meta-parameters after the first correction iteration operation ; According to the meta-parameters , reselect test samples from the test sample set to generate the first support set corrected through the test stage ; According to the meta-parameters and the meta-parameters and their corresponding optimal scaling values , calculate the gradient parameter of the loss function under the first support set ; According to the meta-parameters and the gradient parameters , determine the first adaptive parameter corrected through the testing phase .
2. The prediction method according to claim 1, wherein The training stage includes the following steps: A neural network model with meta-parameters θ as the basic model of the model-agnostic meta-learning framework with the variable scaling step size; For the neural network model perform initialization to determine the initial meta-parameters and the second support set ~ window size ; obtaining the plurality of training samples to form a training sample set; According to the initial meta-parameters select from the training sample set training samples to generate an initial second support set ; Input the multiple training samples into the neural network model to calculate the loss function corresponding to the initial meta-parameters ; ; According to the initial meta-parameters , the loss function , the second support set and the scaling value of the variable scaling step , calculate the meta-parameters after the first training iteration operation ; and According to the meta-parameters calculate the corresponding loss function , if the loss function is greater than or equal to a preset loss function threshold , then according to the meta-parameters and its corresponding second support set recalculate the meta-parameters after the next training iteration operation and its corresponding loss function , and so on until the loss function is less than the loss function threshold , and then determine the meta-parameters after t times of the training iteration operation as the meta-parameters trained through the training stage.
3. The prediction method according to claim 2, wherein The step of calculating the meta-parameters after the first said training iteration operation comprises: Using the loss function for the scaling value take the derivative and use the gradient descent method to calculate the local minimum of the scaling value as the meta-parameter corresponding optimal scaling value , where the optimal scaling value indicates the optimal scaling step for iterating the initial meta-parameter to the meta-parameter ; and According to the initial meta-parameters , the loss function , the initial second support set and the optimal scaling value , calculate the meta-parameters .
4. The prediction method according to claim 3, wherein The training stage further includes the following steps: Meta-parameters obtained according to the calculation Re-select from the training sample set training samples to generate a corresponding second support set , where -1; According to the meta-parameters and the said meta-parameters and their corresponding optimal scaling values calculate the loss function The gradient parameters under the said second support set ; and Based on the meta-parameters and the gradient parameters determine the adaptive parameters after such training iteration operations 5. The prediction method according to claim 4, characterized in that The steps of and so on include: In response to the loss function being greater than or equal to the loss function threshold , according to the meta-parameters and the loss function and the gradient parameters in the second support set , use the gradient descent method to calculate the meta-parameters corresponding to the optimal scaling value after the next training iteration operation; and According to the meta-parameter , the loss function , the second support set and the optimal scaling value , calculate the meta-parameter .
6. The prediction method according to claim 2, wherein The training stage is further provided with a maximum number of iterations M, and the training stage further includes the following steps: judging whether the current number of iterations reaches the maximum number of iterations M; and In response to the current iteration number reaching the maximum iteration number M, it is determined that the training phase is completed, and the meta-parameters after M times of the training iteration operations are determined as the meta-parameters trained through the training phase .
7. The prediction method according to claim 2, wherein The training phase further includes the following steps: according to the input task distribution divide the training sample set to determine a plurality of batch tasks ~ , wherein each of the batch tasks ~ includes a plurality of the training samples The step of selecting according to the initial meta-parameters from the training sample set to generate an initial second support set includes: according to the initial meta-parameters select from multiple training samples of the first batch of tasks to generate an initial second support set and , inputting the plurality of training samples into the neural network model to calculate a loss function corresponding to the initial meta-parameters ; the steps include: inputting the plurality of training samples of the first batch of tasks into the neural network model to calculate the loss function of the first batch of tasks corresponding to the initial meta-parameters ; the loss function , The step of calculating the meta-parameters after the first said training iteration operation includes: According to the initial meta-parameters , the loss function , the second support set and the corresponding scaling value , calculate the meta-parameters after one iteration operation on the first batch of tasks ; and According to the meta-parameters , the loss function , the second support set and the corresponding scaling value , calculate one by one the meta-parameters ~ after performing one iteration operation on the remaining batch tasks , and determine the finally obtained meta-parameters as the meta-parameters after the first training iteration operation , where .
8. The prediction method according to claim 7, wherein The step of calculating one by one for each of the remaining batch tasks ~ The meta-parameters after performing one iteration operation includes: After each of the iterative operations, calculate each of the meta-parameters respectively ~ corresponding loss function value; In response to any loss function value is less than the previous loss function value, determine that the current iteration is valid and record the corresponding meta-parameters ; and In response to any loss function whose value is greater than or equal to the previous loss function determine that the scaling value of the current iteration is too large, halve the scaling value and perform the next iteration until the maximum number of iterations is reached or convergence to a local minimum .
9. The prediction method according to claim 1, characterized in that The training stage and the testing stage further include the following steps: obtaining a plurality of key variable data of a chemical process, wherein the chemical process includes the continuous catalytic naphtha reforming process, and the key variable data includes input variable data and output variable data; preprocessing the plurality of key variable data according to the 3σ criterion to remove outliers and extreme values therein; and dividing the plurality of key variable data after the preprocessing into the plurality of training samples and the plurality of test samples according to a preset ratio.
10. The prediction method according to claim 9, characterized in that, The step of determining the quality index corresponding to the data to be measured includes: In the model-agnostic meta-learning framework with variable scaling step sizes, based on the first adaptive parameter corrected during the testing phase and the first support set corrected during the testing phase , predict the output variable data corresponding to the input variable data.
11. A prediction device for a quality index, characterized in that, including: a memory; and a processor, the processor being connected to the memory and configured to implement the prediction method of the quality index according to any one of claims 1 to 10.
12. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the computer instructions are executed by the processor, the prediction method of the quality index according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Local self-adaptive WNN (Wavelet Neural Network) training system, device and method
CN103676649A
Ash haze predicting system and method based on BP neural network
CN106067079A