Count quality variable prediction method based on variational bayesian gaussian-poisson mixed regression model
By using a variational Bayesian Gaussian-Poisson mixed regression model, the multimodal fitting problem of non-negative discrete count data in industrial production was solved, achieving efficient prediction of count-type quality variables and improving data fitting effect and training efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2023-03-06
- Publication Date
- 2026-04-21
AI Technical Summary
Existing soft measurement modeling methods cannot effectively fit non-negative discrete count data in industrial production processes, especially in multimodal cases, where finite mixture models (FMMs) cannot provide non-negative discrete probability estimates.
A variational Bayesian Gaussian-Poisson mixed regression model is adopted. By mixing multiple probability density functions, the model is trained using variational inference methods, sharing the mixing coefficients to represent the multimodal characteristics of process variables and quality variables, and optimizing the parameter learning process through Laplace approximation.
It improves the fitting ability to multimodal data, realizes discrete probability estimation of count-type quality variables, has a shorter training time than the Markov chain-Monte Carlo method, and has better prediction performance than other methods.
Smart Images

Figure CN116224937B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of industrial process prediction and control, specifically relating to a method for predicting count-type quality variables based on a variational Bayesian-Gaussian-Poisson mixed regression model. Background Technology
[0002] In industrial production, the measurement data generated covers all aspects of "people," "machines," "materials," "methods," and "environment," such as the quality information of raw materials in each production process, process parameters, equipment operating status, and the experience and working status of operators and inspectors. Among these, there are some very important count data, which are non-negative integers like 0, 1, 2, ... , such as the number of equipment failures within a certain period, the number of defects in products, and other quality variables expressed as non-negative integers. For key quality indicators like the number of product defects, soft measurement technology can be used to achieve online real-time prediction, thereby reducing inspection time, improving production efficiency, and enabling production managers to monitor product quality in real time and react more quickly to defective production lines.
[0003] Conventional soft sensor modeling methods mainly include those based on multivariate statistical analysis, statistical learning, deep learning, and probability estimation. Multivariate statistical analysis-based methods, such as ordinary least squares, partial least squares, and multiple linear regression, all rely on the assumption that variables follow a Gaussian distribution, failing to provide support for non-negativity and discreteness in count data. Statistical learning and deep learning-based methods, such as SVR and ANN, lack data interpretability and similarly fail to meet the non-negativity and discreteness requirements of count data. Probability estimation-based methods, such as Gaussian process regression and Gaussian mixture regression, are often based on Gaussian or Gaussian mixture distributions, which are also unsuitable for fitting non-negativity and discrete count data. Therefore, it is necessary to utilize count data models for regression prediction of count data in industrial production processes. Poisson regression, as a typical discrete regression method, is widely used in basic research on count data modeling. Other commonly used count data modeling methods, such as negative binomial regression and Poisson mixture regression, are all based on Poisson regression.
[0004] In industrial production processes, besides the counting characteristics of the data, another important feature is that the values of independent and dependent variables, as well as their mapping relationships, change due to the switching of machines between multiple operating conditions. Therefore, the data exhibits multimodality and no longer generally follows a single Gaussian distribution. Finite Mixture Models (FMMs) can specifically address the multimodal problems in actual industrial processes because they use a weighted sum of multiple probability density functions as the overall probability density of the random variable. Compared to a single probability density function, this allows for learning to fit the actual situation of multiple peaks and troughs under multimodal conditions. Theoretically, a mixture model with infinite components can approximate any unknown random distribution. However, the finite mixture models (FMMs) currently used in industrial production processes cannot provide non-negative discrete probability estimates for count-type quality variables. Summary of the Invention
[0005] To address the issues of nonnegative discreteness of count-type quality variables and multimodal industrial processes, this invention proposes a prediction method for count-type quality variables based on a variational Bayesian-Gaussian-Poisson mixed regression model.
[0006] The specific technical solution of the present invention is as follows:
[0007] A method for predicting count-type quality variables based on a variational Bayesian Gaussian-Poisson mixed regression model includes the following steps:
[0008] (1) Collect data samples of count-type quality variables and related process variables as the training set of the model; perform preprocessing of missing values, outliers and standardization on the data to obtain the processed training set;
[0009] (2) Using variational inference, a variational Bayesian Gaussian-Poisson mixed regression model is trained offline on the processed training set; the expression of the variational Bayesian Gaussian-Poisson mixed regression model is:
[0010]
[0011]
[0012]
[0013] in, z is the set of process variables, quality variables, and latent variables for all samples. ik For z i The value of the k-th dimension; β k Let be the regression coefficient of the k-th Poisson regression component; μ represents the density function of a Gaussian distribution. k Λ k Let be the mean vector and precision matrix of the k-th Gaussian distribution component, respectively; πk These are the mixing coefficients of the mixture model;
[0014] β k μ k Λ k π k Obey the following priors:
[0015] ①β k They all follow a Gaussian distribution, that is Where β0 and Σ0 are β k The mean vector and covariance matrix of the Gaussian distribution, where β is the resultant of all components. k All priors are the same;
[0016] ②μ k and Λ k They all follow a Gaussian-Wishart distribution, i.e.
[0017] Where m0 and γ0 are μ k The mean vector and scale parameter of the distribution. This represents the probability density function of the Wishart distribution. Let v0 > d⁻¹ be the scale matrix of the distribution, and v₀ > d⁻¹ be the number of degrees of freedom of the distribution.
[0018] ③π k The joint distribution follows a Dirichlet distribution, i.e. Where a0 is the parameter of the Dirichlet distribution, a0 > 0 to ensure that the distribution can be normalized.
[0019] (3) Collect the test sample, and after the same preprocessing as in step (1) for missing values, outliers, and standardization, obtain the processed test sample x. q Then, the variational Bayesian Gaussian-Poisson mixed regression model trained in step (2) is used for online prediction.
[0020] Furthermore, in step (2), the specific steps for offline training of the variational Bayesian Gaussian-Poisson mixture regression model on the processed training set are as follows:
[0021] (2.1) Using Bayes' theorem, the latent variables in the model are obtained. and parameter variable β k μ k Λ k π k Variational posterior distribution:
[0022] ①
[0023] ②
[0024] ③
[0025] ④
[0026] Where, ξ ik τ k Σ k m k γ k W k v k and a k Let be the distribution parameters of these variational posterior distributions, where i = 1, 2, ..., N, k = 1, 2, ..., K;
[0027] (2.2) Set the hyperparameters a0, m0, γ0, W0, v0, β0, Σ0 of the prior distribution of the model parameters, the number of mixed components K, and the iteration convergence threshold δ, the maximum number of iterations M, and the current iteration number t = 0;
[0028] (2.3) Random initialization: First, a multinomial distribution is used to generate the initialization process. <z i The initial values of > (i = 1, 2, ..., N), where <z i >for z i The expectation; the parameters of the multinomial distribution are K-dimensional vectors, and the value of the vector is 1 / K; then other expectations are randomly initialized, including: <β k >、<(x i -μ k ) T Λ k (x i -μ k )>、 <ln|Λ k |>、 <lnπ k >, among which <β k > represents the function within parentheses in the variational posterior distribution q. * (β k The expectation on ) <(x) i -μ k ) T Λ k (x i -μ k )>、 <ln|Λ k |> represents the function within the parentheses in q. * (μ k ,Λ k ) on expectations, <lnπ k > represents the function within parentheses in q * Expectation on (π);
[0029] (2.4) Increment the current iteration number by 1, i.e., t = t + 1; for i = 1, 2, ..., N and k = 1, 2, ..., K, calculate the parameters of the variational posterior distribution defined by the model according to the following formula:
[0030] ① in,
[0031]
[0032] ②
[0033] ③
[0034] ④
[0035] ⑤
[0036] ⑥
[0037] ⑦ in, The Newton-Raphson method can be used for optimization. The specific steps are as follows: First, randomly initialize... Set the optimal convergence threshold δ β Using formula renew Until The maximum value in the absolute difference vector before and after the update is less than δ. β ,in,
[0038]
[0039]
[0040] ⑧
[0041] (2.5) Calculate the expected value involved in the model according to the following formula:
[0042] ① <z ik >=ξ ik ,
[0043] ②<β k >=τ k ,
[0044] ③
[0045] ④
[0046] ⑤
[0047] ⑥ Where ψ(·) is the digamma function,
[0048] ⑦
[0049] Where i = 1, 2, ..., N; k = 1, 2, ..., K;
[0050] (2.6) Calculate the lower bound of evidence for the current iteration step using the following formula.
[0051]
[0052] Where <·> represents the expectation of the function enclosed in parentheses on the variational posterior distribution of all parameters involved in the function, and the specific calculation formula is as follows:
[0053] ①
[0054] ②
[0055] ③
[0056] ④
[0057] ⑤ in,
[0058]
[0059] ⑥
[0060] ⑦
[0061] ⑧
[0062] 9 Among them, B(W) k ,v k It has the same functional form as B(W0,v0) above;
[0063] ⑩
[0064] (2.7) Repeat step (2.4) until the maximum number of iterations t = M is reached or in, The lower bound of evidence is the value calculated in the (t-1)th iteration, i.e., the previous iteration. When t=1
[0065] Furthermore, the calculation formula for online prediction using the model trained in step (2) in step (3) is as follows:
[0066]
[0067] in, Let be the predicted value of the sample to be tested, and St(·) be the probability density function of Student's distribution. v k -d+1 is the parameter of the function.
[0068] The beneficial effects of this invention are as follows:
[0069] (1) The prediction method of the present invention mixes multiple probability density functions in a finite mixing manner, i.e. in a weighted manner, by using two mixed models to share the mixing coefficients to represent that the process variable and the quality variable come from the same mode, thereby improving the fitting ability of multimodal data.
[0070] (2) This invention uses the Poisson mixture distribution to characterize count-type quality variables in industrial processes, thereby realizing discrete probability estimation of count-type quality variables.
[0071] (3) This invention uses variational inference to train the variational Bayesian Gaussian-Poisson mixed regression model. In this process, the Laplace approximation is used, that is, the non-conjugate variational posterior of the Poisson regression parameters is approximated as a conjugate Gaussian distribution, and an iterative parameter learning process is derived. Compared with the existing Markov chain-Monte Carlo method, the training time is shorter.
[0072] Furthermore, as an important type of discrete data, count data not only exists in industrial production processes, but also widely exists in other professional fields such as actuarial science, biostatistics, economics, and sociology. The method proposed in this invention is also of certain value for the study of count data in these fields. Attached Figure Description
[0073] Figure 1 This is a flowchart of the count-type mass variable prediction method of the present invention.
[0074] Figure 2 The data is generated by the numerical simulation system (Example 1); where Figure (a) shows the data points of the process variables of real number type, Figure (b) shows the distribution histogram of the count type mass variables, and Figure (c) shows the values of the count type mass variables.
[0075] Figure 3 This is the average value of the training set prediction index of the method proposed in this invention under different initial component numbers (Example 1).
[0076] Figure 4These are the prediction curves of various methods on the test set (Example 1); where Figure (a) represents the PLS prediction result, Figure (b) represents the PR prediction result, Figure (c) represents the NBR prediction result, and Figure (d) represents the GPMR prediction result.
[0077] Figure 5 This is a process flow chart for rolling medium and heavy steel plates.
[0078] Figure 6 It is a histogram showing the distribution of the number of defects in the steel plate.
[0079] Figure 7 This is the average MAE predicted on the training set by the method proposed in this invention under different initial component numbers (Example 2).
[0080] Figure 8 This is the prediction curve of the method proposed in this invention on the test set (Example 2).
[0081] Figure 9 Figure 2 shows partial curves of prediction results for the test set (samples 0-300) by various methods. Figure (a) shows the PLS partial prediction result, Figure (b) shows the PR partial prediction result, Figure (c) shows the NBR partial prediction result, and Figure (d) shows the GPMR partial prediction result. Detailed Implementation
[0082] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. The purpose and effects of the present invention will become clearer. It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0083] The present invention provides a method for predicting count-type quality variables based on a variational Bayesian Gaussian-Poisson mixed regression model. The variational Bayesian Gaussian-Poisson mixed regression model is defined as follows: for the i-th data sample (x... i ,y i ), where x i Represented as a process variable, y i The mass variable y is represented as a count type mass variable, and the design number type mass variable is y. i The process variable x follows a Poisson regression mixture model with K components, and is a real number. i It follows a Gaussian mixture model with the same number of components and the same mixing coefficient, i.e.
[0084]
[0085]
[0086] Where, β k Let be the regression coefficient of the k-th Poisson regression component. Let μ represent the density function of the Gaussian distribution. k Λ k These are the mean vector and precision matrix of the k-th Gaussian distribution component, respectively, and π k It is the mixing coefficient of the mixture model, which can be understood as the probability that the i-th sample comes from the k-th component.
[0087] Introducing the latent variable z i For mixed models, the latent variables follow the parameter π = (π1, π2, ..., π). K Given a multinomial distribution of K-dimensional binary variables (i.e., taking values only {0, 1} and having only one dimension equal to 1), the variational Bayesian Gaussian-Poisson mixed regression model for the entire training set is transformed into the following form:
[0088]
[0089]
[0090]
[0091] in, z is the set of process variables, quality variables, and latent variables for all samples. ik For z i The value of the k-th dimension.
[0092] Model parameter β k μ k Λ k π k The following priors apply, where k = 1, 2, ..., K:
[0093] ①β k They all follow a Gaussian distribution, that is Where β0 and Σ0 are β k The mean vector and covariance matrix of the Gaussian distribution, where β is the resultant of all components. k All priors are the same;
[0094] ②μ k and Λ k They all follow a Gaussian-Wishart distribution, i.e.
[0095] Where m0 and γ0 are μ k The mean vector and scale parameter of the distribution. This represents the probability density function of the Wishart distribution. Let v0 > d⁻¹ be the scale matrix of the distribution, and v₀ > d⁻¹ be the number of degrees of freedom of the distribution.
[0096] ③π kThe joint distribution follows a Dirichlet distribution, i.e. Where a0 is the parameter of the Dirichlet distribution, a0 > 0 to ensure that the distribution can be normalized.
[0097] like Figure 1 As shown, the specific steps of the count-type quality variable prediction method based on the variational Bayesian Gaussian-Poisson mixed regression model of the present invention are as follows:
[0098] Step 1: Collect data samples of count-type quality variables and related process variables as the training set for the model. Preprocess the data by removing missing values, outliers, and standardizing to obtain a processed training set.
[0099] Step 2: Using variational inference methods, train a variational Bayesian Gaussian-Poisson mixture regression model offline on the processed training set; the specific steps are as follows:
[0100] (2.1) Using Bayes' theorem, the latent variables in the model are obtained. and parameter variable β k μ k Λ k π k The variational posterior distribution of (k = 1, 2, ..., K) is as follows:
[0101] ①
[0102] ②
[0103] ③
[0104] ④
[0105] Where, ξ ik τ k Σ k m k γ k W k v k and a k Let be the distribution parameters of these variational posterior distributions, where k = 1, 2, ..., K.
[0106] (2.2) Set the hyperparameters a0, m0, γ0, W0, v0, β0, Σ0 of the prior distribution of the model parameters, the number of mixed components K, and the iteration convergence threshold δ, the maximum number of iterations M, and the current iteration number t = 0;
[0107] (2.3) Random initialization: First, generate <z using a multinomial distribution. i > initial values of (i = 1, 2, ..., N), where < zi > for z i The expectation; the parameters of the multinomial distribution are K-dimensional vectors, and the value of the vector is 1 / K; then other expectations are randomly initialized, including: <β k >、<(x i -μ k ) T Λ k (x i -μ k )>、 <ln|Λ k |>、 <lnπ k >, among which <β k > represents the variational posterior distribution q of the function enclosed in parentheses. * (β k The expectation on ) <(x) i -μ k ) T Λ k (x i -μ k )>、 <ln|Λ k |> represents the function within the parentheses in q. * (μ k ,Λ k ) on expectations, <lnπ k > represents the function within parentheses in q * Expectation on (π).
[0108] (2.4) Increment the current iteration number by 1, i.e., t = t + 1. For i = 1, 2, ..., N and k = 1, 2, ..., K, calculate the parameters of the variational posterior distribution defined by the model according to the following formula:
[0109] ① in,
[0110]
[0111] ②
[0112] ③
[0113] ④
[0114] ⑤
[0115] ⑥
[0116] ⑦ in, The Newton-Raphson method can be used for optimization. The specific steps are as follows: First, randomly initialize... Set the optimal convergence threshold δ β Using formula renew Until The maximum value in the absolute difference vector before and after the update is less than δ. β ,in,
[0117]
[0118]
[0119] ⑧
[0120] (2.5) Calculate the expected value involved in the model according to the following formula:
[0121] ① <z ik >=ξ ik ,
[0122] ②<β k >=τ k ,
[0123] ③
[0124] ④
[0125] ⑤
[0126] ⑥ Where ψ(·) is the digamma function,
[0127] ⑦
[0128] Where i = 1, 2, ..., N; k = 1, 2, ..., K.
[0129] (2.6) Calculate the lower bound of evidence for the current iteration step using the following formula.
[0130]
[0131] Where <·> represents the expectation of the function enclosed in parentheses on the variational posterior distribution of all parameters involved in the function, and the specific calculation formula is as follows:
[0132] ①
[0133] ②
[0134] ③
[0135] ④
[0136] ⑤ in,
[0137]
[0138] ⑥
[0139] ⑦
[0140] ⑧
[0141] 9 Among them, B(W) k ,v k It has the same functional form as B(W0,v0) above.
[0142] ⑩
[0143] (2.7) Repeat step (2.4) until the maximum number of iterations t = M is reached or in, The lower bound of evidence is the value calculated in the (t-1)th iteration, i.e., the previous iteration. When t=1
[0144] Step 3: Collect the test samples on site. After preprocessing and standardizing with the same missing and outlier values as the training data in (1), the processed test samples x are obtained. q Then, the variational Bayesian Gaussian-Poisson mixed regression model trained in (2) is used for online prediction. The specific calculation formula for the predicted value is as follows:
[0145]
[0146] in, Let be the predicted value of the sample to be tested, and St(·) be the probability density function of Student's distribution. v k -d+1 is the parameter of the function.
[0147] The effectiveness of the invention is verified below using a numerical simulation example and a specific example of soft measurement of the number of defects in medium-thick steel plates during the rolling process. The model's prediction performance is evaluated using the MAE and R-squared values of the prediction results on the test set. 2 For quantification, the formula for calculating MAE is:
[0148]
[0149] Where, N t It is the total number of samples to be tested. and yi R represents the predicted and actual values of the quality variable for the i-th sample to be tested, respectively. 2 The calculation formula is:
[0150]
[0151] in, It is the average of all the quality variables of the samples to be tested. The smaller the MAE, the more accurate the prediction result. R 2 The smaller the difference from 1, the better the model fits.
[0152] Example 1:
[0153] This embodiment is a numerical simulation. The input to the numerical simulation system, i.e., the process variable, is two-dimensional real data, and the output, i.e., the quality variable, is one-dimensional count data. The input is generated using a three-component Gaussian mixture distribution, and the output is generated using a three-component Poisson mixture regression distribution with the same mixing coefficients as the three-component Gaussian mixture distribution. The mixing coefficients of the two distributions, i.e., the distribution parameters of the three components, are shown in Table 1.
[0154] Table 1 Parameter Configuration of Numerical Simulation System
[0155]
[0156] 2500 samples were collected from the system, of which 2000 were used for training and 500 were used for testing. The generated samples are as follows: Figure 2 As shown, where, Figure 2 In the diagram, (a) represents the input data points for 2500 samples. Figure 2 (b) in the figure shows the main distribution histogram of the outputs for these samples, which is concentrated between 0 and 150. Figure 2 In the diagram, (c) represents the numerical value of the output for these samples. Figure 2 As can be seen in (c), the numerical distributions of the mass variables of the three components are significantly different.
[0157] In this embodiment, the hyperparameter a0 = 1e-3 of the prior distribution of the model parameters in step (2.2) of the present invention... γ0=1e-3, W0=I, v0=2, β0=0, Σ0=I, iterative convergence threshold δ=1e-6, maximum number of iterations M=2000; except for step (2.3) <z i All expectations listed above are initialized to zero; the Newton-Raphson optimization update in step (2.4) The set convergence threshold δ β =1e-6, initial value is
[0158] The method proposed in this invention is denoted as GPMR. GPMR models with different numbers of mixture components K = 1, 2, ..., 10 are trained. The training set prediction MAE obtained from 10 independent experiments for each component is as follows: Figure 3 As shown. From Figure 3 As can be seen, when K≥3, MAE stabilizes at a certain value. Therefore, K=3 is chosen as the optimal initial number of components for GPMR in this embodiment.
[0159] Poisson regression (PR), negative binomial regression (NBR), and partial least squares (PLS) were used as comparison methods. The optimal discrete parameters of NBR and the optimal number of principal components of PLS were determined by grid search. The discrete parameter of NBR is α = 3e-2, and the number of principal components of PLS is K = 2. The prediction quantification indices of each method are shown in Table 2. GPMR is the method proposed in this invention.
[0160] Table 2. Predictive Quantitative Indicators for Each Method (Example 1)
[0161] PR NBR PLS GPMR MAE 7.54 7.538 8.844 2.688 <![CDATA[R 2 ]]> 0.847 0.849 0.557 0.912
[0162] As can be seen from the prediction quantification metrics, the MAE predicted by the method proposed in this invention on the test set is significantly lower than that predicted by other methods, and the prediction R of the method proposed in this invention is also significantly lower. 2 The value closest to 1 indicates that the proposed method fits the data best. Figure 4 The prediction curves for each method on the test set are shown. Among the prediction curves of the four methods, the prediction curve of GPMR (3.d) best matches the actual data curve of the test set.
[0163] Example 2:
[0164] This embodiment is based on the actual rolling process of medium and heavy steel plates, and all data were collected from the steel plate rolling process of a steel plant. Figure 5 This diagram illustrates the steel plate rolling process, primarily including steps such as steelmaking, casting, rolling, cooling, and cutting. The collected intact data was divided into a training set and a test set, with 6000 samples in the training set and 2000 samples in the test set. 153 relevant process variables were selected, including element content, casting speed, continuous casting temperature, and heating time during steelmaking. The quality variable was the number of defects in the steel plate, and the histogram showing the distribution of defects is shown below. Figure 6 As shown. In this embodiment, the hyperparameter a0 = 1e-3 of the prior distribution of the model parameters in step (2.2) is... γ0=1e-3, W0=I, v0=153, β0=0, Σ0=I, iterative convergence threshold δ=1e-6, maximum number of iterations M=2000; except for step (2.3) <z iAll expectations listed above are initialized to zero; the Newton-Raphson optimization update in step (2.4) The set convergence threshold δ β =1e-6, initial value is
[0165] GPMR models with different numbers of mixture components K = 1, 2, ..., 30 were trained separately. The MAE predictions obtained from the training set of 10 independent experiments for each component are as follows: Figure 7 As shown in the figure, the predicted MAE of GPMR decreases continuously with the increase of the initial number of mixed components, until it tends to level off when K≥24. Considering that an excessively large number of initial components would also increase the complexity of the model, K=24 was chosen as the optimal number of initial components for GPMR in this embodiment.
[0166] Table 3. Predictive Quantitative Indicators for Each Method (Example 2)
[0167] PR NBR PLS GPMR MAE 2.947 2.944 3.242 1.994 <![CDATA[R 2 ]]> 0.822 0.820 0.651 0.863
[0168] Poisson regression (PR), negative binomial regression (NBR), and partial least squares regression (PLS) were used as comparative methods. The optimal discrete parameters for NBR and the optimal number of principal components for PLS were determined by a grid search method. The discrete parameter for NBR was α = 1e-5, and the number of principal components for PLS was K = 46. The prediction quantification indices for each method are shown in Table 3, with GPMR representing the method proposed in this invention. As can be seen from the table, the MAE predicted by the method proposed in this invention on the test set is significantly lower than the MAE of other methods, and the prediction R-squared value of the method proposed in this invention is also significantly lower. 2 The value closest to 1 indicates that the proposed method has the best fit and prediction effect on the data. Figure 8 The GPMR prediction curve for the test set is shown. To more intuitively compare the prediction results of the above methods, the prediction results of each method for the first 300 samples of the test set are plotted, as follows. Figure 9 As shown in the figure, the prediction curve of GPMR best matches the actual defect number curve of the test set, therefore the prediction effect of this method is the best.
[0169] The two embodiments described above verify that the proposed method for predicting count-type quality variables based on a variational Bayesian Gaussian-Poisson mixed regression model is feasible and effective.
[0170] It will be understood by those skilled in the art that the above descriptions are merely preferred examples of the invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art can still modify the technical solutions described in the foregoing examples or make equivalent substitutions for some of the technical features. All modifications and equivalent substitutions made within the spirit and principles of the invention should be included within the scope of protection of the invention.
Claims
1. A method for predicting count-type quality variables based on a variational Bayesian Gaussian-Poisson mixed regression model, characterized in that, Includes the following steps: (1) Collect data samples of count-type quality variables and related process variables as the training set of the model; perform preprocessing of missing values, outliers and standardization on the data to obtain the processed training set; (2) Using variational inference methods, a variational Bayesian Gaussian-Poisson mixed regression model is trained offline on the processed training set; the expression of the variational Bayesian Gaussian-Poisson mixed regression model is: ; ; ; in, , , For all samples, the set of process variables, quality variables, and latent variables, for The The value that a dimension can take; For the first The regression coefficients of each Poisson regression component; ; The density function representing the Gaussian distribution; , The first The mean vector and precision matrix of each Gaussian distribution component; These are the mixing coefficients of the mixture model; , , , Obey the following priors: They all follow a Gaussian distribution, that is ,in, for The mean vector and covariance matrix of the Gaussian distribution are given here for all components. All priors are the same; and They all follow a Gaussian-Wishart distribution, i.e. ,in, and They are The mean vector and scale parameter of the distribution. This represents the probability density function of the Wishart distribution. Let be the scaling matrix of this distribution. denoted as the number of degrees of freedom of the distribution; The joint distribution follows a Dirichlet distribution, i.e. ,in, For the parameters of the Dirichlet distribution, To ensure that the distribution can be normalized, ; (3) Collect the test sample and, after the same preprocessing as in step (1) for missing values, outliers, and standardization, obtain the processed test sample. Then, the variational Bayesian Gaussian-Poisson mixed regression model trained in step (2) is used for online prediction.
2. The method for predicting count-type quality variables based on a variational Bayesian Gaussian-Poisson mixed regression model according to claim 1, characterized in that, In step (2), the specific steps for offline training of the variational Bayesian Gaussian-Poisson mixture regression model on the processed training set are as follows: (2.1) Using Bayes' theorem, the latent variables in the model are obtained. and parameter variables Variational posterior distribution: ; ; ; ; in, , , , , , as well as Let be the distribution parameters of these variational posterior distributions, where , ; (2.2) Set the hyperparameters of the prior distribution of the model parameters , , , , , , Number of mixed components and the iterative convergence threshold Maximum number of iterations M, current number of iterations ; (2.3) Random initialization: First, a multinomial distribution is used to generate the initialization process. The initial value, where for The expected value; the parameters of the multinomial distribution are K-dimensional vectors, and the value of the vector is... Then, other expectations are randomly initialized, including: , , , , ,in, , The function in parentheses is in the variational posterior distribution. Expectations , For the function in parentheses Expectations For the function in parentheses Expectations; (2.4) Increment the current iteration count by 1, that is... ;for and The parameters of the variational posterior distribution defined by the model are calculated using the following formula: ,in, ; ; ; ; ; ; ,in, The result can be obtained using the Newton-Raphson method, with the following steps: First, randomly initialize... Set an optimal convergence threshold Using formula renew until The maximum value in the absolute difference vector before and after the update is less than ,in, , ; ; (2.5) Calculate the expected value involved in the model according to the following formula: , , , , , ,in, For the digamma function, ; in, ; ; (2.6) Calculate the lower bound of evidence for the current iteration step using the following formula. : ;in, The expected value of the function enclosed in parentheses is given by the variational posterior distribution of all parameters involved in the function. The specific calculation formula is as follows: ; ; ; ; ;in, ; ; ; ; ;in, Compared with the above They have the same function form; ; (2.7) Repeat step (2.4) until the maximum number of iterations is reached. or ,in, For the first The next iteration is the lower bound of evidence calculated in the previous iteration, when... hour .
3. The method for predicting count-type quality variables based on a variational Bayesian Gaussian-Poisson mixed regression model according to claim 1, characterized in that, The calculation formula for online prediction using the model trained in step (2) in step (3) is as follows: ; in, The predicted value for the sample to be tested. Let be the probability density function of Student's distribution. It is a parameter of the function.
Citation Information
Patent Citations
Multi-working-condition process adaptive soft measurement modeling method based on local double-weighted probability hidden variable regression model
CN114239400A
Multi-task learning using bayesian model with enforced sparsity and leveraging of task correlations
US20130151441A1