Probability partial least square model modeling method based on scalar likelihood function

By introducing scalar likelihood functions and hierarchical relaxation intrapoint method in PPLS model, the problems of high computational cost and slow convergence during parameter estimation of PPLS model are solved, and more efficient and accurate parameter estimation is achieved, which expands the application scope of the model.

CN119988816APending Publication Date: 2025-05-13BEIJING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411858087.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing probability partial least squares (PPLS) model has problems such as high computational cost, sensitivity to initialization, easy to fall into local minimum values ​​and slow convergence when estimating parameters.

Method used

A PPLS model modeling method based on scalar likelihood functions is proposed. The likelihood function in the matrix form is reduced to a scalar form by using the rank n update technology, and the constraint optimization solution is performed through the hierarchical relaxation inner point method (HR-IPM), the orthogonal constraints of the load matrix are processed, the upper and lower bounds of the parameters are set using prior information, and the hierarchical optimization strategy is used to optimize the key parameters and secondary parameters.

Benefits of technology

It improves the accuracy and efficiency of PPLS model parameter estimation, reduces the computational complexity, enhances the feasibility and accuracy of model solving, and expands the scope of application of PPLS model in practical problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988816A_ABST
    Figure CN119988816A_ABST
Patent Text Reader

Abstract

The invention relates to a high-dimensional data modeling method based on a probability partial least square method, and the method comprises the steps: 1, carrying out probability partial least square modeling, enabling # imgabs0 # and # imgabs1 # to represent two groups of observation variables of a p dimension and a q dimension respectively, and building a probability partial least square model if r hidden variables # imgabs2 # and # imgabs3 # obey Gaussian distribution exist; and step 2, scalar likelihood function derivation: based on joint distribution of observation variables X and Y, calculating a logarithm likelihood function # imgabs4 # containing a model parameter theta, and simplifying # imgabs5 # into a scalar likelihood function # imgabs6 # about matrix elements by using matrix algebraic properties and a rank n updating technology. The method has the advantages that the accuracy of parameter estimation is improved by means of prior noise information and the like, complex matrix operation is converted into scalar operation by means of matrix algebraic properties through the rank n updating method, the calculation complexity is greatly reduced, the modeling efficiency is improved, and the method is suitable for large-scale popularization and application. The number of parameters that need to be optimized is reduced by pre-estimating sample noise.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis, and in particular to a probabilistic partial least squares modeling method based on a scalar likelihood function. Background Art

[0002] Partial Least Squares (PLS) is a statistical learning method widely used in high-dimensional data regression and dimensionality reduction. PLS introduces hidden variables to achieve dimensionality reduction and regression of data while maximizing the covariance of independent and dependent variables. However, the traditional PLS method lacks probabilistic interpretation and is difficult to characterize the inherent uncertainty and noise of the data.

[0003] To overcome the above shortcomings, el Bouhaddani et al. proposed the Probabilistic Partial Least Squares (PPLS) method. PPLS is a statistical modeling technique that combines probabilistic latent variable analysis with partial least squares regression. PPLS introduces probability distribution to describe the relationship between latent variables, observed variables and noise, gives PLS a probabilistic interpretation, and enhances the expressive power of the model. Existing PPLS models usually use the EM algorithm for parameter estimation. However, the EM algorithm has the following limitations when processing PPLS models: (1) EM relies on a complete likelihood function containing latent variables, which makes the computational cost high and sensitive to the initialization of latent variables, (2) EM is prone to fall into local minima and converges slowly, and (3) EM is not sufficient to cope with diverse memory, speed and other requirements.

[0004] Therefore, an efficient and accurate PPLS model parameter estimation method is needed, and an optimization algorithm needs to be designed to deal with the constraints in the model, which is of great value for the application of PPLS in practical problems. Summary of the invention

[0005] The purpose of the present invention is to overcome the defects of the prior art and to propose a high-dimensional data modeling method based on probabilistic partial least squares method.

[0006] The high-dimensional data modeling method based on probabilistic partial least squares method comprises the following steps:

[0007] Step 1. Probabilistic partial least squares modeling, let the observed variables and observed variables Represents two sets of observed variables of p dimension and q dimension respectively, assuming that there are r latent variables and It follows Gaussian distribution and establishes a probabilistic partial least squares model;

[0008] Step 2. Derivation of the scalar likelihood function: Based on the joint distribution of the observed variables X and Y, calculate the log-likelihood function containing the model parameters θ Using matrix algebraic properties and rank n update technology, Reduces to a scalar likelihood function with respect to the matrix elements

[0009] Step 3. Introduce the hierarchical relaxed interior point method (HR-IPM) to perform constrained optimization on the scalar likelihood function, process the orthogonality constraints of the load matrix and solve the model parameters, transform the orthogonality constraints into inequality constraints, use prior information to set the upper and lower bounds of the parameters, and use a hierarchical optimization strategy to optimize the key parameters and secondary parameters respectively to obtain parameter estimates;

[0010] Step 4. Prediction application: Based on the estimated PPLS model parameters, the unknown target data can be predicted by solving the conditional distribution mean of the target variable under the given new observed variables and parameters.

[0011] Furthermore, in step 1, in the PPLS model, it is assumed that x i and i Separate random noise and pollution, and t i and u i All subject to random noise The PPLS model is expressed as follows:

[0012]

[0013] In formula (1), x i , e i ,y i , f i are row vectors of length p, p, q, q respectively, W and C are loading matrices of size p×r and q×r respectively, e i and f i are independent noise components that follow an isotropic normal distribution. t i and u i is a vector of length r, t i Following anisotropic distribution, in, B=diag(b1,...,b r ) is the connection t i and u i The diagonal matrix, h i =[h i1 ,...,h ir ] is the noise part with isotropic normal distribution,

[0014] Furthermore, in step 2, when calculating the log-likelihood function In the process of , where A is an m×n matrix (m>n), and its column vectors are orthogonal to each other, that is In order to reduce computational complexity and improve numerical stability, the determinant and inverse matrix of the expression are directly calculated to obtain The closed-form solution of , introduces new constraints to impose orthogonality on the row vectors of A, in order to obtain The closed-form solution of , where the rank n update method is introduced, is to replace the matrix Decomposed into a superposition of a series of rank-one matrices, as shown in the following formula (2):

[0015]

[0016] In formula (2), a i is the i-th column of A, σ i is the i-th diagonal element of Σ, using And the Sherman-Morrison calculation formula, we get The closed-form solution of the determinant and inverse matrix is ​​as follows (3):

[0017]

[0018] In this step 2, the rank n update method is used to simplify the original matrix form of the likelihood function is the scalar likelihood function By and Substituting into the above formula (3), we get The determinant and analytical expression of the inverse matrix of scalar form of .

[0019] Furthermore, in step 3, the hierarchical relaxed interior point method (HR-IPM) is introduced to perform constrained optimization on the scalar likelihood function:

[0020] In order to simplify the estimation of PPLS model parameters, the original matrix form log-likelihood function is simplified to a scalar form by using matrix algebraic properties and rank n update technology, denoted by is the parameter set in the PPLS model, based on the conditional probability distribution of the observed variables X and Y, the log-likelihood function is written as follows (4):

[0021]

[0022] Omit the constant term and the negative sign, It is expressed as the following formula (5):

[0023]

[0024] In order to Simplifying to scalar form:

[0025] set up k>0, if Then we get the following formula (6):

[0026]

[0027] To prove that (6) holds: Here, let A = [a1, a2, ..., a n ], where a i is the i-th column of A, since Therefore

[0028]

[0029] right Take the determinant and use And the Sherman-Morrison formula, that is:

[0030]

[0031] right Inverse, use and the Sherman-Morrison formula, yielding the same results as above:

[0032]

[0033] Also suppose: Where M is a diagonal matrix, then:

[0034]

[0035] To prove the above equation: Let D = ABM, and let the row vector of A be A i ,i=1,2,...,n, let the column vector of B be B j ,j=1,2,...,r, let the diagonal elements of M be M jj ,j=1,2,...,r, then:

[0036]

[0037] So the (i,j)th element D of D ij It is expressed as the following formula:

[0038]

[0039] Furthermore, using the definition of the trace of the matrix product and the properties of the matrix transpose, we get the following formula:

[0040]

[0041] As above, tr(SΣ -1 )calculate;

[0042] Again, let the parameter set of the PPLS model be Let S be the covariance matrix of the observed samples, and use the above proof steps to transform the log-likelihood function based on Θ into Simplified to a scalar likelihood function As follows:

[0043]

[0044] Among them, K and M are intermediate variables related to the elements of Θ, and are defined as follows:

[0045] make Then M2, M4, and M6 are diagonal matrices, and their diagonal elements are as follows:

[0046]

[0047] The K variable is determined by S and the load matrix W, C:

[0048]

[0049] Compared with the original matrix form, by Perform optimization solution, that is, estimate the parameters of PPLS model;

[0050] To prove the above formula is true: ln|Σ| and tr(SΣ -1 ) are expanded respectively, and To simplify:

[0051] For the ln|Σ| term, we use the properties of the block matrix determinant and simplify it to obtain:

[0052]

[0053] For tr(SΣ -1 ) term, using Σ -1 The block representation and simplification of , we get:

[0054]

[0055] Put ln|Σ| and tr(SΣ -1 ) We still get the scalar likelihood function expression.

[0056] Furthermore, in step 4, based on the estimated PPLS model parameters, the conditional distribution mean of the target variable under the given new observed variables and parameters is solved to predict the unknown target data:

[0057] In order to improve the accuracy and efficiency of PPLS model parameter estimation, a noise distribution pre-estimation method is proposed. By using the statistical characteristics of the observed samples, the likelihood function is optimized. Previously, the variance of the observed sample noise e and f was estimated separately and And use it as prior knowledge to improve the accuracy of parameter estimation. Here, the parameter θ is is a vector with a length of (p+q+2)×r+3. When p, q and r take large values, this vector will become very large and affect the optimization process. and Fixed, substituting its estimated value into the likelihood function as prior knowledge can effectively reduce the complexity of the optimization problem and improve the convergence speed and stability of the algorithm;

[0058] To this end, let the observed sample x i By signal s i and noise i Composition, that is, x i =s i +e i , where the noise e i It has a mean of zero and a variance of Gaussian distribution, given N observation samples Use the mean of the observed sample As a single sample signal s i According to the definition of the PPLS model, when the sample size N is large enough, according to the law of large numbers, the noise mean will converge to zero, and the observed sample mean Mainly by the signal mean Determine, in addition, when the latent variable t i When the variance is small, or the element value of the loading matrix W is small, the signal s of different samples i The difference between them will be reduced accordingly, and the mean value of the signal will be used. A signal s that can approximate a single sample i , thus we get the noise e i The estimation is as follows (7):

[0059]

[0060] According to the maximum likelihood estimation principle, the noise variance Estimated value of The likelihood function should be maximized, as shown in equation (8):

[0061]

[0062] Taking the logarithm of the above equation and taking its derivative, setting the derivative to zero, we get the following equation (9):

[0063]

[0064] Similarly, let the observed sample y i By signal t i and noise f i Composition, that is, y i =t i +f i , where the noise f i It has a mean of zero and a variance of Gaussian distribution, given N observation samples Noise variance The maximum likelihood estimate of is as follows (10):

[0065]

[0066] Compared with the prior art in this technical field, the present invention has the following superior technical effects:

[0067] 1. The probabilistic partial least squares modeling method based on the scalar likelihood function described in the present invention improves the accuracy of parameter estimation by utilizing prior noise information and other means.

[0068] 2. The probabilistic partial least squares modeling method based on the scalar likelihood function described in the present invention uses the matrix algebraic properties to convert complex matrix operations into scalar operations, which greatly reduces the computational complexity and improves the modeling efficiency. Pre-estimating the sample noise also reduces the number of parameters that need to be optimized.

[0069] 3. The probabilistic partial least squares modeling method based on the scalar likelihood function described in the present invention improves the solution accuracy and robustness while ensuring the constraint satisfaction by introducing relaxation variables and inequality constraints, parameter constraints based on prior information and hierarchical optimization strategies, thereby improving the feasibility and accuracy of model solution.

[0070] 4. The probabilistic partial least squares modeling method based on the scalar likelihood function described in the present invention and the prediction method proposed based on the PPLS model parameters provide new ideas for the application of PPLS in practical problems and expand the application scope of the PPLS model. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 It is a flow chart of the probabilistic partial least squares modeling method based on the scalar likelihood method of the present invention;

[0072] Figure 2 It is a schematic diagram of the parameter estimation process of the probabilistic partial least squares modeling method based on the scalar likelihood method described in the present invention;

[0073] Figure 3 It is a schematic diagram of the prediction process of the probabilistic partial least squares modeling method based on the scalar likelihood method described in the present invention. DETAILED DESCRIPTION

[0074] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0075] Example

[0076] The probabilistic partial least squares modeling method based on the scalar likelihood method of the present invention comprises the following steps:

[0077] Step 1. Probabilistic partial least squares modeling, let the observed variables and observed variables Represents two sets of observed variables of p dimension and q dimension respectively, assuming that there are r latent variables and It follows Gaussian distribution and establishes a probabilistic partial least squares model;

[0078] Step 2. Derivation of the scalar likelihood function: Based on the joint distribution of the observed variables X and Y, calculate the log-likelihood function containing the model parameters θ Using matrix algebraic properties and rank n update technology, Reduces to a scalar likelihood function with respect to the matrix elements

[0079] Step 3. Introduce the hierarchical relaxed interior point method (HR-IPM) to perform constrained optimization on the scalar likelihood function, process the orthogonality constraints of the load matrix and solve the model parameters, transform the orthogonality constraints into inequality constraints, use prior information to set the upper and lower bounds of the parameters, and use a hierarchical optimization strategy to optimize the key parameters and secondary parameters respectively to obtain parameter estimates;

[0080] Step 4. Prediction application: Based on the estimated PPLS model parameters, the unknown target data can be predicted by solving the conditional distribution mean of the target variable under the given new observed variables and parameters.

[0081] As a specific step example, Figure 1 As shown, in step 1, in the PPLS model, it is assumed that x i andi Separate random noise and pollution, and t i and u i All subject to random noise The PPLS model is expressed as follows:

[0082]

[0083] In formula (1), x i , e i ,y i , f i are row vectors of length p, p, q, q respectively, W and C are loading matrices of size p×r and q×r respectively, e i and f i are independent noise components that follow an isotropic normal distribution. t i and u i is a vector of length r, t i Following anisotropic distribution, in, B=diag(b1,...,b r ) is the connection t i and u i The diagonal matrix, h i =[h i1 ,...,h ir ] is the noise part with isotropic normal distribution,

[0084] As a specific step example, Figure 2 As shown, in step 2, when calculating the log-likelihood function In the process of , where A is an m×n matrix (m>n), and its column vectors are orthogonal to each other, that is In order to reduce computational complexity and improve numerical stability, the determinant and inverse matrix of the expression are directly calculated to obtain The closed-form solution of , introduces new constraints to impose orthogonality on the row vectors of A, in order to obtain The closed-form solution of , where the rank n update technique is introduced, is to replace the matrix Decomposed into a superposition of a series of rank-one matrices, as shown in the following formula (2):

[0085]

[0086] In formula (2), a i is the i-th column of A, σi is the i-th diagonal element of Σ, using And the Sherman-Morrison calculation formula, we get The closed-form solution of the determinant and inverse matrix is ​​as follows (3):

[0087]

[0088] In this step 2, the rank n update technique is used to simplify the original matrix form of the likelihood function is the scalar likelihood function By and Substitute them into the above formula (3) to obtain their determinants and analytical expressions of inverse matrices, and then get scalar form of .

[0089] As a specific step example, in step 3, the hierarchical relaxed interior point method (HR-IPM) is introduced to perform constrained optimization on the scalar likelihood function:

[0090] In order to simplify the estimation of PPLS model parameters, the original matrix form log-likelihood function is simplified to a scalar form by using matrix algebraic properties and rank n update technology, denoted by is the parameter set in the PPLS model, based on the conditional probability distribution of the observed variables X and Y, the log-likelihood function is written as follows (4):

[0091]

[0092] Omit the constant term and the negative sign, It is expressed as the following formula (5):

[0093]

[0094] In order to Simplifying to scalar form:

[0095] set up k>0, if Then we get the following formula (6)

[0096]

[0097]

[0098] To prove that (6) holds: Let A = [a1, a2, ..., a n ], where a i is the i-th column of A, since Therefore

[0099]

[0100] right Take the determinant and use And the Sherman-Morrison formula, that is,

[0101]

[0102] right Inverse, use And the Sherman-Morrison calculation formula, we get:

[0103]

[0104] Also suppose: Where M is a diagonal matrix, then:

[0105]

[0106] To prove the above equation: Let D = ABM, and let the row vector of A be A i ,i=1,2,...,n, let the column vector of B be B j ,j=1,2,...,r, let the diagonal elements of M be M jj ,j=1,2,...,r, then:

[0107]

[0108] So the (i,j)th element D of D ij It is expressed as:

[0109]

[0110] Furthermore, using the definition of the trace of the matrix product and the properties of the matrix transpose, we get:

[0111]

[0112] As above, tr(SΣ -1 ) calculation;

[0113] In addition, let the parameter set of the PPLS model be Let S be the covariance matrix of the observed samples, and use the above proof steps to transform the log-likelihood function based on Θ into Simplified to a scalar likelihood function

[0114]

[0115] Among them, K and M are intermediate variables related to the elements of Θ, and are defined as follows:

[0116] make Then M2, M4, and M6 are diagonal matrices, and their diagonal elements are:

[0117]

[0118] The K variable is determined by S and the load matrix W, C:

[0119]

[0120] Compared with the original matrix form, the scalar likelihood function has the advantages of simple form and efficient calculation, which lays the foundation for subsequent parameter estimation. Perform optimization solution, that is, estimate the parameters of PPLS model;

[0121] To prove that the above formula is true: ln|Σ| and tr(SΣ -1 ) are expanded respectively, and To simplify:

[0122] For the ln|Σ| term, using the properties of the block matrix determinant and simplifying it, we get the following formula:

[0123]

[0124] For tr(SΣ -1 ) term, using Σ -1 The block representation and the simplification lemma give the following formula:

[0125]

[0126] Put ln|Σ| and tr(SΣ -1 ) That is, we get the scalar likelihood function expression.

[0127] As a specific step example, Figure 3 As shown in Figure 4, based on the estimated PPLS model parameters, the conditional distribution mean of the target variable under the given new observed variables and parameters is solved to predict the unknown target data:

[0128] In order to further improve the accuracy and efficiency of PPLS model parameter estimation, a noise distribution pre-estimation method is proposed. By utilizing the statistical characteristics of the observed samples, the likelihood function is optimized. Previously, the variance of the observed sample noise e and f was estimated separately and Using it as prior knowledge can effectively decouple the interaction between the noise parameter and other parameters, reduce the complexity of the optimization process, and improve the accuracy of parameter estimation. It should be noted that the parameter θ is a vector with a length of (p+q+2)×r+3. When p, q and r take large values, this vector will become very large and affect the optimization process. and The (p+q+2)×r+3 variables may not seem important, but they are important in ln|Σ| and tr(SΣ -1 ) appears in almost every term in the expanded form, indicating that they have important interactions with other parameters. and Fixing these terms can effectively decouple these terms, thereby reducing the complexity of the entire optimization process, alleviating the impact of potential local minima and saddle points, simplifying the calculation of gradients, and reducing the nonlinear characteristics of the optimization problem. Substituting these estimated values ​​into the likelihood function as prior knowledge can effectively reduce the complexity of the optimization problem and improve the convergence speed and stability of the algorithm.

[0129] To this end, let the observed sample x i By signal s i and noise i Composition, that is, x i =s i +e i , where the noise e i It has a mean of zero and a variance of Gaussian distribution, given N observation samples Use the mean of the observed sample As a single sample signal s i According to the definition of the PPLS model, when the sample size N is large enough, according to the law of large numbers, the noise mean will converge to zero, and the observed sample mean Mainly by the signal mean Determine, in addition, when the latent variable t i When the variance is small, or the element value of the loading matrix W is small, the signal s of different samples i The difference between them will be reduced accordingly, and the mean value of the signal will be used. A signal s that can approximate a single sample i , thus we get the noise e i The estimation is as follows (7):

[0130]

[0131] According to the maximum likelihood estimation principle, the noise variance Estimated value of The likelihood function should be maximized, as shown in equation (8):

[0132]

[0133] Taking the logarithm of the above equation and taking its derivative, setting the derivative to zero, we get the following equation (9):

[0134]

[0135] Similarly, let the observed sample y i By signal t i and noise f i Composition, that is, y i =t i +f i , where the noise f i It has a mean of zero and a variance of Gaussian distribution, given N observation samples Noise variance The maximum likelihood estimate of is as follows (10):

[0136]

[0137] The algorithm and strategy adopted by the probabilistic partial least squares modeling method based on the scalar likelihood function of the present invention are further described in detail below:

[0138] Hierarchical Relaxation Interior Point Method

[0139] This paper proposes a hierarchical relaxed interior point method (HR-IPM) to solve the parameter estimation problem of the PPLS model. Based on the traditional interior point method, HR-IPM introduces three main innovations according to the characteristics of the PPLS model to improve the efficiency and reliability of the algorithm:

[0140] The traditional interior point method is an effective method for solving nonlinear optimization problems with inequality constraints. Its basic idea is to transform the original problem into a series of obstacle subproblems, deal with inequality constraints by introducing obstacle functions, and approximate the solution of the original problem by gradually reducing the obstacle parameters:

[0141] Consider the following optimization problem with inequality constraints:

[0142] m θ in f(θ)

[0143] stg i (θ)≤0,i=1,…,m,

[0144] Among them, f(θ) is the objective function, g i (θ) is the inequality constraint function;

[0145] The core of the traditional interior point method is to construct the barrier function B(θ,μ) and transform the above problem into an unconstrained optimization problem:

[0146]

[0147] Among them, μ>0 is the obstacle parameter. When μ approaches 0, the obstacle function B(θ,μ) approaches the original problem.

[0148] The traditional interior point method obtains the solution to the original problem by gradually reducing μ and solving the corresponding obstacle subproblems. In each subproblem, methods such as Newton's method can be used to find the stable point of the obstacle function.

[0149] Slack variables and inequality constraints

[0150] The PPLS model involves some equality constraints, such as the orthogonality constraint of the loading matrix. In order to make full use of the special structure of the PPLS model, HR-IPM transforms these equality constraints into inequality constraints and introduces slack variables;

[0151] Specifically, HR-IPM introduces slack variables w,c , convert the equality constraint into the following inequality constraint:

[0152]

[0153] Among them, δ ij is the Kronecker symbol;

[0154] Through this slack variable technology, HR-IPM transforms the original problem into an optimization problem containing only inequality constraints, which enhances the flexibility and applicability of the algorithm.

[0155] Parameter constraints based on prior information

[0156] In order to make full use of the prior information of the PPLS model and improve the quality of parameter estimation and the stability of the algorithm, HR-IPM further introduces upper and lower bounds for the model parameters.

[0157] Based on prior knowledge, HR-IPM sets the following constraints for model parameters:

[0158]

[0159] These upper and lower bound constraints effectively limit the range of parameter values, eliminate some unreasonable solutions, reduce the difficulty of parameter estimation problems, and improve the estimation quality and algorithm stability.

[0160] Layered optimization strategy

[0161] Considering the importance of the load matrices W and C in the PPLS model, the HR-IPM algorithm introduces a hierarchical optimization strategy. Specifically, this paper divides the optimization problem into two layers: the outer optimization problem only optimizes W and C, and the inner optimization problem optimizes other parameters when W and C are fixed.

[0162] Let θ1=(W,C) represent the outer optimization variable, represents the inner optimization variable, and θ=(θ1,θ2) is the complete parameter vector.

[0163] The outer optimization problem can be expressed as:

[0164]

[0165] in, is the negative log-likelihood function of the PPLS model, is a constraint function related to θ1, including the orthogonality constraint of the load matrix and the upper and lower bound constraints of the parameters. is the current estimate of the inner optimization variable, which is temporarily fixed during the outer optimization process.

[0166] The inner optimization problem can be expressed as:

[0167]

[0168] in, is a constraint function related to θ2, mainly including the upper and lower bound constraints of the parameters, is the current estimate of the outer optimization variable, which is temporarily fixed during the inner optimization process.

[0169] In each outer iteration, HR-IPM first fixes Solve the outer optimization problem and get a new estimate of θ1 Then, in the inner optimization problem, HR-IPM fixes Solve the inner optimization problem and get a new estimate of θ2 This process is repeated alternately until the termination condition of the outer optimization is met.

[0170] The hierarchical optimization strategy makes full use of the importance differences of the PPLS model parameters. By focusing on optimizing the key parameters W and C, it can accelerate the convergence of the algorithm and improve the efficiency of parameter estimation. At the same time, it also simplifies the complexity of the optimization problem at each layer and enhances the flexibility and modularity of the algorithm. In the outer layer optimization, HR-IPM can focus on processing the constraints and properties of the load matrix; while in the inner layer optimization, HR-IPM can use the known estimated values ​​of W and C to optimize other parameters more efficiently. This hierarchical processing idea not only improves the efficiency of the algorithm, but also enhances the interpretability and controllability of the algorithm.

[0171] In summary, the HR-IPM algorithm is a variant of the interior point method designed specifically for the PPLS model parameter estimation problem. It introduces three main innovations on the basis of the traditional interior point method: slack variables and inequality constraints, parameter constraints based on prior information, and hierarchical optimization strategies. These innovations make full use of the special structure and prior information of the PPLS model, effectively improve the efficiency, reliability and applicability of the algorithm, and provide a new and effective solution to the parameter estimation problem of the PPLS model.

[0172] Prediction method based on PPLS model

[0173] After obtaining the parameter estimates of the PPLS model, the model can be used to predict the target data for new observation data. For new observation data, is the corresponding unknown target data, then we can solve y new At a given x new With model parameters The mean of the conditional distribution under the condition is new Predicted value of:

[0174]

[0175] To solve the above conditional distribution, we use the conditional distribution property of the bivariate Gaussian distribution, that is, for the

[0176] The bivariate Gaussian distribution of , given the condition of x, the conditional distribution of y satisfies:

[0177]

[0178] From the definition of the PPLS model, we can see that the joint distribution of the observed data and the target data is:

[0179]

[0180] Replace X and Y in the above formula with x new ,y new , substituting into the mean formula of conditional distribution, we can get:

[0181]

[0182] in: are the parameter estimates of the PPLS model.

[0183] Therefore, using the new observation data x new Parameter estimation with PPLS model The corresponding target data y can be obtained through simple matrix operations newThe mean prediction The matrix operation form of the estimation formula for the objective function is realized. The method is easy to implement and has high computational efficiency.

[0184] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with the technical field within the scope disclosed in the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. A high-dimensional data modeling method based on probabilistic partial least squares method, comprising the following steps: Step 1. Probabilistic partial least squares modeling, let and Represents two sets of observed variables of p dimension and q dimension respectively, assuming that there are r latent variables and It obeys Gaussian distribution, and the probability partial least squares model PPLS is established; Step 2. Derivation of the scalar likelihood function: Based on the joint distribution of the observed variables X and Y, calculate the log-likelihood function containing the model parameters θ Using matrix algebraic properties and rank n update technology, Reduces to a scalar likelihood function with respect to the matrix elements Step 3. Introduce the hierarchical relaxed interior point method to perform constrained optimization on the scalar likelihood function, process the orthogonality constraints of the load matrix and solve the model parameters, transform the orthogonality constraints into inequality constraints, use prior information to set the upper and lower bounds of the parameters, and use the hierarchical optimization strategy to optimize the key parameters and secondary parameters respectively to obtain the parameter estimates; Step 4. Prediction application: Based on the estimated PPLS model parameters, the unknown target data can be predicted by solving the conditional distribution mean of the target variable under the given new observed variables and parameters.

2. According to the high-dimensional data modeling method based on probabilistic partial least squares method as claimed in claim 1, in step 1, it is assumed in the PPLS model that x i and i Separate random noise and pollution, and t i and u i All subject to random noise The PPLS model is expressed as follows: x i =t i W T +e i ,y i =u i C T +f i ,u i =t i B+h i ......(1), In formula (1), x i , e i ,y i , f i are row vectors of length p, p, q, q respectively, W and C are loading matrices of size p×r and q×r respectively, e i and f i are independent noise components that follow an isotropic normal distribution. t i and u i is a vector of length r, t i Following anisotropic distribution, Among them, among them, B=diag(b1,...,b r ) is the connection t i and u i The diagonal matrix of h i =[h i1 ,...,h ir ] is the noise part with isotropic normal distribution, 3. According to the high-dimensional data modeling method based on PPLS as claimed in claim 1, in step 2, when calculating the log-likelihood function In the process, we need to process the form AΣA T +kI m , where A is an m×n matrix (m>n), and its column vectors are orthogonal to each other, that is In order to reduce computational complexity and improve numerical stability, the determinant and inverse matrix of this expression are directly calculated to obtain AΣA T +kI m The closed-form solution of A is to introduce new constraints to impose orthogonality on the row vectors of A. In order to obtain AΣA without introducing new constraints T +kI m The closed-form solution of rank n is introduced here, and the matrix AΣA T +kI m Decomposed into a superposition of a series of rank-one matrices, as shown in the following formula (2): In formula (2), a i is the i-th column of A, σ i is the i-th diagonal element of Σ, using And the Sherman-Morrison calculation formula, we get AΣA T +kI m The closed-form solution of the determinant and inverse matrix is ​​as follows (3): Here, the rank n update method is used to simplify the original matrix form of the likelihood function is the scalar likelihood function By and Substituting into the above formula (3), we get AΣA T +kI m The determinant and analytical expression of the inverse matrix of scalar form of .

4. According to the high-dimensional data modeling method based on probabilistic partial least squares method of claim 1, in step 3, a hierarchical relaxed interior point method (HR-IPM) is introduced to perform constrained optimization solution on the scalar likelihood function: In order to simplify the estimation of PPLS model parameters, the original matrix form log-likelihood function is simplified to a scalar form by using matrix algebraic properties and rank n update technology, denoted by is the parameter set in the PPLS model, based on the conditional probability distribution of the observed variables X and Y, the log-likelihood function is written as follows (4): Omit the constant term and the negative sign, It is expressed as the following formula (5): In order to Simplifying to scalar form: set up k>0, if A T A=I n , then we get the following formula (6):

5. According to the high-dimensional data modeling method based on probabilistic partial least squares method of claim 1, in step 4, based on the estimated PPLS model parameters, the conditional distribution mean of the target variable under the given new observed variables and parameters is solved to achieve the prediction of unknown target data: In order to improve the accuracy and efficiency of PPLS model parameter estimation, a noise distribution pre-estimation method is proposed. By using the statistical characteristics of the observed samples, the likelihood function is optimized. Previously, the variance of the observed sample noise e and f was estimated separately and And use it as prior knowledge to improve the accuracy of parameter estimation. Here, the parameter θ is is a vector with a length of (p+q+2)×r+3. When p, q and r take large values, this vector will become very large and affect the optimization process. and Fixed, substituting its estimated value into the likelihood function as prior knowledge can effectively reduce the complexity of the optimization problem and improve the convergence speed and stability of the algorithm; To this end, let the observed sample x i By signal s i and noise i Composition, that is, x i =s i +e i , where the noise e i It has a mean of zero and a variance of Gaussian distribution, given N observation samples Use the mean of the observed sample As a single sample signal s i According to the definition of the PPLS model, when the sample size N is large enough, according to the law of large numbers, the noise mean will converge to zero, and the observed sample mean Mainly by the signal mean Determine, in addition, when the latent variable t i When the variance is small, or the element value of the loading matrix W is small, the signal s of different samples i The difference between them will be reduced accordingly, and the mean value of the signal will be used. A signal s that can approximate a single sample i , thus we get the noise e i The estimation is as follows (7): According to the maximum likelihood estimation principle, the noise variance Estimated value of The likelihood function should be maximized, as shown in equation (8): Taking the logarithm of the above equation and taking its derivative, setting the derivative to zero, we get the following equation (9): Similarly, let the observed sample y i By signal t i and noise f i Composition, that is, y i =t i +f i , where the noise f i It has a mean of zero and a variance of Gaussian distribution, given N observation samples Noise variance The maximum likelihood estimate of is as follows (10):