Bayesian Tensor Completion Method Based on Multiple Measurements

Through the Bayesian tensor completion method, Gibbs sampling and CP decomposition are used to solve the problem of outliers in multi-measurement tensor data, and more accurate data completion is achieved.

CN114756813BActive Publication Date: 2025-08-01FUDAN UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210331188.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-08-01
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

In the prior art, when processing tensor data of multiple measurement values, it is difficult to effectively utilize all measurement information, especially in the presence of outliers, resulting in inaccurate estimation.

Method used

The Bayesian tensor completion method based on multi-measures is used to estimate the tensor elements through Gibbs sampling and CP decomposition, combined with the conjugated prior distribution, and data completion is performed using all measured value information.

Benefits of technology

More accurate data estimation is achieved, especially in the presence of outliers, providing a more accurate completion effect of missing values.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114756813B_ABST
    Figure CN114756813B_ABST
Patent Text Reader

Abstract

The present invention provides a Bayesian tensor completion algorithm based on multi-measurements. The multi-measurement data is represented by multiple tensors, and it is assumed that each measurement value of each tensor element of the tensor follows a Gaussian distribution. Then, CP decomposition is performed on the tensor to obtain the corresponding factor matrices, and it is assumed that the parameters of the factor matrices follow a conjugate prior distribution. Furthermore, the Gibbs sampling method is used to sample the posterior conditional distributions of the respective parameters, and the estimated value of the tensor is output. The missing values in the multi-measurement data are interpolated based on the estimated value of the tensor, thereby realizing data completion. In summary, the completion method of the present invention is aimed at measurement data with low measurement accuracy, high cost, and repeated measurements in some regions. The Gibbs sampling method combined with CP decomposition is used to realize data completion. Compared with the completion methods in the prior art, since the method of the present invention can utilize the information of all measurement data, it can provide a more accurate estimated value, thereby realizing more accurate data completion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of tensor completion, and particularly relates to a Bayesian tensor completion method based on multiple measurements. Background Art

[0002] A tensor is a high-dimensional array, which is a generalization of vectors and matrices and can be used to represent multi-dimensional data with complex internal structures. A tensor is a natural representation form of high-dimensional data in the real world. For example, a color image can be regarded as a three-dimensional tensor, where one dimension is the color mode and the other two dimensions are spatial variables. A video composed of color images is a four-dimensional tensor, and the fourth dimension is the time variable. Therefore, tensor analysis has become an important tool in multi-dimensional data analysis and has applications in many fields, such as computer vision, data mining, and collaborative filtering. Therefore, the data completion method based on tensor analysis (tensor completion method) has become the most active field in tensor analysis. The tensor completion method attempts to interpolate the missing data or unobserved data in the tensor.

[0003] In some cases of natural science, sometimes the measurement technology is not accurate. Therefore, when fixing certain variables, there are multiple measurements. Similarly, there are some areas that are not measured and thus have no measurement values. Since the measurement cost is very high, it is of great significance to estimate the values of the unmeasured areas. For example, in the scattering data in the field of nuclear physics, the three dimensions are: atomic nucleus mass number, neutron energy, and angle. After discretizing the data, the data can be regarded as a three-dimensional tensor. However, there are multiple measurements at some points and no measurement data in some areas. In the prior art, there is no computer method that can directly process this type of data. One method for processing data with multiple measurements is to first take the average of the data at the same points and then use the tensor completion method to estimate the missing values. However, some of the measurement values deviate significantly from the true values and can be regarded as outliers. If the averaging method is adopted, it will inevitably lead to inaccurate estimation and be affected by the outliers. How to utilize all the measurement value information and consider the weights of the measurement values becomes a key point. Summary of the Invention

[0004] The present invention is made to solve the above problems. For tensor data with low measurement accuracy, high cost, and repeated measurements in some areas for multiple times, an efficient method, MBGCP (Multiple Bayesian Gaussian CANDECOMP / PARAFAC), is provided. The Gibbs sampling method is used in combination with decomposition to estimate the missing tensor element data. By comparing with the existing method that takes the average value of the data at the same point and then uses BGCP (Bayesian Gaussian CANDECOMP / PARAFAC) to estimate the missing values, this method can utilize all data information and provide a more accurate estimated value. The present invention adopts the following technical solutions:

[0005] The present invention provides a Bayesian tensor completion method based on multiple measurements, which is characterized by including the following steps:

[0006] Step S1: Represent the multiple measurement data as several tensors, and set the measured values of the tensor elements of each of the tensors to follow a Gaussian distribution;

[0007] Step S2: Use the CP decomposition method to decompose the tensor to obtain multiple factor matrices, and calculate the outer product of the multiple factor matrices as the estimated value of the tensor;

[0008] Step S3: Set the prior distribution of each parameter used for Gibbs sampling to a conjugate prior distribution, and obtain the conditional posterior distribution of each parameter based on the prior distribution of each parameter;

[0009] Step S4: Use the Gibbs sampling method to sample from the conditional posterior distribution of each parameter;

[0010] Step S5: Determine whether the predetermined number of iterations is reached. When the determination is no, repeat Step S4. When the determination is yes, enter Step S6;

[0011] Step S6: When the determination in Step S5 is yes, output the final estimated value of the tensor, which is used to interpolate the missing measured values in the multiple measurement data, thereby realizing the completion of the missing measured values.

[0012] In the Bayesian tensor completion method based on multiple measurements provided by the present invention, it can also have the following technical feature. In Step S1, the multiple measurement data is represented as a d-dimensional tensor without missing values, including i tensor elements x i , where i = (i1,..., i d ), (1 ≤ i k ≤ n k, for k = 1, …, d), some of the tensor elements have no observed values, i.e., there are missing values, and the remaining observable tensor elements have l observed values using a series of the tensors to represent all the observed values, i.e., the multi-measurement data, where m = max(m i ) corresponding to all i, the first observed values of all the tensor elements are arranged and placed in If a certain tensor element has no observed value, then the corresponding element in is also missing. to Similarly, a series of indicator tensors are used to indicate the situation of the missing values. The element of is p when i ∈ Ω Otherwise where Ω p represents the set of subscripts of the observed elements.

[0013] In step S2, the estimated value of is decomposed as follows:

[0014]

[0015] In the formula, is the j-th column of the factor matrix , and the symbol represents the outer product.

[0016] The Bayesian tensor completion method based on multi-measurements provided by the present invention may further have the following technical feature: the parameters of the factor matrix include the row vectors superparameters μ (k) , Λ (k) and precision τ ∈ .

[0017] The Bayesian tensor completion method based on multi-measurements provided by the present invention may further have the following technical feature: in step S3, the prior distributions of the row vectors are all set to satisfy where

[0018] The conditional posterior distribution of the row vector <00001-- 1> is:

[0019]

[0020] In the formula, L() is the likelihood function, N() is the normal distribution, and there is:

[0021]

[0022] wherein denotes the Hadamard product, and the row vectors to the conditional posterior distribution of can be obtained in the same way. Set the prior distribution of the hyperparameter to satisfy Gaussian-Wishart(μ0, β0, W0, v0),

[0023] i.e., p(μ (k) , Λ (k) | - ) = N(μ (k) |μ (0) , (β0Λ (k) ) -1 ) × Wishart(Λ (k) |W0, μ0),

[0024] When k = 1, the conditional posterior distribution of the hyperparameters μ (k) , Λ (k) is as follows:

[0025]

[0026] wherein:

[0027]

[0028] The variables and S (1) are respectively:

[0029]

[0030] When k = 2,..., d, the conditional posterior distribution of the hyperparameters μ (k) , Λ (k) can be obtained in the same way.

[0031] Set the prior distribution of the precision τ ∈ to satisfy τ ∈ ~Gamma(a0, b0), i.e.,

[0032] The conditional posterior distribution of the precision τ ∈ is as follows:

[0033]

[0034] wherein:

[0035]

[0036] The Bayesian tensor completion method based on multiple measurement values provided by the present invention may further have the following technical feature: in step S5, the estimated value of the final tensor is the average of the estimated values of the N tensors after N iterations.

[0037] Function and effect of the invention

[0038] According to the Bayesian tensor completion method based on multiple measurement values of the present invention, multiple measurement value data is represented by multiple tensors, and it is assumed that each measurement value of each tensor element of the tensor follows a Gaussian distribution; then the tensor is subjected to CP decomposition to obtain corresponding factor matrices, and it is assumed that the parameters follow a conjugate prior distribution; furthermore, the Gibbs sampling method is used to sample the conditional posterior distributions of the respective parameters, and the estimated value of the tensor is output, and the missing values in the multiple measurement value data are interpolated based on the estimated value of the tensor, thereby realizing data completion. In summary, the completion method of the present invention is aimed at measurement data with low measurement accuracy, high cost, and repeated measurement in certain areas multiple times. The Gibbs sampling method is combined with CP decomposition to realize data completion. Compared with the existing completion methods, since the method of the present invention can utilize the information of all measurement data, more accurate estimated values can be provided, thereby realizing more accurate data completion. Description of the drawings

[0039] Figure 1 is a flowchart of the Bayesian tensor completion method based on multiple measurement values in an embodiment of the present invention;

[0040] Figure 2 is a schematic diagram of nuclear physics scattering in an embodiment of the present invention;

[0041] Figure 3 is a schematic diagram of CP decomposition in an embodiment of the present invention;

[0042] Figure 4 is a schematic diagram of the Gibbs sampling process in an embodiment of the present invention;

[0043] Figure 5 is a comparison chart of the completion effects of MGBCP and BGCP on nuclear physics data in an embodiment of the present invention. Detailed implementation manners

[0044] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the following specifically describes the Bayesian tensor completion method based on multiple measurement values of the present invention in conjunction with embodiments and drawings.

[0045] <Embodiment>

[0046] This embodiment provides a Bayesian tensor completion method based on multiple measurements, which combines CP decomposition and Gibbs sampling to perform tensor completion on data with multiple measurements. In this embodiment, the data with multiple measurements is nuclear physics scattering data, which is obtained through multiple measurements. The experimental measurement principle is as Figure 4 shown. Due to reasons such as high measurement costs, there are missing values in the data with multiple measurements.

[0047] At the same time, in this embodiment, the Bayesian tensor completion method based on multiple measurements is implemented using Matlab for experimental analysis, that is, the completion method in this embodiment runs on the computing server as a pre-written program.

[0048] Figure 1 is the flowchart of the Bayesian tensor completion method based on multiple measurements in this embodiment.

[0049] As Figure 1 shown, the Bayesian tensor completion method based on multiple measurements in this embodiment specifically includes the following steps:

[0050] Step S1, represent the data with multiple measurements as several tensors, and set the measured values of the tensor elements of each tensor to follow a Gaussian distribution.

[0051] In this embodiment, use to represent the true value of the d-dimensional tensor, which has no missing values. Use x i to represent the value of an element in the tensor , where i = (i1,..., i d ), (1 ≤ i k ≤ n k , k = 1,..., d) are the subscripts of the tensor elements. In real-world applications, an element of a tensor may have many observed values, and for different tensor elements, the number of observed values may be different, and some may have zero observed values (i.e., missing). For those observable tensor elements, use to represent the l-th observed value of x i . For the element with subscript i, there are a total of m i observed values. Therefore, a series of tensors can be used to represent all the observed values, where m = max(m i ) corresponds to all i. The first observed value of all tensor elements is arranged in . If an element has no observed value, then the corresponding element in is also missing. Next, the second observed value of all tensor elements is arranged in In [a certain context], if an element has no observed value or only one observed value, then the corresponding element in will be missing, and so on. Therefore, the observed values of the tensor can have missing values. Let denote 's estimated value, and let denote the value of one of its elements, which has no missing value. Let Ω p represent the set of subscripts of the observed elements. Define a series of indicator tensors whose elements are when \(i\in\Omega\) p , otherwise,

[0052] By the above-defined method, that is, representing multi-measurement data as multiple tensors, and then setting the measurement values of each tensor element in the tensor to follow a Gaussian distribution For ease of explanation, more conveniently, a matrix \(T\) with \(d + 1\) columns can be used to represent the tensor data with multi-measurements: The first column, the second column, up to the \(d\)-th column represent the values of \(i_1,i_2\) to \(i\) d respectively, and the \((d + 1)\)-th column represents 's value.

[0053] Figure 2 is a schematic diagram of nuclear physics scattering in this embodiment.

[0054] As Figure 2 shown, for the scattering data in the field of nuclear physics, its three dimensions are respectively: the atomic nucleus mass number (\(A\)), the neutron energy (\(E\)) and the angle (\(\theta\)), and the scattering data is probability. Since the experiment for measuring the scattering data is very costly, it is impossible to obtain the scattering probability data for any combination of the three dimensions. However, after discretizing the scattering data, the scattering data can be regarded as a three-dimensional tensor, and the tensor completion method can be used to complete the missing values therein. At the same time, fixing the values of a certain \(A\), \(E\), \(\theta\), there may be multiple measurement values of tensor elements.

[0055] In this embodiment, after discretizing the scattering data, a \(238\times11\times88\) tensor can be used to represent the scattering data, and 10% of the scattering data is artificially taken away for comparison, then the missing rate of the scattering data is 94.87%.

[0056] Step S2, decompose the tensor using the CP decomposition method to obtain the corresponding multiple factor matrices, and calculate the outer product of the multiple factor matrices as the estimated value of the tensor.

[0057] Figure 3It is a schematic diagram of CP decomposition in this embodiment.

[0058] CP decomposition is used to decompose a multi-dimensional tensor into multiple one-dimensional factor matrices. Its basic principle is as Figure 3 shown. In this embodiment, the CP decomposition method is applied to the estimated value of the above-mentioned tensor to decompose the tensor to obtain corresponding multiple factor matrices, and calculate the outer product of the multiple factor matrices as the estimated value of the tensor

[0059]

[0060] In the formula, is the j-th column of the factor matrix , and the symbol represents the outer product.

[0061] In this embodiment, for the convenience of explanation, it is assumed that the measured values of the tensor elements follow a Gaussian distribution In the formula, τ ∈ is the precision of the Gaussian distribution. Correspondingly, it is assumed that the prior distributions of all row vectors of the factor matrix U (k) of the CP decomposition all satisfy where the prior distribution of the hyperparameter is set to the conjugate prior distribution Gaussian-Wishart.

[0062] Step S3: Set the prior distribution of each parameter used for Gibbs sampling to the conjugate prior distribution, and obtain the conditional posterior distribution of each parameter based on the prior distribution of each parameter.

[0063] In this embodiment, the parameters include three parts: the row vectors of the factor matrix the hyperparameters μ (k) , Λ (k) and the precision τ ∈ . For each part of the parameters, first set its prior distribution to the conjugate prior distribution, and then obtain the conditional posterior distribution probability based on its prior distribution.

[0064] Step S4: Use the Gibbs sampling method to sample from the conditional posterior distribution of each parameter of the factor matrix.

[0065] Step S5: Judge whether the predetermined number of iterations is reached. When the judgment is no, repeat Step S4. When the judgment is yes, enter Step S6.

[0066] Step S6: When the judgment in Step S5 is yes, output the final estimated value of the tensor, which is used to interpolate the missing values in the multi-measurement data, so as to complete the filling of the missing values.

[0067] Figure 4It is a schematic diagram of the Gibbs sampling process in this embodiment.

[0068] In this embodiment, according to Figure 4 the step method of Gibbs sampling shown (corresponding to the above steps S3 - S5), the parameters of the above three parts are sampled in sequence:

[0069] (1) Sample the row vectors of the factor matrix Set the prior distributions of all row vectors of the factor matrix U

[0070] obtained by CP decomposition to satisfy (k) Given the observed value Taking k = 1 as an example, the likelihood function can be expressed as:

[0071]

[0072] In the above formula, represents the Hadamard product.

[0073] The conditional posterior distribution (i.e., the conditional posterior distribution probability) of

[0074]

[0075] And in the above formula, there are:

[0076]

[0077] Similarly, the conditional posterior distribution of to can be obtained, and Gibbs sampling is performed separately from these conditional posterior distributions one by one.

[0078] (2) Sample the hyperparameters μ (k) , Λ (k) Sample

[0079] Similarly, set the prior distribution of the hyperparameter to the conjugate prior distribution:

[0080] Gaussian - Wishart(μ0, β0, W0, v0)

[0081] That is, p(μ (k) , Λ (k) | - ) = N(μ (k) |μ (0) , (β0Λ (k) ) -1 ) × Wishart(Λ (k) |W0, μ0).

[0082] Taking \(k = 1\) as an example, \(\mu\) (1) , \(\Lambda\) (1) The conditional posterior distribution of is:

[0083]

[0084] In the above formula, there are:

[0085]

[0086] The variables and \(S\) (1) are respectively:

[0087]

[0088] Similarly, the conditional posterior distributions of \(\mu\) (k) , \(\Lambda\) (k) can be obtained for \(k = 2,\cdots,d\), and Gibbs sampling is performed one by one from these conditional posterior distributions respectively.

[0089] (3) Sampling for the precision \(\tau\) ∈ The likelihood function of all observed values is:

[0090]

[0091]

[0092] Set the prior distribution of the precision \(\tau\) ∈ to follow the conjugate prior distribution \(\tau\) ∈ \(\sim Gamma(a_0,b_0)\), that is The conditional posterior distribution of the precision \(\tau\) ∈ can be obtained as:

[0093]

[0094] In the above formula:

[0095]

[0096] Sample the parameters of the above three parts in sequence and perform iteration (that is, repeat the above step S4 to extract multiple samples). When the number of iterations reaches the predetermined maximum number of iterations \(N\), the iteration stops, and the estimated values of the tensors in the last \(N\) iterations are output The final tensor estimate is the average of the tensor estimates at \(N\) times. Based on the final tensor estimate, the missing values in the multi-measurement data can be interpolated, thereby realizing the tensor completion of the multi-measurement data.

[0097] Figure 5 is the comparison chart of the completion effects of MGBCP and BGCP in the nuclear physics data in the embodiments of the present invention.

[0098] As Figure 5 shown, the experimental results indicate that more accurate completion effects can be achieved through the completion method (i.e., MGBCP) of this embodiment. When the observed values are relatively concentrated, the effects of the two methods are relatively similar; however, when there are observed values deviating from the true values and can be marked as outliers, BGCP is more susceptible to the influence of outliers, while the completion method of this embodiment performs more robustly.

[0099] In addition, in this embodiment, the parts not described in detail are common general knowledge in the art.

[0100] Functions and effects of the embodiment

[0101] According to the Bayesian tensor completion method based on multiple measurements provided in this embodiment, the multiple measurement data is represented by multiple tensors, and it is assumed that each measurement value of each tensor element of the tensor follows a Gaussian distribution; then the tensor is decomposed by CP decomposition to obtain the corresponding factor matrices, and it is assumed that the parameters of the factor matrices follow a conjugate prior distribution; furthermore, the Gibbs sampling method is used to sample the conditional posterior distributions of the respective parameters separately, and the estimated value of the tensor is output, and the missing values in the multiple measurement data are interpolated based on the estimated value of the tensor, thereby realizing data completion. In summary, the completion method of this embodiment is aimed at measurement data with low measurement accuracy, high cost, and repeated measurements in certain areas. By using the Gibbs sampling method combined with CP decomposition to realize data completion, compared with the existing completion methods (such as the above-mentioned BGCP), since the method of this embodiment can utilize the information of all measurement data, more accurate estimated values can be provided, thereby realizing more accurate data completion.

[0102] The above embodiments are only used to illustrate the specific implementation manners of the present invention, and the present invention is not limited to the description scope of the above embodiments.

[0103] In the above embodiment, the Bayesian tensor completion method based on multiple measurements is applied to complete the measured nuclear physics scattering data. In fact, the completion method of this embodiment can also be applied to tensor completion of other types of multiple measurement data.

Claims

1. A Bayesian tensor completion method based on multiple measurement values, characterized in that, It includes the following steps: Step S1: Represent the multi-measurement data as a number of tensors, and set the measurement values of the tensor elements of each of the tensors to follow a Gaussian distribution, where the multi-measurement data is nuclear physics scattering data; Step S2: Decompose the tensors using the CP decomposition method to obtain multiple factor matrices, and calculate the outer product of the multiple factor matrices as the estimated value of the tensors; Step S3: Set the prior distribution of each parameter used for Gibbs sampling to a conjugate prior distribution, and obtain the conditional posterior distribution of each parameter based on the prior distribution of each parameter; Step S4: Sample from the conditional posterior distribution of each parameter using the Gibbs sampling method; Step S5: Determine whether the predetermined number of iterations has been reached. If the determination is no, repeat Step S4; Step S6: When the determination in Step S5 is yes, output the estimated value of the final tensor, which is used to interpolate the missing measurement values in the multi-measurement data, thereby completing the filling of the missing measurement values.

2. The Bayesian tensor completion method based on multi-measurements according to claim 1, wherein: Among them, In step S1, the multi-measurement data is represented as a d-dimensional tensor There are no missing values, including i tensor elements x i , where i = (i1,..., i d ), (1 ≤ i k ≤ n k , k = 1,..., d), Some of the tensor elements have no observed values, that is, there are missing values. The remaining observable tensor elements have l observations using a series of said tensors to represent all of the observations, i.e., the multi-measurement data, where m = max(m i ) corresponding to all i, The first observations of all the said tensor elements are arranged and placed in If there is no observation for a certain said tensor element, the corresponding element in is also missing, to Similarly, Use a series of indicator tensors to indicate the situation of the missing values, and the elements of are p when \(i\in\Omega\) otherwise where \(\Omega\) p represents the set of subscripts of the observed elements In step S2, for the estimated value of perform decomposition: In the formula, is the j-th column of the factor matrix , and the symbol represents the outer product.

3. The Bayesian tensor completion method based on multi-measurements according to claim 2, wherein: Among them, The parameters of the factor matrix include row vectors hyperparameter μ (k) , Λ (k) and precision τ ∈ .

4. The Bayesian tensor completion method based on multi-measurements according to claim 3, wherein: Among them, In step S3, it is set that the prior distributions of the row vectors all satisfy wherein, Row vector The posterior distribution of the condition is as follows: In the formula, L() is the likelihood function, N() is the normal distribution, and there is: wherein, represents the Hadamard product, and the conditional posterior distribution of the row vectors to can be obtained in the same way. Set the hyperparameters The prior distribution of which satisfies Gaussian-Wishart(μ0,β0,W0,v0), That is, p(μ (k) , Λ (k) | -) = N(μ (k) | μ (0) , (β0Λ (k) ) -1 ) × Wishart(Λ (k) | W0, μ0), When k = 1, the conditional posterior distribution of the hyperparameter μ (k) , Λ (k) is as follows: In the formula: Variable and S (1) are respectively: The hyperparameters μ when k = 2, …, d (k) , Λ (k) The conditional posterior distribution of can be obtained in the same way, Set the accuracy τ ∈ The prior distribution of ∈ satisfies τ The precision τ ∈ has the following conditional posterior distribution: In the formula:

5. The Bayesian tensor completion method based on multi-measurements according to claim 1, wherein: Among them, In Step S5, the estimated value of the final tensor is the average of the estimated values of the N tensors in N iterations.

Citation Information

Patent Citations

  • Visual data completion method based on local low-rank tensor estimation

    CN107016649A

  • Missing traffic data repairing method based on Bayesian enhanced tensor

    CN110223509A