Mixed probability distribution parameter estimation method based on deep learning

Through the hybrid probability distribution parameter estimation method based on deep learning, the problem of high computational complexity of mixed probability distribution parameter estimation in the prior art is solved, and fast and accurate parameter estimation under large data volumes are realized.

CN120045833AActive Publication Date: 2025-05-27ZHEJIANG SCI-TECH UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510100412.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art has problems with high computational complexity and long processing time in the estimation of mixed probability distribution parameters. Especially in the case of large data volumes, it is difficult to achieve efficient and accurate parameter estimation.

Method used

The mixed probability distribution parameter estimation method based on deep learning is adopted, and the probability density function and objective function are established through data preprocessing, and the lower bound function is iteratively optimized in combination with deep learning methods to approximate the solution of parameter θ. This method does not need to traverse all data during the iteration process, and only uses representative data for parameter estimation.

Benefits of technology

It realizes fast and accurate estimation of complex mixed probability distribution parameters of large data volumes, improves processing speed and convergence speed, and ensures high accuracy of parameter estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045833A_ABST
    Figure CN120045833A_ABST
Patent Text Reader

Abstract

The invention discloses a mixed probability distribution parameter estimation method based on deep learning, which comprises the following steps: calculating the mean value and variance of all data samples in a given data set X, then normalizing each data sample to obtain a normalized sample X (i), i = [1, 2,..., N], and N being the total number of samples; pk (X (i), theta k) is the probability density function of the k-th component distribution in the mixed probability distribution, wk is the weight of the k-th probability distribution in the overall distribution, theta k is the parameter of the probability density function of the k-th component, k = [1, 2,..., K], and K is the total number of components of the mixed distribution; establishing a target function and a lower bound function of hybrid probability distribution parameter estimation; and iteratively optimizing the lower bound function by adopting a deep learning method, and solving a parameter theta. The method has the advantages of being capable of processing large data volume, high in processing speed and high in estimation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for estimating parameters of a mixed probability distribution, in particular to a method for estimating parameters of a mixed probability distribution based on deep learning. Background Art

[0002] A mixed probability distribution refers to a probability distribution formed by combining a number of independent parametric probability distribution components according to a certain weight. Parameter estimation of a mixed probability distribution refers to the process of inferring or estimating unknown parameters and weights in the probability distribution components, which is widely applied in fields such as engineering simulation, anomaly detection, signal compression, etc. Existing methods for estimating parameters of probability distributions include statistical methods such as the method of moment estimation, maximum likelihood estimation, Bayesian estimation, least squares method, maximum spacing estimation, quantile estimation, EM algorithm, etc.

[0003] Since the mixed probability distribution is more complex in form than ordinary probability distributions, many problems will occur in the practical application of existing methods for estimating parameters of probability distributions in estimating mixed probability distributions. For example, methods such as maximum likelihood estimation and EM algorithm that use optimization methods to find optimal parameters need to calculate the partial derivatives of parameters for the specific probability distribution of each component in each mixed distribution. The calculation method is complex, difficult to directly solve, and prone to errors.

[0004] When using a deterministic iterative method such as Newton's iterative method to calculate the roots of the partial derivative formulas obtained in maximum likelihood estimation and EM algorithm, all data need to be traversed in each iteration, the processing time is long, and the application is relatively limited, unable to handle large amounts of data.

[0005] Therefore, in the prior art, there is no method with a fast processing speed and high estimation accuracy for estimating parameters of complex mixed probability distributions with large amounts of data. Summary of the Invention

[0006] The object of the present invention is to provide a method for estimating parameters of a mixed probability distribution based on deep learning. The present invention has the characteristics of being able to handle large amounts of data, with a fast processing speed and high estimation accuracy.

[0007] The technical solution of the present invention: A method for estimating parameters of a mixed probability distribution based on deep learning includes the following steps:

[0008] 1. Data preprocessing:

[0009] Given a set of data sets X, calculate the mean and variance of all data samples in the data set, and then normalize each data sample to obtain the normalized sample X (i) , i = [1, 2,..., N], where N is the total number of samples;

[0010] 2. Establishment of the probability density function of the mixture probability distribution:

[0011] Probability density function of the mixture probability distribution where p k (X (i) , θ k ) is the probability density function of the k-th component probability distribution in the mixture probability distribution, w k is the weight of the k-th probability distribution in the overall distribution, θ k is the parameter of the probability density function of the k-th component, k = [1, 2,..., K], and K is the total number of components of the mixture distribution; θ = <θ 1 , θ 2 ,..., θ K > is the total parameter vector composed of the probability density function parameters of all components;

[0012] 3. Establishment of the objective function for parameter estimation of the mixture probability distribution:

[0013] Establish the objective function and the lower bound function for parameter estimation of the mixture probability distribution;

[0014] 4. Use the deep learning method to iteratively optimize the lower bound function and solve for the parameter θ:

[0015] 4.1) Sample the parameter θ from a standard normal distribution with a mean of 0 and a variance of 1 as the initial value of the parameter;

[0016] 4.2) Initialize w k = 1 / K, initialize the error count en to 0, and the number of steps step = 1;

[0017] 4.3) Define the latent variable Z as a two-dimensional matrix of N rows and K columns and calculate the value of Z;

[0018] 4.4) Calculate w k according to the value of Z, calculate the Q function according to the value of Z and w k and calculate θ' according to the Q function;

[0019] 4.5) Let the error e = |θ - θ'|. If the error e is less than 0.001, then let en = en + 1; otherwise, let en = 0;

[0020] 4.6) Let θ = θ'. If en > the error termination condition, then output the parameter θ; otherwise, execute step 4.7);

[0021] 4.7) Let step = step + 1. If step > the maximum number of steps, then output the parameter θ; otherwise, execute step 4.3).

[0022] In the above-mentioned method for estimating parameters of a mixed probability distribution based on deep learning, in step 1, the data set X = {x (1) , x (2) ,......, x (N)}; where x (i) is the i-th data sample, i = [1, 2,..., N], N is the total number of samples, and each x (i) is represented as a P-dimensional vector, that is, x (i) = <x i1 , x i2 ,....x iP >, and P is the dimension of the vector.

[0023] In the above-mentioned method for estimating parameters of a mixed probability distribution based on deep learning, in step 1, the mean of the data samples The standard deviation of the data samples

[0024]

[0025] In the above-mentioned method for estimating parameters of a mixed probability distribution based on deep learning, in step 1, the normalized sample X (i) = <(x i1 - u 1 ) / σ 1 , (x i2 - u 2 ) / σ 2 ,...,(x iP - u P ) / σ P >.

[0026] In the above-mentioned method for estimating parameters of a mixed probability distribution based on deep learning, in step 3, the objective function The lower bound function

[0027]

[0028] In the above-mentioned method for estimating parameters of a mixed probability distribution based on deep learning, before solving the parameter θ in step 4, it also includes predefined hyperparameters, and the hyperparameters include: the number of data batches batch_size, the number of execution rounds num_epochs, the learning rate lr, the error termination condition em, and the maximum number of steps maxstep.

[0029] In the above-mentioned method for estimating parameters of a mixed probability distribution based on deep learning, batch_size is 512 - 4096, and the default value is 2048; num_epochs is 20 - 100, and the default value is 30; lr is 1e-3; em is 5, and maxstep is 100.

[0030] In the aforementioned method for estimating parameters of a mixed probability distribution based on deep learning, 4.3) specifically is as follows:

[0031] 4.3.1) Input the entire data set X;

[0032] 4.3.2) Initialize all elements of the latent variable Z to 0;

[0033] 4.3.3) Let i = 1;

[0034] 4.3.4) Obtain the i-th data X in the data set X (i) ;

[0035] 4.3.5) Let sum_Zi = 0;

[0036] 4.3.6) Let k = 1;

[0037] 4.3.7) Z[i,k] = w k *p k (X (i) ,θ k );

[0038] 4.3.8) sum_Zi = sum_Zi + Z[i,k];

[0039] 4.3.9) Let k = k + 1, if k > K, then execute step 5.3.10), otherwise execute step 5.3.7);

[0040] 4.3.10) Let k = 1;

[0041] 4.3.11) Let Z[i,k] = Z[i,k] / sum_Zi;

[0042] 4.3.12) Let k = k + 1, if k > K, then execute step 4.3.13), otherwise execute step 4.3.11);

[0043] 4.3.13) Let i = i + 1, if i > N, then the calculation of Z[i,k] is completed, and execute step 4.4), otherwise execute step 4.3.4).

[0044] In the aforementioned method for estimating parameters of a mixed probability distribution based on deep learning, 4.4) specifically is as follows:

[0045] 4.4.1) Define θ' = θ, let B = N / batch_size, divide the data set into B batches, with the data volume of each batch being batch_size, and randomly batch each data in the data set;

[0046] 4.4.2) Let epoch = 1;

[0047] 4.4.3) Let b = 1. Let X[b] be the b-th batch of data, and X[b, i] be the i-th data in the b-th batch of data. The serial number of this data in the original dataset is idx i ;

[0048] 4.4.4) Calculate the lower bound function Q function;

[0049] 4.4.5) Calculate the gradient of the parameter θ' on the negative function -Q of the Q function and perform backpropagation processing. Among them, the optimizer for backpropagation uses the Nadam method, the optimization step size uses a learning rate of 1e-3. After backpropagation, θ' is updated by the Nadam optimizer to obtain θ';

[0050] 4.4.6) Let b = b + 1. If b > B, execute step 4.4.7), otherwise execute step 4.4.4);

[0051] 4.4.7) Let epoch = epoch + 1. If epoch > num_epochs, execute step 4.4.8), otherwise jump to 4.4.3).

[0052] 4.4.8) Let

[0053] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0054] The present invention uses the expectation maximization process to achieve approximate optimization, and uses the stochastic optimization method in the approximate maximization calculation process, so that each iteration does not require the use of all data, but selects a batch of representative data for parameter estimation, thereby being able to process the probability distribution parameters of a large amount of data, greatly improving the processing speed and convergence speed of the probability distribution parameter estimation, and having high parameter estimation accuracy. Description of the Drawings

[0055] Figure 1 is a set of actual signals.

[0056] Figure 2 is the density map given by the existing statistical non-parametric estimation method.

[0057] Figure 3 is the probability distribution map obtained according to the parameters calculated in the 10th iteration of the present invention.

[0058] Figure 4 is the probability distribution map obtained according to the parameters calculated in the 50th iteration of the present invention.

[0059] Figure 5 is the probability distribution map obtained according to the parameters calculated in the 90th iteration of the present invention. Detailed Embodiments

[0060] The present invention will be further described below in conjunction with embodiments, but it shall not be used as a basis for limiting the present invention.

[0061] Embodiment 1:

[0062] A method for estimating parameters of a mixed probability distribution based on deep learning, comprising the following steps:

[0063] 1. Data preprocessing:

[0064] 1.1) Given a set of data sets X: X = {x (1) , x (2) ,......, x (N)}; where x (i) is the i-th data sample, i = [1, 2,..., N], N is the total number of samples, and each x (i) is represented as a P-dimensional vector, that is, x (i) = <x i1 , x i2 ,....x iP >, and P is the dimension of the vector.

[0065] 1.2) Calculate the mean and variance of all data samples, where the mean of each dimension The standard deviation of each dimension

[0066] 1.3) Normalize each data sample to obtain the normalized sample X (i) : X (i) = <(x i1 - u 1 ) / σ 1 , (x i2 - u 2 ) / σ 2 ,...,(x iP - u P ) / σ P >.

[0067] 2. Establishment of the probability density function of the mixed probability distribution:

[0068] The probability density function pdf of the mixed probability distribution: where p k (X (i) , θ k ) is the probability density function of the k-th component probability distribution in the mixed probability distribution, w k is the weight of the k-th probability distribution in the overall distribution, θ k is the parameter of the probability density function of the k-th component, k = [1, 2,..., K], and K is the total number of components of the mixed distribution; θ = <θ 1 , θ2 ,..., θ K is the total parameter vector composed of the probability density function parameters of all components, that is, the parameter to be estimated.

[0069] 3. Establishment of the objective function for estimating the parameters of the mixture probability distribution:

[0070] Objective function for estimating the parameters of the mixture probability distribution

[0071]

[0072] In practical applications, the objective function usually has a relatively complex form and is difficult to directly calculate precisely. Therefore, in the present invention, a deep learning method is used to iteratively approximate the solution of the objective function. The objective function LMLE has a complex form in practical applications and is difficult to directly optimize for precise calculation. Therefore, the present invention establishes a lower bound function of the objective function and iteratively optimizes the lower bound function through a deep learning method to achieve the effect of approximately solving the objective function.

[0073] 4. Predefined hyperparameters:

[0074] 4.1) The number of data batches batch_size is 512 - 4096, and the default value is 2048;

[0075] 4.2) The number of execution rounds num_epochs is 20 - 100, and the default value is 30;

[0076] 4.3) The learning rate lr is 1e - 3;

[0077] 4.4) The error termination condition em is 5, and the maximum number of steps maxstep is 100.

[0078] 5. Use a deep learning method to iteratively optimize the Q function and solve for the parameter θ:

[0079] 5.1) Sample the parameter θ from a standard normal distribution with a mean of 0 and a variance of 1 as the initial value of the parameter.

[0080] 5.2) Initialize w k = 1 / K, initialize the error count en to 0, and the number of steps step = 1.

[0081] 5.3) Define the latent variable Z as a two-dimensional matrix of N rows and K columns, representing the posterior probability that X (i) comes from the k-th component in the mixture probability distribution, and perform the E-step to calculate the value of Z. The specific process is as follows:

[0082] 5.3.1) Input the entire dataset X;

[0083] 5.3.2) Initialize all elements of the latent variable Z to 0;

[0084] 5.3.3) Let i = 1;

[0085] 5.3.4) Obtain the i-th data Xi in the data set X (i) ;

[0086] 5.3.5) Let sum_Zi = 0;

[0087] 5.3.6) Let k = 1;

[0088] 5.3.7) Z[i, k] = w k *p k (X (i) , θ k );

[0089] 5.3.8) sum_Zi = sum_Zi + Z[i, k];

[0090] 5.3.9) Let k = k + 1. If k > K, then execute step 5.3.10); otherwise, execute step 5.3.7);

[0091] 5.3.10) Let k = 1;

[0092] 5.3.11) Let Z[i, k] = Z[i, k] / sum_Zi;

[0093] 5.3.12) Let k = k + 1. If k > K, then execute step 5.3.13); otherwise, execute step 5.3.11);

[0094] 5.3.13) Let i = i + 1. If i > N, then the calculation of Z[i, k] is completed, and execute step 5.4); otherwise, execute step 5.3.4);

[0095] 5.4) Execute random M steps to calculate θ', and the specific process is as follows:

[0096] 5.4.1) Define θ' = θ. Let B = N / batch_size, divide the data set into B batches, with the data volume of each batch being batch_size, and randomly batch each data in the data set;

[0097] 5.4.2) Let epoch = 1;

[0098] 5.4.3) Let b = 1, let X[b] be the b-th batch of data, X[b, i] be the i-th data in the b-th batch of data, and the serial number of this data in the original data set is idx i ;

[0099] 5.4.4) Calculate the Q function. The Q function represents the expectation of the posterior of the joint probability distribution of the data and the latent variable Z with respect to the latent variable under the current parameters:

[0100]

[0101] 5.4.5) Take the gradient of the parameter θ' with respect to the negative function -Q of the Q function and perform backpropagation, where the optimizer for backpropagation uses the Nadam method, the optimization step size uses a learning rate of 1e-3, and after backpropagation, θ' is updated by the Nadam optimizer to obtain θ';

[0102] 5.4.6) Let b = b + 1. If b > B, then execute step 5.4.7); otherwise, execute step 5.4.4);

[0103] 5.4.7) Let epoch = epoch + 1. If epoch > num_epochs, then execute step 5.4.8); otherwise, jump to 5.4.3).

[0104] 5.4.8) Let

[0105] 5.5) Let the error e = |θ - θ'|. If the error e is less than 0.001, then let en = en + 1; otherwise, let en = 0.

[0106] 5.6) Let θ = θ'. If en > em, then the parameter learning process ends, and the learned parameter θ is output; otherwise, execute step 5.7);

[0107] 5.7) Let step = step + 1. If step > maxstep, then the parameter learning process ends, and the learned parameter θ is output; otherwise, execute step 5.3).

[0108] Example 2:

[0109] In a communication system, a signal is represented as a two-dimensional vector X = {X (i)} = {X i1 , X i2}, i = [1, 2,..., N]. According to known experience, the signals in this communication system follow an asymmetric binary mixture generalized normal probability distribution, and its probability density function is as follows:

[0110]

[0111] where

[0112] where θ = <θ 1 , θ 2 ,... θ K >, θ k = <α kd , μ kd , σ kd1, σ kd2 , k = [1, 2,...., K], d = [1, 2].

[0113] In the above probability density function, θ is an unknown parameter, that is, the parameter θ to be estimated is a four-dimensional vector. In order to analyze these signals, it is often necessary to first determine the probability distribution parameters of these signals and then perform subsequent signal processing.

[0114] Figure 1 A set of samples of actual signals is given, where the blue dots are two-dimensional signal points; Figure 2 is the density map given by the kernel density estimation method (KDE), which can be regarded as an approximation of the true distribution; Figures 3 - 5 Are the probability distribution graphs drawn with the parameters calculated by this method at the 10th iteration (step = 10), the 50th iteration, and the 90th iteration respectively. It can be seen that as the number of iterations increases, the probability distribution graph becomes more and more similar to the true distribution. By obtaining the mathematical representation of the mixed probability distribution, that is, the mathematical form and parameters of the mixed probability distribution, in practical engineering applications, signal compression, simulation, etc. can be further performed.

[0115] It should be understood that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Those skilled in the art can modify the technical solutions recorded in the above embodiments or equivalently replace some of the technical features; and all such modifications and replacements should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for estimating parameters of a mixed probability distribution based on deep learning, characterized in that: The following steps are involved:

1. Data preprocessing: Given a set of data sets X, calculate the mean and variance of all data samples in the data set, and then normalize each data sample to obtain the normalized sample X (i) , i=[1,2,...,N], N is the total number of samples; 2. Establishment of probability density function of mixed probability distribution: Probability density function of a mixed probability distribution where p k (X (i) ,θ k ) is the probability density function of the kth component probability distribution in the mixed probability distribution, w k is the weight of the kth probability distribution in the overall distribution, θ k is the parameter of the probability density function of the kth component, k = [1, 2, ..., K], K is the total number of components of the mixed distribution; θ = <θ1, θ2, ..., θ K > is the total parameter vector consisting of the probability density function parameters of all components; 3. Establishment of the objective function for mixed probability distribution parameter estimation: Establish the objective function and lower bound function for the estimation of the parameters of the mixed probability distribution; 4. Use deep learning methods to iteratively optimize the lower bound function and solve the parameter θ: 4.1) Sample the parameter θ from a standard normal distribution with a mean of 0 and a variance of 1 as the initial value of the parameter; 4.2) Initialize w k =1 / K, initialize the error count en to 0, and the number of steps step = 1; 4.3) Define the latent variable Z as a two-dimensional matrix with N rows and K columns, and calculate the value of Z; 4.4) Calculate w based on Z value k , according to the Z value and w k Calculate the Q function and obtain θ' based on the Q function; 4.5) Let error e = |θ-θ'|. If error e is less than 0.001, let en = en+1, otherwise let en = 0. 4.6) Let θ = θ', if en> error termination condition, then output parameter θ; otherwise, execute step 4.7); 4.7) Let step = step + 1. If step > maximum number of steps, output parameter θ; otherwise, execute step 4.3).

2. The method for estimating parameters of a mixed probability distribution based on deep learning according to claim 1, characterized in that: In step 1, the data set X = {x (1) ,x (2) ,......,x (N) }; where x (i) is the i-th data sample, i=[1,2,...,N], N is the total number of samples, and each x (i) Represented as a P-dimensional vector, that is, x (i) = <x i1 ,x i2 ,....x iP >, P is the dimension of the vector.

3. The method for estimating parameters of a mixed probability distribution based on deep learning according to claim 2, characterized in that: In step 1, the mean of the data sample is The standard deviation of the data sample 4. The method for estimating parameters of a mixed probability distribution based on deep learning according to claim 3, characterized in that: In step 1, normalize the sample X (i) =<(x i1 -u1) / σ1,(x i2 -u2) / σ2,...,(x iP -u P ) / σ P >.

5. The method for estimating parameters of a mixed probability distribution based on deep learning according to claim 1, characterized in that: In step 3, the objective function Lower bound function 6. The method for estimating parameters of a mixed probability distribution based on deep learning according to claim 1, characterized in that: Before solving the parameter θ in step 4, predefined hyperparameters are also included, including: batch size of data, number of execution rounds, learning rate, learning rate, error termination condition, error termination condition, and maximum number of steps, maxstep.

7. The method for estimating parameters of a mixed probability distribution based on deep learning according to claim 6, characterized in that: batch_size is 512-4096, the default value is 2048; num_epochs is 20-100, the default value is 30; lr is 1e-3; em is 5, and maxstep is 100.

8. The method for estimating parameters of a mixed probability distribution based on deep learning according to claim 5, characterized in that: 4.3) Specifically: 4.3.1) Input the entire data set X; 4.3.2) All elements of the hidden variable Z are initialized to 0; 4.3.3) Let i = 1; 4.3.4) Get the i-th data X in the data set X (i) ; 4.3.5) Let sum_Zi = 0; 4.3.6) Let k = 1; 4.3.7)Z[i,k]=w k *p k (X (i) ,θ k ); 4.3.8)sum_Zi=sum_Zi+Z[i,k]; 4.3.9) Let k = k + 1, if k>K, then execute step 5.3.10), otherwise execute step 5.3.7); 4.3.10) Let k = 1; 4.3.11) Let Z[i,k]=Z[i,k] / sum_Zi; 4.3.12) Let k = k + 1, if k>K, then execute step 4.3.13), otherwise execute step 4.3.11); 4.3.13) Let i=i+1. If i>N, then Z[i,k] is calculated and step 4.4) is executed. Otherwise, step 4.3.4) is executed.

9. The method for estimating parameters of a mixed probability distribution based on deep learning according to claim 5, characterized in that: 4.4) Specifically: 4.4.1) Define θ' = θ, set B = N / batch_size, divide the data set into B batches, the amount of data in each batch is batch_size, and randomly divide each data item in the data set into batches; 4.4.2) Set epoch = 1; 4.4.3) Let b = 1, let X[b] be the bth batch of data, X[b,i] be the i-th data in the bth batch of data, and the serial number of this data in the original data set is idx i ; 4.4.4) Calculate the lower bound function Q function; 4.4.5) The gradient of the parameter θ' is calculated on the negative function -Q of the Q function and back-propagated, wherein the back-propagation optimizer adopts the Nadam method, and the optimization step size adopts a learning rate of 1e-3. After back-propagation, θ' is updated by the Nadam optimizer to obtain θ'; 4.4.6) Let b = b + 1, if b>B, then execute step 4.4.7), otherwise execute step 4.4.4); 4.4.7) Let epoch = epoch + 1. If epoch > num_epochs, execute step 4.4.8), otherwise jump to 4.4.3). 4.4.8) Order

Citation Information

Patent Citations

  • Complex equipment reliability hybrid model and construction method thereof

    CN111814342A

  • Atmospheric turbulence channel fading parameter estimation method based on mixed distribution model

    CN112468229A

  • Adaptive weighted stochastic gradient descent

    US20130325401A1