Improved estimation method of distributed random parameters under the condition of probability density overlap
By adjusting the probability density of Gaussian distribution components through the improved EM algorithm, the parameter estimation error problem of the mixed Gaussian distribution model in the case of overlapping probability densities is solved, and higher data processing accuracy is achieved.
Patent Information
- Application Number
- CN202411883535.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-19
AI Technical Summary
In the dynamic analysis of engineering structures, existing technologies have large errors in parameter estimation of mixed Gaussian distribution models when probability densities overlap, affecting the accuracy of data processing.
An improved expectation maximization algorithm (EM algorithm) is used to adjust the probability density of sample points belonging to different Gaussian distribution components by introducing an appropriate adjustment factor α, thereby improving the parameter estimation method of the mixed Gaussian distribution model and improving the estimation accuracy.
It effectively reduces the error in parameter estimation of the mixed Gaussian distribution model and improves the accuracy of data processing, especially the accuracy of parameter estimation in areas with overlapping probability densities.
Smart Images

Figure CN119761030B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing, and relates to a method for estimating distributed random parameters, and in particular to an improved method for estimating distributed random parameters under the condition of overlapping probability densities. Background Art
[0002] In the dynamic analysis of actual engineering structures, accurate measurement of parameters such as structural mass, moment of inertia, and operating speed is crucial for predicting and subsequently evaluating the structural operational state. However, due to differences in structural component production batches and operating environments, the measured parameters exhibit randomness. Furthermore, the probability density peaks of data measured from different batches are often closely spaced, resulting in overlapping overall probability density distributions. Therefore, accurately estimating the distribution of structural parameters in the presence of overlapping probability densities is crucial for operational evaluation of engineering structures. The Gaussian distribution model (normal distribution) is the most common and widely used random distribution model, characterized by a single peak. However, in the measured data of structural state parameters in actual engineering applications, the heterogeneity of structural products makes their true random distribution difficult to accurately fit using a single normal distribution. The Gaussian mixture distribution model (GMM) is a further generalization of the Gaussian distribution model (normal distribution). However, when the probability density peaks of the measured data are close together, probability density aliasing occurs. Conventional Gaussian mixture distribution model parameter estimation based on expectation maximization produces large errors in the overlapping region, necessitating further improvement.
[0003] A single Gaussian distribution model has the following form:
[0004] T~N(μ,σ2)(1)
[0005] in:
[0006] μ represents the mean of the Gaussian distribution; σ 2 Represents the variance of the Gaussian distribution.
[0007] The mixed Gaussian distribution model has the following form:
[0008]
[0009] Among them, represents the probability of the Gaussian distribution component and satisfies:
[0010]
[0011] μ i is the mean of the ith Gaussian distribution,
[0012] σ i is the standard deviation of the ith Gaussian distribution.
[0013] When processing data based on the Gaussian mixture model, it is assumed that the measured data is x={x1,x2,…x N}, is the number of measurement data, then the mixed Gaussian model can be expressed as follows:
[0014]
[0015] in:
[0016] x i is the true value of the measured data;
[0017] K represents the number of Gaussian distributions in the mixed Gaussian distribution;
[0018] represents the kth Gaussian distribution in the mixture model;
[0019] represents the mean and variance of the k-th Gaussian distribution.
[0020] If two Gaussian distributions are required, then K = 2. π k is the weight coefficient of the kth Gaussian distribution, which represents the weight of the kth component in the mixed Gaussian model. The sum of all weights should be 1, that is:
[0021]
[0022] When using a Gaussian mixture model to determine the random distribution of measurement data, it is necessary to determine the relevant parameters of the model, including the mean, variance, and weights of each component. However, research has shown that in existing technologies, when data aliasing is severe, the estimated parameters of the Gaussian mixture distribution have relatively large errors, typically exceeding 10%, affecting the accuracy of data processing. Summary of the Invention
[0023] In order to solve the above technical problems existing in the background technology, the present invention provides an improved estimation method for distributed random parameters under probability density overlap, which can reduce estimation errors and improve data processing accuracy.
[0024] In order to achieve the above object, the present invention adopts the following technical solutions:
[0025] An improved method for estimating distributed random parameters under overlapping probability densities, characterized in that the improved method for estimating distributed random parameters under overlapping probability densities comprises the following steps:
[0026] 1) Obtaining dynamic parameters of the engineering structure; the dynamic parameters include structural mass, moment of inertia, and operating speed;
[0027] 2) performing distributed random parameter estimation on the kinetic parameters obtained in step 1) to improve the accuracy of the probability density of the kinetic parameter distribution when the probability density overlaps; the specific implementation method of step 2) is:
[0028] 2.1) Obtaining measurement data of structural mass and moment of inertia of dynamic parameters;
[0029] 2.2) Based on the mean and variance of the measured data obtained in step 2.1), the probability density parameter θ of the distributed parameter to be estimated is initialized, and the initial value of the estimate is taken as θ0.
[0030] 2.3) For each strategy data sample of structural mass and moment of inertia, estimate the generation probability wi of each measurement data sample xi in each category based on the expectation maximization method k ;
[0031] 2.4) According to the sample point x i The numerical distribution of and the 3σ criterion in the normal distribution are used to determine the generation probability w estimated in step 2.3). ik Whether to update, if so, update the probability density And the updated probability density As the new probability w i,k If not, keep the generation probability w estimated in step 2.3) ik ;
[0032] 2.5) The probability w determined in step 2.4) ik , re-estimate the parameters of each Gaussian distribution Determine whether the estimated value converges. If so, complete the distributed random parameter estimation. If not, repeat steps 2.3) to 2.4) until the estimated value obtained in step 2.5) converges, and complete the distributed random parameter estimation.
[0033] 3) Based on the result obtained in step 2), a stability analysis is performed on the engineering structure, wherein the engineering structure stability analysis is performed in a manner of predicting or evaluating the operating status of the engineering structure.
[0034] Each measurement data sample x in step 2.3) above i The expression of the generation probability in each category is:
[0035] w ik =P(z=k|x i ,θ)
[0036] in:
[0037] wik represents the probability that the i-th measurement data belongs to the k-th classification;
[0038] Z is a K-dimensional random variable, z = k represents the k-th classification, indicating the category implied by the measurement data, and the value range of Z is: z = {1, 2, ..., K};
[0039] =P(z=k|x i ,θ) is calculated as follows:
[0040]
[0041] φ(x i |θ k ) indicates that is the Gaussian distribution function with mean and variance;
[0042] xi is the experimental measurement data, xi obeys Gaussian distribution, that is, p(x i |z i =k) represents the conditional probability density of the kth Gaussian distribution, satisfying the Gaussian distribution
[0043] The joint probability density of zi and xi should satisfy p(x i ,z i )=p(x i |z i )p(z i ).
[0044] The above random parameter π k The expression is:
[0045]
[0046] The random parameter μ k The expression is:
[0047]
[0048] The random parameters The expression is:
[0049]
[0050] The updated conditions in step 2.4) above are:
[0051] When x i Satisfy μ k+1 -3σ k+1 <x i <μ k +3σ k When w i,k and w i,k+1 , otherwise no update;
[0052] in:
[0053] μ k and μ k+1 are the means of the kth and k+1th Gaussian distributions respectively;
[0054] σ k and σ k+1 are the standard deviations of the kth and k+1th Gaussian distributions respectively.
[0055] The specific implementation method of the update in the above step 2.4) is:
[0056]
[0057] in:
[0058] w i,k and w i,k+1 are the probability densities of two adjacent components respectively;
[0059] as well as are the probability densities of two adjacent components after updating;
[0060] α is about (w i,k -w i,k+1 ) adjustment factor.
[0061] The value of α mentioned above varies with (w i,k -w i,k+1 )The absolute value of α changes. When the difference between the two is large, α becomes smaller, and when the difference between the two is small, α becomes larger.
[0062] The above α is a constant, and the α is 0.5.
[0063] The advantages of the present invention are:
[0064] The present invention provides a distributed random parameter estimation method based on an improved EM algorithm and a mixture Gaussian distribution. The method includes: a modified expectation-maximization (EM) algorithm to improve the parameter estimation accuracy of a mixture Gaussian distribution model (GM) in the presence of overlapping probability densities. When estimating parameters of a mixture Gaussian model with a relatively large overlapping region, the traditional EM algorithm is used to increase the probability that a sample point belongs to the component with the larger probability density. This method proposes introducing an appropriate adjustment factor for sample points in the overlapping probability region to adjust the probability density of the sample points belonging to different components when calculating the probability density of the sample points, thereby improving the parameter estimation accuracy of the traditional EM algorithm in the presence of overlapping probability densities. When estimating parameters of a mixture Gaussian model with a relatively large overlapping region, the method uses the traditional EM algorithm to increase the probability that a sample point belongs to the component with the larger probability density. When estimating parameters of a mixture Gaussian model with a relatively large overlapping region, the method introduces an appropriate adjustment factor for sample points in the overlapping probability region to adjust the probability density of the sample points belonging to different components when calculating the probability density of the sample points belonging to different components, thereby improving the parameter estimation accuracy of the traditional EM algorithm in the presence of overlapping probability densities. A comparison of the traditional EM algorithm with the modified EM algorithm is conducted through a numerical example. The results demonstrate that the improved EM algorithm can improve the accuracy of GM model parameter estimation, verifying the effectiveness of the proposed method. The present invention mainly aims at the random distribution and overlapping probability density characteristics of the measurement data of parameters such as structural mass, moment of inertia, and operating speed caused by different production batches and operating environments of components during the dynamic analysis of engineering structures. An improved estimation method for distributed random parameters under the condition of overlapping probability density is proposed. Through this method, the stability of engineering structures can be analyzed. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is a schematic diagram of the probability density estimation of the mixed Gaussian distribution;
[0066] Figure 2 Parameter estimation results based on the distributed random parameter estimation method provided by the present invention;
[0067] Figure 3 is the parameter estimation result of the traditional EM method. DETAILED DESCRIPTION
[0068] The present invention provides an improved method for estimating distributed random parameters in the case of overlapping probability densities. The core idea is: first, based on the given observed data, the values of the model parameters are estimated; then, based on the parameter values estimated in the previous step, the values of the missing data are estimated; then, based on the estimated missing data plus the previously observed data, the parameter values are re-estimated; and then the iteration is repeated until convergence is achieved and the iteration ends. Specifically, the improved method for estimating distributed random parameters in the case of overlapping probability densities provided by the present invention includes the following steps:
[0069] 1) Initialize the parameters to be estimated Take the initial value as θ0;
[0070] 2) Estimate the generation probability w of each measurement data sample in each Gaussian distribution based on the expectation maximization method ik ;
[0071] 3) Based on the generation probability distribution obtained in step 2) and the location of the sample collection, and according to the 3σ criterion in the normal distribution, determine whether to update the random parameters in the mixed Gaussian distribution;
[0072] 4) According to the judgment result obtained in step 3), according to the update criterion, the new random parameter w is obtained. i,k , w i,k+1 The updated value of .
[0073] 5) According to the updated w ik , estimate the parameters for each class Determine whether the estimated value converges. Otherwise, repeat 2)-5) until the estimated value of the random parameter of each classification obtained in step 3) converges, thereby completing the distributed random parameter estimation.
[0074] Among them, in step 2), in order to obtain the generation probability, the present invention adopts the following method: when determining the generation probability of each parameter value in each category, it can actually be understood as a two-step implementation. First, one is selected from the k components, and the probability of each component being selected is its coefficient π k After selecting a component, we consider its probability distribution in this component distribution separately. At this time, its distribution is a normal Gaussian distribution. Figure 1 , a comparison between Gaussian distribution and mixed Gaussian distribution. When the probability density is calculated based on data, it belongs to the category of probability density estimation. When the form of the probability density function is known (or assumed) and only the parameters need to be estimated, the problem is transformed into parameter estimation. The parameters to be estimated in formula (3) are:
[0075]
[0076] This is the parameter estimation problem of the Gaussian mixture distribution model, where Bayesian estimation and expectation maximization algorithm (EM) are needed.
[0077] Among them, π k It can be regarded as the probability of the kth category being selected. A new K-dimensional random variable z={1,2,…,K} is introduced, where z=k represents the kth category, indicating the category implied by the data. i Classification of z i After, x i Obey the Gaussian distribution, that is Zi, xi are both unknown variables, so their joint distribution can be obtained and satisfies
[0078] p(x i ,z i )=p(x i |z i )p(z i )(5)
[0079] The maximum likelihood estimation method is used to estimate the unknown parameters. The logarithmic maximum likelihood function is:
[0080]
[0081] This formula (6) cannot be derived to 0 to determine the estimated formula of the parameters, so the EM method needs to be used to solve it.
[0082] The first step is to assume that the zi value of each sample is known, and the above formula can be simplified to:
[0083]
[0084] Here, we can further derive the parameters of formula (7):
[0085]
[0086]
[0087]
[0088] Previously, it was assumed that zi was known, but in reality zi is unknown. In practice, it is necessary to adopt the EM idea to guess the implicit variable zi, and in the second step update other variables to obtain the maximum likelihood estimate.
[0089] In the parameter results obtained by equation (7), it is assumed that the sample point can only belong to one component. However, in practice, since the probability densities of the components may overlap, it is more appropriate to assume that a sample point has a certain probability of belonging to each component. Here, a new variable wik is introduced to represent the probability that the i-th data belongs to the k-th classification, that is, w ik =P(z=k|x i ,θ), according to the posterior probability density of Bayesian estimation, we can get:
[0090]
[0091] The Bayesian estimation and total probability formula are used in the derivation of formula (8). The above calculation of the generation probability of each data in each category is based on each data, and each category is a Gaussian distribution. According to the definition of Gaussian distribution, it can be estimated that The expression is:
[0092]
[0093]
[0094]
[0095] In the above process, the update conditions are:
[0096] For the sample points in the probability overlap part, the factor α is introduced to adjust the probability density w of the sample points belonging to different components in the calculation. i,k When α is taken as i,k and w i,k+1 When they are close, the smaller value is appropriately increased. In most cases, it is not desirable to change w too much which is not near the sample point. i,k Value, only update the value with overlapping area, according to the 3σ criterion in normal distribution, only when x i Satisfy μ k+1 -3σ k+1 <x i <μ k +3σ k When w i,k and w i,k+1 , otherwise no update.
[0097] The specific implementation of the update is:
[0098] If the probability density w of the sample point belongs to its two adjacent components i,k , w i,k+1 , the updated probability density is Satisfies the following expression:
[0099]
[0100] Generally, α can be a certain value about (w i,k -w i,k+1 ) function, whose value varies with (w i,k -w i,k+1 ) changes in their absolute values. When the difference between the two is large, α becomes smaller, and when the difference is small, α becomes larger. The simplest case is to choose α as a constant function, with a value of 0.5.
[0101] Example 1: To verify the accuracy of the proposed improved EM algorithm, the traditional EM algorithm and the improved EM algorithm are used to analyze the generated measurement data. The data contains a mixture Gaussian distribution of three univariate normal distribution components. The model is as follows:
[0102]
[0103] Table 1 The true parameter values of each Gaussian component function
[0104]
[0105] In practice, the sample data to be analyzed can be generated by model (8). 4000 samples are randomly generated, of which 2000 are generated by component 1, 1200 are generated by component 2, and 800 are generated by component 3. The traditional EM algorithm and the improved EM algorithm are used for parameter estimation respectively, and the parameter values after sample estimation can be obtained:
[0106] Table 2 Parameter estimates of each Gaussian component function
[0107]
[0108]
[0109] As can be seen from Table 1 and Table 2, compared with the traditional EM algorithm, the improved EM algorithm can better estimate the parameters in the mixed Gaussian distribution model. In order to demonstrate the effectiveness of the estimation method, the measured data are statistically analyzed, and the statistical probability density of the data can be obtained by the Monte Carlo method (MCS). Then, the mixed Gaussian distribution obtained by the EM method is compared with the simulated distribution obtained by the MCS method and the true solution. Figure 2 It can be seen that the probability density of the third Gaussian mixture distribution component overlaps less with the first two. Therefore, both the traditional EM algorithm and the improved EM algorithm can fit the probability density of the sample well. However, since the means of the first Gaussian distribution component and the second Gaussian distribution are close, the probability density overlaps more. The mixed Gaussian distribution model estimated by the traditional EM algorithm (see Figure 3 ), with large errors in these two regions, and the maximum error value of the probability density is 0.044. As can be seen from the figure, the mixed probability density obtained based on the improved EM method is in good agreement with both the theoretical solution and the simulation solution, which can increase the accuracy of the estimated parameters. The maximum error value of the probability density is 0.03, thus better fitting the probability density of the sample.
Claims
1. An improved method for estimating distributed random parameters in the case of overlapping probability densities, characterized by: The improved estimation method of distributed random parameters under the condition of overlapping probability densities comprises the following steps: 1) Obtaining dynamic parameters of the engineering structure; the dynamic parameters include structural mass, moment of inertia, and operating speed; 2) performing distributed random parameter estimation on the kinetic parameters obtained in step 1) to improve the accuracy of the probability density of the kinetic parameter distribution when the probability density overlaps; the specific implementation method of step 2) is: 2.1) Obtaining measurement data of structural mass and moment of inertia of dynamic parameters; 2.2) Based on the mean and variance of the measured data obtained in step 2.1), the probability density parameter θ of the distributed parameter to be estimated is initialized, and the initial value of the estimate is taken as θ0. 2.3) For each strategy data sample of structural mass and moment of inertia, estimate each measurement data sample x based on the expectation maximization method i The generation probability w in each category ik ; 2.4) According to the sample point x i The numerical distribution of and the 3σ criterion in the normal distribution are used to determine the generation probability w estimated in step 2.3). ik Whether to update, if so, update the probability density And the updated probability density As the new probability w i,k If not, keep the generation probability w estimated in step 2.3) ik ; 2.5) The probability w determined in step 2.4) i k , re-estimate the parameters of each Gaussian distribution Determine whether the estimated value converges. If so, complete the distributed random parameter estimation. If not, repeat steps 2.3) to 2.4) until the estimated value obtained in step 2.5) converges, and complete the distributed random parameter estimation. 3) Based on the result obtained in step 2), a stability analysis is performed on the engineering structure, wherein the engineering structure stability analysis is performed in a manner of predicting or evaluating the operating status of the engineering structure.
2. The improved method for estimating distributed random parameters under overlapping probability densities according to claim 1, characterized in that: In step 2.3), each measurement data sample x i The expression of the generation probability in each category is: w ik =P(z=k|x i ,θ) in: w ik Represents the probability that the i-th measurement data belongs to the k-th classification; Z is a K-dimensional random variable, z = k represents the k-th probability density category, indicating the probability density category implied by the measurement data, and the value range of Z is: z = {1, 2, ..., K}; P(z=k|x i ,θ) is calculated as follows: φ(x i |θ k ) indicates that is the Gaussian distribution function with mean and variance; x i is the experimental measurement data, x i Obey Gaussian distribution, that is, p(x i |z i =k) represents the conditional probability density of the kth Gaussian distribution, satisfying the Gaussian distribution The z i and x i The joint probability density should satisfy p(x i ,z i )=p(x i |z i )p(z i ).
3. The improved method for estimating distributed random parameters under overlapping probability densities according to claim 2, characterized in that: The random parameter π k The expression is: The random parameter μ k The expression is: The random parameters The expression is:
4. The improved method for estimating distributed random parameters under overlapping probability densities according to claim 3, characterized in that: The conditions updated in step 2.4) are: When x i Satisfy μ k+1 -3σ k+1 <x i <μ k +3σ k When w i,k and w i,k+1 , otherwise no update; in: μ k and μ k+1 are the means of the kth and k+1th Gaussian distributions respectively; σ k and σ k+1 are the standard deviations of the kth and k+1th Gaussian distributions respectively.
5. The improved method for estimating distributed random parameters under overlapping probability densities according to claim 4, characterized in that: The specific implementation method of the update in step 2.4) is: in: w i,k and w i,k+1 are the probability densities of two adjacent components respectively; as well as are the probability densities of two adjacent components after updating; α is about (w i,k -w i,k+1 ) adjustment factor.
6. The improved method for estimating distributed random parameters under overlapping probability densities according to claim 5, characterized in that: The value of α varies with (w i,k -w i,k+1 )The absolute value of α changes. When the difference between the two is large, α becomes smaller, and when the difference between the two is small, α becomes larger.
7. The improved method for estimating distributed random parameters under overlapping probability densities according to claim 6, characterized in that: The α is usually selected as a constant, and the α is 0.5.
Citation Information
Patent Citations
Data-driven complex electromechanical system service quality state evaluation method
CN106682835A
Distributed high-dimensional uncertainty quantification method based on depth flow model
CN113128100A