Time series data stochastic simulation method based on greedy learning Gaussian mixture model
By introducing a greedy learning strategy into the Gaussian mixture model, gradually increasing the Gaussian components and optimizing the parameter set, the problem of low simulation accuracy in the existing technology is solved, efficient and accurate random simulation of time series data is achieved, and the simplicity and interpretability of the model are improved.
Patent Information
- Application Number
- CN202511105749.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-08-08
AI Technical Summary
Existing stochastic simulation methods have problems such as low simulation accuracy, difficulty in parameter optimization, sensitivity to initial values, and local optimality in time series data simulation, making it difficult to effectively capture the actual needs of power system management and optimization.
A greedy learning strategy is adopted to gradually add Gaussian components to the Gaussian mixture model. Through the split candidate component generation mechanism and double threshold constraints, the parameter set of the Gaussian mixture model is optimized, the probability density function and conditional cumulative distribution function of the Gaussian mixture model are constructed, and random simulation of time series data is performed.
It improves the simulation accuracy, ensures the simplicity and robustness of the model, can accurately describe the autocorrelation characteristics of the time series, reduces the risk of data generalization, and enhances the interpretability and uncertainty quantification capabilities of the model.
Smart Images

Figure CN120611537A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to data modeling and random simulation, and more specifically, relates to a random simulation method for time series data based on a greedy learning Gaussian mixture model. Background Art
[0002] Stochastic simulation is an important technical approach for power system planning and operation, and is widely used in stochastic modeling of time-series data such as runoff, wind speed, power output, and load. However, due to limitations in observation technology and the uneven spatial distribution of observation stations, historical observational data struggles to capture the long-term variability of real-world conditions, making it difficult to effectively meet the practical needs of power system management and optimization. Therefore, generating a large number of simulation sequences through efficient stochastic simulation methods for time-series data is crucial for addressing uncertain events that may occur in the future but have not been observed historically, and for enhancing the scientific nature of power system management.
[0003] Currently, existing stochastic simulation methods primarily include statistically driven methods, copula functions, and generative adversarial networks. Statistically driven methods offer simple structures and clear concepts, but they suffer from a strong reliance on linear and Gaussian distribution assumptions or an inability to extrapolate to extreme scenarios. While copula functions and generative adversarial networks have made progress in nonlinear representation, the former still requires a pre-defined marginal distribution and is difficult to model in high dimensions. The latter, due to its black-box nature, struggles to explicitly analyze the generated distribution. Against this backdrop, the Gaussian mixture model (GMM) has emerged as a new breakthrough in stochastic simulation methods due to its multimodal fitting capabilities and explicit probability structure. However, its parameter estimation relies on the traditional expectation-maximization (EM) algorithm, which suffers from limitations such as sensitivity to initial values, the need for a pre-defined number of components, and a tendency to fall into local optima, limiting its applicability to complex time series data. For GMM parameter optimization, existing techniques primarily focus on the latent variable handling mechanism of the EM algorithm. Variants such as classified EM, random EM, and Monte Carlo EM have been proposed, achieving improvements in computational efficiency, convergence speed, and model complexity. However, these variant algorithms each have their own focus and still face common problems such as sensitivity to initial values, local optimality, and accuracy-efficiency trade-offs, lacking systematic solutions. Therefore, it is necessary to propose a more accurate, reliable, and efficient GMM parameter optimization method that can properly capture the characteristics of time series data and further improve the accuracy of GMM in random simulations of time series data. Summary of the Invention
[0004] In response to the above defects or improvement needs of the prior art, the present invention provides a random simulation method for time series data based on a greedy learning Gaussian mixture model to solve the problem of low simulation accuracy of the existing random simulation method.
[0005] To achieve the above object, according to one aspect of the present invention, a method for constructing a Gaussian mixture model based on greedy learning is provided, the method comprising the following steps: S1 uses the data to be processed to construct a two-dimensional vector corresponding to each moment of the data to be processed, and constructs a two-dimensional vector corresponding to the two-dimensional vector at each moment. K Gaussian mixture model of Gaussian components and initialize the parameter set of the Gaussian mixture model; S2: Divide the two-dimensional vector into components corresponding to the Gaussian components. K disjoint first-generation subsets; for any first-generation subset A k , randomly select two data points in the subset, and use the two data points to transform the subset A k Divide into two disjoint second-generation subsets, calculate the parameter sets of each second-generation subset respectively, and use the second-generation subsets as candidate components; Screening the candidate components to obtain screened candidate components; Calculate the difference between any selected candidate components and the K Gaussian components form K+ The log-likelihood function of the Gaussian mixture model with 1 Gaussian component, the maximum value of the log-likelihood function corresponds to K+ 1 Gaussian mixture model as a new Gaussian mixture model, and updating the parameter set of the new Gaussian mixture model; S3 K = K + 1. Return to step S2 until the termination condition is met, and the Gaussian mixture model obtained is the optimal Gaussian mixture model.
[0006] Further preferably, the Gaussian mixture model is as follows: ; in, For multidimensional vectors x The probability density function of is a multidimensional Gaussian probability density function; is the set of parameters to be optimized for GMM, where It is the first k The mixture weights of Gaussian components, It is k The mean vector of the components, It is k The covariance matrix of the components; K is the total number of Gaussian components of the current model; D For multidimensional vectors x dimension.
[0007] Further preferably, the second-generation subset is as follows: ; in, 、 They are two disjoint second-generation subsets respectively; a 、 b are two randomly selected data point vectors; x Represents a subset of the generation A k The data points in .
[0008] Further preferably, the parameter set includes a mean vector, a covariance matrix and a mixing weight; the mean and covariance of the second-generation subset are calculated using the elements in the second-generation subset, and the mixing weight of the second-generation subset is 50% of the mixing weight of the first-generation subset.
[0009] Further preferably, the step of screening the candidate components is as follows: (1) Preliminary screening of candidate components; the condition for the preliminary screening is that the number of elements in the candidate components or the mixed weight of the candidate components is greater than or equal to a preset threshold; (2) Update the parameter set of the candidate components obtained by preliminary screening until the preset conditions are met; (3) Further screening the updated candidate components and retaining the candidate components that meet the screening conditions; the further screening condition is that the weight of the updated candidate component is greater than or equal to the preset component weight threshold.
[0010] Further preferably, the formula for updating the parameter set of the candidate components obtained by preliminary screening is as follows: ; in, It is the first K+ 1 Gaussian component of the mixture weight; represents the mixing weight of the candidate component; It is K+ 1-component mean vector, It is K+ 1-component covariance matrix; K+ 1 is the total number of Gaussian components in the model after adding candidate components; is the sample data point x i For the first K+ A latent variable with 1 component, representing x i Is it the first K+ 1 serving; is the posterior probability of the latent variable, indicating the K+ 1 component pair x i The responsiveness of the first generation subset A k The number of data points in ; Yuan K Probability density function of the component Gaussian mixture model.
[0011] Further preferably, in step (2), the preset conditions are as follows: ; in, Yuan K The Gaussian mixture model adds a A k The split S After candidate components, the parameter optimization iteration K+ Log-likelihood function of 1-component Gaussian mixture model; Update the convergence conditions for the preset parameters; represents the mixing weight of the candidate component; is the parameter set of the candidate component; Yuan K Probability density function of component Gaussian mixture model; is the probability density function of the Gaussian mixture model after adding candidate components.
[0012] Further preferably, in step S2, the formula for updating the parameter set of the new Gaussian mixture model is as follows: ; in, It is the first k The mixture weights of the Gaussian components; It is k The mean vector of the components; It is k The covariance matrix of the components; is the sample data point x i For the first k The latent variables of the components are expressed as x i Is it the first k Quantity is the posterior probability of the latent variable, indicating the k Component pairs x i responsiveness; M is the total number of sample data points.
[0013] According to another aspect of the present invention, a method for stochastic simulation of time series data based on greedy learning Gaussian mixture model is provided, the method comprising the following steps: The Gaussian mixture model is used to establish the probability density function and conditional cumulative distribution function of the two-dimensional vector of the data to be simulated; The conditional cumulative distribution function is used to calculate the simulated value at each moment of each year to obtain the random simulated data of the target year.
[0014] Further preferably, the probability density function and conditional cumulative distribution function are as follows: ; The said i Year t Simulated value at the moment as follows: ; in, x 1, x 2 refers to two-dimensional vectors x t The first and second column elements of ; p ( x 2 | x 1 )and F ( x 2 | x 1 ) are known x 1 hour x 2 The conditional probability density function and conditional cumulative distribution function of the Gaussian mixture model; It is the first k The mixture weights of the Gaussian components; are known x 1 The mean vector and covariance matrix of ; are known x 1 hour x 2 The conditional mean vector and conditional covariance matrix of ; is a uniform random number.
[0015] In general, the above technical solutions conceived by the present invention have the following beneficial effects compared with the prior art: 1. The present invention gradually adds Gaussian components to the Gaussian mixture model. This method introduces the greedy learning strategy into the parameter optimization process of the Gaussian mixture model. Compared with the innovative design of split candidate component generation mechanism, two-stage optimization and double threshold constraint, the present invention ensures the simplicity of the model (the optimal number of components is 100%). K Lower) improves the ability to fit the data, achieves the coordinated optimization of computational efficiency and model robustness, and improves the simulation accuracy of random simulation.
[0016] 2. The greedy learning Gaussian mixture model employed in this invention accurately describes both the intra-annual and inter-annual autocorrelation characteristics of time series, while ensuring ease of model construction and optimization. Its simulated data maintains the statistical characteristics of the observed time series (mean, standard deviation, skewness coefficient, etc.) with higher simulation accuracy than comparable methods. Furthermore, the simulated data of this invention comprehensively considers coverage and interval width, achieving precise coverage of the core data distribution at varying confidence levels. Compared to comparable methods, the interval width is minimized, thus avoiding the risk of overgeneralization and ensuring superior uncertainty quantification.
[0017] 3. The present invention does not require a preset data distribution or rely on linear assumptions. Its explicit probability structure enhances the interpretability of the model and improves the accuracy and reliability of the present invention's random simulation of time series data. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flowchart of a method for stochastic simulation of time series data based on greedy learning Gaussian mixture model constructed according to a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0020] like Figure 1 As shown, the present invention proposes a random simulation method for time series data based on the greedy learning Gaussian mixture model (GL-GMM), and the specific steps are as follows: S1: Obtain a time series dataset of historical observations and input the preset parameters of the greedy learning algorithm.
[0021] Get a time series dataset of historical observations ,in M is the dataset size, N is the length of time. t- 1 moment data Q t-1 Hedi t Time data Q t Two-dimensional vector x t .variable X The expression is:
[0022] Where: Q t Represents the historical observation time series data sequence t A vector of time; if the time length is 1 day, then N= 365; if the time length is 10 days, then N= 36; If the time length is one month, then N= 12; Two-dimensional vector represents the dimension of sample data D= 2.
[0023] Enter the preset parameters of the greedy learning algorithm and set the maximum number of components , the number of candidate components to be split for each existing component S , sample size threshold , weight threshold , stop iteration threshold , regularization coefficient .
[0024] Construct a 1-component Gaussian mixture model based on the observed data and initialize the model parameters.
[0025] GMM is a composite probability model. K The linear superposition of a single Gaussian distribution is mathematically described as follows:
[0026] in, For multidimensional vectors x The probability density function of is a multidimensional Gaussian probability density function; is the set of parameters to be optimized for GMM, where It is the first k The mixture weights of Gaussian components, It is k The mean vector of the components, It is k The covariance matrix of the components; K is the total number of Gaussian components of the current model; D For multidimensional vectors x dimension.
[0027] Based on the variables described in S1 X , for every moment t Two-dimensional vector x t Construct a Gaussian mixture model with one component, that is K = 1; Initialize the mixing weights of its components as =1; mean vector of components is a two-dimensional vector x t The mean of the components is 2D vector x t The covariance of plus the regularization term; the above constitutes the initial parameter set of the model θ .
[0028] According to the initial parameter set θ , calculate the log-likelihood function of the initial model L 1;
[0029] S3: Perform a global parameter space search on each component of the existing Gaussian mixture model and use a split candidate component generation mechanism to obtain the expected number of candidate components.
[0030] S31: Soft segmentation based on posterior probability, the two-dimensional vector x t Divided into K disjoint subsets , where each subset A k Corresponding to a Gaussian component in the existing mixture model. ;
[0031] Where: is the current number of Gaussian components; For the i Sample data points x i For the first k The hidden variables of the components indicate the selection of each sub-component; is the posterior probability of the latent variable.
[0032] S32: The split candidate component generation mechanism is A k Select two data points uniformly randomly from the global range a and b , based on the Euclidean distance metric A k Divide into two disjoint subsets and ;
[0033] S33: Subset and The two new candidate components are used to calculate the mean vector and covariance matrix of the elements in the subset as the parameters of the candidate components, and inherit 50% of the parent component mixing weight as the initial weight.
[0034] S34: Repeat S32 and S33 to obtain the next two candidate components until each subset A k Generate the expected S Group candidate components; obtain K × S candidate components.
[0035] S4: Perform double threshold constraints and local optimization on the candidate components of S3 to optimize and update the parameters of the candidate components.
[0036] right K × S The candidate components are preliminarily screened and constrained by the component sample size threshold and weight threshold. When the number of samples belonging to the candidate component or the candidate component weight is lower than the double threshold constraint, the component is eliminated.
[0037] EM local optimization is performed on each candidate component that has been preliminarily screened. The initial parameters of EM local optimization are the mean vector and covariance matrix of each candidate component. The parameters of each component of the existing Gaussian mixture model are fixed, and only the mixing weights of the candidate components are optimized. and parameter sets , that is, simplifying the multi-component parameter optimization problem into a two-component parameter optimization problem. ; ; Where: Parameter set of the optimized candidate component , For the k The first split of the set S candidate subsets, for A k The number of data points in .
[0038] The iteration stop condition is set to ;
[0039] Where: For existing K The probability density function of the component mixture model, is the probability density function of the newly added component, is the probability density function of the mixture model after adding the new component, is the probability density function of the mixture model after adding the new component. To express the In the next iteration, add k The first S The log-likelihood function of the mixture model after the candidate components.
[0040] The candidate components after parameter optimization are further screened and constrained by the component weight threshold. If the candidate component weight is lower than the threshold, it will be eliminated.
[0041] S5: Optimizing the candidate components after the parameter optimization described in S4, adding the selected components to the existing Gaussian mixture model described in S2, and performing global optimization to update the parameters of the Gaussian mixture model.
[0042] Based on the maximum likelihood principle, the candidate component that maximizes the log-likelihood function of the mixed model is selected from the candidate components optimized and screened in step S4 as a new one. K The component mixture models together form a new K+ 1-component mixed model.
[0043] The weights and parameter sets of the hybrid model after the newly added components are used as initial values, and a global EM optimization is performed to complete the parameter update;
[0044] S6: Repeat S3-S5 until the termination condition of the greedy learning algorithm is met to obtain the final optimized Gaussian mixture model.
[0045] Specifically, repeat steps S3-S5 until the number of components reaches the preset maximum number of components. K max Or the log-likelihood function satisfies L K >L K+1 The termination condition of the model is the ideal number of components. K *, the corresponding log-likelihood function is the global maximum.
[0046] Output model ideal number of components K *, the mixture weight set and parameter set of each component of the model to obtain the final optimized Gaussian mixture model.
[0047] S7: Based on the optimized Gaussian mixture model described in S6, perform random simulation of the time series data to obtain a simulation sequence.
[0048] S71: Based on the final optimized Gaussian mixture model described in S6, construct a two-dimensional vector x t = [x 1 , x 2 ]’s conditional probability relationship:
[0049] Where: x 1 , x 2 Respectively refer to two-dimensional vectors x t The first and second column elements of ; p ( x 1 , x 2 ) represents a vector x 1 With vector x 2 The probability density function of F ( x t ) represents a vector x 1 With vector x 2 The two-dimensional joint distribution of represents the bivariate Gaussian probability density function.
[0050] S72: According to the total probability formula and conditional probability formula, deduce the known x 1 hour x 2 The conditional probability density function of the Gaussian mixture distribution is p ( x 2 | x 1 ) and the conditional cumulative distribution function F ( x 2 | x 1 ):
[0051] Where: p ( x 1 )for x 1 The marginal probability density function of Respectively x 1 The mean vector and covariance matrix of ; are known x 1 hour x2 The conditional mean vector and conditional covariance matrix of are calculated as follows:
[0052] S73: When performing the first year of random simulation, first let the simulated runoff (in M is the number of years of sample data, For the i Year N observed runoff at each moment), generate a uniform random number and assign it to the conditional CDF, i.e. , thereby further obtaining the first year t Simulated value at the moment , t = t + 1. Repeat the above process until t = N , then the first year of random simulation is completed.
[0053] S74: In progress i When simulating the randomness of years, we first generate a uniform random number and assign it to the conditional CDF, i.e. , thereby further obtaining the i Year t Simulated value at the moment, t = t +1. It should be noted that Repeat the above process until t = N , then complete the i Random simulation of the year.
[0054] S75: Order i = i + 1. Repeat step S74 until , then complete stochastic simulation of the year (where is the target number of years to be generated).
[0055] The present invention will be further described below with reference to specific embodiments.
[0056] The monthly, ten-day and daily historical runoff time series data of Pingshan Hydrological Station in the lower reaches of Jinsha River from 1940 to 2022 were selected as the implementation objects to test the simulation effect of the random simulation method of time series data based on greedy learning Gaussian mixture model.
[0057] The present invention, the traditional EM algorithm-optimized GMM method, the Copula method, the seasonal autoregressive method (SAR), and the conditional generative adversarial network method (CGAN) were used to randomly simulate and generate 1000-year monthly, ten-day, and daily runoff series, and a comparative analysis was performed. The randomly simulated runoff series required high performance in terms of preserving statistical features, characterizing autocorrelation structures, and quantifying uncertainty. The embodiment used metrics such as mean, variance, skewness coefficient, root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE), autocorrelation coefficient (ACF), coverage ratio (CR), and interval width (IW) to measure the quality of the generated series. ; ; .
[0058] Tables 1, 2, and 3 show the quantitative evaluation results of the statistical characteristics of the random simulation sequences generated by each method at the monthly, ten-day, and daily scales, respectively. The results show that the GL-GMM method shows significant advantages in maintaining the statistical characteristics of runoff sequences at different time scales. At the monthly and ten-day scales, GL-GMM achieved optimal levels of RMSE, MAE, and MAPE for the three statistical indicators of mean, variance, and skewness coefficient, among which the error reduction of the skewness coefficient was particularly prominent. However, in terms of maintaining the mean runoff at the daily scale, although the MAPE of GL-GMM (0.87%) was slightly inferior to that of the Copula method (0.54%), the difference was still within the high-precision range and did not affect its overall advantages in the other three indicators. The bold numbers in Tables 1, 2, and 3 indicate that the method has the smallest error; .
[0059] Table 4 shows the average autocorrelation coefficients and their normalized values for runoff series generated by different methods. The data show that as the time scale decreases from month to day, the normalized ACF values of the runoff series generated by each method show a significant increasing trend, indicating that the short-term runoff process has stronger time-related characteristics. Among them, the ACF values of the GL-GMM method at all three scales reached the highest level (monthly 0.6330, ten-day 0.8868, daily 0.9898), verifying its superior performance in maintaining the time series dependency structure. The bold numbers in Table 4 indicate that the method has the largest ACF value; .
[0060] Table 5 shows the accuracy of runoff series generated by different methods, specifically the evaluation of coverage ratio (CR) and interval width (IW). The results show that with increasing confidence levels, the CR and IW of the series generated by the GL-GMM method exhibit a synchronous and uniform growth trend. This indicates that the generated series can steadily converge to a reasonable probability distribution at the corresponding confidence level as the probability constraints are strengthened, avoiding the anomalies of other methods (such as Copula and SAR) such as stagnant CR growth or sudden increases in IW at specific confidence levels. While achieving more accurate coverage, the GL-GMM method uses stricter intervals to represent the true distribution of the runoff series, avoiding the loss of accuracy caused by excessive interval expansion. Compared to traditional methods, it demonstrates superior uncertainty modeling advantages.
[0061] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for constructing a Gaussian mixture model based on greedy learning, characterized in that: The method comprises the following steps: S1 uses the data to be processed to construct a two-dimensional vector corresponding to each moment of the data to be processed, and constructs a two-dimensional vector corresponding to the two-dimensional vector at each moment. K Gaussian mixture model of Gaussian components and initialize the parameter set of the Gaussian mixture model; S2: Divide the two-dimensional vector into components corresponding to the Gaussian components. K disjoint first-generation subsets; for any first-generation subset A k , randomly select two data points in the subset, and use the two data points to transform the subset A k Divide into two disjoint second-generation subsets, calculate the parameter sets of each second-generation subset respectively, and use the second-generation subsets as candidate components; Screening the candidate components to obtain screened candidate components; Calculate the difference between any selected candidate components and the K Gaussian components form K+ The log-likelihood function of the Gaussian mixture model with 1 Gaussian component, the maximum value of the log-likelihood function corresponds to K+ 1 Gaussian mixture model as a new Gaussian mixture model, and updating the parameter set of the new Gaussian mixture model; S3 K=K+ 1. Return to step S2 until the termination condition is met, and the Gaussian mixture model obtained is the optimal Gaussian mixture model.
2. A method for constructing a Gaussian mixture model based on greedy learning according to claim 1, characterized in that: The Gaussian mixture model is as follows: ; in, For multidimensional vectors x The probability density function of is a multidimensional Gaussian probability density function; is the set of parameters to be optimized for GMM, where It is the first k The mixture weights of Gaussian components, It is k The mean vector of the components, It is k The covariance matrix of the components; K is the total number of Gaussian components of the current model; D For multidimensional vectors x dimension.
3. A method for constructing a Gaussian mixture model based on greedy learning according to claim 1 or 2, characterized in that: The second generation subsets are as follows: ; in, 、 They are two disjoint second-generation subsets respectively; a 、 b are two randomly selected data point vectors; x Represents a subset of the generation A k The data points in .
4. A method for constructing a Gaussian mixture model based on greedy learning according to claim 3, characterized in that: The parameter set includes a mean vector, a covariance matrix and a mixing weight; the mean and covariance of the second-generation subset are calculated using the elements in the second-generation subset, and the mixing weight of the second-generation subset is 50% of the mixing weight of the first-generation subset.
5. The method for constructing a Gaussian mixture model based on greedy learning according to claim 1, wherein: The steps of screening the candidate components are as follows: (1) Preliminary screening of candidate components; the condition for the preliminary screening is that the number of elements in the candidate components or the mixed weight of the candidate components is greater than or equal to a preset threshold; (2) Update the parameter set of the candidate components obtained by preliminary screening until the preset conditions are met; (3) Further screening the updated candidate components and retaining the candidate components that meet the screening conditions; the further screening condition is that the weight of the updated candidate component is greater than or equal to the preset component weight threshold.
6. A method for constructing a Gaussian mixture model based on greedy learning according to claim 5, characterized in that: The formula for updating the parameter set of candidate components obtained by preliminary screening is as follows: ; in, It is the first K+ 1 Gaussian component of the mixture weight; represents the mixing weight of the candidate component; It is K+ 1-component mean vector, It is K+ 1-component covariance matrix; K+ 1 is the total number of Gaussian components in the model after adding candidate components; is the sample data point x i For the first K+ A latent variable with 1 component, representing x i Is it the first K+ 1 serving; is the posterior probability of the latent variable, indicating the K+ 1 component pair x i The responsiveness of the first generation subset A k The number of data points in ; Yuan K Probability density function of the component Gaussian mixture model.
7. A method for constructing a Gaussian mixture model based on greedy learning according to claim 5 or 6, characterized in that: In step (2), the preset conditions are as follows: ; in, Yuan K The Gaussian mixture model adds a A k The split S After candidate components, the parameter optimization iteration K+ Log-likelihood function of 1-component Gaussian mixture model; Update the convergence conditions for the preset parameters; represents the mixing weight of the candidate component; is the parameter set of the candidate component; Yuan K Probability density function of component Gaussian mixture model; is the probability density function of the Gaussian mixture model after adding candidate components.
8. The method for constructing a Gaussian mixture model based on greedy learning according to claim 1, wherein: In step S2, the formula for updating the parameter set of the new Gaussian mixture model is as follows: ; in, It is the first k The mixture weights of the Gaussian components; It is k The mean vector of the components; It is k The covariance matrix of the components; is the sample data point x i For the first k The latent variables of the components are expressed as x i Is it the first k Quantity is the posterior probability of the latent variable, indicating the k Component pairs x i responsiveness; M is the total number of sample data points.
9. A random simulation method for time series data based on greedy learning Gaussian mixture model, characterized in that: The method comprises the following steps: Using the Gaussian mixture model described in any one of claims 1 to 8 to establish a probability density function and a conditional cumulative distribution function of the two-dimensional vector of the data to be simulated; The conditional cumulative distribution function is used to calculate the simulated value at each moment of each year to obtain the random simulated data of the target year.
10. The method for stochastic simulation of time series data based on greedy learning Gaussian mixture model according to claim 9, characterized in that: The probability density function and conditional cumulative distribution function are as follows: ; The said i Year t Simulated value at the moment as follows: ; in, x 1, x 2 refers to two-dimensional vectors x t The first and second column elements of ; p ( x 2 | x 1 )and F ( x 2 | x 1 ) are known x 1 hour x 2 The conditional probability density function and conditional cumulative distribution function of the Gaussian mixture model; It is the first k The mixture weights of the Gaussian components; are known x 1 The mean vector and covariance matrix of ; are known x 1 hour x 2 The conditional mean vector and conditional covariance matrix of ; is a uniform random number.
Citation Information
Patent Citations
Time sequence construction method and system for output of multiple wind power plants
CN111027790A
Runoff stochastic simulation method and system based on Gaussian mixture model
CN113191561A
Gaussian mixture model clustering machine learning method under condition of missing features
WO2022179241A1