A new energy nonlinear correlation data enhancement method for limited observation

CN122620490BActive Publication Date: 2026-09-22SICHUAN AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611096583.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-23
Publication Date
2026-09-22
Estimated Expiration
2046-07-23

AI Technical Summary

Technical Problem

[0007]本发明的目的在于提供一种面向有限观测的新能源非线性相关数据增强方法,解决现有技术在有限观测场景下,光伏功率建模参数估计偏差大、模型易受异常样本干扰、误差耦合放大引发潮流分析失真的问题

Benefits of technology

1、本发明采用非参数建模思路,经对数变换、核密度估计、D-Vine Copula模型复现光伏功率多维非线性相关特征,搭配插值微扰动逆变换消除样本分层、指数变换保障功率物理约束,将样本集的分布距离显著压缩,能够在有限信息条件下有效提升生成稳定性与分布连续性,同时避免传统后验截断约束所导致的分布失真与相关结构破坏问题,有效缓解有限信息条件下分布估计不稳定的问题,实现对原始数据分布的平滑重构与扩展。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122620490B_ABST
    Figure CN122620490B_ABST
Patent Text Reader

Abstract

The application discloses a new energy nonlinear related data enhancement method facing limited observation, and belongs to the field of interval probability power flow estimation, and comprises the following steps: preprocessing and logarithmic space mapping, continuous edge distribution modeling, probability integral transformation and boundary clipping, D-Vine Copula model hierarchical fitting, probability space layer-by-layer conditional sampling, interpolation micro-disturbance composite inverse mapping reconstruction, exponential inverse transformation to restore original power space, and Euclidean distribution distance distribution consistency check.The application has the beneficial effects that: a non-parametric modeling idea is adopted, logarithmic transformation, kernel density estimation, D-Vine Copula model are used to reproduce the multi-dimensional nonlinear correlation characteristics of photovoltaic power, interpolation micro-disturbance inverse transformation is used to eliminate sample stratification, and exponential transformation is used to guarantee power physical constraints, so that the generation stability and distribution continuity can be effectively improved under the condition of limited information, and the problem of unstable distribution estimation under the condition of limited information can be effectively relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interval probabilistic power flow estimation, and in particular to a method for enhancing nonlinear correlation data of new energy sources with limited observations. Background Technology

[0002] In large-scale photovoltaic grid-connected systems, when the statistical data on photovoltaic power is limited, how to effectively model the uncertain photovoltaic power, propagate this uncertainty and reflect it in the output variables of the power flow, and ultimately establish a theoretical framework for power flow quantitative analysis that takes into account the correlation of photovoltaic power rate under incomplete information has become a key problem that urgently needs to be solved in the field of uncertainty quantitative analysis of power system operating points.

[0003] Current mainstream modeling techniques include a fusion framework of ellipsoidal theory and evidence theory, D-Vine Copula, coupled models of evidence theory, and an integration method of partitioning around medoids (a clustering approach) and evidence theory. These approaches can characterize the nonlinear correlations between photovoltaic power outputs, but their application typically requires relatively complete probabilistic statistical information or sufficient historical sample data. In practical engineering, due to factors such as short photovoltaic farm commissioning times, missing measurement data, communication failures, and data recording interruptions caused by severe weather conditions, available photovoltaic power statistics are often highly incomplete or even extremely scarce. Conventional modeling frameworks will face identification ambiguities due to insufficient prior information, leading to overly conservative or distorted quantification results of uncertainties.

[0004] Currently, there are no similar technical solutions for enhancing nonlinear correlation data of new energy sources with limited observations. When information is limited, the confidence interval for parameter estimation of the marginal distribution of photovoltaic power expands significantly. The construction of basic probability assignments in evidence theory lacks objective basis, and relying on subjective assignment will introduce additional artificial uncertainty, resulting in the focal element structure being too coarse or too fine, making it difficult to balance the model's expressive accuracy and computational complexity.

[0005] A few abnormal observations can easily lead to misjudgment of the tail characteristics of the marginal distribution, while the focal element boundary is extremely sensitive to power sample perturbations, resulting in improper aggregation or dispersion of probability quality, causing a pseudo-amplification of cognitive uncertainty and the masking of real uncertainty.

[0006] When using multi-method fusion evidence theory for uncertainty propagation, errors in each stage are amplified through nonlinear coupling, leading to an excessive expansion of the confidence intervals of output quantities such as node voltage and branch power. This masks the true operational risks and causes conservative redundancy or risk underestimation in scheduling decisions. Summary of the Invention

[0007] The purpose of this invention is to provide a new energy nonlinear correlation data enhancement method for limited observation, which solves the problems of large deviation in photovoltaic power modeling parameter estimation, model susceptibility to abnormal sample interference, and power flow analysis distortion caused by error coupling amplification in the case of limited observation.

[0008] The objective of this invention is achieved through the following technical solution: A method for enhancing nonlinear correlation data of new energy sources with limited observations includes the following steps: S1. The acquired raw photovoltaic power data is preprocessed and logarithmic transformation constant is combined to complete the logarithmic space mapping, transforming the physically constrained photovoltaic power data into a continuous unconstrained real number space to obtain logarithmic space power data. S2. For the power data corresponding to each photovoltaic power station in the logarithmic space, the stable kernel density estimation method is used to construct the marginal probability density function of each dimension, and the continuous cumulative distribution function of each dimension is obtained by integral operation. S3. Based on the cumulative distribution function, perform probability integral transformation to map the logarithmic space power data to a unified probability space, and perform linear smoothing correction and boundary clipping on the transformed data to obtain probability space samples. S4. Based on the probability space samples, construct a D-Vine Copula model, decompose the joint dependency of three-dimensional photovoltaic power into two-level binary pairwise connection functions, and use Gaussian connection functions to fit the correlation structure between variables in each level, and solve to obtain all parameters of the D-Vine Copula model. S5. Based on the fitted D-Vine Copula model and parameters, perform layer-by-layer conditional sampling in the probability space to generate new probability space samples that retain the original multivariate correlation characteristics and conditional dependencies. S6. Perform edge inverse mapping reconstruction on the new probability space samples. The inverse mapping reconstruction includes inverse cumulative distribution transformation, combined with continuous interpolation and adaptive Gaussian micro-perturbation, to obtain enhanced data in logarithmic space. S7. Perform an inverse exponential transformation on the enhanced data to restore it to the original data space, and obtain a photovoltaic power enhancement dataset that satisfies the non-negative physical constraints. S8. Using Euclidean distance, perform consistency verification between the photovoltaic power enhancement dataset obtained in step S7 and the original benchmark dataset, and output the enhanced photovoltaic power data.

[0009] Furthermore, in step S1, preprocessing includes outlier checking, data normalization, and numerical stabilization; the expression for the logarithmic space mapping is:

[0010] In the formula, Indicates the first The first sample j Original photovoltaic power data, Represents the logarithmic transformation constant. Describe the th in the logarithmic space Dimensional variables.

[0011] Furthermore, in step S2, the stable kernel density estimation method uses a Gaussian kernel function to perform nonparametric kernel density estimation, and the bandwidth parameter of the kernel density estimation is determined based on the standard deviation of the logarithmic space samples in each dimension.

[0012] Furthermore, in step S3, the expressions for linear smoothing correction and boundary clipping are as follows:

[0013] In the formula, This represents the initial data obtained after the probability integral transformation. This represents the probability space sample obtained after linear smoothing correction and boundary clipping. This represents the global boundary clipping constant, whose magnitude is greater than that of the logarithmic transformation constant.

[0014] Furthermore, in step S4, the first level of the two-level binary pairwise connection function is the marginal pairwise connection function, which represents the direct dependency relationship between each adjacent dimension, and the second level is the conditional pairwise connection function, which represents the conditional residual dependency relationship between the remaining two dimensions after fixing the second dimension variable. The Gaussian link function maps uniform marginal variables to a normal space through a standard normal quantile transformation, and then quantifies the dependence strength between variables with a linear correlation coefficient. The Pearson correlation coefficient obtained in the normal space is used as the parameter of the Gaussian link function. When fitting the conditional pairwise link function, the marginal uniform variables of the corresponding dimension are selected as inputs, skipping the explicit conditional transformation step, and the approximate correlation coefficient used to characterize the residual dependence is calculated in the normal space.

[0015] Furthermore, in step S5, the conditional sampling layer by layer defines the sampling interval based on the global boundary pruning constant. The intermediate variable is extracted as the root node of the conditional sampling within the uniform distribution corresponding to the sampling interval. The root node variable is first mapped to the normal space, and then the conditional normal sample is generated layer by layer by combining the Gaussian connection function and introducing independent standard normal random quantities. Subsequently, the conditional normal sample is mapped back to the uniform space to obtain new uniform variables in each dimension. Finally, the new uniform variables in each dimension are spliced ​​together to form a three-dimensional uniform sample vector.

[0016] Furthermore, in step S6, when performing the inverse cumulative distribution transformation, a single-dimensional uniform variable is extracted from the three-dimensional uniform sample vector in the new probability space, its continuous position in the sorted sequence is calculated, the left and right neighborhood indices and interpolation weights are determined, and linear interpolation is performed using the adjacent order statistics to obtain the estimated value of the variable mapped back to the logarithmic space. Then, Gaussian random perturbation is superimposed to complete the adaptive micro-perturbation processing.

[0017] Furthermore, in step S7, the expression for the inverse exponential transform is:

[0018] In the formula, This indicates the first step after applying a Gaussian perturbation. i The first sample j Data is generated in a logarithmic space of dimension. This represents the photovoltaic power value after exponential restoration.

[0019] The present invention has the following advantages: 1. This invention adopts a nonparametric modeling approach, using logarithmic transformation, kernel density estimation, and D-Vine Copula model to reproduce the multidimensional nonlinear correlation characteristics of photovoltaic power. Combined with inverse interpolation micro-perturbation transformation to eliminate sample stratification and exponential transformation to ensure power physical constraints, the distribution distance of the sample set is significantly compressed. This can effectively improve the generation stability and distribution continuity under limited information conditions, while avoiding the distribution distortion and related structure destruction problems caused by traditional posterior truncation constraints. It effectively alleviates the problem of unstable distribution estimation under limited information conditions and achieves smooth reconstruction and expansion of the original data distribution.

[0020] 2. By mapping the restricted non-negative photovoltaic power to the global real domain through logarithmic transformation, and constructing a continuous smooth edge cumulative distribution with adaptive bandwidth Gaussian kernel density estimation, the traditional discrete step-type empirical distribution is abandoned. At the same time, a simplified fitting strategy is adopted for the conditional pairwise connection function, omitting complex explicit conditional transformations, which greatly suppresses the error amplification in the finite sample iteration process. It can accurately decompose and replicate the multi-layer nonlinear coupling dependence relationship between three-dimensional photovoltaic power.

[0021] 3. Data restoration is achieved by using a reversible exponential transformation with offset compensation, which is strictly mathematically inverse of the preceding logarithmic transformation. Relying on the inherently positive property of the exponential function, the generation of illegal samples with negative power is prevented from the mathematical mechanism. There is no need to add additional post-processing such as outlier filtering, iterative truncation, and secondary correction, thus simplifying the entire data augmentation project implementation process. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the process of the present invention.

[0023] Figure 2The original photovoltaic power scatter plots of the three photovoltaic power plants selected for the example are shown below.

[0024] Figure 3 This is a scatter plot of different numbers of photovoltaic power samples taken from a baseline distribution according to the present invention.

[0025] Figure 4 This is a scatter plot of photovoltaic power for different original sample numbers after enhancement using this method. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0027] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0028] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other.

[0029] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0030] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0031] refer to Figure 1 As shown, one embodiment of the present invention is as follows: A data augmentation method for nonlinear correlations in new energy sources, oriented towards limited observations, is proposed. This method maps the non-negative right-skewed data from photovoltaic power plants into a logarithmic space. In the transformed logarithmic space, a stable kernel density estimation method is first used to nonparametrically model the marginal distributions of each dimension, and a smoothed cumulative distribution function is used to achieve a probabilistic mapping of the data to the connection function space. Subsequently, a D-Vine Copula structure is used to decompose and model the nonlinear dependencies among multiple variables, and a Gaussian pairwise connection function is used to achieve conditional correlation structure learning and layer-by-layer joint sampling. Finally, the original data space is restored through continuous inverse cumulative distribution function mapping and inverse exponential transformation, thereby generating new photovoltaic power sample data that satisfies marginal statistical characteristics, nonlinear correlation structure, and strict non-negativity constraints. The method is validated using a constructed Euclidean distribution distance to form a closed-loop feedback loop, outputting the final augmented dataset. Thus, this scheme achieves synergistic optimization between maintaining the statistical characteristics of marginal distributions, accurately characterizing multidimensional correlation structures, eliminating the discrete layer perception of synthetic samples, and strictly satisfying non-negativity physical constraints. It can effectively improve the generation stability and distribution continuity under limited information conditions, while avoiding the distribution distortion and correlation structure destruction problems caused by traditional posterior truncation constraints. It effectively alleviates the problem of unstable distribution estimation under limited information conditions, and achieves smooth reconstruction and expansion of the original data distribution.

[0032] This embodiment assumes that the power output of the photovoltaic power plant is an uncertain renewable energy source, and takes three-dimensional photovoltaic power data as the research object. It assumes that the original finite photovoltaic power dataset is... The original data baseline sample consisted of 5000 groups, which were considered the baseline. Figure 2 As shown. Data augmentation experiments were conducted using four finite sample groups of 30, 50, 100, and 300, as illustrated. Figure 3 As shown. After data augmentation using the method of this application, the limited observations of photovoltaic power were expanded to 5000 sets of photovoltaic power samples to study its accuracy. The results are as follows. Figure 4 As shown. In the following discussion, each variable in one dimension represents the photovoltaic power data of a photovoltaic power plant.

[0033] The specific steps are as follows: S1. Data Preprocessing and Logarithmic Space Mapping: The acquired raw photovoltaic power data is preprocessed and logarithmic space mapping is completed by combining the logarithmic transformation constant. The photovoltaic power data with physical constraints is transformed into a continuous unconstrained real number space to obtain logarithmic space power data, thereby reducing the impact of skewed distribution and extreme values ​​on the subsequent probability modeling process and providing a basis for the subsequent generated results to meet the physical rationality constraints.

[0034] In step S1, preprocessing includes outlier detection, data normalization, and numerical stabilization. Outlier detection is used to remove invalid samples that are out of range, have abrupt changes, or have missing data. Numerical normalization is used to eliminate differences in dimensions. Numerical stabilization preprocessing is used to avoid numerical singularities caused by subsequent logarithmic operations approaching zero.

[0035] Since photovoltaic power data is non-negativity, direct kernel density estimation can easily introduce boundary bias near zero. This embodiment uses logarithmic transformation to map the data to an approximately symmetric real space, thereby improving the accuracy and numerical stability of kernel density estimation. The expression for the logarithmic space mapping is as follows:

[0036] In the formula, Indicates the first The first sample Original photovoltaic power data, Represents the logarithmic transformation constant. Describe the th in the logarithmic space Dimensional variables.

[0037] In this embodiment, This invention employs a hierarchical numerical stabilization strategy: a minimal positive constant is introduced during the logarithmic transformation stage. This allows for the maximum preservation of the original data's precision information during the logarithmic mapping stage, avoiding the introduction of excessively large systematic offsets. The minimal positive constant is much smaller than the global boundary pruning constant. The functional separation and magnitude differentiation design of the two-level constants achieve hierarchical optimization that maintains precision and ensures numerical stability.

[0038] S2. Continuous marginal distribution modeling: For the power data corresponding to each photovoltaic power station in the logarithmic space, the stable kernel density estimation method is used to construct the marginal probability density function of each dimension. This enables the model to learn the statistical characteristics of the data with limited information without relying on prior distribution assumptions. The continuous cumulative distribution function of each dimension is obtained through integral operation, which provides the marginal distribution basis for subsequent joint probability modeling.

[0039] The stable kernel density estimation method uses a Gaussian kernel function for nonparametric kernel density estimation. The bandwidth parameter for kernel density estimation is determined based on the standard deviation of the logarithmic space samples in each dimension. The probability density function expression of the Gaussian kernel is as follows:

[0040] In the formula, Indicates the first Estimates of the probability density function of a dimensional variable in logarithmic space; This represents any point in the logarithmic space (the position to be estimated). Indicates the first The bandwidth parameter for kernel density estimation. , Indicates the first The standard deviation of the logarithmic space samples is calculated, and the bandwidth is adaptively adjusted according to the number of samples and the data dispersion, making it suitable for scenarios with limited samples. Let represent the Gaussian kernel function, where To standardize the distance, .

[0041] Based on the above The kernel density estimation results are used to construct the cumulative distribution function for each dimension of the variable by integration:

[0042] In the formula, Indicates the first dimensional cumulative distribution function, The cumulative distribution function represents the standard normal distribution.

[0043] S3. Probability Integral Transformation and Boundary Clipping: Based on the cumulative distribution function, a probability integral transformation is performed to map the logarithmic space power data to a unified probability space. Using the cumulative distribution mapping corresponding to the power data of each photovoltaic power station, the original photovoltaic power data is converted into probability variables following a unified probability distribution, thereby eliminating differences in the dimensions and distribution forms of different variables and providing a unified expression space for multivariate correlation structure learning. Specifically, the logarithmic space power data is mapped to a unit hypercube, as shown in the following expression:

[0044] Further linear smoothing and boundary clipping are performed on the transformed data to obtain the probability space samples, as shown in the following expression:

[0045] In the formula, This represents the initial data obtained after the probability integral transformation; This represents the probability space sample obtained after linear smoothing correction and boundary clipping. This represents the global boundary clipping constant. The magnitude is larger than that of the logarithmic transformation constant, so a relatively large magnitude is used. To ensure the numerical safety of the standard inverse normal transform, functional separation and magnitude-differentiated design of the two-level constants achieve hierarchical optimization that preserves accuracy and stabilizes numerical values.

[0046] The aforementioned linear smoothing correction and boundary clipping are key stability designs for photovoltaic power data enhancement scenarios: standard normal inverse cumulative distribution function. The presence of vertical asymptotes near 0 and 1 means that directly mapping the cumulative probability of edge samples to the vicinity of the boundary will result in extreme outliers in the subsequent inverse transformation, undermining the physical plausibility of the synthesized data. This invention achieves an optimal engineering balance between numerical stability and distribution fidelity by linearly shrinking all samples away from the singular boundary while maintaining the relative ordering of samples.

[0047] It should be noted that traditional connection function methods typically use empirical cumulative distribution functions (step functions based on sample ranking) for probability integral transformation. While this approach is computationally simple, it is essentially a discrete mapping, which can easily introduce a step-like pseudo-structure with a limited number of samples, and there is hard truncation at the boundaries, causing the samples to only take values ​​at a finite number of discrete points during the subsequent inverse transformation. This invention uses kernel density estimation analytical integration to obtain a continuous cumulative distribution function. Its output is a smooth continuous value, which fundamentally avoids the discrete ladder defect of empirical distribution, and provides a continuous basis for the subsequent construction of D-Vine Copula model and continuous interpolation inverse transformation, significantly improving the naturalness of the synthetic samples and the distribution approximation accuracy.

[0048] S4. D-Vine Copula Model Construction and Parameter Fitting: To characterize the nonlinear correlation structure between different dimensions of the logarithmic space, a D-Vine Copula model is further constructed after obtaining the probability space samples. The D-Vine Copula model decouples the joint distribution of multivariate random variables from their marginal distributions, thereby independently modeling the dependency structure between variables without changing the nonparametric features of each dimension's marginal distributions, achieving a stable expression of the multidimensional nonlinear correlation structure. Specifically, in this embodiment, the D-Vine Copula model will... The joint dependency of three-dimensional photovoltaic power is decomposed into two-level binary pairwise connection functions to reduce the complexity of high-dimensional parameter estimation and improve numerical stability.

[0049] Specifically, record the first The uniform space vector of each sample is The D-Vine Copula model approximates the three-dimensional joint distribution through the following two-level sequence of pairwise connection functions: The first level consists of edge pairwise join functions: containing 、 , To characterize the interdependence between the first and second dimensions, Describe the interdependence between the second and third dimensions.

[0050] The second level consists of conditional pairwise join functions: containing , Characterizes the residual dependence between the first and third dimensions given the second dimension.

[0051] For both levels of pairwise connection functions, a Gaussian connection function is used to fit the correlation structure between variables at each level. The Gaussian connection function maps uniform marginal variables to a normal space through standard normal quantile transformation, and then quantifies the dependence strength between variables using linear correlation coefficients. Taking the first level pairwise connection function as an example... For example, the fitting process is as follows: For the The first and second dimensions of uniform variables for each sample , Perform inverse normal quantile transform:

[0052] In the formula, Indicates the first The first-dimensional marginal uniform variable of each sample is mapped to the first-level transformed value in the standard normal space after undergoing inverse normal quantile transformation. Indicates the first The second-dimensional marginal uniform variable of each sample is mapped to the first-level transformed value of the standard normal space after the inverse normal quantile transformation; This represents the inverse cumulative distribution function of the standard normal distribution.

[0053] The estimated Pearson linear correlation coefficient obtained in the normal space is used as the parameter of the pairwise connection function:

[0054] In the formula, These are the sample means for the corresponding dimensions. Similarly, the same operation is performed on the second and third dimensions to obtain the parameters of the pairwise connection function for the other edge of the first layer. .

[0055] For the second level, it is a conditional pairwise join function. Conventional processing requires first constructing conditional pseudo-observations through conditional distribution function transformation, and then fitting a pairwise connection function. However, in the small-sample or sparse-sampling scenario of photovoltaic power data in this application, strict inverse operation of the conditional distribution function easily introduces numerical errors and amplifies the marginal estimation bias. Therefore, this application adopts a stable fitting strategy based on simplified conditional approximation: directly using the first-dimensional and third-dimensional marginal uniform variables... and As input, the explicit conditional transformation step is skipped, and the approximate correlation coefficient of its residual dependence is estimated in the normal space. This strategy sacrifices theoretical rigor for numerical robustness in small sample scenarios, and is particularly suitable for practical engineering scenarios such as photovoltaic power data, which have non-negative constraints, sharp peaks and heavy tails, and limited original observation information.

[0056] For the The first-dimensional uniform variable of each sample The third dimension of uniform variable Perform inverse normal quantile transform:

[0057] In the formula, Indicates the first The first-dimensional marginal uniform variable of each sample is mapped to the second-level transformed value in the standard normal space after inverse normal quantile transformation. Indicates the first The third-dimensional marginal uniform variable of each sample is mapped to the second-level transformed value in the standard normal space after inverse normal quantile transformation.

[0058] Estimating the second-level conditional pairwise connection function Approximate dependency parameters:

[0059] In the formula, , These are the sample means for the corresponding dimensions of the second layer. The simplification strategy described above maintains the interpretability of the D-Vine Copula model while avoiding numerical instability in the conditional quantile iteration, making it particularly suitable for scenarios like photovoltaic power data, which exhibit non-negative constraints and leptokurtic characteristics.

[0060] Through the above layer-by-layer fitting, the complete parameter set of the D-Vine Copula model is obtained:

[0061] S5. Layer-by-layer conditional sampling in probability space: Based on the fitted D-Vine Copula model and parameters, layer-by-layer conditional sampling is performed in the probability space to generate new probability space samples that retain the original multivariate correlation characteristics and conditional dependencies.

[0062] The sampling theory of the classic D-Vine Copula model assumes accurate marginal distribution estimation and no boundary singularities. However, directly applying this theory to the photovoltaic power data enhancement scenario of this application faces three challenges: First, with small samples, the cumulative probability of the margins tends to fall near the boundary, triggering a numerical explosion in the inverse normal transformation; second, errors accumulate layer by layer in the conditional sampling chain, causing the synthesized samples to deviate from the true joint distribution; and third, the generated uniform variables, after undergoing the standard PPF inverse transformation, directly map back to a finite number of discrete sample points, producing visually identifiable stepped layered pseudo-samples. To address these challenges, this invention introduces a triple stabilization mechanism based on the classic theory: boundary pruning constraints, simplified conditional approximations, and continuous interpolation micro-perturbation inverse transformation in subsequent step S6, forming an engineering-adaptive sampling scheme for photovoltaic scenarios.

[0063] Specifically, considering the small sample sparsity and non-negativity physical constraints of photovoltaic power data, this scheme adopts the following stabilized layer-by-layer conditional sampling process in the probability space: From a distance away from the boundary (to prevent the singularity of the standard normal inverse cumulative distribution function at the boundary, based on the global boundary pruning constant defined above). Extraction from a uniform distribution (used for boundary clipping constraints during sampling and transformation processes). The first intermediate variable is used as the root node for conditional sampling. Let the newly generated sampling variable be the first intermediate variable. The uniform variable of the second-dimensional connection function for each sample is:

[0064] Based on the first level as an edge pairwise connection function ,by Conditional variables are used to generate the first-dimensional uniform sample for each sample.

[0065] First, the condition variable is mapped to the standard normal space to obtain the first... The second-dimensional normal space mapping value of each sample:

[0066] Then, the Gaussian conditional distribution is generated. A new sample in the first dimension of the normal space of the samples:

[0067] In the formula, For the first The first dimension of each sample is an independent random number; This represents a standard normal distribution with a mean of 0 and a standard deviation of 1.

[0068] Then, the new samples in the first-dimensional normal space are mapped back to the uniform space to obtain the... The newly generated uniform variable in the first dimension of each sample:

[0069] Similarly, based on the first layer Gaussian pairwise connection function With the same For conditional variables, generate new samples in the third-dimensional normal space and new uniform variables that are mapped to the uniform space for each sample:

[0070]

[0071] In the formula, These are independent standard normal random numbers.

[0072] Finally, by combining the above layer-by-layer sampling results, we obtain the first... Three-dimensional uniform sample vector of each sample:

[0073] The layer-by-layer conditional sampling mechanism in this application uses intermediate variables Using the Gaussian conditional distribution in normal space as the hub, the hierarchical dependency structure fitted above is strictly maintained by spreading outward through the normal space, while the independent random noise introduced in each step is also eliminated. , This ensured the diversity of the generated samples.

[0074] S6. Edge Inverse Mapping Reconstruction: The new probability space samples are subjected to edge inverse mapping reconstruction. The inverse mapping reconstruction includes inverse cumulative distribution transformation, and continuous interpolation and adaptive Gaussian micro-perturbation are introduced in the inverse cumulative distribution transformation (to reduce the discrete layered structure caused by empirical inverse transformation and improve the continuity and distribution smoothness of the generated data) to obtain enhanced data in logarithmic space.

[0075] Record No. Samples in logarithmic space The order statistics after sorting in ascending order are:

[0076] In the formula, Indicates the first The first sample in the logarithmic space is sorted in ascending order. The order statistic corresponds to the th order statistic after sorting. +1 smaller value, and satisfy In the formula, For the first The ascending order permutation of dimensions is used to change the sort order. Map to the original sample index.

[0077] For the three-dimensional uniform sample vector in the new probability space, the first... The first newly generated sample 2D uniform variable First, calculate its consecutive positions in the sorted sequence:

[0078] Determine the left neighbor index Right Neighbor Index And calculate the interpolation weights. By linear interpolating the edge samples in the logarithmic space, the first... The first sample The estimated value after linear interpolation in the logarithmic space of dimension:

[0079] In the formula, Indicates the first The first sample The estimated value after linear interpolation in the logarithmic space of dimension.

[0080] Existing technologies typically return the sorted original sample points directly as the inverse transformation result, causing the synthesized samples to only take values ​​from a finite number of discrete values, forming a distinct "ladder-like" layered distribution, severely reducing the naturalness and diversity of data augmentation. This invention constructs a continuous transition between adjacent order statistics through linear interpolation, allowing the synthesized samples to take any value within the convex hull of the original samples, fundamentally breaking the discrete layered structure. Furthermore, the effectiveness of the continuous linear interpolation mechanism depends on the continuous cumulative distribution function in step S2. The smooth output is achieved by using a traditional empirical step distribution, where the probability integral transformation result has only a finite number of discrete values. No matter how refined the subsequent interpolation is, it cannot break through the discrete support set of the original sample points. This invention obtains a continuous cumulative distribution function through kernel density estimation analytical integral, fundamentally extending the range continuity of the inverse transform.

[0081] To further eliminate the layered distribution defects caused by sample discreteness, a minimal Gaussian perturbation is applied to the interpolation results, resulting in enhanced data in the logarithmic space after the perturbation:

[0082] In the formula, This indicates the first step after applying a Gaussian perturbation. i The first sample j Data is generated in a logarithmic space of dimension. Indicates a Gaussian perturbation. Indicates the first The standard deviation of the logarithmic space sample. The intensity of the above Gaussian perturbation. After optimization and calibration: if the perturbation is too large, it will destroy the statistical consistency of the marginal distribution; if the perturbation is too small, it will not be able to effectively break the layered structure. This invention adopts an adaptive perturbation strength that is proportional to the standard deviation of the data itself, ensuring that the perturbation amplitude is always on the order of one-thousandth of the variability of the original data. While maintaining the distribution fidelity, it achieves sample continuity. This adaptive calibration strategy does not require manual parameter tuning and has a wide range of application scenarios.

[0083] S7. Inverse Exponential Transform to Restore the Original Space: For Augmented Data Perform an inverse exponential transform to restore the original data space. The expression for the inverse exponential transform is:

[0084] In the formula, This represents the exponentially reduced photovoltaic power value. This exponential reduction step is not a simple mathematical inversion, but a customized design addressing the physical constraints of photovoltaic power. The introduction of a logarithmic transformation maps the original non-negative data to the entire real number space, and after synthesis, it must strictly revert to the non-negative domain.

[0085] This invention adopts Instead of directly This ensures the precise reversibility of the transformation. It compensates for the small shift during logarithmic transformation, and ensures [the stability of the system] through the monotonic positive value property of the exponential function. The constant validity eliminates the generation of illegal samples with negative power from a mathematical perspective, eliminating the need for subsequent filtering or iterative supplementation and simplifying the engineering implementation process.

[0086] S8. Distribution Consistency Verification and Output: To quantitatively verify the accuracy of the data distribution of photovoltaic observations after data augmentation using this method, the distribution difference between the augmented dataset and the original dataset is evaluated using the Euclidean distance D, and the augmented photovoltaic power data is output.

[0087] The expression for the Euclidean distance D is:

[0088] In the formula, d This represents the number of photovoltaic power plants. The data for one photovoltaic power plant represents one dimension in spatial geometry, which is also represented in the formula. d 3D space, when each photovoltaic power range is divided into When dividing into uniform segments, based on the original scatter plot, we have d The entire uncertainty hypercube constructed from the dimensional photovoltaic interval is divided into: A hypercube, denoted as _____. .and Represents a sub-hypercube The volume. This indicates that the data from the original scatter plot falls into the sub-hypercube. The number of power samples, This indicates that the data set of the comparison scheme falls into the sub-hypercube. The number of power samples. Furthermore... This represents the total number of power samples from the original scatter plot of this invention. This represents the total number of power samples compared to the other scheme.

[0089] Determine the upper and lower bounds of the power intervals for each dimension based on the original dataset. Calculate the minimum value of each dimension With the maximum value , will the interval Divided into equal parts The space of the hypercube is defined and its volume is calculated. The number of samples in each hypercube, both from the original dataset and the augmented dataset, is then used to calculate the Euclidean distance D. The closer the D value is to 0, the closer the augmented dataset distribution is to the original distribution, and the better the data augmentation effect.

[0090] The results are analyzed from two perspectives: the effectiveness of the Euclidean distance and the advantages of data augmentation. The effectiveness analysis of the Euclidean distance is as follows: Original sample sets and baseline distributions under different finite observation scales ( Figure 2 The Euclidean distribution distances of ) are shown in Table 1.

[0091] Table 1. Euclidean distance between the original sample set and the baseline distribution under different finite observation scales

[0092] Combination Figure 3 The photovoltaic power distribution of the three photovoltaic power plants under different finite observation values ​​shows that when the number of photovoltaic power observation values ​​is 30, 50, 100, and 300, the corresponding distance value is 364.71. 289.10 186.89 104.60 This decreasing trend aligns with the convergence of empirical distributions in statistical learning theory: as... Figure 3With increasing limited observational information, the distribution of the limited samples becomes increasingly closer to the baseline distribution. Specifically, when the number of observations is 30, the sample is extremely sparse, with very few scattered points and a highly uneven spatial distribution. The overall outline of the three-dimensional joint distribution is almost indistinguishable, with large gaps at the boundary regions and modal gaps. When the number of observations is 50, the sample size increases slightly, but the scattered points remain sparse. The main clustering areas are only vaguely visible, and the distribution edges and tails are severely under-covered, which would introduce significant bias when used directly for statistical inference. When the number of observations is 100, the sample density increases, and the scattered point distribution in the core clustering area begins to show a clustering trend similar to the baseline distribution. However, there are still significant gaps in the edge transition regions and high-power tails, making it difficult to fully characterize the tail features of the joint distribution. When the number of observations is 300, the sample is relatively abundant, and the overall outline of the scattered point distribution is close to the baseline distribution. Figure 2 The baseline distribution is established, but due to the inherent complex nonlinear correlation of photovoltaic power and multi-peak operating modes, insufficient coverage still exists in local sparse regions and transition regions between modes, ultimately leading to a monotonically decreasing Euclidean distance. However, when the number of observations is 30, the distance value is still as high as 364.71. This indicates that in the limited observation scenarios commonly seen in engineering practice, directly using the original sample set for modeling will introduce significant distribution bias, and it is urgent to correct and expand the distribution structure through data augmentation techniques.

[0093] The advantages of data augmentation are analyzed as follows: To verify the improvement effect of the proposed method on finite observation scenarios, after obtaining photovoltaic power observations of 30, 50, 100, and 300, the following steps were performed sequentially: logarithmic space transformation, kernel density estimation for nonparametric modeling, three-dimensional joint dependency deconstruction, logarithmic space linear interpolation and inverse transformation, followed by exponential reduction to generate an enhanced sample set with the same dimension as the original finite sample. The enhanced sample set is shown below. Figure 4 As shown. Specifically, when the number of observations is 30, although the original input consists of only 30 extremely sparse samples, after enhancement by this method, Figure 4 The scattered points are dense and the three-dimensional joint distribution outline is clearly discernible, effectively filling the gaps. Figure 3 The large blank sparse areas and boundary gaps in the scatter plot with 30 observations, the scatter clustering pattern, nonlinear correlation trend and non-negative boundary characteristics are all consistent with... Figure 2 The baseline distributions are highly similar, validating the effectiveness of adaptive bandwidth kernel density estimation in nonparametric edge modeling, and the D-Vine Copula model's ability to stably preserve the correlation structure under sparse conditions. When the number of observations is 50, the enhanced scatter distribution is uniform and full, with significant recovery of the main operating modes and distribution tails; compared to... Figure 3 The sparse scatter plot has 50 observations. Figure 4The scatter plot with 50 observations did not exhibit obvious stepped layered discreteness defects. This is attributed to the inverse cumulative distribution transformation strategy of continuous linear interpolation in logarithmic space superimposed with minimal Gaussian micro-perturbations, which successfully eliminated the pseudo-sample layered structure caused by direct sampling of a finite number of samples. When the number of observations reached 100, the enhanced scatter plot further approximated the desired result. Figure 2 The baseline distribution has sufficient edge coverage, and the transition regions between modes are smooth and coherent; with Figure 3 Compared to when the number of observations is 100, the sparse region at the tail is effectively extrapolated and filled, and the empirical covariance structure more closely resembles the true multidimensional correlation characteristics. When the number of observations is 300, the enhanced scatter plot and... Figure 2 The baseline distribution is almost identical in visual morphology, and the tails, boundaries, and multimodal modes of the joint distribution are all completely reproduced; compared to Figure 3 The scatter plot with 300 observations shows a particularly noticeable effect in filling in local sparse regions, demonstrating that even with moderate sample input, this method can further improve distribution accuracy through continuous interpolation and micro-perturbation mechanisms, achieving refined preservation of the joint distribution. Furthermore, the Euclidean distance between the enhanced sample set and the baseline distribution is calculated. Table 2 below shows the Euclidean distance between the enhanced sample set and the baseline distribution.

[0094] Table 2. Euclidean distance between the enhanced sample set and the baseline distribution using this method.

[0095] As can be seen from Tables 1 and 2, the distribution distance of the sample set enhanced by this invention is significantly smaller than that of the corresponding finite observation sample set, and this advantage becomes more prominent as the original sample size decreases. Figure 4 Visualization shows that, with the increase of observational information, the augmented sample set gets closer and closer to the baseline distribution. When the number of observations is 30, the finite sample distance is 364.71. After enhancement, it decreased to 127.05. The reduction reached 65.2%. In this extremely small sample scenario, traditional parametric generation models are difficult to apply directly due to the mismatch between the preset distribution form and the non-negative right-skewed characteristic of photovoltaic power. At the same time, direct resampling easily destroys the multidimensional correlation structure and produces discrete ladder-like pseudo-samples. This method first maps the non-negative power data to an approximately symmetric real number space through logarithmic transformation, and combines adaptive bandwidth kernel density estimation to perform nonparametric modeling of the marginal distribution of each dimension, avoiding the loss of distribution structure caused by over-smoothing. Then, the three-dimensional joint dependency is deconstructed into a hierarchical combination of binary Gaussian pairwise connection functions through the D-Vine Copula model, accurately capturing the nonlinear dependence between variables. Finally, the ladder-like defects caused by direct sampling of finite sample points are fundamentally eliminated by continuous linear interpolation in logarithmic space superimposed with minimal Gaussian micro-perturbation. After exponential transformation and strict guarantee of non-negativity, the empirical distribution of the enhanced sample set significantly approximates the original baseline distribution while maintaining physical legitimacy.

[0096] When the number of observations is 50 and 100, the enhanced distance decreases to 118.34 km, respectively. With 100.85 Compared to finite samples, the errors were reduced by 59.1% and 46.0%, respectively. Within this range, the original samples already possessed a preliminary distribution profile, but due to sample sparsity, the estimation of the correlation structure between variables was distorted, and boundary coverage was insufficient. This method utilizes hierarchical conditional sampling of the D-Vine Copula model to perform joint distribution extrapolation filling on the locally sparse regions of the original data. Simultaneously, by using adaptive bandwidth to maintain compactness in dense regions and moderate smoothness in sparse regions, a trade-off between edge fidelity and correlation structure preservation is achieved, thereby significantly reducing the estimation error of the joint distribution.

[0097] When the number of observations is 300, the finite sample distance has decreased to 104.60. After enhancement, it further decreased to 88.32. The reduction was 15.6%. Although the original sample was relatively sufficient, due to the complex nonlinear correlation and multi-peak operating modes of photovoltaic power data, the limited sample still could not fully characterize the tail and transition regions between modes of the joint distribution. This invention generates enhanced samples through nonparametric edge modeling of kernel density estimation and conditional sampling of the D-Vine Copula model, which effectively fills the sparse regions of the original sample at the distribution edges and mode gaps, making the empirical covariance matrix closer to the real structure, thus still achieving considerable distribution improvement.

[0098] Without assuming any specific parameter distribution, this method simultaneously addresses the multi-coupling challenges of preserving marginal distributions, high-dimensional correlation structures, sample continuity, and non-negativity constraints. Regardless of the abundance of observational information, the proposed logarithmic transformation-adaptive kernel density estimation-D-Vine Copula model fitting-continuous interpolation inverse transformation-exponential restoration enhancement method significantly compresses the distribution distance of the sample set. Moreover, the scarcer the original information, the more significant the relative benefit of the enhancement, fully demonstrating the effectiveness and engineering practical value of this method in solving the problem of sparsity in photovoltaic power data with limited information.

[0099] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for enhancing nonlinear correlation data of new energy sources with limited observations, characterized in that: Includes the following steps: S1. Data preprocessing and logarithmic space mapping: The acquired raw photovoltaic power data is preprocessed and logarithmic space mapping is completed in combination with the logarithmic transformation constant. The photovoltaic power data with physical constraints is transformed into a continuous unconstrained real number space to obtain logarithmic space power data. S2. Continuous marginal distribution modeling: For the power data corresponding to each photovoltaic power station in the logarithmic space, the stable kernel density estimation method is used to construct the marginal probability density function of each dimension, and the continuous cumulative distribution function of each dimension is obtained through integral operation; S3. Probability integral transformation and boundary clipping: Based on the cumulative distribution function, a probability integral transformation is performed to map the logarithmic space power data to a unified probability space. The transformed data is then subjected to linear smoothing correction and boundary clipping to obtain probability space samples. S4. Construction and parameter fitting of D-Vine Copula model: Based on the probability space samples, a D-Vine Copula model is constructed. The joint dependency of three-dimensional photovoltaic power is decomposed into two-level binary pairwise connection functions. The binary pairwise connection functions are fitted with Gaussian connection functions to fit the correlation structure between variables in each level. All parameters of the D-Vine Copula model are obtained by solving the problem. S5. Layer-by-layer conditional sampling in probability space: Based on the fitted D-Vine Copula model and parameters, layer-by-layer conditional sampling is performed in the probability space to generate new probability space samples that retain the original multivariate correlation characteristics and conditional dependencies. S6. Edge Inverse Mapping Reconstruction: The new probability space samples are subjected to edge inverse mapping reconstruction. The inverse mapping reconstruction includes inverse cumulative distribution transformation, combined with continuous interpolation and adaptive Gaussian micro-perturbation, to obtain enhanced data in logarithmic space. S7. Inverse exponential transform to restore the original space: Perform an inverse exponential transform on the enhanced data to restore it to the original data space, and obtain a photovoltaic power enhancement dataset that satisfies the non-negative physical constraints. S8. Distribution Consistency Verification and Output: Using Euclidean distribution distance, the consistency of the photovoltaic power enhancement dataset obtained in step S7 with the original benchmark dataset is verified, and the enhanced photovoltaic power data is output.

2. The method for enhancing nonlinear correlation data of new energy sources based on limited observations according to claim 1, characterized in that: In step S1, the preprocessing includes outlier detection, data normalization, and numerical stabilization. The expression for the logarithmic space mapping is: In the formula, Indicates the first The first sample j Original photovoltaic power data, Represents the logarithmic transformation constant. Describe the th in the logarithmic space Dimensional variables.

3. The method for enhancing nonlinear correlation data of new energy sources based on limited observations according to claim 1, characterized in that: In step S2, the stable kernel density estimation method uses a Gaussian kernel function to perform nonparametric kernel density estimation, and the bandwidth parameter of the kernel density estimation is determined based on the standard deviation of the logarithmic space samples in each dimension.

4. The method for enhancing nonlinear correlation data of new energy sources based on limited observations according to claim 1, characterized in that: In step S3, the expressions for linear smoothing correction and boundary clipping are: In the formula, This represents the initial data obtained after the probability integral transformation. This represents the probability space sample obtained after linear smoothing correction and boundary clipping. This represents the global boundary clipping constant, whose magnitude is greater than that of the logarithmic transformation constant.

5. The method for enhancing nonlinear correlation data of new energy sources based on limited observations according to claim 1, characterized in that: In step S4, the first level of the two-level binary pairwise connection function is the edge pairwise connection function, which represents the direct dependency relationship between each adjacent dimension, and the second level is the conditional pairwise connection function, which represents the conditional residual dependency relationship of the remaining two dimensions after fixing the second dimension variable. The Gaussian link function maps uniform marginal variables to a normal space through a standard normal quantile transformation, and then quantifies the dependence strength between variables with a linear correlation coefficient. The Pearson correlation coefficient obtained in the normal space is used as the parameter of the Gaussian link function. When fitting the conditional pairwise link function, the marginal uniform variables of the corresponding dimension are selected as inputs, skipping the explicit conditional transformation step, and the approximate correlation coefficient used to characterize the residual dependence is calculated in the normal space.

6. The method for enhancing nonlinear correlation data of new energy sources based on limited observations according to claim 4, characterized in that: In step S5, the conditional sampling layer by layer defines the sampling interval based on the global boundary pruning constant. The intermediate variable is extracted as the root node of the conditional sampling within the uniform distribution corresponding to the sampling interval. The root node variable is first mapped to the normal space, and then the conditional normal sample is generated layer by layer by combining the Gaussian connection function and introducing independent standard normal random quantities. Subsequently, the conditional normal sample is mapped back to the uniform space to obtain new uniform variables in each dimension. Finally, the new uniform variables in each dimension are spliced ​​together to form a three-dimensional uniform sample vector.

7. The method for enhancing nonlinear correlation data of new energy sources based on limited observations as described in claim 6, characterized in that: In step S6, when performing the inverse cumulative distribution transformation, a single-dimensional uniform variable is extracted from the three-dimensional uniform sample vector in the new probability space, its continuous position in the sorted sequence is calculated, the left and right neighborhood indices and interpolation weights are determined, and linear interpolation is performed using the adjacent order statistics to obtain the estimated value of the variable mapped back to the logarithmic space. Then, Gaussian random perturbation is superimposed to complete the adaptive micro-perturbation processing.

8. The method for enhancing nonlinear correlation data of new energy sources based on limited observations according to claim 2, characterized in that: In step S7, the expression for the inverse exponential transform is: In the formula, This indicates the first step after applying a Gaussian perturbation. i The first sample j Data is generated in a logarithmic space of dimensionality. This represents the photovoltaic power value after exponential restoration.

Citation Information

Patent Citations

  • Photovoltaic output probability modeling method and system for describing photovoltaic multi-dimensional correlation based on R-Vine Copula

    CN119358192A

  • Unit commitment optimization method of multivariable joint probability distribution model based on D-vine Copula

    CN121145605A