Sample adaptive iterative training method and system for existing model tuning
By constructing a net load-irradiance response intensity index and a photovoltaic bistable constraint term, and optimizing the sample training weights, the problem of decreased model sensitivity caused by inverter limiting was solved, thereby improving the accuracy and stability of distribution network net load forecasting.
Patent Information
- Application Number
- CN202610106495.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-27
- Publication Date
- 2026-03-03
AI Technical Summary
Existing sample adaptive iterative training methods, in high-penetration distributed energy scenarios, suffer from inverter limiting actions that reduce the model's sensitivity to solar irradiance, making it unable to accurately capture the driving effect of meteorological changes on net load. This results in the prediction model underestimating the impact of sunlight under unlimited conditions, leading to underfitting.
By acquiring net load and solar irradiance data of the distribution network, a net load-irradiance response intensity index is constructed to determine the baseline values for inverter limiting and normal power generation. The non-truncation probability is calculated, photovoltaic bistable constraint terms and sample training weights are constructed, and the power prediction model parameters are optimized.
It effectively overcomes the gradient collapse problem of the model for meteorological input, improves the prediction accuracy and robustness under complex meteorological and voltage regulation control coupling conditions, avoids the pseudo-optimal trap, and adapts to the complexity of high-penetration distributed energy scenarios.
Smart Images

Figure CN121598089A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power prediction technology for power systems, and more specifically, to a sample adaptive iterative training method and system for optimizing existing models. Background Technology
[0002] As the penetration rate of distributed energy in power distribution networks continues to increase, the load forecasting target on the grid side has shifted from the traditional simple electricity load to the net load that includes the impact of distributed generation. In this scenario, the power generation of downstream photovoltaic (PV) power cannot often be directly measured, but is implied in the net load data.
[0003] In practical engineering operations, to address the risk of voltage exceeding limits caused by high penetration rates, modern power distribution networks widely adopt intelligent inverter control strategies (such as Volt-Watt control) that comply with standards like IEEE 1547. When the grid voltage is too high or the equipment reaches its rated power, the inverter actively limits the active power output, causing the power generation that should increase with increasing solar irradiance to be clipped or interrupted. Furthermore, influenced by atmospheric cloud field movements, the surface solar irradiance itself exhibits intermittent and abrupt changes.
[0004] Existing prediction model tuning methods typically employ sample-adaptive iterative training (such as boosting algorithms or reweighting methods based on hard example mining), which forces the model to reduce residuals by assigning larger training weights to high-error samples. However, in the scenario described above where inverter limiting exists, the following problems arise: Specifically, when the inverter triggers a limiting action, although the external solar irradiance is high, the actual grid-connected power is forcibly reduced. In this situation, if the model predicts a higher power output based on solar irradiance (consistent with normal photovoltaic conversion logic), a large prediction residual will occur. Existing adaptive training mechanisms classify such samples as hard examples and assign them extremely high weights, forcing the model to fit these truncated power values. This leads the model to incorrectly learn the mapping relationship that increased solar irradiance does not necessarily lead to increased power during iteration, causing an abnormal decrease in the model's sensitivity to the solar irradiance parameter (i.e., gradient effectiveness decay).
[0005] Ultimately, this pseudo-optimal state caused by the existing training mechanism will cause the model to underfit under normal weather conditions without amplitude limiting, due to underestimating the impact of sunlight, and thus fail to accurately capture the driving effect of meteorological changes on net load. Summary of the Invention
[0006] This invention provides a sample adaptive iterative training method and system for optimizing existing models, which solves the technical problems mentioned in the background art.
[0007] The first aspect is the sample-adaptive iterative training method for optimizing existing models, including: Obtain net load data and solar irradiance data of the distribution network as time series samples; Based on the response slope of the net load data relative to the solar irradiance data within a local window, the net load-irradiance response intensity index is determined. The statistical distribution of the response intensity index is used to determine the cutoff mode reference value representing inverter limiting and the response mode reference value representing normal power generation, and the non-cutoff probability of the sample being in normal power generation state is calculated accordingly. Calculate the marginal response strength of the power prediction model to the solar irradiance data; The critical factor for capturing the inverter's operating boundary is calculated based on the non-truncation probability, and the sample training weights are constructed by combining the prediction residuals. The critical factor takes a maximum value when the non-truncation probability indicates that the inverter is switching between limiting and normal power generation states. A photovoltaic bistable constraint term is constructed, and the non-truncation probability is used to constrain the marginal response intensity to approach the response mode reference value under normal power generation conditions and to approach the truncation mode reference value under inverter limiting conditions. The parameters of the power prediction model are updated based on the optimization objective function composed of the power prediction error term weighted by the sample training weights and the photovoltaic bistable constraint term.
[0008] Secondly, a sample adaptive iterative training system for optimizing existing models, applied to any of the sample adaptive iterative training methods for optimizing existing models, includes: The data acquisition module is used to acquire net load data and solar irradiance data of the distribution network as time series samples; The index determination module is used to determine the net load-irradiance response intensity index based on the response slope of the net load data relative to the solar irradiance data within a local window. The state probability calculation module is used to determine the cutoff mode reference value representing inverter limiting and the response mode reference value representing normal power generation by using the statistical distribution of the response intensity index, and calculate the non-cutoff probability of the sample being in the normal power generation state accordingly. The model response calculation module is used to calculate the marginal response strength of the power prediction model to the solar irradiance data; The weight construction module is used to calculate the critical factor for capturing the inverter operation boundary based on the non-truncation probability, and to construct sample training weights in combination with the prediction residual, wherein the critical factor takes a maximum value when the non-truncation probability indicates the inverter switching between limiting and normal power generation states. The constraint construction module is used to construct photovoltaic bistable constraint terms. Using the non-truncation probability, the marginal response intensity is constrained to approach the response mode reference value under normal power generation conditions and to approach the truncation mode reference value under inverter limiting conditions. The parameter update module is used to update the parameters of the power prediction model based on the optimization objective function composed of the power prediction error term weighted by the sample training weights and the photovoltaic bistable constraint term.
[0009] The beneficial effects of this invention include: It effectively overcomes the gradient collapse problem in high-penetration distributed energy scenarios, where the predictive model loses its responsiveness to meteorological inputs due to inverter limiting actions. By introducing photovoltaic bistable constraint terms and boundary reinforcement weights during training, this invention forces the model to maintain high sensitivity to solar irradiance under normal power generation conditions, while adapting to a saturation cutoff mechanism under limiting conditions, thus avoiding the algorithm falling into the pseudo-optimal trap of ignoring environmental factors. Furthermore, this method requires no additional hardware measurement devices; it can automatically identify device operating boundaries using only existing data, significantly improving the accuracy and robustness of distribution network net load prediction under complex meteorological and voltage regulation control coupling conditions. Attached Figure Description
[0010] Figure 1 This is a flowchart of the sample adaptive iterative training method for optimizing existing models according to the present invention; Figure 2 This is a schematic diagram illustrating a specific implementation of the present invention. Detailed Implementation
[0011] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.
[0012] Example 1: As Figure 1 As shown, the sample-adaptive iterative training method for optimizing existing models includes: Obtain net load data and solar irradiance data of the distribution network as time series samples; Based on the response slope of the net load data relative to the solar irradiance data within a local window, the net load-irradiance response intensity index is determined. The statistical distribution of the response intensity index is used to determine the cutoff mode reference value representing inverter limiting and the response mode reference value representing normal power generation, and the non-cutoff probability of the sample being in normal power generation state is calculated accordingly. Calculate the marginal response strength of the power prediction model to the solar irradiance data; The critical factor for capturing the inverter's operating boundary is calculated based on the non-truncation probability, and the sample training weights are constructed by combining the prediction residuals. The critical factor takes a maximum value when the non-truncation probability indicates that the inverter is switching between limiting and normal power generation states. A photovoltaic bistable constraint term is constructed, and the non-truncation probability is used to constrain the marginal response intensity to approach the response mode reference value under normal power generation conditions and to approach the truncation mode reference value under inverter limiting conditions. The parameters of the power prediction model are updated based on the optimization objective function composed of the power prediction error term weighted by the sample training weights and the photovoltaic bistable constraint term.
[0013] Preferably, acquiring net load data and solar irradiance data of the distribution network as time series samples includes: Define time index Obtaining input data and building a model: The net load data of the distribution network is represented as follows: ; The solar irradiance data is expressed as: ; Construct a feature vector containing ambient temperature data and calendar information. ; Establish the power prediction model: in, These are the model's predicted values. For model parameters A defined nonlinear mapping function, is the dimension of the feature vector.
[0014] Distribution network net load data is the difference between the actual electricity consumed by the distribution network per unit time and the power generated by distributed power sources. It can be collected in real time through smart meters and load monitoring terminals deployed along the distribution network, or historical and real-time data can be obtained through the SCADA (Supervisory Control and Data Acquisition) system of the power dispatch center.
[0015] Solar irradiance data is the solar radiation energy projected onto a unit area per unit time. It can be collected through irradiance sensors at weather stations, satellite remote sensing data inversion, ground-based irradiance observation networks, or data provided by commercial irradiance data service platforms.
[0016] Ambient temperature data refers to the air temperature in the area where the power distribution network is located. It can be obtained through weather stations, temperature sensors deployed in power distribution stations, and IoT temperature acquisition devices.
[0017] Calendar information records the time attributes corresponding to a date. It can be directly obtained using date functions in a computer system, including whether it is a weekday, holiday, season, and time period.
[0018] A feature vector is a multidimensional data array formed by combining solar irradiance data, ambient temperature data, and calendar information in a fixed order. It is preferably 5 to 15 dimensions to comprehensively consider the number of key influencing factors in net load forecasting and model computational efficiency, avoiding information loss due to too few dimensions or overfitting due to too many dimensions.
[0019] Power prediction models are based on input feature vectors and output predicted values of the distribution network net load. Deep learning models such as Long Short-Term Memory networks and gradient boosting tree models such as extreme gradient boosting trees are preferred because both types of models have strong generalization ability in time series prediction and can adapt to the nonlinear characteristics of net load.
[0020] Model parameters are internal variables that determine the output of the power prediction model. The preferred values are between -0.01 and 0.01, and random or Xavier initialization should be used to avoid gradient explosion due to excessively large initial parameters or gradient vanishing due to excessively small initial parameters.
[0021] Since solar irradiance directly affects the power generation of distributed photovoltaics, it is the core driving factor for net load fluctuations; ambient temperature affects the conversion efficiency of photovoltaic modules and the electricity demand of loads such as air conditioning, indirectly affecting the net load; calendar information can distinguish the differences in load patterns between weekdays and holidays, and between peak and off-peak periods. The combination of these three factors can comprehensively cover the key dimensions affecting the net load.
[0022] The selection and initialization of the power prediction model need to be adapted to subsequent sample adaptive iterative training to avoid excessive deviation of the initial model, which would lead to difficulty in optimization. For example, when choosing a long short-term memory network, the number of hidden layer neurons should preferably be 32 to 128 to match the dimension of the feature vector, so as to ensure a balance between the model's fitting ability and computational efficiency.
[0023] The time step for data acquisition is preferably 15 minutes to 1 hour, based on the regular acquisition frequency of distribution network load and irradiance data. The 15-minute step is suitable for high-precision forecasting scenarios, while the 1-hour step is suitable for day-ahead forecasting and other scenarios.
[0024] The dimensions of the feature vector can be determined by recursive feature elimination or Pearson correlation coefficient screening. For example, by calculating the correlation coefficient between each feature and the net load, features with an absolute value of correlation coefficient greater than 0.3 are retained, and the final dimension is determined to be between 5 and 15.
[0025] Model parameters are initialized using Xavier initialization first. For the ReLU (Rectified Linear Unit) activation function, the initialization range is adjusted to -0.02 to 0.02 to adapt to the gradient characteristics of the activation function and improve the stability of the initial model.
[0026] It should be noted that by collecting data on the core influencing factors of net load, a feature vector comprehensively reflecting the changing patterns of net load is formed, providing sufficient information support for the model. A model suitable for time series prediction is selected, and parameters are initialized appropriately to avoid abnormal initial states affecting subsequent training. The benefits of this approach are that the model can quickly adapt to the nonlinear and fluctuating characteristics of net load, laying a stable foundation for adaptive iterative training of samples, reducing convergence difficulties during training, and the targeted combination of feature vectors avoids interference from redundant information. Appropriate initialization values for the model can improve training efficiency.
[0027] Preferably, the net load-irradiance response intensity index is determined based on the response slope of the net load data relative to the solar irradiance data within a local window, including: For each moment Define a local time window : in, To smooth the window radius, This represents the total number of time steps. Calculate the window mean: in, The number of samples within the window. For a moment Solar irradiance data, For a moment Net load data of the distribution network; Calculate centralized irradiance data With centralized net load data : Calculate the local regression slope : in, It is a numerically stable term; Calculate the net load-irradiance response intensity index : A local time window is a dataset defined around the current moment, containing several time steps before and after it, used for local regression calculations.
[0028] The smoothing window radius is the number of time steps that the local time window extends forward and backward. It is preferably 3 to 7 to balance the correlation of local data and the interference of outliers. 3 to 7 time steps can effectively capture short-term response patterns, avoiding both excessive data fluctuations caused by a window that is too small and the introduction of irrelevant data by a window that is too large.
[0029] The total number of time steps is the total number of time series samples of net load and solar irradiance in the distribution network, and it is a data acquisition parameter. It can be obtained by counting the total number of complete time series data collected.
[0030] The mean value of solar irradiance data within a local window is the arithmetic mean of all solar irradiance data within the local time window.
[0031] The mean value of the net load data window of the distribution network within a local window is the arithmetic mean of all net load data of the distribution network within the local time window.
[0032] Centralized irradiance data is the result of subtracting the mean value of the irradiance data within a window from each solar irradiance data point within a local time window.
[0033] Centralized net load data is the result of subtracting the window average of net load data within a local time window from the net load data of each distribution network within that window.
[0034] The local regression slope is a linear regression coefficient calculated from the centered irradiance data and the centered net load data within a local window. It reflects the degree of linear response of the net load to irradiance within a local area.
[0035] The numerical stability term is a small constant introduced to avoid numerical anomalies caused by the denominator approaching zero in local regression calculations. It is preferably between 10 to the power of -6 and 10 to the power of -8, because small values in this range will not affect the accuracy of the calculation results, while effectively avoiding the calculation risk of the denominator being zero.
[0036] The net load-irradiance response intensity index is a parameter obtained by inverting the local regression slope and is used to quantify the degree of response of net load to changes in solar irradiance.
[0037] Traditional methods often employ simple difference or global regression, which are ill-suited for scenarios involving the coupling of sudden irradiance changes and inverter cutoff in high-penetration distributed energy environments. Symmetrical windows can balance the impact of time steps before and after the current moment, avoiding the lag bias caused by unidirectional windows; centralized processing can eliminate interference from overall data offset within the window, focusing on local fluctuations; local linear regression is more robust than the difference method, reducing the impact of outliers caused by sudden irradiance changes; and the inversion operation transforms the negative correlation between net load and irradiance into a positive response indicator, facilitating subsequent bistable feature identification. For example, when the time step is 15 minutes, a smoothing window radius of 5, calculated around the current moment and including 11 data points across 5 time steps before and after, can effectively smooth irradiance fluctuations caused by fragmented clouds.
[0038] The boundary handling methods for local windows include: when the current time is less than the smoothing window radius, forward padding is used, extending the window from the first time step forward by twice the smoothing window radius for several time steps; when the current time is greater than the total number of time steps minus the smoothing window radius, backward padding is used, extending the window from the total number of time steps minus twice the smoothing window radius for several time steps to the last time step. The specific selection of the smoothing window radius can be combined with the data time step size; 5 is preferred for a time step size of 15 minutes, and 3 is preferred for a time step size of 1 hour. The specific value of the numerical stability term can be adjusted according to the data precision; 10 to the power of -6 is used when the data precision is 6 decimal places, and 10 to the power of -8 is used when the data precision is higher.
[0039] It should be noted that, through local window analysis, an intensity index is constructed that accurately reflects the local response of net load to irradiation, adapting to the nonlinear response characteristics caused by sudden irradiation changes and inverter cutoff under high-penetration distributed energy. By using symmetrical window design, centralization, and local linear regression, the impact of outliers and overall data offset is reduced, and a numerical stability term is introduced to ensure the reliability of the calculation process. As a result, the obtained response intensity index clearly presents the local sensitivity of net load to irradiation.
[0040] Preferably, the statistical distribution of the response intensity index is used to determine the cutoff mode reference value representing inverter limiting and the response mode reference value representing normal power generation, and the non-cutoff probability of the sample being in a normal power generation state is calculated accordingly, including: Set the cutoff modal reference value : Calculate the global median With absolute median : in, This is a net load-irradiance response intensity index for the entire time series. It is a numerically stable term; Calculate the response modal reference value : in, It is a logical stethoscope function; Define uniform scale parameters Calculate time Cut-off state energy value With response state energy value : Calculate the non-truncated probability : in, Softly assign temperature parameters for the state.
[0041] The cutoff mode reference value is a parameter representing the standard value of the net load-irradiance response intensity under inverter limiting conditions. It is preferably 0 because the photovoltaic output exhibits saturation characteristics when the inverter is limited, and the net load response to irradiance changes approaches 0, which conforms to the operating mechanism of power electronic equipment.
[0042] The response mode reference value is a parameter representing the standard value of net load-irradiance response intensity under normal power generation conditions, reflecting the stable response level without inverter limiting.
[0043] The global median is the value in the middle position after the net load-irradiance response intensity index is sorted by numerical value at all times, and is used to characterize the central tendency of the response intensity.
[0044] The absolute median difference is the absolute value of the difference between the net load-irradiance response intensity index at all times and the global median, and then the median is taken as the value. It is used to characterize the dispersion of the response intensity.
[0045] The sigmoid activation function is used to construct weighted coefficients, achieving a smooth weighting of response intensity indicators and avoiding interference from outliers.
[0046] The cutoff state energy value is the square of the standardized Euclidean distance between the current response intensity index and the cutoff mode reference value, and is used to quantify the degree to which the current sample belongs to the inverter's limiting state.
[0047] The response state energy value is the square of the standardized Euclidean distance between the response intensity index and the response mode reference value at the current moment, and is used to quantify the degree to which the current sample belongs to the normal power generation state.
[0048] The state soft assignment temperature parameter is a custom parameter that adjusts the smoothness of the non-truncation probability calculation. It is preferably between 0.5 and 2.0, because this range can balance the sensitivity and stability of state division. If the value is too small, it will easily lead to over-segmentation of states, and if the value is too large, it will easily blur the boundary between the two states.
[0049] The non-truncated probability is the probability value that the current sample is in a normal power generation state. It is used to achieve a soft division between inverter limiting and normal power generation state, and its value ranges from 0 to 1.
[0050] Traditional state partitioning methods often employ unfounded free learning or hard thresholding, which are difficult to adapt to high-penetration distributed energy scenarios. This solution, leveraging the saturation characteristics of inverter limiting, presets the truncated mode benchmark value to 0, establishing a direct link between the mechanism and the benchmark. Weighted coefficients are constructed using the logistic function and the absolute median difference to calculate the response mode benchmark value, effectively suppressing outlier interference from irradiance mutations and ensuring the robustness of the benchmark value. Furthermore, the deviation of the sample from the two benchmark values is quantified by energy values, and non-truncated probabilities are obtained using exponential normalization, achieving smooth, soft state partitioning and avoiding boundary mutation problems caused by hard thresholding. For example, when the response intensity index is concentrated around 0 and 0.8, the global median is 0.4, and the absolute median difference is 0.2. After weighted calculation using the logistic function, the response mode benchmark value is approximately 0.75, accurately reflecting the true response level of the normal power generation state.
[0051] No additional outlier removal is needed when calculating the absolute median, as the median itself has strong outlier resistance and can be directly calculated for all response intensity indices. The standardization of the squared Euclidean distance is based on the absolute median. The calculation first divides the difference between the current response intensity index and the baseline value by the absolute median, then squares the result, using a robust metric to avoid interference from extreme values in the distance calculation. The specific value of the soft-assignment temperature parameter can be adjusted based on data fluctuations: 0.5 to 1.0 is used when irradiance changes frequently and data fluctuations are large; 1.5 to 2.0 is used when the weather is stable and data fluctuations are small. For example, 0.8 is used in areas with frequent scattered clouds, and 1.6 is used in areas with clear skies and few clouds.
[0052] It should be noted that, based on the statistical regularities of inverter operating characteristics and response intensity indicators, a bistable benchmark system is constructed. A soft partitioning method is used to quantify the probability of a sample belonging to a particular state, adapting to the switching characteristics of the two operating states under high-penetration distributed energy. A preset truncated mode benchmark value aligns with the equipment's operating mechanism, robustly calculating the response mode benchmark value avoids outlier interference, and energy value and exponent normalization ensure the smoothness of state partitioning. The resulting bistable benchmark value accurately reflects the essential characteristics of the two states, and the non-truncated probability allows for continuous quantification of state attribution.
[0053] Preferably, calculating the marginal response strength of the power prediction model to the solar irradiance data includes: If the power prediction model Differentiable, calculate the marginal response strength : in, For the current number A power prediction model based on round-iteration iterations. For solar irradiance data, For feature vectors; If the power prediction model is not differentiable, calculate the marginal response strength. : in, This is a preset fixed difference step size.
[0054] Marginal response strength is a parameter that quantifies the sensitivity of the power prediction model output to changes in solar irradiance data, reflecting the magnitude of the predicted net load change caused by small changes in irradiance.
[0055] The small perturbation step size is the tiny change applied to the solar irradiance data when calculating the marginal response intensity of the non-differentiable model. It is preferably between 10^-3 and 10^-2, because this range ensures the accuracy of the difference calculation while preventing excessive perturbation that deviates from the actual range of irradiance variation, thus avoiding affecting the reasonableness of the calculation results.
[0056] Traditional calculations often employ single methods, making it difficult to adapt to different types of power prediction models. This solution directly calculates partial derivatives for differentiable models, such as deep learning neural networks, fully utilizing the model's differentiability to improve computational efficiency. For non-differentiable models, such as gradient boosting trees and random forests, it uses the central difference method instead of forward or backward difference methods. By symmetrically perturbing the irradiance data, it offsets the systematic errors caused by perturbations in a single direction, improving computational accuracy. For example, for a gradient boosting tree model, when the actual solar irradiance is 800, a perturbation step size of 10 to the power of -3 is applied to calculate the model outputs corresponding to irradiances of 800.001 and 799.999. The marginal response intensity is then obtained through difference, resulting in a result closer to the actual sensitivity characteristics.
[0057] When calculating partial derivatives of a differentiable model, numerical stabilization is required. A numerical stabilization term with values ranging from 10⁻⁶ to 10⁻⁸ can be used to avoid numerical overflow or anomalies during calculation. The boundary handling method for central difference calculation is as follows: if the irradiance is too low (below 0) or too high (above 2000) after applying a perturbation, a clamping method is used to limit the perturbed irradiance to a physically reasonable range of 0 to 2000 before performing the difference calculation. The specific value of the small perturbation step size can be adjusted according to the magnitude of the solar irradiance. When the typical irradiance range is 0 to 2000, a value of 10⁻³ is used; when the irradiance data accuracy is low, it can be adjusted to 5 multiplied by 10⁻³.
[0058] It should be noted that, considering the characteristics of different types of power prediction models, a universal and accurate method for calculating marginal response intensity is provided, adapting to the diversity of model selection in net load forecasting. Differentiable models are calculated directly using partial derivatives, taking efficiency into account; non-differentiable models employ the central difference method to ensure accuracy, while boundary treatment and numerical stabilization measures ensure computational reliability. The beneficial effect of this approach is that, regardless of whether the model is differentiable or not, reliable data reflecting the model's radiation sensitivity characteristics can be obtained.
[0059] Preferably, a critical factor for capturing the inverter's operating boundary is calculated based on the non-truncation probability, and sample training weights are constructed by combining the prediction residuals. The critical factor reaches its maximum value when the non-truncation probability indicates the inverter's switching between limiting and normal power generation states, including: Calculate the first Predicted residuals from round iterations With global robust scale : in, For the current model, For net load label, For median operations, It is a numerically stable term; Calculate the critical factor : in, The non-truncated probability, Emphasizing the index for the preset boundary and ; Calculate the iterative scheduling coefficients : in, For the maximum weighted range, For the current iteration round, This represents the total number of iterations. Calculate the training weights of the samples : The prediction residual of the k-th iteration is the difference between the output value of the power prediction model and the actual data of the net load of the distribution network in the k-th iteration, which is used to quantify the current prediction deviation of the model.
[0060] The global robustness metric for the k-th iteration is the result of adding the median of the absolute values of all predicted residuals in the k-th iteration to the numerical stability term. This metric is used to standardize the residuals and reduce the impact of outliers.
[0061] The critical factor is a parameter constructed based on the non-truncated probability and its complementary probability, used to enhance the weight of samples in the critical region of inverter state switching.
[0062] The boundary emphasis index is a custom parameter that adjusts the degree of emphasis the critical factor places on the state transition region. The preferred value is between 1 and 3. The value of 1 to 3 is chosen because this range effectively highlights samples in the critical region, avoiding both values that are too small (resulting in unclear boundaries) and values that are too large (resulting in excessive weight concentration).
[0063] The iterative scheduling coefficient is a coefficient that is dynamically adjusted with each iteration round and is used to control the degree of reinforcement of sample weights as the training process progresses.
[0064] The maximum weighting magnitude is the upper limit of the iterative scheduling coefficient. It is preferably between 1 and 5. The basis for this value is that this range can balance the adjustment intensity of the weights, avoiding the model overfitting due to excessive weights or the adjustment being ineffective due to excessively small weights.
[0065] The current iteration round is the sequence number of the training round that is currently in progress during the adaptive iterative training of the samples.
[0066] The total number of iterations is the preset total number of complete training iterations for adaptive iterative training of samples. It is preferably between 50 and 200, because this range can ensure that the model converges sufficiently, while avoiding excessive iterations that lead to low training efficiency.
[0067] The sample training weights are parameters used to weight the model prediction error, and are dynamically adjusted in combination with residual difficulty, state boundary, and iteration process.
[0068] Traditional sample weighting is often based solely on residual magnitude, making it susceptible to misleading high residual samples caused by inverter limiting. In this scheme, the critical factor is constructed by multiplying the non-truncated probability with its complementary probability, reaching its maximum at a non-truncated probability of 0.5, accurately capturing the state transition zone between inverter limiting and normal power generation. The iterative scheduling coefficient increases with each round, achieving a balance between global fitting in the early stages and dynamic adjustment of key samples in the later stages. Furthermore, the globally robust scale-standardized residuals prevent outliers from dominating the weights. For example, when the non-truncated probability is 0.5, the critical factor reaches its maximum. If, at this point, the standardized prediction residual is 2, the iterative scheduling coefficient is 0.8, the maximum weighting amplitude is 3, and the boundary emphasis index is 2, then the training weights of the samples will be significantly higher than those of ordinary samples, highlighting the importance of difficult examples in the critical region.
[0069] When calculating the globally robust scalar, the numerical stability term ranges from 10⁻⁶ to 10⁻⁸ to ensure computational stability. No additional outlier removal is needed during the globally robust scalar calculation, as the median itself has outlier resistance properties and can be directly calculated on all prediction residuals in the k-th round. The total number of iterations can be determined in conjunction with the model convergence criterion: training can be terminated early when the change in prediction error on the validation set is less than 0.01 over 10 consecutive rounds; otherwise, training continues for the preset total number of iterations. The specific value of the boundary emphasis index can be adjusted according to the frequency of state transitions in the data: 2 to 3 for frequent transitions, and 1 to 2 for infrequent transitions.
[0070] It's important to note that we construct dynamically adaptable training weights for high-penetration distributed energy scenarios, thus avoiding the shortcomings of traditional static weights that only focus on residuals and ignore state transition regions. We strengthen key samples for state transitions through critical factors, iteratively schedule coefficients to adapt to the training process, and standardize residuals to reduce outlier interference, achieving both structured and dynamic weights. This allows sample weights to accurately focus on key samples with high residuals within the state transition critical region, preventing the model from being misled by false high-residual samples.
[0071] Preferably, a photovoltaic bistable constraint term is constructed, using the non-truncation probability to constrain the marginal response intensity to approach the response mode reference value under normal power generation conditions and to approach the truncation mode reference value under inverter limiting conditions, including: Calculate single-point sensitivity loss : in, The non-truncated probability, The marginal response strength, The reference value for the response mode is... The cut-off modal reference value; Calculate the photovoltaic bistable constraint term : in, This represents the total number of time steps. For the first Model parameters for each iteration.
[0072] Single-point sensitivity loss is a loss value that quantifies the degree of deviation between the marginal response strength of the power prediction model and the corresponding state baseline value at a single time step.
[0073] The normal power generation response loss is the square of the difference between the marginal response strength and the baseline value of the response mode, weighted by non-truncation probability, and is used to measure the sensitivity deviation under normal power generation conditions.
[0074] The inverter limiting response loss is the square of the difference between the marginal response strength and the cutoff mode reference value, weighted by the complementary probabilities of the non-cutoff probabilities, and is used to measure the sensitivity deviation of the inverter under limiting conditions.
[0075] The photovoltaic bistable constraint term is the sum of all single-point sensitivity losses over the entire time series. It is used to impose bistable constraints during model training to guide the model to adapt to two operating states.
[0076] Traditional constraints often employ a globally uniform approach, which is ill-suited to the bistable characteristics of high-penetration distributed energy resources. This scheme uses non-truncation probabilities and their complementary probabilities to weight the sensitivity deviations between the two states, achieving a differentiated design that strengthens irradiance sensitivity constraints during normal power generation and relaxes them during limited-amplitude states. For example, when the non-truncation probability is 0.9, the marginal response strength of the key constraints approximates the baseline value of the response mode; when the non-truncation probability is 0.1, the marginal response strength of the key constraints approximates the baseline value of the truncated mode, allowing the constraints to accurately match the actual operating state of the sample.
[0077] The calculation scope of the photovoltaic bistable constraint term is the entire time series, covering all time steps without omissions or selective calculations. If extreme outliers occur in the marginal response strength, no additional handling is required, as the calculation process for the marginal response strength has already ensured robustness through designs such as central difference and numerical stability terms, and the squared form of the single-point sensitivity loss can suppress the impact of extreme values to a certain extent. The weighting logic of the two response losses does not need to be adjusted; the sum of the non-truncated probability and its complementary probability is always 1, which naturally balances the constraint strength of the two states.
[0078] It should be noted that, considering the bistable characteristics of high-penetration distributed energy, state-based constraint terms are constructed to guide the model to exhibit reasonable irradiance sensitivity under different operating conditions, avoiding adaptation imbalances caused by global constraints. Through non-truncated probability weighting, the constraint terms accurately match the state assignment of samples, specifically correcting sensitivity biases. This allows the photovoltaic bistable constraint terms to directly combat the irradiance gradient collapse problem, enabling the model to maintain its irradiance response capability under normal power generation conditions and adapt to saturation characteristics under limited conditions.
[0079] Preferably, the power prediction model is updated with parameters based on an optimization objective function composed of a power prediction error term weighted by the sample training weights and a photovoltaic bistable constraint term, including: Calculate the power prediction error term : in, This represents the total number of time steps. For the first The training weights of the samples in each iteration, For the first The predicted residual of each iteration; Calculate the constraint strength coefficient : in, The preset maximum constraint strength, For the current iteration round, This represents the total number of iterations. Construct the optimization objective function : in, This refers to the photovoltaic bistable constraint term; Update the parameters of the power prediction model: in, For the updated model parameters, These are the current model parameters. The preset learning rate, The gradient of the optimization objective function is given.
[0080] The power prediction error term is a value obtained by weighting the squares of the prediction residuals with the sample training weights. It is used to quantify the degree of deviation between the model prediction results and the actual net load.
[0081] The constraint strength coefficient is a coefficient that is dynamically adjusted as the iteration process progresses. It is used to control the weight of the photovoltaic bistable constraint term in the optimization objective function.
[0082] The maximum constraint strength is the upper limit of the constraint strength coefficient. It is preferably between 0.1 and 1.0, because this range can balance the optimization of prediction error and structural constraints, avoiding excessively strong constraints that would increase prediction bias or excessively weak constraints that would fail to function.
[0083] The objective function, including the power prediction error term and the weighted photovoltaic bistable constraint term, serves as the basis for optimizing the model parameters.
[0084] The learning rate is a custom parameter that controls the step size of model parameter updates. It is preferably between 0.001 and 0.01, because this range can ensure stable convergence of parameter updates and avoid oscillations caused by excessively large step sizes or inefficient training caused by excessively small step sizes.
[0085] The current model parameters are internal variables of the power prediction model in the current iteration process, and are the basis for parameter updates.
[0086] The updated model parameters are new variables obtained by adjusting the current model parameters using gradient descent, and are used for the next iteration or final prediction.
[0087] Traditional optimization techniques often employ fixed constraint strengths or single objective functions, which are ill-suited to the training requirements of high-penetration distributed energy scenarios. In this approach, the constraint strength coefficient monotonically increases with each iteration, enabling the model to focus on optimizing prediction errors in the early stages, allowing it to grasp the basic load patterns; and then strengthening bistable constraints in later stages, guiding the model to adapt to dynamic training logic that adapts to state characteristics. The optimization objective function integrates weighted prediction errors with bistable constraints, preventing the model from falling into a pseudo-optimal state by solely pursuing minimum error.
[0088] Gradient calculations should preferably use full-batch gradient descent. If the data volume is too large, mini-batch gradient descent can be used, with a batch size of 32 to 128, balancing computational efficiency and gradient estimation accuracy. Gradient pruning should be introduced during gradient calculations, with a pruning threshold of 1 to 5 to avoid gradient explosion leading to abnormal parameter updates. The learning rate can be used in conjunction with momentum gradient descent, with a momentum coefficient of 0.9 to improve convergence stability. The specific value of the maximum constraint strength can be adjusted according to the model type: 0.3 to 1.0 for deep learning models and 0.1 to 0.5 for traditional statistical models, adapting to the fitting characteristics of different models.
[0089] It's important to note that a dynamically adapted optimization objective function is constructed to integrate data-driven prediction error optimization with structure-driven bistable constraints. This guides the model to balance accuracy and rationality during iterations, avoiding pseudo-optimal approaches that ignore irradiance characteristics. The dynamic constraint strength coefficients adapt to the model's learning patterns, while the learning rate ensures stable updates. This allows the model to rapidly reduce prediction bias in the early stages of training and gradually strengthen the bistable structure adaptation in later stages. This ensures both the accuracy of net load prediction and avoids irradiance gradient collapse, guaranteeing that the model exhibits reasonable response characteristics under different operating conditions.
[0090] like Figure 2 As shown, Figure 2 This demonstration showcases the connection architecture of user-side equipment, data transmission, and model optimization systems in a high-penetration distributed energy scenario: On the left is the distribution network feeder / transformer area, which connects to multiple user units. Each unit includes downstream photovoltaic (PV) power, user load, inverters with limiting / reduction functions, and smart meters. Downstream PV provides distributed generation, the inverter is responsible for AC-DC conversion of PV power and can limit power according to operational needs, and the smart meter collects the net load data of the unit. Simultaneously, solar irradiance sensors and ambient temperature sensors are deployed within the scenario to collect irradiance and ambient temperature data, respectively. These net load, irradiance, and ambient temperature data are transmitted to the prediction and model optimization server on the right via an edge gateway / communication network. This server performs net load prediction and corresponding model optimization based on this data, forming a complete link from user-side equipment to data acquisition, transmission, and prediction optimization.
[0091] Example 2: A sample adaptive iterative training system for optimizing existing models, applied to any of the sample adaptive iterative training methods for optimizing existing models, including: The data acquisition module is used to acquire net load data and solar irradiance data of the distribution network as time series samples; The index determination module is used to determine the net load-irradiance response intensity index based on the response slope of the net load data relative to the solar irradiance data within a local window. The state probability calculation module is used to determine the cutoff mode reference value representing inverter limiting and the response mode reference value representing normal power generation by using the statistical distribution of the response intensity index, and calculate the non-cutoff probability of the sample being in the normal power generation state accordingly. The model response calculation module is used to calculate the marginal response strength of the power prediction model to the solar irradiance data; The weight construction module is used to calculate the critical factor for capturing the inverter operation boundary based on the non-truncation probability, and to construct sample training weights in combination with the prediction residual, wherein the critical factor takes a maximum value when the non-truncation probability indicates the inverter switching between limiting and normal power generation states. The constraint construction module is used to construct photovoltaic bistable constraint terms. Using the non-truncation probability, the marginal response intensity is constrained to approach the response mode reference value under normal power generation conditions and to approach the truncation mode reference value under inverter limiting conditions. The parameter update module is used to update the parameters of the power prediction model based on the optimization objective function composed of the power prediction error term weighted by the sample training weights and the photovoltaic bistable constraint term.
[0092] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.
Claims
1. A sample-adaptive iterative training method for optimizing existing models, characterized in that, include: Obtain net load data and solar irradiance data of the distribution network as time series samples; Based on the response slope of the net load data relative to the solar irradiance data within a local window, the net load-irradiance response intensity index is determined. The statistical distribution of the response intensity index is used to determine the cutoff mode reference value representing inverter limiting and the response mode reference value representing normal power generation, and the non-cutoff probability of the sample being in normal power generation state is calculated accordingly. Calculate the marginal response strength of the power prediction model to the solar irradiance data; The critical factor for capturing the inverter's operating boundary is calculated based on the non-truncation probability, and the sample training weights are constructed by combining the prediction residuals. The critical factor takes a maximum value when the non-truncation probability indicates that the inverter is switching between limiting and normal power generation states. A photovoltaic bistable constraint term is constructed, and the non-truncation probability is used to constrain the marginal response intensity to approach the response mode reference value under normal power generation conditions and to approach the truncation mode reference value under inverter limiting conditions. The parameters of the power prediction model are updated based on the optimization objective function composed of the power prediction error term weighted by the sample training weights and the photovoltaic bistable constraint term.
2. The sample adaptive iterative training method for optimizing existing models according to claim 1, characterized in that, Obtain net load data and solar irradiance data of the distribution network as time series samples, including: Acquire the net load data of the distribution network collected in time steps, and the solar irradiance data; Construct a feature vector that includes the solar irradiance data, ambient temperature data, and calendar information; A power prediction model is established with the feature vector as input and the net load data of the distribution network as the prediction target, and the model parameters of the power prediction model are initialized.
3. The sample adaptive iterative training method for optimizing existing models according to claim 1, characterized in that, Based on the response slope of the net load data relative to the solar irradiance data within a local window, a net load-irradiance response intensity index is determined, including: A local time window is defined with the current time as the center, and the window mean of the solar irradiance data and the net load data of the distribution network within the local time window is calculated; Subtract the corresponding window mean from each data point within the local time window to obtain the centralized irradiance data and the centralized net load data, respectively. The sum of the products of the centralized irradiance data and the centralized net load data within the local time window is calculated and divided by the sum of the squares of the centralized irradiance data and the sum of the preset numerical stability terms to obtain the local regression slope. By inverting the local regression slope, the net load-irradiance response intensity index is obtained.
4. The sample adaptive iterative training method for optimizing existing models according to claim 1, characterized in that, The statistical distribution of the response intensity index is used to determine the cutoff mode reference value representing inverter limiting and the response mode reference value representing normal power generation, and the non-cutoff probability of the sample being in normal power generation state is calculated accordingly, including: Based on the saturation characteristics of the inverter under limiting conditions, the cutoff mode reference value is set to zero; Calculate the global median and absolute median difference of the net load-irradiance response intensity index; Using the logistic function and the absolute median difference to construct weighting coefficients, the robust weighting center of the net load-irradiance response intensity index is calculated and used as the reference value of the response mode; Calculate the squared normalized Euclidean distance between the net load-irradiance response intensity index at the current moment and the cutoff mode reference value and the response mode reference value, respectively, to obtain the cutoff state energy value and the response state energy value; Based on the cutoff state energy value and the response state energy value, the non-cutoff probability is calculated using an exponential normalization function that includes a temperature parameter.
5. The sample adaptive iterative training method for optimizing existing models according to claim 1, characterized in that, Calculating the marginal response strength of the power prediction model to the solar irradiance data includes: When the power prediction model is differentiable with respect to the solar irradiance data, the partial derivative of the output of the power prediction model with respect to the solar irradiance data is calculated as the marginal response strength; When the power prediction model is not differentiable with respect to the solar irradiance data, the central difference of the power prediction model at the current moment is calculated using a preset small perturbation step size, which is used as the marginal response strength.
6. The sample adaptive iterative training method for optimizing existing models according to claim 1, characterized in that, The critical factor for capturing the inverter's operating boundary is calculated based on the non-truncation probability, and the sample training weights are constructed by combining the prediction residuals. The critical factor reaches its maximum value when the non-truncation probability indicates the inverter switching between limiting and normal power generation states, including: The difference between the output value of the power prediction model and the net load data of the distribution network is calculated to obtain the prediction residual, and the global robustness scale of the prediction residual is calculated. The critical factor is obtained by multiplying the non-truncated probability and its complementary probability, and by performing a power operation using a preset boundary emphasis exponent, such that the critical factor reaches its maximum when the non-truncated probability is 50%. The iteration scheduling coefficient is calculated based on the ratio of the current iteration round to the total number of iteration rounds. The sample training weights are obtained by multiplying the iterative scheduling coefficients, the critical factor, and the absolute value of the prediction residuals standardized by the global robust scale, and then adding the basic weight constants.
7. The sample adaptive iterative training method for optimizing existing models according to claim 1, characterized in that, A photovoltaic bistable constraint term is constructed, and using the non-truncation probability, the marginal response intensity is constrained to approach the response mode reference value under normal power generation conditions and to approach the truncation mode reference value under inverter limiting conditions, including: Calculate the square of the difference between the marginal response intensity and the response mode reference value, and use the non-truncated probability for weighting to obtain the normal power generation response loss; The square of the difference between the marginal response strength and the cut-off mode reference value is calculated, and the complementary probabilities of the non-cut-off probabilities are used for weighting to obtain the inverter limiting response loss. The normal power generation response loss is added to the inverter limiting response loss to obtain the single-point sensitivity loss; The photovoltaic bistable constraint term is obtained by summing the single-point sensitivity loss over the entire time series.
8. The sample adaptive iterative training method for optimizing existing models according to claim 1, characterized in that, Based on the optimization objective function composed of the power prediction error term weighted by the sample training weights and the photovoltaic bistable constraint term, the parameters of the power prediction model are updated, including: The power prediction error term is obtained by weighting and summing the squares of the prediction residuals using the sample training weights. Based on the ratio of the current iteration round to the total number of iteration rounds, and combined with the preset maximum constraint strength, calculate the constraint strength coefficient that monotonically increases with the iteration process. The power prediction error term is added to the photovoltaic bistable constraint term weighted by the constraint strength coefficient to obtain the optimization objective function; Calculate the gradient of the optimization objective function with respect to the current model parameters, and update the parameters of the power prediction model in the opposite direction of the gradient using a preset learning rate.
9. A sample adaptive iterative training system for optimizing existing models, applied in the sample adaptive iterative training method for optimizing existing models as described in any one of claims 1-8, characterized in that, include: The data acquisition module is used to acquire net load data and solar irradiance data of the distribution network as time series samples; The index determination module is used to determine the net load-irradiance response intensity index based on the response slope of the net load data relative to the solar irradiance data within a local window. The state probability calculation module is used to determine the cutoff mode reference value representing inverter limiting and the response mode reference value representing normal power generation by using the statistical distribution of the response intensity index, and calculate the non-cutoff probability of the sample being in the normal power generation state accordingly. The model response calculation module is used to calculate the marginal response strength of the power prediction model to the solar irradiance data; The weight construction module is used to calculate the critical factor for capturing the inverter operation boundary based on the non-truncation probability, and to construct sample training weights in combination with the prediction residual, wherein the critical factor takes a maximum value when the non-truncation probability indicates the inverter switching between limiting and normal power generation states. The constraint construction module is used to construct photovoltaic bistable constraint terms. Using the non-truncation probability, the marginal response intensity is constrained to approach the response mode reference value under normal power generation conditions and to approach the truncation mode reference value under inverter limiting conditions. The parameter update module is used to update the parameters of the power prediction model based on the optimization objective function composed of the power prediction error term weighted by the sample training weights and the photovoltaic bistable constraint term.
Citation Information
Cited By
Method for optimizing minimum fuel consumption performance of an aero-engine based on flight speed closed loop
CN122215944A