A method for generating extreme power load samples
By identifying and generating models for industrial load disturbances, the problem of identifying and generating extreme power load patterns under small sample conditions is solved, generating realistic extreme load samples and improving the quality of training data and the interpretability of scheduling decisions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 成都亿成科技有限公司
- Filing Date
- 2025-12-24
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies struggle to effectively identify extreme power load patterns and generate reasonable extreme load samples under small sample conditions. In particular, in small power grid scenarios, traditional methods struggle to distinguish industrial disturbance components, resulting in data that lacks specificity and interpretability.
An industrial load disturbance identification model is used to separate the base load and industrial disturbance components. Combined with multi-scale time-frequency features and cluster analysis, a diffusion generation model is used to generate realistic extreme load samples. Multi-scale identification, constraint generation and disturbance separation are integrated into a unified process.
Generating realistic extreme load samples under small sample conditions improves the quality of training data and the interpretability of scheduling decisions, demonstrating good engineering usability and potential for widespread application.
Smart Images

Figure CN122087445A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power load extreme sample generation technology, and specifically relates to a method for generating power load extreme samples. Background Technology
[0002] As power systems evolve towards intelligence and precision, load forecasting models need to maintain high accuracy in increasingly complex and variable scenarios. However, for urban or regional power grids with relatively small load volumes and significant random fluctuations, load curves are frequently affected by holiday effects, extreme weather, and fluctuations in industrial user production activities, exhibiting extreme change patterns far removed from normal. These extreme load events occur infrequently and for short periods, resulting in scarce historical data samples. Traditional forecasting models are prone to a sharp decline in accuracy when encountering such situations. For example, sudden drops in electricity consumption during major holidays, temporary load declines caused by heavy rain, or load jumps triggered by sudden shutdowns of large industrial installations may all exceed the model's empirical range, leading to forecasting biases and scheduling risks.
[0003] In existing technologies, researchers are attempting to improve models' ability to capture abnormal fluctuations by performing frequency domain decomposition or feature extraction on historical load sequences. On the other hand, in the field of data augmentation, techniques such as Generative Adversarial Networks (GANs) are used to generate and repair power data to compensate for insufficient data. However, these methods still have limitations: decomposition-based prediction methods primarily focus on improving prediction accuracy and do not address the lack of extreme samples; while generative models such as GANs can synthesize data, their training process is unstable and often lacks consideration for physical constraints and event interpretability, potentially generating unreasonable load curves. Especially in small power grid scenarios, fluctuations in a single large industrial load can significantly affect the total load. Existing models struggle to distinguish the industrial disturbance components within the total load, resulting in data lacking specificity and thus limiting the model's effectiveness in extreme situations.
[0004] In summary, a new technical solution is urgently needed to effectively identify extreme power load patterns and generate more reasonable extreme load samples under small sample conditions. On the one hand, it is necessary to integrate knowledge from the power sector (such as load cyclical patterns and physical constraints) with artificial intelligence generation models to ensure that the generated data is both diverse and physically plausible. On the other hand, it is necessary to introduce a specialized processing mechanism for industrial load disturbances to improve the interpretability of the generated samples, enabling them to clearly reflect the impact of specific disturbance sources on the load. This will provide more comprehensive training and exercise data support for load forecasting models and dispatching decisions. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method for generating extreme power load samples, so as to at least solve some of the above-mentioned technical problems.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: like Figure 1 As shown, a method for generating extreme power load samples includes the following steps: Step 1: Obtain and preprocess the original load sequence to obtain the preprocessed original load sequence; use the industrial load disturbance identification model to separate the preprocessed original load sequence to obtain the base load sequence and industrial disturbance component; slice the base load sequence by day to obtain multiple daily base load sequences; the multiple daily base load sequences form a daily base load sequence set; construct external covariates; Step 2: Process the daily basic load sequence set using the discrete wavelet decomposition method to obtain the low-frequency smoothed trend sequence and multiple scale detail sequences under each daily basic load sequence; construct a multi-scale time-frequency feature vector based on the multiple scale detail sequences under each daily basic load sequence. Step 3: Perform density peak clustering analysis on the daily basic load sequence set based on multi-scale time-frequency feature vectors to obtain extreme load pattern samples; Step 4: Use dynamic density clustering algorithm to perform fine-grained clustering on the time segments corresponding to the extreme load pattern samples to obtain the initial time segment set of extreme events; use Markov random field to optimize the initial time segment set to obtain the spatiotemporal template of extreme events. Step 5: Design a diffusion generation model based on the spatiotemporal template of extreme events, and train the diffusion generation model using extreme load pattern samples; use the trained diffusion generation model to generate multiple load sequences, and superimpose the multiple load sequences with their corresponding industrial disturbance components to obtain extreme load samples.
[0007] Further, step 1 includes: The original load sequence is acquired and preprocessed to obtain the preprocessed original load sequence. The preprocessing includes missing value imputation, outlier removal, and normalization. An industrial load disturbance identification model is used to separate the preprocessed original load sequence to obtain the base load sequence. and industrial disturbance components i(t) ; Base load sequence Slicing by day yields multiple daily basic load sequences; these multiple daily basic load sequences form a daily basic load sequence set. , ; in, For extreme day indexes, This represents the total number of extreme days. This represents the number of time segments contained in a single day's basic load sequence. For the first The base load value for the first time segment in a single-day base load sequence. For the first The first of the daily basic load sequences Baseline load values for each time segment; Constructing external covariates External covariates include calendar and weather information.
[0008] Furthermore, in step 1, the industrial load disturbance identification model adopts a Transformer architecture based on a multi-layer self-attention structure, and the minimization objective function value of the industrial load disturbance identification model is... L The solution formula is: ; in, ; ; A set of discrete-time indices within a sample window. For time step index; To observe the total load; Background load component; Industrial disturbance component; , These are the time series vectors of the background load component and the industrial disturbance component, respectively; This is the set of trainable parameters for an industrial load disturbance identification model. These are the weighting coefficients for the industrial disturbance smoothing regularization term. ; These are the weighting coefficients of the relevance constraint term. ; The number of selected industrial disturbance indicators; For the first Time series of industrial disturbance indicators; for and Correlation loss between them , They are respectively , exist The sample mean above; Pearson correlation coefficient; Industrial disturbance component i(t) The dynamic constraint formula that is satisfied is: ;in, This is the upper limit of the rate. For industrial load disturbance identification model in t- Industrial disturbance components extracted at time 1.
[0009] Further, step 2 includes: Step 21: Perform analysis on the daily basic load sequence set. Layered discrete wavelet decomposition yields the low-frequency smoothed trend sequence and J-scale detail sequences for each daily basic load sequence; the discrete wavelet decomposition formula is: ; in, For the first d A low-frequency smoothed trend sequence under a single-day basic load series. For the first d The first daily basic load sequence Layer-scale detail sequence, , For time step index; The number of discrete wavelet decomposition layers; For wavelet decomposition scale index; Step 22: Based on J scale detail sequences under each daily base load sequence, calculate the wavelet energy entropy under each daily base load sequence. and energy percentage ; Calculate wavelet energy entropy The formula is: ;in, For the first Energy percentage of each scale detail sequence It is a logarithmic function; The calculation formula is: ;in, For the first Energy of a sequence of details at a scale k The index variable for summation; The calculation formula is: ; Step 23: Based on wavelet energy entropy and energy percentage Construct multi-scale time-frequency feature vectors ; ,in, This represents the energy percentage of the detail sequence at scales 1 to J. For the first d The energy entropy value of a single-day basic load sequence. For dimension The real vector space.
[0010] Furthermore, step 3 includes: Step 31: Define the composite distance metric Suppose the size of the set of daily basic load sequences to be analyzed is... The index set of the daily basic load sequence is ;in, These represent the indices of two single-day basic load sequences; For the first A multi-scale time-frequency feature vector of a single-day basic load sequence; For the first External covariates of a single-day basic load sequence; Composite distance metric for: ;in, For the first A single-day basic load sequence. The Mahalanobis distance, It is a diagonal matrix. As the first weighting coefficient, This is the second weighting coefficient. For the first External covariates of a single-day basic load sequence It is the Euclidean norm; Step 32: Based on composite distance metric Density peak clustering algorithm is used to... Analyzing a single-day basic load sequence, the number of output clusters is: , No. The cluster number to which each daily basic load sequence belongs is: , .
[0011] Furthermore, step 3 also includes: Step 33: Select the cluster with the most basic load sequences on a single day as the master cluster. , morphological feature center for: For any cluster ( ), its morphological center for: ; morphological distance Defined as: The mean and standard deviation of the morphological distances of all clusters were calculated as follows: and ,when At that time, it was considered that the cluster The curve shape deviates significantly from the main cluster, among which, These are empirical parameters; Step 34, Primary Cluster External covariate feature center for: For any cluster , External covariate feature center for: external covariate distance for: The mean and standard deviation of the external covariate distances for all clusters were calculated as follows: and ;when At that time, it was considered that the cluster The distribution of external covariates differs significantly from that of the main cluster, among which, These are empirical parameters; Step 35: Simultaneously satisfying and The cluster number set is used as the sample set of extreme load patterns. , Extreme load pattern sample index set for: .
[0012] Furthermore, step 4 includes: Step 41, the process of using dynamic density clustering algorithm to perform fine-grained clustering of time series segments corresponding to extreme load pattern samples to obtain the initial set of time series segments for extreme events includes: on each extreme day Calculate the first-order difference. The formula is: ; with a preset threshold Forming a candidate mutation set , To obtain the initial time series fragment set of extreme events. , ,in, Represents the norm, Representing different subtypes; Step 42, the process of optimizing the initial time series segment set using Markov random fields to obtain the spatiotemporal template of extreme events includes: optimizing the initial time series segment set... Modeling the sequence as a one-dimensional Markov random field and minimizing the energy yields a smoothed set of initial time series segments; for the same subtype The events are time-normalized and aligned to obtain the template curve. The event amplitude, downlink rate, recovery rate, and duration are collected into a parameter vector. ; and The spatiotemporal template that constitutes extreme events; energy The calculation formula is: ,in, It is a single-point potential. They are adjacent potentials.
[0013] Furthermore, in step 5, the forward diffusion process of the diffusion generation model is defined using a Markov chain; the reverse diffusion process of the diffusion generation model uses a multi-layer deep neural network as a noise estimator. The forward diffusion process is as follows: ; in, Let be the probability density function. For diffusion time step index, For the first The load sequence state at each time step. For the first The load sequence state at each time step. The mean is Covariance is The multivariate Gaussian distribution, This is the noise variance hyperparameter. It is the identity matrix; ;in, The cumulative scaling factor. , ; To obtain from the standard normal distribution Random noise in the mid-sample, This is a sample of extreme load patterns.
[0014] Furthermore, the noise estimator uses Predict injected noise; For conditional vectors, , A one-hot indicator vector for the event type; Objective function of diffusion generation model for: ; in, For mean square error loss, ; [] represents extreme load mode samples ,noise and time step The joint distribution expectation; This is the third weighting coefficient. It is the fourth weighting coefficient. This is the first penalty term, used to suppress the generation of negative load values. The second penalty term is used to constrain the ramp rate of the generated load sequence from being too fast, ensuring that the load change conforms to the laws of engineering physics. The load ramp-up increments between adjacent time steps. The maximum allowed ramp threshold; After the diffusion generation model training converges, sampling starts from... By gradually working backward, we can obtain The extreme load samples are spliced together according to the event positions on the timeline within the extreme day, and the transition between segments is smoothed by the first-order difference continuity criterion. The corresponding industrial disturbance components are then superimposed to obtain the extreme load samples.
[0015] Compared with the prior art, the present invention has the following beneficial effects: This invention employs a diffusion-generative model to learn the distribution of extreme events and synthesizes realistic load curves under small sample conditions. Compared with traditional generative adversarial networks, it has more stable training and more comprehensive pattern coverage.
[0016] This invention employs an industrial load disturbance identification model to separate the preprocessed original load sequence and obtain industrial disturbance components; these components are then recovered in a superposition manner during the generation stage to preserve traceable industrial characteristics.
[0017] This invention integrates multi-scale identification, constraint generation, and disturbance separation into a unified process that supports scheduling simulation, model training, and risk assessment. It has good engineering usability and prospects for promotion, and solves the technical problem that existing technologies cannot generate a large number of reasonable extreme load samples. Attached Figure Description
[0018] Figure 1 This is a flowchart of the steps of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0020] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0021] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; of course, they can also refer to a mechanical connection or an electrical connection; furthermore, they can refer to a direct connection, an indirect connection through an intermediate medium, or a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0022] like Figure 1 As shown, a method for generating extreme power load samples includes the following steps: Step 1: Obtain and preprocess the original load sequence to obtain the preprocessed original load sequence; use the industrial load disturbance identification model to separate the preprocessed original load sequence to obtain the base load sequence and industrial disturbance component; slice the base load sequence by day to obtain multiple daily base load sequences; the multiple daily base load sequences form a daily base load sequence set; construct external covariates; Step 2: Process the daily basic load sequence set using the discrete wavelet decomposition method to obtain the low-frequency smoothed trend sequence and multiple scale detail sequences under each daily basic load sequence; construct a multi-scale time-frequency feature vector based on the multiple scale detail sequences under each daily basic load sequence. Step 3: Perform density peak clustering analysis on the daily basic load sequence set based on multi-scale time-frequency feature vectors to obtain extreme load pattern samples; Step 4: Use dynamic density clustering algorithm to perform fine-grained clustering on the time segments corresponding to the extreme load pattern samples to obtain the initial time segment set of extreme events; use Markov random field to optimize the initial time segment set to obtain the spatiotemporal template of extreme events. Step 5: Design a diffusion generation model based on the spatiotemporal template of extreme events, and train the diffusion generation model using extreme load pattern samples; use the trained diffusion generation model to generate multiple load sequences, and superimpose the multiple load sequences with their corresponding industrial disturbance components to obtain extreme load samples.
[0023] Preferably, step 1 includes: The original load sequence is acquired and preprocessed to obtain the preprocessed original load sequence. The preprocessing includes missing value imputation, outlier removal, and normalization. An industrial load disturbance identification model is used to separate the preprocessed original load sequence to obtain the base load sequence. and industrial disturbance components i(t) ; Base load sequence Slicing by day yields multiple daily basic load sequences; these multiple daily basic load sequences form a daily basic load sequence set. , ; in, For extreme day indexes, This represents the total number of extreme days. This represents the number of time segments contained in a single day's basic load sequence. For the first The base load value for the first time segment in a single-day base load sequence. For the first The first of the daily basic load sequences Baseline load values for each time segment; Constructing external covariates External covariates include calendar and weather information.
[0024] Preferably, in step 1, the industrial load disturbance identification model adopts a Transformer architecture based on a multi-layer self-attention structure, and the minimization objective function value of the industrial load disturbance identification model is... L The solution formula is: ; in, ; ; A set of discrete-time indices within a sample window. For time step index; To observe the total load; Background load component; Industrial disturbance component; , These are the time series vectors of the background load component and the industrial disturbance component, respectively; This is the set of trainable parameters for an industrial load disturbance identification model. These are the weighting coefficients for the industrial disturbance smoothing regularization term. ; These are the weighting coefficients of the relevance constraint term. ; The number of selected industrial disturbance indicators; For the first Time series of industrial disturbance indicators; for and Correlation loss between them , They are respectively , exist The sample mean above; Pearson correlation coefficient; Industrial disturbance component i(t) The dynamic constraint formula that is satisfied is: ;in, This is the upper limit of the rate. For industrial load disturbance identification model in t- Industrial disturbance components extracted at time 1.
[0025] Preferably, step 2 includes: Step 21: Perform analysis on the daily basic load sequence set. Layered discrete wavelet decomposition yields the low-frequency smoothed trend sequence and J-scale detail sequences for each daily basic load sequence; the discrete wavelet decomposition formula is: ; in, For the first d A low-frequency smoothed trend sequence under a single-day basic load series. For the first d The first daily basic load sequence Layer-scale detail sequence, , For time step index; The number of discrete wavelet decomposition layers; For wavelet decomposition scale index; Step 22: Based on J scale detail sequences under each daily base load sequence, calculate the wavelet energy entropy under each daily base load sequence. and energy percentage ; Calculate wavelet energy entropy The formula is: ;in, For the first Energy percentage of each scale detail sequence It is a logarithmic function; The calculation formula is: ;in, For the first Energy of a sequence of details at a scale k The index variable for summation; The calculation formula is: ; Step 23: Based on wavelet energy entropy and energy percentage Construct multi-scale time-frequency feature vectors ; ,in, This represents the energy percentage of the detail sequence at scales 1 to J. For the first d The energy entropy value of a single-day basic load sequence. For dimension The real vector space.
[0026] Preferably, step 3 includes: Step 31: Define the composite distance metric Suppose the size of the set of daily basic load sequences to be analyzed is... The index set of the daily basic load sequence is ;in, These represent the indices of two single-day basic load sequences; For the first A multi-scale time-frequency feature vector of a single-day basic load sequence; For the first External covariates of a single-day basic load sequence; Composite distance metric for: ;in, For the first A single-day basic load sequence. The Mahalanobis distance, It is a diagonal matrix. As the first weighting coefficient, This is the second weighting coefficient. For the first External covariates of a single-day basic load sequence It is the Euclidean norm; Step 32: Based on composite distance metric Density peak clustering algorithm is used to... Analyzing a single-day basic load sequence, the number of output clusters is: , No. The cluster number to which each daily basic load sequence belongs is: , .
[0027] Step 33: Select the cluster with the most basic load sequences on a single day as the master cluster. , morphological feature center for: For any cluster ( ), its morphological center for: ; morphological distance Defined as: The mean and standard deviation of the morphological distances of all clusters were calculated as follows: and ,when At that time, it was considered that the cluster The curve shape deviates significantly from the main cluster, among which, These are empirical parameters; Step 34, Primary Cluster External covariate feature center for: For any cluster , External covariate feature center for: external covariate distance for: The mean and standard deviation of the external covariate distances for all clusters were calculated as follows: and ;when At that time, it was considered that the cluster The distribution of external covariates differs significantly from that of the main cluster, among which, These are empirical parameters; Step 35: Simultaneously satisfying and The cluster number set is used as the sample set of extreme load patterns. , Extreme load pattern sample index set for: .
[0028] Preferably, step 4 includes: Step 41, the process of using dynamic density clustering algorithm to perform fine-grained clustering of time series segments corresponding to extreme load pattern samples to obtain the initial set of time series segments for extreme events includes: on each extreme day Calculate the first-order difference. The formula is: ; with a preset threshold Forming a candidate mutation set , To obtain the initial time series fragment set of extreme events. , ,in, Represents the norm, Representing different subtypes; Step 42, the process of optimizing the initial time series segment set using Markov random fields to obtain the spatiotemporal template of extreme events includes: optimizing the initial time series segment set... Modeling the sequence as a one-dimensional Markov random field and minimizing the energy yields a smoothed set of initial time series segments; for the same subtype The events are time-normalized and aligned to obtain the template curve. The event amplitude, downlink rate, recovery rate, and duration are collected into a parameter vector. ; and The spatiotemporal template that constitutes extreme events; energy The calculation formula is: ,in, It is a single-point potential. They are adjacent potentials.
[0029] Preferably, in step 5, the forward diffusion process of the diffusion generation model is defined using a Markov chain; the reverse diffusion process of the diffusion generation model uses a multi-layer deep neural network as a noise estimator. The forward diffusion process is as follows: ; in, Let be the probability density function. For diffusion time step index, For the first The load sequence state at each time step. For the first The load sequence state at each time step. The mean is Covariance is The multivariate Gaussian distribution, This is the noise variance hyperparameter. It is the identity matrix; ;in, The cumulative scaling factor. , ; To obtain from the standard normal distribution Random noise in the mid-sample, This is a sample of extreme load patterns.
[0030] Noise estimator Predict injected noise; For conditional vectors, , A one-hot indicator vector for the event type; Objective function of diffusion generation model for: ; in, For mean square error loss, ; [] represents extreme load mode samples ,noise and time step The joint distribution expectation; This is the third weighting coefficient. It is the fourth weighting coefficient. This is the first penalty term, used to suppress the generation of negative load values. The second penalty term is used to constrain the ramp rate of the generated load sequence from being too fast, ensuring that the load change conforms to the laws of engineering physics. The load ramp-up increments between adjacent time steps. The maximum allowed ramp threshold; After the diffusion generation model training converges, sampling starts from... By gradually working backward, we can obtain The extreme load samples are spliced together according to the event positions on the timeline within the extreme day, and the transition between segments is smoothed by the first-order difference continuity criterion. The corresponding industrial disturbance components are then superimposed to obtain the extreme load samples.
[0031] Finally, it should be noted that the above embodiments are merely preferred embodiments of the present invention used to illustrate the technical solutions of the present invention, and are not intended to limit the invention, nor are they intended to limit the patent scope of the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention. That is to say, any changes or refinements made to the main design concept and spirit of the present invention that are not of substantial significance, but whose technical problems are still consistent with the present invention, should be included within the protection scope of the present invention. In addition, the direct or indirect application of the technical solutions of the present invention to other related technical fields are similarly included within the patent protection scope of the present invention.
Claims
1. A method for generating extreme power load samples, characterized in that, Includes the following steps: Step 1: Obtain and preprocess the original load sequence to obtain the preprocessed original load sequence; use the industrial load disturbance identification model to separate the preprocessed original load sequence to obtain the base load sequence and the industrial disturbance component; The basic load series is sliced by day to obtain multiple daily basic load series; the multiple daily basic load series are combined into a daily basic load series set; external covariates are constructed. Step 2: Process the daily basic load sequence set using the discrete wavelet decomposition method to obtain the low-frequency smoothed trend sequence and multiple scale detail sequences under each daily basic load sequence; construct a multi-scale time-frequency feature vector based on the multiple scale detail sequences under each daily basic load sequence. Step 3: Perform density peak clustering analysis on the daily basic load sequence set based on multi-scale time-frequency feature vectors to obtain extreme load pattern samples; Step 4: Use dynamic density clustering algorithm to perform fine-grained clustering on the time segments corresponding to the extreme load pattern samples to obtain the initial time segment set of extreme events; use Markov random field to optimize the initial time segment set to obtain the spatiotemporal template of extreme events. Step 5: Design a diffusion generation model based on the spatiotemporal template of extreme events, and train the diffusion generation model using extreme load pattern samples; Multiple load sequences are generated using a trained diffusion generation model. These multiple load sequences are then superimposed with their corresponding industrial disturbance components to obtain extreme load samples.
2. The method for generating extreme power load samples according to claim 1, characterized in that, Step 1 includes: The original load sequence is acquired and preprocessed to obtain the preprocessed original load sequence. The preprocessing includes missing value imputation, outlier removal, and normalization. An industrial load disturbance identification model is used to separate the preprocessed original load sequence to obtain the base load sequence. and industrial disturbance components i(t) ; Base load sequence Slicing by day yields multiple daily basic load sequences; these multiple daily basic load sequences form a daily basic load sequence set. , ; in, For extreme day indexes, This represents the total number of extreme days. This represents the number of time segments contained in a single-day basic load sequence. For the first The base load value for the first time segment in a single-day base load sequence. For the first The first of the daily basic load sequences Baseline load values for each time segment; Constructing external covariates External covariates include calendar and weather information.
3. A method for generating extreme power load samples according to claim 2, characterized in that, In step 1, the industrial load disturbance identification model adopts a Transformer architecture based on a multi-layer self-attention structure. The minimum objective function value of the industrial load disturbance identification model is... L The solution formula is: ; in, ; ; A set of discrete-time indices within a sample window. For time step index; To observe the total load; Background load component; Industrial disturbance component; , These are the time series vectors of the background load component and the industrial disturbance component, respectively; This is the set of trainable parameters for an industrial load disturbance identification model. These are the weighting coefficients for the industrial disturbance smoothing regularization term. ; These are the weighting coefficients of the relevance constraint term. ; The number of selected industrial disturbance indicators; For the first Time series of industrial disturbance indicators; for and Correlation loss between them , They are respectively , exist The sample mean above; The Pearson correlation coefficient; Industrial disturbance component i(t) The dynamic constraint formula that is satisfied is: ;in, This is the upper limit of the rate. For industrial load disturbance identification model in t- Industrial disturbance components extracted at time 1.
4. The method for generating extreme power load samples according to claim 2, characterized in that, Step 2 includes: Step 21: Perform analysis on the daily basic load sequence set. Layered discrete wavelet decomposition yields the low-frequency smoothed trend sequence and J-scale detail sequences for each daily basic load sequence; the discrete wavelet decomposition formula is: ; in, For the first d A low-frequency smoothed trend sequence under a single-day basic load series. For the first d The first daily basic load sequence Layer-scale detail sequence, , For time step index; The number of discrete wavelet decomposition layers; For wavelet decomposition scale index; Step 22: Based on J scale detail sequences under each daily base load sequence, calculate the wavelet energy entropy under each daily base load sequence. and energy percentage ; Calculate wavelet energy entropy The formula is: ;in, For the first Energy percentage of each scale detail sequence It is a logarithmic function; The calculation formula is: ;in, For the first Energy of a sequence of details at a scale k The index variable for summation; The calculation formula is: ; Step 23: Based on wavelet energy entropy and energy percentage Construct multi-scale time-frequency feature vectors ; ,in, This represents the energy percentage of the detail sequence at scales 1 to J. For the first d The energy entropy value of a single-day basic load sequence. For dimension The real vector space.
5. The method for generating extreme power load samples according to claim 4, characterized in that, Step 3 includes: Step 31: Define the composite distance metric Suppose the size of the set of daily basic load sequences to be analyzed is... The index set of the daily basic load sequence is ;in, These represent the indices of two single-day basic load sequences; For the first A multi-scale time-frequency feature vector of a single-day basic load sequence; For the first External covariates of a single-day basic load sequence; Composite distance metric for: ;in, For the first A single-day basic load sequence. The Mahalanobis distance, It is a diagonal matrix. As the first weighting coefficient, This is the second weighting coefficient. For the first External covariates of a single-day basic load sequence It is the Euclidean norm; Step 32: Based on composite distance metric Density peak clustering algorithm is used to... Analyzing a single-day basic load sequence, the number of output clusters is: , No. The cluster number to which each daily basic load sequence belongs is: , .
6. The method for generating extreme power load samples according to claim 5, characterized in that, Step 3 also includes: Step 33: Select the cluster with the most basic load sequences on a single day as the master cluster. , morphological feature center for: For any cluster ( ), its morphological center for: ; morphological distance Defined as: The mean and standard deviation of the morphological distances of all clusters were calculated as follows: and ,when At that time, it was considered that the cluster The curve shape deviates significantly from the main cluster, among which, These are empirical parameters; Step 34, Primary Cluster External covariate feature center for: For any cluster , External covariate feature center for: external covariate distance for: The mean and standard deviation of the external covariate distances for all clusters were calculated as follows: and ;when At that time, it was considered that the cluster The distribution of external covariates differs significantly from that of the main cluster, among which, These are empirical parameters; Step 35: Simultaneously satisfying and The cluster number set is used as the sample set of extreme load patterns. , Extreme load pattern sample index set for: .
7. The method for generating extreme power load samples according to claim 6, characterized in that, Step 4 includes: Step 41, the process of using dynamic density clustering algorithm to perform fine-grained clustering of time series segments corresponding to extreme load pattern samples to obtain the initial set of time series segments for extreme events includes: on each extreme day Calculate the first-order difference. The formula is: ; with a preset threshold Forming a candidate mutation set , To obtain the initial time series fragment set of extreme events. , ,in, Represents the norm, Representing different subtypes; Step 42, the process of optimizing the initial time series segment set using Markov random fields to obtain the spatiotemporal template of extreme events includes: optimizing the initial time series segment set... Modeling the sequence as a one-dimensional Markov random field and minimizing the energy yields a smoothed set of initial time series segments; for the same subtype The events are time-normalized and aligned to obtain the template curve. The event amplitude, downlink rate, recovery rate, and duration are collected into a parameter vector. ; and The spatiotemporal template that constitutes extreme events; energy The calculation formula is: ,in, It is a single-point potential. They are adjacent potentials.
8. The method for generating extreme power load samples according to claim 7, characterized in that, In step 5, the forward diffusion process of the diffusion generation model is defined using a Markov chain; the reverse diffusion process of the diffusion generation model uses a multi-layer deep neural network as a noise estimator. The forward diffusion process is as follows: ; in, Let be the probability density function. For diffusion time step index, For the first The load sequence state at each time step. For the first The load sequence state at each time step. The mean is Covariance is The multivariate Gaussian distribution, This is the noise variance hyperparameter. It is the identity matrix; ;in, The cumulative scaling factor. , ; To obtain from the standard normal distribution Random noise in the mid-sample, This is a sample of extreme load patterns.
9. A method for generating extreme power load samples according to claim 8, characterized in that, Noise estimator Predict injected noise; For conditional vectors, , A one-hot indicator vector for the event type; Objective function of diffusion generation model for: ; in, For mean square error loss, ; [] represents extreme load mode samples ,noise and time step The joint distribution expectation; This is the third weighting coefficient. It is the fourth weighting coefficient. This is the first penalty term, used to suppress the generation of negative load values. The second penalty term is used to constrain the ramp rate of the generated load sequence from being too fast. The load ramp-up increments between adjacent time steps. The maximum allowed ramp threshold; After the diffusion generation model training converges, sampling starts from... By gradually working backward, we can obtain The extreme load samples are spliced together according to the event positions on the timeline within the extreme day, and the transition between segments is smoothed by the first-order difference continuity criterion. The corresponding industrial disturbance components are then superimposed to obtain the extreme load samples.