Load shock type identification method and system for distributed power grid
By expanding the training data through the preset generative adversarial network of the local curvature perception mechanism, the problems of insufficient accuracy and generalization ability in load shock type identification in distributed power grids are solved, and high-precision identification of load shock types is achieved.
Patent Information
- Application Number
- CN202510886423.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-30
AI Technical Summary
Existing technologies have difficulty in effectively identifying the types of load shocks in distributed power grids, which leads to challenges in the stability, reliability and security of the power grid, and the model generalization ability and classification accuracy are insufficient.
A preset generative adversarial network with a local curvature perception mechanism is used to expand the samples of the initial training set to generate power data samples that are close to real power data. The target recognition model is trained by expanding the training set to improve the recognition accuracy and generalization ability.
The training data scale has been significantly increased, which has improved the model's recognition accuracy and generalization ability for load shock types, reduced the risk of misclassification, and can accurately identify normal fluctuations, minor shocks, moderate shocks, and severe shocks.
Smart Images

Figure CN120408323B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of power grid fault identification and artificial intelligence technology, and in particular to a load impact type identification method and system for distributed power grids. Background Art
[0002] With the widespread integration of distributed energy resources and the increasing diversification of electricity loads, the operating environment of the power grid has become more complex and dynamic. In this context, load shocks of various sizes can occur within the grid. These shocks can be caused by a variety of factors, such as the startup and shutdown of large-scale equipment, load fluctuations caused by changing weather conditions, or failures of internal grid equipment. These load shocks not only pose challenges to the stability, reliability, and security of the grid, but can also trigger a series of problems, such as voltage fluctuations, frequency deviations, and even localized power outages. Therefore, effectively identifying the types of load shocks in distributed power grids has become a pressing issue. Summary of the Invention
[0003] The purpose of the embodiments of the present invention is to provide a load shock type identification method and system for distributed power grids. By introducing a preset generative adversarial network with a local curvature perception mechanism, generated power data samples with a geometric structure close to real power data are obtained, which can achieve the expansion of training samples and thereby improve the recognition accuracy and generalization ability of the target identification model for load shock types.
[0004] A first embodiment of the present invention provides a method for identifying load impact types in a distributed power grid, comprising:
[0005] Obtain an initial training set of the target power system;
[0006] Using a preset generative adversarial network that introduces a local curvature perception mechanism, the initial training set is expanded to obtain an expanded training set; wherein the expanded training set includes a plurality of power data samples;
[0007] The initial model is trained by using the expanded training set to obtain a target recognition model; wherein the target recognition model is used to identify the power data to be detected and output the corresponding load impact type.
[0008] Optionally, the generator loss in the preset generative adversarial network includes: adversarial loss, multi-scale local curvature loss and high-order curvature alignment loss.
[0009] Optionally, the multi-scale local curvature loss is calculated using the following formula:
[0010] ;
[0011] in, is the multi-scale local curvature loss function; For the generator; A set of real power data samples input to the preset generative adversarial network for the current batch; For the The generated power data samples are Second-order derivatives at different scales; For the The real power data samples are Second-order derivatives at different scales; For the The third derivative of the generated power data samples; For the The third-order derivative of a real power data sample; For the Generate power data samples; For the Real power data samples; For the The first step to generate power data samples Features For the The real power data samples Features is the curvature loss balance coefficient; is the total number of scales; The number of samples input to the preset generative adversarial network for the current batch; is a positive integer; is a positive integer and ; is the L2 norm.
[0012] Optionally, the generator loss is calculated using the following formula:
[0013] ;
[0014] in, is the loss function of the generator; A set of real power data samples input to the preset generative adversarial network for the current batch; is the set of latent space vectors corresponding to the real power data sample set; is the generator function; is the discriminator function; is the latent variable interaction term; The generator input is corresponding expectations; is the multi-scale local curvature loss function; For the A high-order curvature term that generates power data samples; For the The higher-order curvature terms of real power data samples; is the loss weight coefficient of the first generator; is the loss weight coefficient of the second generator; The number of samples input to the preset generative adversarial network for the current batch.
[0015] Optionally, the method further includes:
[0016] During the training process of the initial model, the gradient update amount of the current weight matrix is calculated based on the expanded training set; wherein the gradient update amount is composed of at least the following three components: an interference elimination term, a disturbance term, and a nonlinear adaptive adjustment term; the interference elimination term is used to correct the gradient update direction according to the historical gradient vector of the power data sample; the disturbance term is used to adjust the gradient update direction according to the deviation between the power data sample and the sample mean vector; the nonlinear adaptive adjustment term is used to adjust the gradient update direction according to the nonlinear relationship in the power data sample.
[0017] Optionally, the interference elimination term is calculated using the following formula:
[0018] ;
[0019] in, is the interference elimination term; is the learning rate adjustment coefficient of the initial model; is the model output corresponding to the i-th power data sample; represents the gradient update direction of the i-th power data sample; is the historical gradient vector of the i-th power data sample; The number of samples input into the initial model for training for the current batch.
[0020] Optionally, the disturbance term is calculated using the following formula:
[0021] ;
[0022] in, is the disturbance term; is the i-th power data sample; is the sample mean vector; is the model output corresponding to the i-th power data sample; represents the gradient update direction of the i-th power data sample; The number of samples input into the initial model for training for the current batch.
[0023] Optionally, the nonlinear adaptive adjustment item is obtained by the following steps:
[0024] Performing high-order feature expansion on the power data sample to obtain a corresponding nonlinear feature vector; wherein the nonlinear feature vector includes: the original features of the power data sample, the square terms of the original features, and the combined cross terms between the original features;
[0025] Calculating the response intensity of the current weight matrix to the nonlinear eigenvector to obtain a nonlinear adjustment factor; wherein the nonlinear adjustment factor is used to characterize the adaptability of the initial model to the nonlinear relationship;
[0026] Based on the nonlinear adjustment factor and the influence coefficient factor, the gradient update direction is adjusted to obtain the nonlinear adaptive adjustment item.
[0027] Optionally, the method further includes:
[0028] Update the current weight matrix based on the gradient update amount to obtain a transition weight matrix;
[0029] Based on the weight matrix difference and the viscosity factor of adjacent iteration cycles, the transition weight matrix is smoothly corrected to obtain an updated weight matrix.
[0030] A second embodiment of the present invention provides a load impact type identification system for a distributed power grid, comprising:
[0031] An initial training set acquisition module, used to acquire an initial training set of a target power system;
[0032] An expanded training set acquisition module is configured to expand the samples of the initial training set using a preset generative adversarial network that introduces a local curvature perception mechanism to obtain an expanded training set; wherein the expanded training set includes a plurality of power data samples;
[0033] The target recognition model construction module is used to train the initial model through the expanded training set to obtain the target recognition model; wherein, the target recognition model is used to identify the power data to be detected and output the corresponding load impact type.
[0034] Compared with the prior art, an embodiment of the present invention provides a method and system for identifying load impact types for distributed power grids. The method includes: obtaining an initial training set for the target power system; using a preset generative adversarial network that introduces a local curvature perception mechanism to perform sample expansion on the initial training set of gradient update quantities to obtain an expanded training set; wherein the gradient update quantity expanded training set includes several power data samples; using the gradient update quantity expanded training set to train the initial model to obtain a target recognition model; wherein the gradient update quantity target recognition model is used to identify the power data to be detected and output the corresponding load impact type. The embodiment of the present invention obtains generated power data samples whose geometric structure approximates real power data by introducing a preset generative adversarial network that introduces a local curvature perception mechanism, thereby achieving the expansion of training samples and thereby improving the target recognition model's recognition accuracy and generalization ability for load impact types. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a flow chart of an embodiment of a method for identifying load impact types for a distributed power grid provided by the present invention;
[0036] Figure 2 This is a flow chart of an embodiment of training a preset generative adversarial network provided by the present invention;
[0037] Figure 3 It is a structural diagram of an embodiment of a load impact type identification system for distributed power grids provided by the present invention. DETAILED DESCRIPTION
[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this technical field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0039] See also Figure 1 , is a flow chart of an embodiment of a load impact type identification method for a distributed power grid provided by the present invention.
[0040] The first embodiment of the present invention provides a method for identifying load impact types for a distributed power grid, including steps S1 to S3, as follows:
[0041] Step S1: Obtain an initial training set of the target power system;
[0042] Step S2: using a preset generative adversarial network that introduces a local curvature perception mechanism to perform sample expansion on the initial training set to obtain an expanded training set; wherein the expanded training set includes a plurality of power data samples;
[0043] Step S3: The initial model is trained using the expanded training set to obtain a target recognition model; wherein the target recognition model is used to identify the power data to be detected and output the corresponding load impact type.
[0044] It should be noted that in traditional methods, the collection of training data is often limited by factors such as time, space, and equipment. This not only leads to insufficient sample size, but may also cause uneven data distribution, thus adversely affecting the model's generalization ability and classification accuracy.
[0045] This embodiment of the present invention introduces a local curvature perception mechanism into a generative adversarial network (GAN). By calculating the curvature difference between generated power data and real power data, it ensures that the geometric structure of the generated power data approximates the distribution characteristics of real power data (such as voltage fluctuation trends, current mutation patterns, and meteorological fluctuation trends), thus avoiding the generation of "pseudo-power data."
[0046] In addition, the embodiment of the present invention can significantly increase the scale of training data by presetting the generative adversarial network to expand the samples, so that the model can learn the characteristics of various load shocks more robustly and comprehensively during training, thereby accurately identifying the four types of normal fluctuations, slight shocks, moderate shocks and severe shocks, and reducing the risk of misclassification due to insufficient samples.
[0047] In step S1, the embodiment of the present invention constructs an initial training set by collecting multi-source monitoring data from the target power system and performing data preprocessing. In other words, the initial training set is constructed by collecting data such as current, voltage, equipment status, and meteorological parameters from the target power system and performing data cleaning, feature extraction, discretization, and expert labeling.
[0048] The data collected by the embodiments of the present invention comes from a wide range of sources, covering a variety of real-time monitoring devices in the power system, including but not limited to current sensors, voltage sensors, load monitoring equipment, equipment status monitoring systems, and external meteorological data sources. The following are the main methods and sources of data collection:
[0049] (1) Real-time monitoring: Sensors in power substations and distribution networks collect current, voltage, load fluctuations, and other values in real time and transmit them to the data acquisition system. Each acquisition device collects data at a fixed time interval, and the data cycle can be flexibly adjusted according to demand, usually 1 to 10 seconds.
[0050] (2) Equipment status information: With the help of the real-time status monitoring system of the equipment, the switch status, operating mode and fault information of various equipment (such as transformers, circuit breakers, etc.) can be obtained.
[0051] (3) Meteorological data: By accessing meteorological monitoring stations, external meteorological data such as local temperature, humidity, wind speed, and precipitation are obtained; this information has a certain correlation with the load fluctuations of the power grid, so it needs to be collected synchronously.
[0052] All collected data is stored in a unified cloud platform data warehouse and timestamped to ensure the raw data retains time series properties. Data is stored in a structured JSON format to improve the efficiency of subsequent processing and analysis.
[0053] The collected raw data needs to go through multiple steps of pre-processing to ensure its quality and applicability. The main processing flow is as follows:
[0054] (1) Data cleaning and outlier detection: The raw data may contain noise, missing values, or outliers, which can adversely affect subsequent analysis and modeling. First, outliers are detected through statistical analysis methods (such as the Z-score method), and then the raw data is cleaned by interpolation or direct deletion of outliers.
[0055] (2) Feature conversion and data standardization: The collected time series data such as current, voltage, load, and meteorological parameters are subjected to feature conversion and standardization. Specifically, the values of current and voltage are standardized to ensure that different features are calculated on the same scale. In addition, the sliding window technology is used to extract statistical features such as the maximum, minimum, mean, and variance within different time windows.
[0056] (3) Discretization and feature construction: For time series data, discretization methods are used to convert them into discrete features. For example, the current value is divided into different intervals (such as 0–10A, 10–20A, 20–30A, etc.) through discretization methods, and the corresponding discrete feature values are generated according to the interval to which the current belongs.
[0057] After data preprocessing, the collected training data is labeled. Specifically, experts manually label the load impact type corresponding to each moment of data based on characteristics such as current and voltage fluctuation amplitudes and load change rates. Load impact types include the following four types: normal fluctuation, slight impact, moderate impact, and severe impact.
[0058] It is understandable that the collection, labeling and preprocessing of training data are time-consuming and labor-intensive, and insufficient training samples (i.e., power data samples) can easily lead to poor model generalization ability and affect the accuracy of the model.
[0059] To address the shortcomings of power data samples in terms of scale, distribution, and diversity, embodiments of the present invention employ a pre-configured generative adversarial network based on local curvature perception to generate samples, thereby expanding the power data samples and obtaining an expanded training set. In the expanded training set, the generated power data samples and the real power data samples are collectively referred to as power data samples. Specifically, by employing a local curvature perception mechanism and multi-scale curvature modeling, the geometric structure of the generated power data samples is optimized, improving their authenticity, diversity, and adaptability to complex power data structures.
[0060] The embodiments of the present invention not only effectively expand the scale of power data samples and alleviate the problem of insufficient samples, but also improve the diversity and distribution balance of samples in the expanded training set, providing a solid data foundation for subsequent model training and optimization.
[0061] In an optional embodiment, the generator loss in the preset generative adversarial network includes: adversarial loss, multi-scale local curvature loss and high-order curvature alignment loss.
[0062] Furthermore, the multi-scale local curvature loss is calculated by the following formula:
[0063] ;
[0064] in, is the multi-scale local curvature loss function; For the generator; A set of real power data samples input to the preset generative adversarial network for the current batch; For the The generated power data samples are Second-order derivatives at different scales; For the The real power data samples are Second-order derivatives at different scales; For the The third derivative of the generated power data samples; For the The third-order derivative of a real power data sample; For the Generate power data samples; For the Real power data samples; For the The first step to generate power data samples Features For the The real power data samples Features is the curvature loss balance coefficient; is the total number of scales; The number of samples input to the preset generative adversarial network for the current batch; is a positive integer; is a positive integer and ; is the L2 norm.
[0065] Furthermore, the generator loss is calculated by the following formula:
[0066] ;
[0067] in, is the loss function of the generator; A set of real power data samples input to the preset generative adversarial network for the current batch; is the set of latent space vectors corresponding to the real power data sample set; is the generator function; is the discriminator function; is the latent variable interaction term; The generator input is corresponding expectations; is the multi-scale local curvature loss function; For the A high-order curvature term that generates power data samples; For the The higher-order curvature terms of real power data samples; is the loss weight coefficient of the first generator; is the loss weight coefficient of the second generator; The number of samples input to the preset generative adversarial network for the current batch.
[0068] It should be noted that if Figure 2 The figure is a flow chart of an embodiment of the present invention for training a preset generative adversarial network. Figure 2In the process, the generator and discriminator are first initialized, and then the iterative training phase begins. In each iteration, the generator generates data and the discriminator determines the authenticity of the data; then, the geometric structure differences of the generated data are evaluated through multi-scale local curvature loss, and the geometric structure of the generated data is optimized accordingly. Subsequently, the losses of the generator and discriminator are calculated separately; among them, the loss of the generator combines the discrimination results and the geometric structure optimization goal, while the loss of the discriminator is based on its ability to classify real and fake data. Using these losses, the parameters of the generator and discriminator are updated through backpropagation. The training process continues to cycle until the maximum number of iterations is reached, and finally the training of the preset generative adversarial network (i.e., the power data augmentation model) is completed. The above training process not only covers the basic training steps of GAN, but also introduces multi-scale geometric structure optimization to improve the authenticity of the generated power data samples.
[0069] In specific implementation, the main steps in the training process of the preset generative adversarial network with the local curvature perception mechanism are as follows:
[0070] Step 1: Based on the framework of the generative adversarial network, the generator adopts a multi-layer neural network structure, with inputs of randomly sampled latent space vectors and feature vectors of the original power data (i.e., real power data samples), and outputs generated power data (i.e., generated power data samples). The discriminator is built based on a multi-layer convolutional network and is used to output the authenticity probability after inputting power data.
[0071] Let the generator be , the input of the generator is (i.e., the potential space vector set corresponding to the real power data sample set), the real power data is (i.e., the current batch of real power data samples input to the preset generative adversarial network), the generated power data output by the generator is (i.e. the generator for the current batch The generated power data sample set outputted by the discriminator is , the input of the discriminator is or ; The output probability of the discriminator is ,and , represents the probability that the discriminator determines that the power data is true; the weight parameters of the generator and the discriminator are and , initialized by random assignment.
[0072] Step 2: Optimize the geometric structure of the generated power data through the local curvature perception mechanism. The local curvature perception mechanism means that during the training process of the generator, the curvature difference between the generated power data and the real power data at different scales is calculated, and the complex geometric features are captured by combining high-order derivative information to constrain the generator output to approximate the geometric distribution of the real power data in the local area. At this stage, the optimization of the local curvature perception mechanism is achieved by calculating the multi-scale local curvature loss function, which is expressed as:
[0073] ;
[0074] in, is the multi-scale local curvature loss function; For the generator; A set of real power data samples input to the preset generative adversarial network for the current batch; For the generator Output The generated power data samples are The second-order derivative at this scale is used to characterize the The local curvature of the generated power data samples; For the The real power data samples are The second-order derivative at this scale is used to characterize the The local curvature of a real power data sample; For the generator Output The third-order derivative of the generated power data sample is used to characterize the A high-order geometric variation of generated power data samples; For the The third-order derivative of a real power data sample is used to characterize the High-order geometric changes of real power data samples; For the Generate power data samples; For the Real power data samples; For the The first step to generate power data samples Features For the The real power data samples Features is the curvature loss balance coefficient; is the total number of scales; The number of samples input to the preset generative adversarial network for the current batch; is a positive integer; is a positive integer and ; is the L2 norm. Preferably, Set to 0.2.
[0075] It is worth noting that the total number of scales It refers to the scope of modeling the local geometric structure of power data samples at different time windows and feature combination levels. For example, =1 means that in a short time window (such as 1 second), the basic features are used to calculate the curvature to capture micro mutations; As the time window length and feature combination complexity increase, they are successively expanded to model the trend changes and high-order structural features at the meso- and macro-levels, thereby achieving accurate simulation and structural preservation of load disturbance behaviors at different levels.
[0076] Step 3: Enhance the generator’s ability to capture complex feature relationships through latent variable interaction modeling. Perform nonlinear mapping between latent space vectors and input features to enhance the deep interaction between the two and improve the diversity of generated power data. The calculation method of latent variable interaction terms is expressed as:
[0077] ;
[0078] Where, is the latent variable interaction term, is the ReLU activation function, indicating that the ReLU activation function is used for nonlinear feature extraction; is the Sigmoid activation function; is the weight parameter of the generator; is the bias parameter of the generator; ⊙ is the element multiplication; The input to the generator and real power data Vector concatenation operation; is the weight coefficient of the polynomial term; The first k values input to the generator; are the first k values of the real power data. Preferably, Set to 0.2.
[0079] Furthermore, based on the latent variable interaction term, the final output of the generator is expressed as:
[0080] ;
[0081] Where ⊙ is element-wise multiplication; is a generator function.
[0082] Step 4: Based on the traditional adversarial loss, the generator loss function is constructed based on the multi-scale local curvature loss function to balance the optimization objectives of power data distribution and local structure. The calculation method is expressed as:
[0083] ;
[0084] in, is the loss function of the generator, which constrains the training process of the generator; A set of real power data samples input to the preset generative adversarial network for the current batch; is the set of latent space vectors corresponding to the real power data sample set; is the generator function; is the discriminator function; is the latent variable interaction term; express expectations; The generator input is expectations when is the multi-scale local curvature loss function; The output of the generator is The high-order curvature terms of the generated power data samples are expressed as ; For the The third derivative of the generated power data samples, For the The second-order derivative of the generated power data sample is calculated. When calculating the second-order derivative and the third-order derivative of the generated power data sample, the value of each feature dimension is derived separately and then accumulated; For the The calculation method of the high-order curvature term of the real power data sample is similar to that of the generated power data sample, that is, ; is the loss weight coefficient of the first generator; is the loss weight coefficient of the second generator; The number of samples input to the preset generative adversarial network for the current batch.
[0085] Step 5: The discriminator training process is constrained by the discriminator loss function to enhance sensitivity to subtle differences in the generated power data. The calculation method is expressed as:
[0086] ;
[0087] Where, is the loss function of the discriminator; Express expectations, Indicates that the input is expectations of the time, Indicates that the input is expectations of the time.
[0088] Step 6: Update the parameters of the generator and discriminator using an alternating iterative strategy. The generator minimizes the total loss through gradient descent, and the discriminator maximizes the adversarial loss through gradient ascent. The calculation method is expressed as:
[0089] Update of generator parameters: ;
[0090] Update of discriminator parameters: ;
[0091] Where, To preset the learning rate of the generative adversarial network; The loss function of the generator is gradient; is the loss function of the discriminator gradient; and are the weight parameters of the generator and discriminator respectively; Indicates a parameter update operation. Preferably, Set to 0.001.
[0092] Step 7: Repeat the above steps until a preset stopping condition is met, indicating that training of the preset generative adversarial network (i.e., the power data augmentation model) is complete. Typically, the stopping condition can be set to reach a preset maximum number of iterations. In this embodiment of the present invention, the maximum number of iterations is preferably set to 1000 to ensure sufficient model convergence and good generalization capabilities.
[0093] Step 8: After the power data augmentation model is trained, the trained model is used to expand the number of samples. For example, if the original number of collected samples is 800, and the power data augmentation model generates 200 new samples, the expanded power data sample set (i.e., the augmented training set) will contain 1000 samples.
[0094] In an optional embodiment, the method further includes:
[0095] During the training process of the initial model, the gradient update amount of the current weight matrix is calculated based on the expanded training set; wherein the gradient update amount is composed of at least the following three components: an interference elimination term, a disturbance term, and a nonlinear adaptive adjustment term; the interference elimination term is used to correct the gradient update direction according to the historical gradient vector of the power data sample; the disturbance term is used to adjust the gradient update direction according to the deviation between the power data sample and the sample mean vector; the nonlinear adaptive adjustment term is used to adjust the gradient update direction according to the nonlinear relationship in the power data sample.
[0096] Furthermore, the interference elimination term is calculated by the following formula:
[0097] ;
[0098] in, is the interference elimination term; is the learning rate adjustment coefficient of the initial model; is the model output corresponding to the i-th power data sample; represents the gradient update direction of the i-th power data sample; is the historical gradient vector of the i-th power data sample; The number of samples input into the initial model for training for the current batch.
[0099] Furthermore, the disturbance term is calculated by the following formula:
[0100] ;
[0101] in, is the disturbance term; is the i-th power data sample; is the sample mean vector; is the model output corresponding to the i-th power data sample; represents the gradient update direction of the i-th power data sample, The number of samples input into the initial model for training for the current batch.
[0102] Furthermore, the nonlinear adaptive adjustment term is obtained by the following steps:
[0103] Performing high-order feature expansion on the power data sample to obtain a corresponding nonlinear feature vector; wherein the nonlinear feature vector includes: the original features of the power data sample, the square terms of the original features, and the combined cross terms between the original features;
[0104] Calculating the response intensity of the current weight matrix to the nonlinear eigenvector to obtain a nonlinear adjustment factor; wherein the nonlinear adjustment factor is used to characterize the adaptability of the initial model to the nonlinear relationship;
[0105] Based on the nonlinear adjustment factor and the influence coefficient factor, the current gradient update direction is adjusted to obtain the nonlinear adaptive adjustment item.
[0106] Furthermore, the method further comprises:
[0107] Update the current weight matrix based on the gradient update amount to obtain a transition weight matrix;
[0108] Based on the weight matrix difference and the viscosity factor of adjacent iteration cycles, the transition weight matrix is smoothly corrected to obtain an updated weight matrix.
[0109] It should be noted that after obtaining the expanded training set, the embodiment of the present invention will perform model training on the initial model based on the expanded training set to obtain the final target identification model. The selection of the initial model is highly flexible and includes, but is not limited to, neural network models, support vector machines, and extreme learning machines. Preferably, the initial model is set to an extreme learning machine. In addition, by combining the expanded training set with the gradient update optimization algorithm, the embodiment of the present invention enables the target identification model to significantly improve the accuracy and robustness in high-dimensional power data classification tasks.
[0110] The following takes the extreme learning machine (ELM) as an example to explain in detail the training process of obtaining the target recognition model.
[0111] It should be noted that in order to solve a series of problems in high-dimensional power data classification, such as noise interference, redundant features, and inconsistent feature scales, the embodiment of the present invention optimizes the extreme learning machine algorithm to improve the accuracy of the classifier (i.e., the target recognition model) and accelerate the training process. Specifically, by introducing interference elimination terms, disturbance terms, and nonlinear adaptive adjustment terms, the gradient update direction is dynamically adjusted to reduce the impact of noise and redundant features.
[0112] The embodiment of the present invention introduces an interference elimination mechanism based on gradient direction during the training process of the extreme learning machine; wherein, the interference elimination mechanism based on gradient direction first calculates the gradient value of each training sample and adjusts the weight of the hidden layer according to the gradient direction.
[0113] In the forward propagation process of the extreme learning machine, the activation method is expressed as:
[0114] ;
[0115] Where, is the Sigmoid activation function; is the weight matrix of the hidden layer of the extreme learning machine, is the bias term of the extreme learning machine; is the output of the extreme learning machine (i.e., the output set of the extreme learning machine under the current batch); is the training sample input to the extreme learning machine (i.e. the set of power data samples input to the extreme learning machine in the current batch). In other words, in the current batch training, each power data sample Through the same and , and obtain the corresponding extreme learning machine output .
[0116] The embodiment of the present invention corrects the gradient update direction through an interference cancellation mechanism. This is achieved by calculating the interference cancellation term. The calculation method is based on the weighted sum of the absolute value of the current gradient and the historical gradient vector to reduce the gradient interference between samples, thereby stabilizing the training direction. The details are as follows:
[0117] ;
[0118] Where, is the interference elimination term; is the learning rate adjustment coefficient of the extreme learning machine; Represents the loss function of the extreme learning machine right gradient; are the parameters of the extreme learning machine, including the weight matrix and bias term; is the loss function of the extreme learning machine; It is the historical gradient vector used to adjust the gradient direction, which represents the accumulation of the historical gradient influence of the sample.
[0119] It is worth noting that the chain rule can be used to obtain , gradient direction It is the gradient update direction of the batch, reflecting the fastest descent direction of the loss function L under the current parameter θ in the current training batch. Represents the collection of historical gradient information of all samples in the current training batch. It is an aggregate representation used to adjust the current gradient direction.
[0120] To further illustrate the loss function of the extreme learning machine right The gradient of , and the method of obtaining the historical gradient vector after gradient adjustment, with samples as the granularity, the expanded calculation formula is expressed as:
[0121] ;
[0122] Where, is the historical gradient vector of the i-th power data sample, is the model (extreme learning machine) output corresponding to the i-th power data sample, The number of samples input to the extreme learning machine for the current batch.
[0123] It is worth noting that the historical gradient vector It is the cumulative average of the gradient vectors corresponding to the i-th sample in historical training, that is, the gradient values corresponding to the i-th sample in previous trainings are accumulated one by one and divided by the number of rounds in which it participates in training, which is used to represent the historical optimization experience of the sample. Represents the gradient update direction of the i-th power data sample.
[0124] During the training process, in order to avoid overfitting and improve training efficiency, the robustness of the extreme learning machine training process is enhanced by the perturbation factor to ensure the best training effect under different types of power data. The calculation formula is expressed as:
[0125] ;
[0126] Where, is the learning rate of the extreme learning machine, is the current weight matrix The gradient update amount (i.e. weight update amount), is the weight matrix of the hidden layer of the extreme learning machine at the tth iteration (i.e., the current weight matrix), is the transition weight matrix of the hidden layer of the extreme learning machine at the t+1th iteration. Preferably, Set to 0.01.
[0127] In the embodiment of the present invention, the gradient update amount It consists of at least three components: interference elimination term, disturbance term, and nonlinear adaptive adjustment term. This can comprehensively consider the effects of gradient correction, disturbance, and nonlinear relationship, optimize the update direction, and balance the generalization ability. The calculation formula is:
[0128] ;
[0129] Where, is the first adjustment parameter, is the second adjustment parameter, is the disturbance term (i.e., disturbance factor), is the interference elimination term, is a nonlinear adaptive adjustment term. Preferably, Set to 0.2, Set to 0.3.
[0130] In this embodiment of the present invention, the disturbance term (disturbance factor) is calculated based on the dot product sum of the difference between the input sample and the mean and the gradient, so as to enhance the robustness of the model to changes in the input distribution and prevent overfitting. Its calculation formula is expressed as follows:
[0131] ;
[0132] Where, is the i-th power data sample input to the extreme learning machine, is the sample mean vector input to the extreme learning machine, The number of samples input to the extreme learning machine for the current batch.
[0133] This embodiment of the present invention also introduces a nonlinear adaptive adjustment term during the gradient update process. First, a nonlinear feature vector for the power sample is constructed. A new high-order feature vector is formed by combining the original input vector, the squared terms of its components, and the cross-products between the features. This constructs a nonlinear relationship between the input features, thereby improving the model's ability to perceive complex patterns in power data.
[0134] After completing the nonlinear feature construction, combined with the calculation mechanism of the nonlinear adjustment factor, that is, by performing a weighted product of the above nonlinear feature vector and the extreme learning machine weight matrix under the current iteration round, and normalizing it through the Sigmoid function, an adjustment factor between 0 and 1 is obtained. It is used to quantify the response degree of the model to the nonlinear structure in the current sample, thereby providing an adaptive adjustment basis for parameter updating, ensuring that the model has higher sensitivity and adaptability to nonlinear features.
[0135] After obtaining the nonlinear adjustment factor, a nonlinear adaptive adjustment term is further constructed. By weightedly combining this factor with the current gradient information, an additional gradient term that can dynamically adjust the update direction and amplitude is generated. In the weight update process, a compensation mechanism for the nonlinear structure of the data is combined, so that the training process can not only follow the guidance of minimizing the error, but also effectively capture the high-order structural features in the input space, enhance the generalization ability of the model, and improve the stability and accuracy of complex power data classification.
[0136] In specific implementation, the nonlinear adaptive adjustment term is defined as a correction term related to the nonlinear relationship of the power data. This term is calculated based on the high-order features of the input power data (such as quadratic terms, cross terms, etc.). Assume that each sample of the power data contains several features, and its nonlinear feature vector is the high-order transformation form of the power data input to the extreme learning machine, expressed as:
[0137] ;
[0138] Where, is the matrix composed of the nonlinear eigenvectors of each sample in the current batch of power data sample sets, The matrix consisting of the square terms of the power data features corresponding to each sample is represented. Represents the cross term between the power data features corresponding to each sample (such as , wait, 、 、 is a matrix composed of three power data features of the power data sample.
[0139] Taking samples as the granularity, the calculation formula is expressed as:
[0140] ;in, is the nonlinear feature vector of the i-th power data sample, represents the square term of the power data feature in the i-th power data sample, represents the cross term between the power data features in the i-th power data sample.
[0141] Furthermore, a nonlinear adjustment factor is defined to quantify the nonlinear relationship in the power data sample and is calculated based on the internal structure of the nonlinear eigenvector. Specifically, the nonlinear adjustment factor is obtained by calculating the response strength of the current weight matrix to the nonlinear eigenvector. The corresponding calculation formula is as follows:
[0142] ;
[0143] Where, is the nonlinear adjustment factor, which represents the adaptability of the model to the nonlinear relationship in the power data sample. The number of samples input to the extreme learning machine for the current batch.
[0144] Taking samples as the granularity, the calculation formula is expressed as:
[0145] .
[0146] Furthermore, the nonlinear adaptive adjustment term It is obtained by adjusting the current gradient update direction based on the nonlinear adjustment factor and the influence coefficient factor. The corresponding calculation formula is as follows:
[0147] ;
[0148] Taking samples as the granularity, the calculation formula is expressed as:
[0149] ;
[0150] Where, It is the weight hyperparameter of the nonlinear adaptive adjustment term, which is used to control the impact of the nonlinear adjustment term on the total weight update. Represents the current gradient calculation, indicating how the model adjusts the weights based on the error; ; is the number of samples input to the extreme learning machine in the current batch. Preferably, Set to 0.2.
[0151] It is worth noting that, in the weight update process, the embodiment of the present invention improves the robustness of the model, prevents overfitting, and optimizes the update direction through the disturbance factor and the nonlinear adaptive adjustment term.
[0152] In addition, in order to avoid the negative impact of excessive parameter update amplitude on model stability, the embodiment of the present invention introduces a sticky control strategy. During each parameter update, the update stride is adjusted according to the historical weight changes to avoid oscillation or too fast convergence.
[0153] Specifically, the current weight matrix is updated based on the gradient update amount to obtain the transition weight matrix. Then, based on the weight matrix difference and viscosity factor of adjacent iteration cycles, the transition weight matrix is smoothed and corrected to obtain the updated weight matrix. The corresponding calculation formula is as follows:
[0154] ;
[0155] Where, is the viscosity factor; is the weight matrix of the hidden layer of the extreme learning machine at the t+1th iteration (i.e., the updated weight matrix), which serves as the training parameter for the next iteration; is the weight matrix of the hidden layer of the extreme learning machine at the t-1th iteration; is the transition weight matrix of the hidden layer of the extreme learning machine at the t+1th iteration.
[0156] It is worth noting that the stickiness factor is calculated based on the average amplitude of historical weight changes to dynamically control the update step size, thereby suppressing oscillations and ensuring stable convergence. Its calculation formula is expressed as:
[0157] ;
[0158] Where T is the total number of iterations of the extreme learning machine. Preferably, T is set to 400.
[0159] After multiple rounds of training, the classifier finally outputs the classification result of each sample. This result is calculated by the output layer of the extreme learning machine and the activation function to obtain the final classification result. The calculation formula is expressed as:
[0160] ;
[0161] Where, is the Softmax activation function, is the classification output of the extreme learning machine.
[0162] After the extreme learning machine model (initial model) is trained, a target recognition model is generated. This target recognition model is used to identify the power data to be tested and output the corresponding load impact type. Specifically, the collected power data (power data to be tested) is input into the trained target recognition model for classification, thereby obtaining a classification result. In this embodiment of the present invention, the classification categories include normal fluctuation, slight impact, moderate impact, and severe impact.
[0163] See also Figure 3 , is a structural diagram of an embodiment of a load impact type identification system for distributed power grids provided by the present invention.
[0164] A second embodiment of the present invention provides a load impact type identification system for a distributed power grid, comprising:
[0165] An initial training set acquisition module 11 is used to acquire an initial training set of a target power system;
[0166] An expanded training set acquisition module 12 is configured to expand the samples of the initial training set using a preset generative adversarial network that introduces a local curvature perception mechanism to obtain an expanded training set; wherein the expanded training set includes a plurality of power data samples;
[0167] The target identification model construction module 13 is used to train the initial model through the expanded training set to obtain a target identification model; wherein the target identification model is used to identify the power data to be detected and output the corresponding load impact type.
[0168] It should be noted that the load impact type identification system for distributed power grids provided in the embodiment of the second aspect of the present invention can implement all the processes of the load impact type identification method for distributed power grids described in any embodiment of the first aspect above. The functions of each module and unit in the system and the technical effects achieved are respectively the same as the functions and technical effects achieved by the load impact type identification method for distributed power grids described in any embodiment of the first aspect above, and will not be repeated here.
[0169] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for identifying load impact types in a distributed power grid, characterized in that: include: Obtaining an initial training set of the target power system; wherein the initial training set is standardized; A preset generative adversarial network that introduces a local curvature perception mechanism is used to expand the sample of the initial training set to obtain an expanded training set; wherein the expanded training set includes a plurality of power data samples; the generator loss in the preset generative adversarial network includes: a multi-scale local curvature loss; The initial model is trained using the expanded training set to obtain a target recognition model; wherein the target recognition model is used to identify the power data to be detected and output the corresponding load impact type; The multi-scale local curvature loss is calculated by the following formula: ; in, is the multi-scale local curvature loss function; For the generator; A set of real power data samples input to the preset generative adversarial network for the current batch; For the The generated power data samples are Second-order derivatives at different scales; For the The real power data samples are Second-order derivatives at different scales; For the The third derivative of the generated power data samples; For the The third-order derivative of a real power data sample; For the Generate power data samples; For the Real power data samples; For the The first step to generate power data samples Features For the The real power data samples Features is the curvature loss balance coefficient; is the total number of scales; The number of samples input to the preset generative adversarial network for the current batch; is a positive integer; is a positive integer and ; is the L2 norm; The method further comprises: During the training of the initial model, a gradient update amount of the current weight matrix is calculated based on the expanded training set; wherein the gradient update amount is composed of at least the following three components: an interference elimination term, a disturbance term, and a nonlinear adaptive adjustment term; the interference elimination term is used to correct the gradient update direction based on the historical gradient vector of the power data sample; the disturbance term is used to adjust the gradient update direction based on the deviation between the power data sample and the sample mean vector; the nonlinear adaptive adjustment term is used to adjust the gradient update direction based on the nonlinear relationship in the power data sample; The nonlinear adaptive adjustment term is obtained by the following steps: Performing high-order feature expansion on the power data sample to obtain a corresponding nonlinear feature vector; wherein the nonlinear feature vector includes: the original features of the power data sample, the square terms of the original features, and the combined cross terms between the original features; Calculating the response intensity of the current weight matrix to the nonlinear eigenvector to obtain a nonlinear adjustment factor; wherein the nonlinear adjustment factor is used to characterize the adaptability of the initial model to the nonlinear relationship; Based on the nonlinear adjustment factor and the influence coefficient factor, the gradient update direction is adjusted to obtain the nonlinear adaptive adjustment item.
2. The method for identifying load impact types for distributed power grids according to claim 1, wherein: The generator loss in the preset generative adversarial network also includes: adversarial loss and high-order curvature alignment loss.
3. The method for identifying load impact types for distributed power grids according to claim 2, wherein: The generator loss is calculated by the following formula: ; in, is the loss function of the generator; A set of real power data samples input to the preset generative adversarial network for the current batch; is the set of latent space vectors corresponding to the real power data sample set; is the generator function; is the discriminator function; is the latent variable interaction term; The generator input is corresponding expectations; is the multi-scale local curvature loss function; For the A high-order curvature term that generates power data samples; For the The higher-order curvature terms of real power data samples; is the loss weight coefficient of the first generator; is the loss weight coefficient of the second generator; The number of samples input to the preset generative adversarial network for the current batch.
4. The method for identifying load impact types for distributed power grids according to claim 1, wherein: The interference elimination term is calculated by the following formula: ; in, is the interference elimination term; is the learning rate adjustment coefficient of the initial model; is the model output corresponding to the i-th power data sample; represents the gradient update direction of the i-th power data sample; is the historical gradient vector of the i-th power data sample; The number of samples input into the initial model for training for the current batch.
5. The method for identifying load impact types for distributed power grids according to claim 1, wherein: The disturbance term is calculated by the following formula: ; in, is the disturbance term; is the i-th power data sample; is the sample mean vector; is the model output corresponding to the i-th power data sample; represents the gradient update direction of the i-th power data sample; The number of samples input into the initial model for training for the current batch.
6. The method for identifying load impact types for distributed power grids according to claim 1, wherein: The method further comprises: Update the current weight matrix based on the gradient update amount to obtain a transition weight matrix; Based on the weight matrix difference and the viscosity factor of adjacent iteration cycles, the transition weight matrix is smoothly corrected to obtain an updated weight matrix.
7. A load impact type identification system for distributed power grids, characterized in that: A method for identifying load impact types for a distributed power grid applicable to any one of claims 1 to 6, comprising: An initial training set acquisition module, used to acquire an initial training set of a target power system; An expanded training set acquisition module is configured to expand the samples of the initial training set using a preset generative adversarial network that introduces a local curvature perception mechanism to obtain an expanded training set; wherein the expanded training set includes a plurality of power data samples; The target recognition model construction module is used to train the initial model through the expanded training set to obtain the target recognition model; wherein, the target recognition model is used to identify the power data to be detected and output the corresponding load impact type.
Citation Information
Patent Citations
Cloth super-resolution method based on generative adversarial network
CN112419157A
Centrifugal pump fault diagnosis system based on artificial intelligence
CN119026038A