Power transmission line multi-class fault waveform data generation method and system, medium and computer program product

Through GAN-based data generation model, fault waveform data that meets characteristics is generated, the problems of insufficient data and imbalance are solved, and the performance and robustness of the fault diagnosis model are improved.

CN120011808APending Publication Date: 2025-05-16WUHAN NARI LIABILITY OF STATE GRID ELECTRIC POWER RES INST +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411899910.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing fault diagnosis methods rely on a large amount of fault waveform data, but it is difficult to obtain these data, especially inadequate data of rare fault types, resulting in limited improvement in diagnostic model performance, especially in multi-category fault diagnosis tasks.

Method used

Using a data generation model based on a generative adversarial network (GAN), a generator and discriminator are constructed, and real fault waveform data is used for training, to generate synthetic data that meets the characteristics of a few types of data, expand the data set, and balance the data distribution.

Benefits of technology

Effectively generate high-quality multi-category fault waveform data, expanding a few-category data sets, making the distribution of various categories more balanced, improving the performance and robustness of the fault diagnosis model, and avoiding the problem of the model biasing towards most classes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011808A_ABST
    Figure CN120011808A_ABST
Patent Text Reader

Abstract

The invention discloses a power transmission line multi-class fault waveform data generation method and system, a medium and a computer program product. The method comprises the following steps: constructing a data generation model based on a generative adversarial network; training the data generation model by using real power transmission line multi-class fault waveform data to obtain a trained data generation model; and inputting randomly generated noise data to the trained data generation model to obtain simulated multi-class fault waveform data of the power transmission line. According to the method, real minority class samples are effectively simulated, a minority class data set is expanded, and data distribution of all classes is relatively balanced; according to the method, the Wasserstein distance is utilized to improve the model training process to solve the gradient disappearance problem; a gradient penalty stability training process is introduced to ensure that the output of the generator is smoother and more stable, a data set containing various variables and complex relationships can be processed, high-dimensional multi-coupling complex data can be adapted, and high-quality data conforming to the category of fault features can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power system fault diagnosis, and in particular relates to a method, system, medium and computer program product for generating multi-category fault waveform data of a power transmission line. Background Art

[0002] During the operation of a transmission line, a variety of faults may occur. According to the frequency of occurrence, they can be generally divided into lightning strikes and non-lightning strikes. Lightning strikes can be further divided into two categories: backlash and backlash. Non-lightning strikes can be further divided into wind deflection, bird damage, ice damage, external force damage and others. However, existing fault diagnosis methods usually rely on a large amount of fault waveform data, but it is very difficult to obtain sufficient fault waveform data in practice, especially for some rare fault types. This lack of data limits the improvement of the performance of the fault diagnosis model, especially when dealing with multi-category fault diagnosis tasks, the problem of data imbalance is more prominent. Therefore, how to effectively generate sufficient and high-quality fault waveform data has become the key to improving the performance of the fault diagnosis system.

[0003] The oversampling technology used in traditional technology is difficult to adapt to high-dimensional complex data, and the actual transmission line fault waveform data is coupled with a variety of environmental factors. Simple random generation of data points cannot effectively reflect the data characteristics of this category. There are also many data enhancement methods in deep learning, but the methods that can effectively process high-dimensional complex time series data are limited, and can only modify the original data, which cannot effectively solve the problem of small data volume. Summary of the invention

[0004] In order to overcome the shortcomings of traditional technologies and improve the performance of fault diagnosis systems, the present invention proposes a method, system, medium and computer program product for generating multi-category fault waveform data for power transmission lines.

[0005] A method for generating multi-category fault waveform data of a power transmission line to achieve one of the purposes of the present invention comprises:

[0006] Build a data generation model based on generative adversarial networks;

[0007] Using real multi-category fault waveform data of power transmission lines to train the data generation model, to obtain a trained data generation model;

[0008] The randomly generated noise data is input into the trained data generation model to obtain the simulated multi-category fault waveform data of the transmission line.

[0009] Furthermore, the data generation model includes a generator and a discriminator, the generator is used to generate fault waveform data consistent with the characteristics of real transmission line multi-category fault waveform data; the discriminator is used to distinguish whether the input data is real transmission line multi-category fault waveform data or fault waveform data generated by the generator.

[0010] Furthermore, the method for training the discriminator of the data generation model includes:

[0011] Sampling from the real distribution of multi-category fault waveform data of transmission lines to obtain real samples;

[0012] Input the noise data sampled from the prior distribution into the generator to obtain fake samples;

[0013] An interpolation method is used to obtain multiple interpolation points from real samples and fake samples;

[0014] The real samples, the fake samples and the multiple interpolation points are input into the constructed loss function based on the Wasserstein distance, and the parameters of the discriminator are iteratively updated according to the loss function.

[0015] Furthermore, the method for training the data generation model includes:

[0016] Initialize model parameters: including batch size, learning rate, gradient penalty coefficient

[0017] Training the discriminator: In each training cycle, the discriminator will perform multiple parameter updates to ensure that the discriminator can better approximate the Wasserstein distance. For each batch of data, the training method includes:

[0018] 1) Sampling real samples from the real distribution of multi-category fault waveform data of transmission lines;

[0019] 2) Sample noise from the prior distribution and input the noise into the generator to obtain false samples;

[0020] 3) Perform linear interpolation between real samples and fake samples to obtain multiple interpolation points;

[0021] 4) Calculating the loss function of the discriminator, the loss function includes three parts: a loss function for discriminating the score of real samples, a loss function for discriminating the score of fake samples, and a loss function for the gradient penalty of the interpolation point; the gradient penalty ensures that the discriminator function has appropriate smoothness with respect to its input;

[0022] Training the generator: After the discriminator is updated multiple times, the generator is trained and its parameters are updated; the goal of the generator is to generate fake samples that can cause the discriminator to misjudge, and the goal is achieved by minimizing the difference in the discriminator's scores for fake samples and real samples; the method for updating the generator includes: inputting the generated random noise into the generator to obtain generated data, calculating the loss function after the generated data passes through the discriminator, and updating the generator parameters according to the loss function;

[0023] Repeat the above steps of training the discriminator and training the generator until the predetermined number of iterations is reached or the model performance no longer improves.

[0024] After training is complete, the generator can be used to generate new fault data samples.

[0025] Furthermore, the loss function includes: a score for distinguishing real samples, a score for distinguishing fake samples, and a gradient penalty for interpolation points. The introduction of gradient penalty can better meet the Lipschitz constraint and improve the stability of the model.

[0026] Furthermore, the loss function includes: min G max D L(D,G)

[0027]

[0028] in, is the distribution of real fault waveform data, is the fault waveform data distribution generated by the generator, x represents the sample input to the discriminator, is the generated data output by the generator, Tong is the mixed distribution of the real fault waveform data and the generated fault waveform data, λ is the weight of the gradient penalty, D(x) represents the score of the discriminator for the input data x, Represents the discriminator to generate data score.

[0029] Furthermore, it also includes data preprocessing of real transmission line multi-category fault waveform data, the preprocessing method including data cleaning and / or data standardization;

[0030] The data cleaning method includes: counting the data length of the sampling rate of the multi-category fault waveform data of the transmission line; comparing the data length of the multi-category fault waveform data of the transmission line with the same sampling rate, and eliminating the multi-category fault waveform data of the transmission line with abnormal data length;

[0031] Methods for data normalization include:

[0032] When the length of the multi-category fault waveform data of the transmission line is greater than the data length specified by the sampling rate:

[0033] If the multi-category fault waveform data of the power transmission line is continuous waveform data, the end of the multi-category fault waveform data of the power transmission line is truncated to reach the standard length; otherwise, the discontinuous part of the waveform is interpolated so that the data length reaches the standard length of the sampling rate;

[0034] When the data length of the multi-category fault waveform of the transmission line is less than the data length specified by the sampling rate:

[0035] If the transmission line multi-category fault waveform data is continuous waveform data, interpolation is performed at the end of the transmission line multi-category fault waveform data so that the data length reaches the standard length; otherwise, interpolation is performed at the discontinuous part of the waveform.

[0036] A system for generating multi-category fault waveform data of a power transmission line to achieve the second objective of the present invention comprises:

[0037] Model building module: used to build a data generation model based on generative adversarial networks;

[0038] Model training module: using real multi-category fault waveform data of power transmission lines to train the data generation model to obtain a trained data generation model;

[0039] Data generation module: used to input randomly generated noise data into the trained data generation model to obtain simulated transmission line multi-category fault waveform data.

[0040] Furthermore, the data generation model includes a generator and a discriminator, the generator is used to generate fault waveform data consistent with the characteristics of real transmission line multi-category fault waveform data; the discriminator is used to distinguish whether the input data is real transmission line multi-category fault waveform data or fault waveform data generated by the generator.

[0041] Furthermore, a discriminator training module is included, which is used to train the discriminator, and the training method includes:

[0042] Sampling from the real distribution of multi-category fault waveform data of transmission lines to obtain real samples;

[0043] Input the noise data sampled from the prior distribution into the generator to obtain fake samples;

[0044] An interpolation method is used to obtain multiple interpolation points from real samples and fake samples;

[0045] The real samples, the fake samples and the multiple interpolation points are input into the constructed loss function based on the Wasserstein distance, and the parameters of the discriminator are iteratively updated according to the loss function.

[0046] Furthermore, the loss function includes: a score for distinguishing real samples, a score for distinguishing fake samples, and a gradient penalty for interpolation points.

[0047] Furthermore, it also includes a data generation model training module for training the data generation model, and the training method includes:

[0048] Initialize model parameters: including batch size, learning rate, and gradient penalty coefficient;

[0049] Training the discriminator: In each training cycle, the discriminator performs multiple parameter updates to ensure that the discriminator can better approximate the Wasserstein distance. For each batch of data, the training method includes:

[0050] The real samples are obtained by sampling from the real distribution of multi-category fault waveform data of transmission lines;

[0051] Sample noise from the prior distribution and input the noise into the generator to obtain fake samples;

[0052] Perform linear interpolation between real samples and fake samples to obtain multiple interpolation points;

[0053] Calculate the loss function of the discriminator, the loss function includes three parts: a score for discriminating real samples, a score for discriminating fake samples, and a gradient penalty for the interpolation point; the gradient penalty ensures that the discriminator function has appropriate smoothness with respect to its input;

[0054] Training the generator: After the discriminator is updated multiple times, the generator is trained and its parameters are updated; the goal of the generator is to generate fake samples that can cause the discriminator to misjudge, and the goal is achieved by minimizing the difference in the discriminator's scores for fake samples and real samples; the method for updating the generator includes: inputting the generated random noise into the generator to obtain generated data, calculating the loss function after the generated data passes through the discriminator, and updating the generator parameters according to the loss function;

[0055] Repeat the above steps of training the discriminator and training the generator until the predetermined number of iterations is reached or the model performance no longer improves.

[0056] Furthermore, it also includes a data preprocessing module for performing data preprocessing on real multi-category fault waveform data of power transmission lines, and the preprocessing method includes data cleaning and / or data standardization;

[0057] The data cleaning method includes: counting the data length of the sampling rate of the multi-category fault waveform data of the transmission line; comparing the data length of the multi-category fault waveform data of the transmission line with the same sampling rate, and eliminating the multi-category fault waveform data of the transmission line with abnormal data length;

[0058] Methods for data normalization include:

[0059] When the length of the multi-category fault waveform data of the transmission line is greater than the data length specified by the sampling rate:

[0060] If the multi-category fault waveform data of the power transmission line is continuous waveform data, the end of the multi-category fault waveform data of the power transmission line is truncated to reach the standard length; otherwise, the discontinuous part of the waveform is interpolated so that the data length reaches the standard length of the sampling rate;

[0061] When the data length of the multi-category fault waveform of the transmission line is less than the data length specified by the sampling rate:

[0062] If the transmission line multi-category fault waveform data is continuous waveform data, interpolation is performed at the end of the transmission line multi-category fault waveform data so that the data length reaches the standard length; otherwise, interpolation is performed at the discontinuous part of the waveform.

[0063] The beneficial effects of the present invention include:

[0064] The present invention generates synthetic data that conforms to the characteristics of minority class data through WGAN-GP. These data can effectively imitate real minority class samples, expand the minority class data set, and make the distribution of data of each category relatively balanced, thereby helping the model to better learn the characteristics of these categories during the training process, avoiding the situation that when the sample size of one or several categories is significantly lower than that of other categories, the model often tends to be biased towards the majority class, resulting in weak recognition ability for the minority class, thereby allowing the classifier to play a role and achieve better results.

[0065] Traditional oversampling methods may cause the model to overfit the minority class, while simple data enhancement techniques (such as rotation, scaling, etc.) may not be sufficient to capture the multi-dimensional coupling characteristics of complex data. Compared with traditional oversampling methods and data enhancement methods, the present invention uses Wasserstein distance to improve the training process of GAN, solves the gradient vanishing problem that may occur in traditional GAN, and further introduces gradient penalty on this basis to stabilize the training process, ensuring that the output of the generator is smoother and more stable. It can handle data sets containing multiple variables and complex relationships, can adapt to high-dimensional multi-coupled complex data, and generate high-quality data that meets the characteristics of this category of faults. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a schematic flow diagram of the method of the present invention;

[0067] Figure 2 This is a schematic diagram of the structure of the data generation model based on the generative adversarial network;

[0068] Figure 3 It is a comparison chart of the original fault waveform data and the generated fault waveform data. DETAILED DESCRIPTION

[0069] The following specific implementations are used to explain the technical solutions of the claims of the present invention so that those skilled in the art can understand the claims. The protection scope of the present invention is not limited to the following specific implementation structures. The technical solutions of the claims of the present invention made by those skilled in the art but different from the following specific implementations are also within the protection scope of the present invention.

[0070] Example 1

[0071] (1) Transmission line fault waveform data processing part, including: data reading and data preprocessing:

[0072] like Figure 1 As shown, the original fault waveform data is read to obtain three types of data contained in the original waveform data set, namely, waveform sampling rate, fault waveform data, and fault type label.

[0073] The data preprocessing includes:

[0074] Data cleaning

[0075] Since some data in the data set are missing fault type labels, and some waveform data have missing data and garbled symbols, this type of data does not meet the standards and cannot be used. Therefore, the waveform data and label data are cleaned, the unusable data are eliminated, and the data that meets the experimental conditions is screened out.

[0076] Because the waveform data may be collected by different sensor devices, and the sampling rates of different sensor devices are different, different sampling rates will lead to inconsistent lengths of waveform data. When modeling, data of different lengths will make the model unable to be trained. Therefore, the data that meets the experimental conditions obtained in the above steps must be further sorted out, the waveform sampling rate of each data is counted, the data length of each data is calculated, and the statistical data length and sampling rate are integrated. The data lengths of data with the same sampling rate (the same sampling rate means the same or similar data length) are compared, and data with abnormal data lengths are eliminated. The abnormal data length includes: the difference between the actual length of the data and the data length corresponding to the sampling rate is greater than a set ratio. The set ratio is set according to actual needs and is not limited here.

[0077] Data Standardization

[0078] Then the data is normalized according to the corresponding relationship between sampling rate and data length:

[0079] When the data length is longer than the data length specified by the sampling rate, first determine whether the data is a continuous waveform through visualization operations. If the waveform is continuous, the end of the data sequence can be truncated to reach the standard length; if the waveform is discontinuous, the disconnected part is subjected to a local cubic spline interpolation algorithm. This algorithm can be combined with the waveform trend for interpolation so that the data length reaches the standard length of the sampling rate.

[0080] When the data length is shorter than the data length specified by the sampling rate, we also first check whether the data is a continuous waveform through visualization operations. If the waveform is continuous, a local cubic spline interpolation algorithm is performed at the end of the data to make the overall waveform data reach the standard length; if the waveform is discontinuous, a local cubic spline interpolation algorithm is performed at the break.

[0081] Data Normalization

[0082] In order to improve the training efficiency of the model and the quality of generated data, the filtered data is further normalized so that the distribution range of the data is unified to [0, 1] or [-1, 1].

[0083] (2) Use the pre-processed fault waveform data to train the WGAN-GP (Wasserstein Generative Adversarial Network with Gradient Penalty) model

[0084] First, build the model, initialize the environment, and build the WGAN-GP model structure with gradient penalty, such as Figure 2 As shown, it includes a generator and a discriminator. This model includes a convolution layer, a leaky ReLU layer, a layer normalization, and a transposed convolution layer. The generator is used to generate fault waveform data that is consistent with the characteristics (such as statistical characteristics) of the real fault data; the discriminator is used to distinguish whether the input data is from the fault waveform data in the real data set or the fault waveform data generated by the generator. The optimizer in the model training usually uses Adaptive Moment Estimation (Adam) to optimize the network parameters of the generator and the discriminator. For the WGAN-GP model, the Wasserstein distance is usually used as the loss function. In addition, the gradient penalty (GradientPenalty) term is introduced to meet the Lipschitz constraint and improve the model stability. The algorithm implementation process is as follows:

[0085] ①First, initialize the model parameters. You need to initialize the parameters of the generator G and the discriminator D.

[0086] ② Secondly, train the discriminator. In each training cycle, the discriminator will be updated multiple times to ensure that the discriminator can better approximate the Wasserstein distance. For each batch of data:

[0087] 1) From the distribution of real fault waveform data The real samples are obtained by sampling.

[0088] 2) Sample noise from the prior distribution and generate fake samples through the generator G.

[0089] 3) Calculate the gradient penalty and sample interpolation points between real samples and fake samples The interpolation points are generated by linear interpolation.

[0090] 4) The loss function of the discriminator consists of three parts: the score for discriminating real samples, the score for discriminating fake samples, and the gradient penalty of the interpolation point. The purpose of the gradient penalty is to ensure that the discriminator's response to the input data is smoother and more stable.

[0091] ③Then, update the generator. After the discriminator is updated several times, the generator is updated once. The goal of the generator is to generate fake samples that can cause the discriminator to misjudge, which is achieved by minimizing the discriminator's score for the generated fake samples. Specifically, the generator optimizes its parameters through back propagation so that the discriminator scores the fake samples as high as possible, thereby "deceiving" the discriminator and making it difficult to distinguish between real samples and fake samples.

[0092] ④ Then, repeat the training process of steps ② and ③ above until the predetermined number of iterations is reached or the model performance no longer improves.

[0093] ⑤Finally, model evaluation and application, after the training is completed, the generator can be used to generate new fault waveform data samples and evaluate their quality and diversity.

[0094] The loss function is defined as:

[0095] min G max D L(D,G)

[0096] This formula can be interpreted as: the generator G hopes that the generated samples can minimize the score difference between the real samples and the fake samples of the discriminator D, while the discriminator D hopes to maximize the score difference between the real samples and the fake samples. Through this adversarial process, the generator and the discriminator improve each other, thereby improving the performance of the generated model.

[0097] maxD Indicates: For the discriminator D, its goal is to maximize the loss function L(D,G). The discriminator hopes to maximize the score difference between the real sample and the generated sample, that is, to maximize the score of the discriminator for the real sample minus the score for the generated sample.

[0098] minG means: For the generator G, its goal is to minimize the loss function L(D,G). The generator hopes to minimize the score of the discriminator on the generated samples, that is, the generator generates realistic fake samples so that the discriminator cannot distinguish between these fake samples and real samples, thereby minimizing the score of the discriminator.

[0099] in:

[0100]

[0101] in, is the distribution of real fault waveform data, is the fault waveform data distribution generated by the generator G, x represents the sample input to the discriminator, is the generated data output by the generator, Tong is the mixed distribution of the real fault waveform data and the generated fault waveform data, λ is the weight of the gradient penalty; D(x) represents the score of the discriminator for the input data x, Represents the discriminator to generate data score.

[0102] Input the preprocessed and normalized fault waveform data into the constructed WGAN-GP model for training, and initialize the training parameters: including batch size, learning rate, and gradient penalty coefficient. The training iteration process involves iteratively updating the parameters of the generator and the discriminator. When training the discriminator, the generator parameters are fixed and the discriminator is updated. Using a batch of real fault waveform data and a batch of fault waveform data generated by the generator, the loss function of the discriminator is calculated, and the discriminator parameters are updated according to the loss function. When training the generator, the discriminator parameters are fixed and the generator parameters are updated. Using a batch of random noise, the random noise is input into the generator to obtain the generated data, the loss function after the generated data passes through the discriminator is calculated, and the generator parameters are updated according to the loss function. The discriminator is trained more than the generator to maintain a balance between the two, and the parameters of the generator and the discriminator are continuously updated. After a sufficient number of iterative training, when the generator can generate high-quality data that is close to the characteristics of the real data, the training can be stopped. The trained data generation model is obtained for subsequent generation of new fault waveform data or further research.

[0103] (3) Use the trained data generation model to generate fault waveform data:

[0104] First, a batch of random noise data needs to be generated as the input of the generator. These random noises usually come from a certain distribution, such as uniform distribution or Gaussian distribution. The dimension of the random noise should match the noise dimension used in training. Determine the number of generated data, that is, the size of the random noise batch, which depends on the number of fault waveform data that need to be generated. Then input the prepared random noise into the trained generator. The generator will generate new fault waveform data based on the input noise. These data should simulate the distribution characteristics of the training data. For each batch of random noise, the generator will output the corresponding batch of fault waveform data. Repeat the process until a sufficient amount of data is generated. Then, the generated data is denormalized and restored to the original numerical range to ensure the practicality and comparability of the data. Output the comparison chart of the data generated by the generator and the original data, which intuitively reflects the quality of the generated data, so as to facilitate targeted tuning of the model. Save the generated high-quality fault waveform data to the file system or database for subsequent use. As needed, organize the generated data, including adjusting the data format, assigning labels, etc., for subsequent analysis and application.

[0105] Through the above steps, the data generation model based on WGAN-GP can be used to generate multi-category transmission line fault waveform data, such as Figure 3 As shown in the figure, it can be seen that the real multi-category fault waveform data of the transmission line is very close to the simulated multi-category fault waveform data of the transmission line generated by the model. These simulated multi-category fault waveform data of the transmission line can not only help solve the problem of scarcity of actual data, but also improve the performance and robustness of the fault diagnosis model.

[0106] Example 2

[0107] A method for generating multi-category fault waveform data of a power transmission line, comprising:

[0108] Build a data generation model based on generative adversarial networks;

[0109] Using real multi-category fault waveform data of power transmission lines to train the data generation model, to obtain a trained data generation model;

[0110] The randomly generated noise data is input into the trained data generation model to obtain the simulated multi-category fault waveform data of the transmission line.

[0111] In some embodiments, the data generation model includes a generator and a discriminator, the generator is used to generate fault waveform data consistent with the characteristics of real transmission line multi-category fault waveform data; the discriminator is used to distinguish whether the input data is real transmission line multi-category fault waveform data or fault waveform data generated by the generator.

[0112] In some embodiments, a method for training a discriminator of the data generation model includes:

[0113] Sampling from the real distribution of multi-category fault waveform data of transmission lines to obtain real samples;

[0114] Input the noise data sampled from the prior distribution into the generator to obtain fake samples;

[0115] An interpolation method is used to obtain multiple interpolation points from real samples and fake samples;

[0116] The real samples, the fake samples and the multiple interpolation points are input into the constructed loss function based on the Wasserstein distance, and the parameters of the discriminator are iteratively updated according to the loss function.

[0117] In some embodiments, a method for training a data generation model includes:

[0118] Initialize model parameters; including batch size, learning rate, gradient penalty coefficient

[0119] Train the discriminator; in each training cycle, the discriminator will perform multiple parameter updates to ensure that the discriminator can better approximate the Wasserstein distance. For each batch of data, the training method includes:

[0120] 1) Sampling real samples from the real fault waveform data distribution;

[0121] 2) Sample noise from the prior distribution and input the noise into the generator to obtain a fake sample;

[0122] 3) Perform linear interpolation between real samples and fake samples to obtain multiple interpolation points;

[0123] 4) Calculate the loss function of the discriminator, the loss function includes three parts: the score of discriminating real samples, the score of discriminating fake samples and the gradient penalty of the interpolation point; the gradient penalty ensures that the discriminator function has appropriate smoothness with respect to its input;

[0124] Training the generator: After the discriminator is updated multiple times, the generator is trained and its parameters are updated; the goal of the generator is to generate fake samples that can cause the discriminator to misjudge, and the goal is achieved by minimizing the difference in the discriminator's scores for fake samples and real samples; the method for updating the generator includes: inputting randomly generated random noise into the generator to obtain generated data, calculating the loss function after the generated data passes through the discriminator, and updating the generator parameters according to the loss function;

[0125] Repeat the above training process of the discriminator and generator until the predetermined number of iterations is reached or the model performance no longer improves;

[0126] After training is complete, the generator can be used to generate new fault data samples.

[0127] In some embodiments, the loss function includes: a score for distinguishing real samples, a score for distinguishing fake samples, and a gradient penalty for interpolation points. The introduction of gradient penalty can better meet the Lipschitz constraint and improve the stability of the model.

[0128] In some embodiments, the loss function includes: min G max D L(D,G)

[0129]

[0130] in, is the distribution of real fault waveform data, is the fault waveform data distribution generated by the generator, is the generated data output by the generator, Tong is the mixed distribution of the real fault waveform data and the generated fault waveform data, λ is the weight of the gradient penalty; D(x) represents the score of the discriminator for the input data x, Represents the discriminator to generate data score.

[0131] In some embodiments, it also includes performing data preprocessing on real transmission line multi-category fault waveform data, and the preprocessing method includes data cleaning and / or data standardization;

[0132] The data cleaning method includes: counting the data length of the sampling rate of the multi-category fault waveform data of the transmission line; comparing the data length of the multi-category fault waveform data of the transmission line with the same sampling rate, and eliminating the multi-category fault waveform data of the transmission line with abnormal data length;

[0133] Methods for data normalization include:

[0134] When the length of the multi-category fault waveform data of the transmission line is greater than the data length specified by the sampling rate:

[0135] If the multi-category fault waveform data of the power transmission line is continuous waveform data, the end of the multi-category fault waveform data of the power transmission line is truncated to reach the standard length; otherwise, the discontinuous part of the waveform is interpolated so that the data length reaches the standard length of the sampling rate;

[0136] When the data length of the multi-category fault waveform of the transmission line is less than the data length specified by the sampling rate:

[0137] If the transmission line multi-category fault waveform data is continuous waveform data, interpolation is performed at the end of the transmission line multi-category fault waveform data so that the data length reaches the standard length; otherwise, interpolation is performed at the discontinuous part of the waveform.

[0138] Example 3

[0139] A system for generating multi-category fault waveform data of a power transmission line, comprising:

[0140] Model building module: used to build a data generation model based on generative adversarial networks;

[0141] Model training module: using real multi-category fault waveform data of power transmission lines to train the data generation model to obtain a trained data generation model;

[0142] Data generation module: used to input randomly generated noise data into the trained data generation model to obtain simulated transmission line multi-category fault waveform data.

[0143] In some embodiments, the data generation model includes a generator and a discriminator, the generator is used to generate fault waveform data consistent with the characteristics of real transmission line multi-category fault waveform data; the discriminator is used to distinguish whether the input data is real transmission line multi-category fault waveform data or fault waveform data generated by the generator.

[0144] In some embodiments, a discriminator training module is further included, for training the discriminator, and the training method includes:

[0145] Sampling from the real distribution of multi-category fault waveform data of transmission lines to obtain real samples;

[0146] Input the noise data sampled from the prior distribution into the generator to obtain fake samples;

[0147] An interpolation method is used to obtain multiple interpolation points from real samples and fake samples;

[0148] The real samples, the fake samples and the multiple interpolation points are input into the constructed loss function based on the Wasserstein distance, and the parameters of the discriminator are iteratively updated according to the loss function.

[0149] In some embodiments, the loss function includes: a score for distinguishing real samples, a score for distinguishing fake samples, and a gradient penalty for interpolation points.

[0150] In some embodiments, a data generation model training module is further included, which is used to train the data generation model. The training method includes:

[0151] Initialize model parameters: including batch size, learning rate, and gradient penalty coefficient;

[0152] Training the discriminator: In each training cycle, the discriminator performs multiple parameter updates to ensure that the discriminator can better approximate the Wasserstein distance. For each batch of data, the training method includes:

[0153] The real samples are obtained by sampling from the real distribution of multi-category fault waveform data of transmission lines;

[0154] Sample noise from the prior distribution and input the noise into the generator to obtain fake samples;

[0155] Perform linear interpolation between real samples and fake samples to obtain multiple interpolation points;

[0156] Calculate the loss function of the discriminator, the loss function includes three parts: a score for discriminating real samples, a score for discriminating fake samples, and a gradient penalty for the interpolation point; the gradient penalty ensures that the discriminator function has appropriate smoothness with respect to its input;

[0157] Training the generator: After the discriminator is updated multiple times, the generator is trained and its parameters are updated; the goal of the generator is to generate fake samples that can cause the discriminator to misjudge, and the goal is achieved by minimizing the difference in the discriminator's scores for fake samples and real samples; the method for updating the generator includes: inputting the generated random noise into the generator to obtain generated data, calculating the loss function after the generated data passes through the discriminator, and updating the generator parameters according to the loss function;

[0158] Repeat the above steps of training the discriminator and training the generator until the predetermined number of iterations is reached or the model performance no longer improves.

[0159] In some embodiments, a data preprocessing module is further included, which is used to perform data preprocessing on real multi-category fault waveform data of power transmission lines, and the preprocessing method includes data cleaning and / or data standardization;

[0160] The data cleaning method includes: counting the data length of the sampling rate of the multi-category fault waveform data of the transmission line; comparing the data length of the multi-category fault waveform data of the transmission line with the same sampling rate, and eliminating the multi-category fault waveform data of the transmission line with abnormal data length;

[0161] Methods for data normalization include:

[0162] When the length of the multi-category fault waveform data of the transmission line is greater than the data length specified by the sampling rate:

[0163] If the multi-category fault waveform data of the power transmission line is continuous waveform data, the end of the multi-category fault waveform data of the power transmission line is truncated to reach the standard length; otherwise, the discontinuous part of the waveform is interpolated so that the data length reaches the standard length of the sampling rate;

[0164] When the data length of the multi-category fault waveform of the transmission line is less than the data length specified by the sampling rate:

[0165] If the transmission line multi-category fault waveform data is continuous waveform data, interpolation is performed at the end of the transmission line multi-category fault waveform data so that the data length reaches the standard length; otherwise, interpolation is performed at the discontinuous part of the waveform.

[0166] Example 4

[0167] A computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the various steps of the method of the present invention are implemented, which will not be described in detail herein.

[0168] The computer-readable storage medium may be the data transmission device provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (smartmedia card, SMC), a secure digital (secure digital, SD) card, a flash card (flash card), etc., provided on the computer device.

[0169] Furthermore, the computer-readable storage medium may include both an internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data to be output or has been output.

[0170] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0171] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0172] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0173] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0174] Example 5

[0175] A computer program product comprises a computer program / instruction, which implements any step of the method for generating multi-category fault waveform data of a power transmission line when the computer program / instruction is executed by a processor.

[0176] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0177] The contents not described in detail in this specification belong to the prior art known to professional and technical personnel in this field.

Claims

1. A method for generating multi-category fault waveform data of a transmission line, characterized in that: include: Build a data generation model based on generative adversarial networks; Using real multi-category fault waveform data of power transmission lines to train the data generation model, to obtain a trained data generation model; The randomly generated noise data is input into the trained data generation model to obtain the simulated multi-category fault waveform data of the transmission line.

2. The method for generating multi-category fault waveform data of a power transmission line according to claim 1, characterized in that: The data generation model includes a generator and a discriminator, wherein the generator is used to generate fault waveform data consistent with the characteristics of multi-category fault waveform data of real power transmission lines; The discriminator is used to distinguish whether the input data is real multi-category fault waveform data of the power transmission line or fault waveform data generated by the generator.

3. The method for generating multi-category fault waveform data of a power transmission line according to claim 2, characterized in that: The method for training the discriminator of the data generation model includes: Sampling from the real distribution of multi-category fault waveform data of transmission lines to obtain real samples; Input the noise data sampled from the prior distribution into the generator to obtain fake samples; An interpolation method is used to obtain multiple interpolation points from real samples and fake samples; The real samples, the fake samples and the multiple interpolation points are input into the constructed loss function based on the Wasserstein distance, and the parameters of the discriminator are iteratively updated according to the loss function.

4. The method for generating multi-category fault waveform data of a power transmission line according to claim 3, characterized in that: The loss function includes: a score for distinguishing real samples, a score for distinguishing fake samples, and a gradient penalty for interpolation points.

5. The method for generating multi-category fault waveform data of a power transmission line according to claim 2, characterized in that: Methods for training data generation models include: Initialize model parameters: including batch size, learning rate, and gradient penalty coefficient; Training the discriminator: In each training cycle, the discriminator performs multiple parameter updates to ensure that the discriminator can better approximate the Wasserstein distance. For each batch of data, the training method includes: The real samples are obtained by sampling from the real distribution of multi-category fault waveform data of transmission lines; Sample noise from the prior distribution and input the noise into the generator to obtain fake samples; Perform linear interpolation between real samples and fake samples to obtain multiple interpolation points; Calculate the loss function of the discriminator, the loss function includes three parts: a score for discriminating real samples, a score for discriminating fake samples, and a gradient penalty for the interpolation point; the gradient penalty ensures that the discriminator function has appropriate smoothness with respect to its input; Training the generator: After the discriminator is updated multiple times, the generator is trained and its parameters are updated; the goal of the generator is to generate fake samples that can cause the discriminator to misjudge, and the goal is achieved by minimizing the difference in the discriminator's scores for fake samples and real samples; the method for updating the generator includes: inputting the generated random noise into the generator to obtain generated data, calculating the loss function after the generated data passes through the discriminator, and updating the generator parameters according to the loss function; Repeat the above steps of training the discriminator and training the generator until the predetermined number of iterations is reached or the model performance no longer improves.

6. The method for generating multi-category fault waveform data of a power transmission line according to claim 2, characterized in that: It also includes data preprocessing of real transmission line multi-category fault waveform data, where the preprocessing method includes data cleaning and / or data standardization; The data cleaning method includes: counting the data length of the sampling rate of the multi-category fault waveform data of the transmission line; comparing the data length of the multi-category fault waveform data of the transmission line with the same sampling rate, and eliminating the multi-category fault waveform data of the transmission line with abnormal data length; Methods for data normalization include: When the length of the multi-category fault waveform data of the transmission line is greater than the data length specified by the sampling rate: If the multi-category fault waveform data of the power transmission line is continuous waveform data, the end of the multi-category fault waveform data of the power transmission line is truncated to reach the standard length; otherwise, the discontinuous part of the waveform is interpolated so that the data length reaches the standard length of the sampling rate; When the data length of the multi-category fault waveform of the transmission line is less than the data length specified by the sampling rate: If the transmission line multi-category fault waveform data is continuous waveform data, interpolation is performed at the end of the transmission line multi-category fault waveform data so that the data length reaches the standard length; otherwise, interpolation is performed at the discontinuous part of the waveform.

7. A system for generating multi-category fault waveform data of a power transmission line, characterized in that: include: Model building module: used to build a data generation model based on generative adversarial networks; Model training module: using real multi-category fault waveform data of power transmission lines to train the data generation model to obtain a trained data generation model; Data generation module: used to input randomly generated noise data into the trained data generation model to obtain simulated transmission line multi-category fault waveform data.

8. The system for generating multi-category fault waveform data of a power transmission line according to claim 7, characterized in that: The data generation model includes a generator and a discriminator, wherein the generator is used to generate fault waveform data consistent with the characteristics of multi-category fault waveform data of real power transmission lines; The discriminator is used to distinguish whether the input data is real multi-category fault waveform data of the power transmission line or fault waveform data generated by the generator.

9. The system for generating multi-category fault waveform data of a power transmission line according to claim 8, characterized in that: It also includes a discriminator training module for training the discriminator, and the training method includes: Sampling from the real distribution of multi-category fault waveform data of transmission lines to obtain real samples; Input the noise data sampled from the prior distribution into the generator to obtain fake samples; An interpolation method is used to obtain multiple interpolation points from real samples and fake samples; The real samples, the fake samples and the multiple interpolation points are input into the constructed loss function based on the Wasserstein distance, and the parameters of the discriminator are iteratively updated according to the loss function.

10. The system for generating multi-category fault waveform data of a power transmission line according to claim 8, characterized in that: The loss function includes: a score for distinguishing real samples, a score for distinguishing fake samples, and a gradient penalty for interpolation points.

11. The system for generating multi-category fault waveform data of a power transmission line according to claim 7, characterized in that: It also includes a data generation model training module for training the data generation model. The training method includes: Initialize model parameters: including batch size, learning rate, and gradient penalty coefficient; Training the discriminator: In each training cycle, the discriminator performs multiple parameter updates to ensure that the discriminator can better approximate the Wasserstein distance. For each batch of data, the training method includes: The real samples are obtained by sampling from the real distribution of multi-category fault waveform data of transmission lines; Sample noise from the prior distribution and input the noise into the generator to obtain fake samples; Perform linear interpolation between real samples and fake samples to obtain multiple interpolation points; Calculate the loss function of the discriminator, the loss function includes three parts: a score for discriminating real samples, a score for discriminating fake samples, and a gradient penalty for the interpolation point; the gradient penalty ensures that the discriminator function has appropriate smoothness with respect to its input; Training the generator: After the discriminator is updated multiple times, the generator is trained and its parameters are updated; the goal of the generator is to generate fake samples that can cause the discriminator to misjudge, and the goal is achieved by minimizing the difference in the discriminator's scores for fake samples and real samples; the method for updating the generator includes: inputting the generated random noise into the generator to obtain generated data, calculating the loss function after the generated data passes through the discriminator, and updating the generator parameters according to the loss function; Repeat the above steps of training the discriminator and training the generator until the predetermined number of iterations is reached or the model performance no longer improves.

12. The system for generating multi-category fault waveform data of a power transmission line according to claim 7, characterized in that: It also includes a data preprocessing module for preprocessing real transmission line multi-category fault waveform data, and the preprocessing method includes data cleaning and / or data standardization; The data cleaning method includes: counting the data length of the sampling rate of the multi-category fault waveform data of the transmission line; comparing the data length of the multi-category fault waveform data of the transmission line with the same sampling rate, and eliminating the multi-category fault waveform data of the transmission line with abnormal data length; Methods for data normalization include: When the length of the multi-category fault waveform data of the transmission line is greater than the data length specified by the sampling rate: If the multi-category fault waveform data of the power transmission line is continuous waveform data, the end of the multi-category fault waveform data of the power transmission line is truncated to reach the standard length; otherwise, the discontinuous part of the waveform is interpolated so that the data length reaches the standard length of the sampling rate; When the data length of the multi-category fault waveform of the transmission line is less than the data length specified by the sampling rate: If the transmission line multi-category fault waveform data is continuous waveform data, interpolation is performed at the end of the transmission line multi-category fault waveform data so that the data length reaches the standard length; otherwise, interpolation is performed at the discontinuous part of the waveform.

13. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for generating multi-category fault waveform data for a power transmission line according to any one of claims 1 to 6 are implemented.

14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the method for generating multi-category fault waveform data of a power transmission line as claimed in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Fault sample generation method and system of DCGAN with gradient penalty

    CN115270953A

  • Data-driven power distribution network fault data enhancement method based on WGAN-GP

    CN118013272A