Low-voltage active transformer area line loss rate prediction method and system based on improved GAN
By improving the Generative Adversarial Network (GAN) and combining data constraints and a segmentation generation strategy, the problem of insufficient sample data in the prediction of line loss rate in low-voltage active transformer areas was solved, and more accurate and stable line loss rate prediction was achieved.
Patent Information
- Application Number
- CN202511332144.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-09-18
AI Technical Summary
Existing technologies lack sample data for predicting line loss rates in low-voltage active transformer areas, leading to inaccurate predictions. Traditional GAN models struggle to handle the complexity and dynamic characteristics of time series data and lack effective methods for processing multivariate time series data with constraints.
An improved Generative Adversarial Network (GAN) is constructed to generate multiple historical sample sequences through a sliding time window. The generator, feature extractor, and discriminator are used for training. The sample sequences are generated in combination with data constraints. The generation strategy is subdivided into a first generation strategy and a second generation strategy to ensure the quality and coherence of the generated samples. The line loss rate is predicted by a temporal neural network.
It improves the accuracy and stability of line loss rate prediction, maintains consistency of generated samples over time series, covers more data features and pattern changes, and enhances the robustness and reliability of the network.
Smart Images

Figure CN120822672A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data prediction technology, and in particular to a method and system for predicting line loss rate in a low-voltage active substation area based on an improved GAN. Background Art
[0002] Low-voltage active substations, as a crucial component of the power system, are characterized by the widespread integration of distributed power sources. However, this integration introduces increased uncertainty and randomness into the distribution system, making the source and load characteristics of low-voltage active substations more complex and variable, and the system's operation more complex, posing new challenges to the calculation and management of line losses. Low-voltage line loss is a key indicator of power supply efficiency and grid operation quality, directly impacting the economic benefits and resource utilization efficiency of power supply companies. Therefore, how to efficiently and accurately predict and control low-voltage line loss using existing data has become a pressing issue.
[0003] Currently, the most commonly used techniques for low-voltage line loss management include traditional machine learning algorithms such as statistical analysis, artificial neural networks, and support vector machines. However, these methods rely heavily on large amounts of high-quality historical data. In practice, data collection for low-voltage distribution networks is limited by factors such as sensor layout and communication capabilities. Consequently, data is often incomplete or noisy, and it is difficult to cover all possible operating conditions. This is especially true when data samples of abnormal grid operation (such as equipment failures and illegal electricity theft) are lacking. This reduces the accuracy of line loss assessment and prediction in low-voltage active substations, impacting the effectiveness of decision-making.
[0004] While existing generative adversarial networks (GANs) have been widely used for sample expansion, they face numerous challenges when applied to low-voltage active substation line loss. First, low-voltage active substation line loss samples are time series data with complex dynamic characteristics and temporal dependencies. Traditional GAN models, which primarily target non-time series data, struggle to effectively handle the complexity and dynamic characteristics of time series data. Second, low-voltage active substation data encompasses a combination of multiple electrical quantities, including substation power supply, power consumption, and distributed power output. These electrical quantities are tightly constrained, and traditional GAN models lack effective methods for processing this type of multivariate, constrained time series data. Summary of the Invention
[0005] In view of the above analysis, the embodiments of the present invention aim to provide a method and system for predicting line loss rate in low-voltage active substations based on an improved GAN, so as to solve the problem of inaccurate line loss rate prediction due to the lack of sample data.
[0006] On the one hand, an embodiment of the present invention provides a method for predicting line loss rate in a low-voltage active area based on an improved GAN, comprising the following steps: Collect historical operation data of low-voltage active substations, construct historical samples for each period, and form multiple historical sample sequences through sliding time windows, and put them into the historical sample set; A generative adversarial network consisting of a generator, a feature extractor, and a discriminator is constructed and trained using a historical sample set. The generator is used to obtain multiple generated sample sequences based on each historical sample sequence and the sample generation strategy. The feature extractor is used to extract real features and multiple generated features based on each historical sample sequence and its corresponding multiple generated sample sequences. The discriminator is used to distinguish all real features from generated features. After training, the generated sample sequence corresponding to the generated features judged as true by the discriminator is obtained and put into the historical sample set. Together with the equipment rated capacity, a prediction data set is constructed. The time series neural network model is trained to obtain a line loss rate prediction model. The operating data and equipment rated capacity used for prediction are collected, and the line loss rate at the prediction time is obtained using the line loss rate prediction model.
[0007] Based on a further improvement of the above method, the generator is used to obtain multiple generated sample sequences according to each historical sample sequence and the sample generation strategy. It uses the historical samples of the first time period in each historical sample sequence and uses the sample generation strategy to generate multiple initial samples for each time period in sequence, and obtains the valid samples that meet the data constraints of each time period as the multiple generated samples of the corresponding time period, thereby obtaining multiple generated sample sequences with the same time window length.
[0008] Based on the further improvement of the above method, the data constraints include: initial period data constraints, adjacent period data constraints, adjacent period change constraints and valid sample quantity constraints; the sample generation strategy includes: a first generation strategy and a second generation strategy in sequence; the first generation strategy is used to obtain multiple generated samples of the first period that meet the initial period data constraints and the valid data quantity constraints based on the historical samples of the first period in each historical sample sequence; the second generation strategy is used to iteratively obtain multiple generated samples of subsequent consecutive periods that meet the adjacent period data constraints, adjacent period change constraints and valid data quantity constraints based on each generated sample of the first period according to the length of the time window, to form multiple generated sample sequences.
[0009] Based on the further improvement of the above method, the historical samples and generated samples in each period include multiple operating data: distribution and transformation load in the substation, distributed power generation output, charging load and user conventional load; The first generation strategy performs the following steps: According to each operating data, fluctuation value and random number in the historical sample of the first period in each historical sample sequence, the corresponding generated data are obtained multiple times as the multiple initial samples of the first period, and the valid samples that meet the data constraints of the initial period are obtained from the multiple initial samples of the first period. The first generation strategy is repeatedly executed until the valid samples of the first period meet the valid sample quantity constraint, and the multiple generated samples of the first period are obtained.
[0010] Based on the further improvement of the above method, the second generation strategy performs the following steps: The first time period is taken as the current time period, and the generated data corresponding to the next time period is obtained according to each generated data, fluctuation value and random number in each generated sample of the current time period, and the initial sample of the next time period corresponding to each generated sample of the current time period is obtained. The valid samples that meet the data constraints of adjacent time periods and the change constraints of adjacent time periods are obtained from them. When the valid samples do not meet the valid sample quantity constraints, the second generation strategy is repeatedly executed; otherwise, the obtained valid samples are used as the generated samples of the next time period, and after the next time period is taken as the current time period, the second generation strategy is executed again until multiple generated samples of multiple consecutive time periods of the time window length are obtained to form multiple generated sample sequences.
[0011] Based on a further improvement of the above method, the corresponding generated data is obtained multiple times based on each operating data, fluctuation value and random number in the historical sample of the first period in each historical sample sequence. The generated data is generated by adding a first random perturbation value to each operating data multiple times. The first random perturbation value is obtained by multiplying the random perturbation ratio with the corresponding operating data. Obtaining generated data corresponding to the next period based on each generated data, fluctuation value, and random number in each generated sample of the current period, by adding a second random perturbation value to each generated data in each generated sample, the second random perturbation value being obtained by multiplying the random perturbation ratio by the generated data of the corresponding generated sample; The random disturbance ratio is obtained based on the fluctuation value and the random number.
[0012] Based on a further improvement of the above method, the initial period data constraint is that the absolute value of the difference between the line loss rate of the initial sample of the first period and the line loss rate of the historical sample of the first period is not greater than a first threshold; The constraint on data in adjacent time periods is that the absolute value of the difference in line loss rate of initial samples in adjacent time periods is not greater than a first threshold; The adjacent time period change constraint is that the change value of the distribution transformer load of the initial sample in the adjacent time period is greater than the change value of the line loss; The effective sample quantity constraint is that the number of effective samples is greater than or equal to a second threshold.
[0013] Based on the further improvement of the above method, the real features and multiple generated features are extracted according to each historical sample sequence and its corresponding multiple generated sample sequences, including: According to multiple preset indicators, the features of each indicator of each operating data are extracted from each historical sample sequence and spliced together to obtain the real features corresponding to each historical sample sequence; the features of each indicator of each operating data are extracted from each generated sample sequence and spliced together to obtain the generated features corresponding to each generated sample sequence.
[0014] Based on the further improvement of the above method, the fluctuation value is a trainable parameter in the generator, and the fluctuation value is automatically updated and optimized through gradient descent during the training process of the generative adversarial network.
[0015] On the other hand, an embodiment of the present invention provides a low-voltage active area line loss rate prediction system based on an improved GAN, comprising: The sample collection module is used to collect historical operation data of low-voltage active substations, construct historical samples for each period, and form multiple historical sample sequences through a sliding time window and put them into a historical sample set; The sample generation module is used to construct a generative adversarial network including a generator, a feature extractor, and a discriminator, and train it using a historical sample set. The generator is used to obtain multiple generated sample sequences based on each historical sample sequence and the sample generation strategy. The feature extractor is used to extract real features and multiple generated features based on each historical sample sequence and its corresponding multiple generated sample sequences. The discriminator is used to distinguish all real features from generated features. The model training module is used to obtain the generated sample sequence corresponding to the generated features judged as true by the discriminator after training, put it into the historical sample set, and construct a prediction data set together with the rated capacity of the equipment. The time series neural network model is trained to obtain the line loss rate prediction model; The line loss rate prediction module is used to collect operating data and equipment rated capacity for prediction, and use the line loss rate prediction model to obtain the line loss rate at the prediction time.
[0016] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects: 1. By improving the generator of the generative adversarial network, the generated samples are kept consistent with the historical samples in terms of time series continuity. Moreover, after screening with data constraints, the quality of the generated samples in each time period is guaranteed, and the rationality of the generated samples is improved. By using each historical sample sequence to obtain multiple generated sample sequences, more data features and pattern changes are covered, avoiding data homogeneity caused by a single generation path, and increasing sample diversity, it is easier for the line loss rate prediction model to learn more comprehensive substation operation characteristics, so that the line loss rate prediction model can more accurately capture the laws of line loss changes, thereby significantly improving the prediction accuracy of the line loss rate.
[0017] 2. The sample generation strategy is divided into two. The first generation strategy focuses on sample generation in the first period. Based on the data constraints and effective data quantity constraints of the initial period, it lays a good foundation for subsequent generation and avoids the impact of initial sample quality issues on the entire sequence. The second generation strategy considers the data constraints and change constraints of adjacent periods. Based on the high-quality samples of the first period, it generates samples for subsequent periods, ensuring the rationality and coherence of the generated sample sequence in the time dimension, making the generated entire time series data more realistic and reliable.
[0018] 3. The features of multiple indicators are extracted from the historical sample sequence and the generated sample sequence generated by the generator, so that the discriminator can judge the comprehensive and implicit deep features in the sample, and has stronger robustness when facing data with different distributions or noise interference, thereby improving the stability and reliability of the network.
[0019] In the present invention, the above-mentioned technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of the present invention will be described in the following description, and some advantages will become apparent from the description or be learned through practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings are only used for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. Like reference symbols denote like components throughout the accompanying drawings. Figure 1 This is a flow chart of a method for predicting line loss rate in a low-voltage active area based on an improved GAN in Example 1 of the present invention. DETAILED DESCRIPTION
[0021] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, and are not used to limit the scope of the present invention.
[0022] Example 1 A specific embodiment of the present invention discloses a method for predicting line loss rate in low-voltage active area based on improved GAN, such as Figure 1 As shown, the following steps are included: S1. Collect historical operation data of low-voltage active substations, construct historical samples for each time period, and form multiple historical sample sequences through sliding time windows, and put them into the historical sample set.
[0023] It should be noted that the historical operating data of the low-voltage active substations collected include: substation distribution transformer load, distributed power supply online output, charging load and user conventional load; among them, the substation distribution transformer load is measured according to the metering cabinet next to the substation distribution transformer, and the distributed power supply online output refers to the output of the distributed power supply connected to the low-voltage substation to supply power to the grid, which is measured according to the meter installed at the distributed power supply access point; the charging load is measured according to the meter installed at the electric vehicle charging pile access point; the user conventional load is measured according to the household meter.
[0024] Collect historical operating data at each measurement moment and perform preprocessing, such as standardization or normalization, and then perform time alignment, that is, obtain the substation distribution transformer load, distributed power supply online output, charging load and user conventional load at each identical measurement moment.
[0025] Furthermore, each measurement moment is divided according to the preset time period length, and based on the measured values of the substation distribution transformer load, distributed power supply online output, charging load and user conventional load at each measurement moment in each time period, the respective average values are calculated as the values of each operating data in each time period to obtain the historical samples of each time period.
[0026] Furthermore, according to the total number of time periods and the sliding step in the preset time window, historical samples of multiple time periods are sequentially combined into a historical sample sequence through the sliding time window and put into the historical sample set.
[0027] S2. Construct a generative adversarial network including a generator, a feature extractor and a discriminator, and train it using a historical sample set; the generator is used to obtain multiple generated sample sequences based on each historical sample sequence and the sample generation strategy; the feature extractor is used to extract real features and multiple generated features based on each historical sample sequence and its corresponding multiple generated sample sequences; the discriminator is used to discriminate all real features and generated features.
[0028] S21. Generator.
[0029] It should be noted that in the generative adversarial network constructed in this embodiment, the generator is not a neural network model, but obtains multiple generated sample sequences by executing a sample generation strategy based on each historical sample sequence, which not only increases the number of data samples, but also improves the diversity of samples.
[0030] Specifically, based on the historical samples of the first time period in each historical sample sequence, the sample generation strategy is used to generate multiple initial samples for each time period in turn, and the valid samples that meet the data constraints of each time period are obtained as the multiple generated samples of the corresponding time period, thereby obtaining multiple generated sample sequences with the same time window length.
[0031] Data constraints include: initial period data constraints, adjacent period data constraints, adjacent period change constraints, and effective sample quantity constraints. Sample generation strategies include: the first generation strategy and the second generation strategy. The first generation strategy focuses on generating samples for the first period, laying a solid foundation for subsequent generation and preventing initial sample quality issues from affecting the entire sequence. The second generation strategy generates samples for subsequent periods based on high-quality samples from the first period, ensuring the rationality and consistency of the generated sample sequence in the temporal dimension, making the generated time series data more authentic and reliable.
[0032] Specifically, the first generation strategy is used to obtain multiple generated samples of the first time period that meet the initial time period data constraints and the valid data quantity constraints based on the historical samples of the first time period in each historical sample sequence; the second generation strategy is used to iteratively obtain multiple generated samples of subsequent consecutive time periods that meet the adjacent time period data constraints, adjacent time period change constraints and valid data quantity constraints based on each generated sample of the first time period according to the length of the time window, to form multiple generated sample sequences.
[0033] That is, first, based on the historical sample sequence, a first generation strategy is used to obtain multiple generated samples for the first time period. Then, based on the multiple generated samples obtained for the first time period, a second generation strategy is used to obtain multiple generated samples for subsequent consecutive time periods in the time window, excluding the first time period, thereby obtaining multiple generated sample sequences of the same length as the historical sample sequence. During the generation process, different data constraints are used to screen valid samples, and a constraint on the number of valid samples is used to control the minimum number of generated samples. Therefore, the number of generated sample sequences obtained by the sample generation strategy for each historical sample sequence is not necessarily the same.
[0034] Furthermore, the first generation strategy performs the following steps: ①According to each running data, fluctuation value and random number in the historical sample of the first period in each historical sample sequence, the corresponding generated data is obtained multiple times as multiple initial samples of the first period.
[0035] It should be noted that the corresponding generated data are obtained multiple times by adding the first random disturbance value to each operating data in the historical sample of the first period in each historical sample sequence multiple times, wherein the first random disturbance value is obtained by multiplying the random disturbance ratio by the corresponding operating data, and the random disturbance ratio is obtained based on the fluctuation value and the random number.
[0036] Specifically, the generated data for the first period is calculated by the following formula: (1), in, 、 、 and Respectively represent The values of the distribution transformer load, distributed generation output, charging load and user conventional load in the historical samples of the first period in the historical sample sequence; 、 、 and Represents the first period The values of the distribution and transformation load, distributed generation output, charging load and user conventional load in the initial sample; represents a random number in the (0,1) interval, Indicates the fluctuation value, initially set to 0.1; represents the random perturbation ratio.
[0037] From formula (1), we can see that the fluctuation value determines the maximum deviation ratio of the disturbance relative to the operating data of the historical sample. The random number is used to generate uniformly distributed random disturbances within the upper and lower deviation ranges for the operating data center, through dynamic scaling of the operating data of the historical sample.
[0038] It should be noted that each time these four data are generated using formula (1), an initial sample of the first period is formed; executing formula (1) multiple times will result in multiple initial samples, and the number of initial samples is controlled by the preset number of samples.
[0039] ② Obtain valid samples that meet the data constraints of the initial period from multiple initial samples of the first period, repeat the first generation strategy until the valid samples of the first period meet the valid sample quantity constraint, and obtain multiple generated samples of the first period.
[0040] It should be noted that the initial period data constraint is that the absolute value of the difference between the line loss rate of the initial sample of the first period and the line loss rate of the historical sample of the first period is not greater than the first threshold. The line loss rate is calculated based on the distribution transformer load, distributed power supply output, charging load and user conventional load in the first period. The formula is as follows: (2), in, represents the initial line loss rate of the initial sample in the first period, represents the line loss rate of the historical sample in the first period; Indicates the first threshold, which is set to 5% in this embodiment.
[0041] In order to improve the accuracy of the data, the initial sample is checked by the initial period data constraint. If the initial sample meets the initial period data constraint, it is regarded as a valid sample, otherwise it is removed. At the same time, the effective sample quantity constraint is used to ensure that the generated samples reach a certain number. Among them, the effective sample quantity constraint is that the number of effective samples is greater than or equal to the second threshold. That is, if the number of valid samples obtained is greater than or equal to the second threshold , then the valid sample is used as the generated sample of the first period and passed into the second generation strategy; otherwise, the valid sample is retained and the first generation strategy is repeated to obtain new valid samples according to the above process until the number of all valid samples (new valid samples and previously retained valid samples) is greater than or equal to the second threshold , as the generated sample of the first period, is passed into the second generation strategy.
[0042] It should be noted that, at this time, each generated sample in the first period corresponds to a generated sample sequence.
[0043] Furthermore, the second generation strategy performs the following steps: ① Take the first period as the current period, and obtain the generated data corresponding to the next period according to each generated data, fluctuation value and random number in each generated sample of the current period, and obtain the initial sample of the next period corresponding to each generated sample of the current period.
[0044] It should be noted that the generated data corresponding to the next time period are obtained by adding a second random disturbance value to each generated data in each generated sample of the current time period, wherein the second random disturbance value is obtained by multiplying the random disturbance ratio by the generated data of the corresponding generated sample; the random disturbance ratio is obtained based on the fluctuation value and the random number.
[0045] Specifically, the following formula is used to calculate the initial sample of the next period corresponding to each generated sample in the current period: (3), in, 、 、 and Respectively represent Period The values of the distribution and transformation load, distributed power supply output, charging load and user conventional load in the generated samples are 、 、 and Respectively represent Period The values of the distribution and transformation load, distributed generation output, charging load and user conventional load in the initial sample; , Indicates the total number of time periods in the time window.
[0046] It can be seen from formula (3) that each time the four data in the generated sample of the current period are used to generate the four data of the next period, the initial sample of the next period corresponding to the generated sample of the current period is formed.
[0047] ② Obtain valid samples that meet the data constraints and change constraints of adjacent time periods from the initial samples of the next time period. When the valid samples do not meet the valid sample quantity constraint, repeat the second generation strategy; otherwise, use the obtained valid samples as the generated samples of the next time period, and after taking the next time period as the current time period, execute the second generation strategy again until multiple generated samples of multiple consecutive time periods of the time window length are obtained to form multiple generated sample sequences.
[0048] It should be noted that the constraint on data in adjacent time periods is that the absolute value of the difference in line loss rate between initial samples in adjacent time periods is not greater than the first threshold. The formula is as follows: (4), in, and Indicates the The generated sample corresponds to Period and The line loss rates of the time periods are respectively Period and The distribution and transformation load of the substation, the online output of distributed power generation, the charging load and the conventional load of users in each period are calculated.
[0049] Furthermore, valid samples that meet the adjacent period data constraints are further checked to see if they meet the adjacent period change constraints. The adjacent period change constraint is that the change value of the distribution transformer load of the initial sample in the adjacent period is greater than the change value of the line loss, as shown in the following formula: (5).
[0050] The initial samples that satisfy both the adjacent period data constraints and the adjacent period change constraints are taken as the first If the number of valid samples does not meet the valid sample quantity constraint, these valid samples are retained and the second generation strategy is repeated, that is, the first generation strategy is repeated. Each generated sample in the time period generates the corresponding The new initial samples of the period are verified by the adjacent period data constraints and adjacent period change constraints to obtain new valid samples. If the number of new valid samples and the number of valid samples retained before meet the valid sample number constraints, all valid samples are taken as the first The generated samples of the time period will be After the period is taken as the current period, the second generation strategy is executed again, and the first Multiple generated samples of the time period, and so on, until the Multiple generated samples for a period of time.
[0051] It should be noted that, in the process of repeatedly executing the second generation strategy, The first The generated sample Multiple initial samples of a time period may not meet the data constraints, or may meet the data constraints partially or completely. Therefore, for each generated sample, the association relationship between time periods should be established, with a complete data set from the first time period to the Generate samples for each period.
[0052] For example, the second threshold value in the valid sample quantity constraint is 40. For the first historical sample sequence, based on its historical samples in the first period, 80 initial samples for the first period are generated by executing the first generation strategy. After verification of the initial period data constraint, 60 valid samples are obtained. 60 is greater than the second threshold value 40. Then, the 60 valid samples are passed as the generated samples for the first period to the second generation strategy, and a corresponding initial sample for the second period is generated for each of them. That is, 60 initial samples for the second period are generated. After verification of the adjacent period data constraint and the adjacent period change constraint: If 50 valid samples are obtained, the 50 valid samples are used as the generated samples for the second period, and it is ensured that these 50 generated samples are associated with the generated samples for the first period. After the second period is taken as the current period, the second generation strategy is executed to obtain 50 initial samples for the third period. If 30 valid samples are obtained, these 30 valid samples are retained, and 60 new initial samples for the second period are generated again based on the 60 generated samples of the first period. These 60 new initial samples are verified. If 20 new valid samples are obtained, plus the 30 valid samples retained before, a total of 50 valid samples are obtained, which is greater than the second threshold of 40. These 50 valid samples are used as the generated samples for the second period, and it is ensured that these 50 generated samples are associated with the generated samples of the first period. After taking the second period as the current period, the second generation strategy is executed to obtain 50 initial samples for the third period. The 50 initial samples of the third period are processed according to the process of the second period, and so on, to obtain the first The valid samples of more than 40 times in a period are used as the generated samples.
[0053] Based on the correlation between each time period of each generated sample, multiple generated sample sequences are formed in the order of the time periods, and correspond to the actual collected historical sample sequences. That is, using the generator of this embodiment, multiple generated samples with data-constrained time series are obtained for each historical sample sequence.
[0054] S22, feature extractor.
[0055] This step extracts data features from each historical sample sequence and its corresponding generated sample sequence. Each historical sample sequence is constructed based on the actual collected data, and the corresponding data features are used as real features. The generated sample sequence obtained by the generator uses the corresponding data features as generated features.
[0056] Specifically, based on multiple preset indicators, the features of each indicator of each operating data point are extracted from each historical sample sequence and concatenated to obtain the true features corresponding to each historical sample sequence. The features of each indicator of each operating data point are extracted from each generated sample sequence and concatenated to obtain the generated features corresponding to each generated sample sequence. The preset indicators include maximum value, minimum value, mean, variance, frequency, and frequency amplitude.
[0057] It should be noted that when extracting the frequency and frequency amplitude of each operating data, Fourier transform is used to obtain multiple frequencies and their frequency amplitudes, from which the maximum frequency amplitude and the second largest frequency amplitude and their corresponding frequencies are selected. If the second largest frequency amplitude is less than 10% of the maximum frequency amplitude, only the maximum frequency amplitude and its corresponding frequency are selected.
[0058] S23, discriminator.
[0059] The discriminator in this embodiment is a binary classifier that uses a neural network model, such as a recurrent neural network, a multilayer perceptron, or a convolutional neural network. The discriminator distinguishes between real features and generated features, and outputs a probability distribution that the input feature is a real feature.
[0060] S24. Train a generative adversarial network.
[0061] When training the generative adversarial network using historical sample sets, the learning rate is initially set to , and after every 5 epochs of training, the learning rate is multiplied by 0.95 to achieve decay.
[0062] It should be noted that the fluctuation value used in the sample generation strategy in the generator, that is, the fluctuation value in formula (1) and formula (3), is a trainable parameter. During the training process of the generative adversarial network, the fluctuation value is automatically updated and optimized through gradient descent, so that the discriminator is more inclined to believe that the generated data is real.
[0063] The adjustment of each parameter in the generative adversarial network follows the following formula: (6), in, and Respectively represent the first Parameter No. Second and The value of the iteration, and Represents the first Parameter No. Second and The value of the iteration; represents the learning rate; and Represent the generator loss function and the discriminator loss function respectively, Represents the generator loss function with respect to the generator The gradient of the parameters; Denotes the discriminator loss function with respect to the discriminator The gradient of the parameters.
[0064] When the loss values of the generator and discriminator no longer decrease / increase significantly and the difference values continue to fluctuate within a preset small range, the training ends and a trained generative adversarial network is obtained.
[0065] S3. After the training is completed, the generated sample sequence corresponding to the generated features judged as true by the discriminator is obtained, put into the historical sample set, and together with the rated capacity of the equipment, a prediction data set is constructed. The time series neural network model is trained to obtain a line loss rate prediction model.
[0066] It should be noted that the generated sample sequences corresponding to the generated features that the discriminator judges as true are placed in the historical sample set and are collectively referred to as sample sequences together with the historical sample sequences. The equipment rated capacity collected as static features includes: the rated capacity of the distribution transformer in the substation, the rated capacity of the distributed power supply, and the charging capacity of the charging pile.
[0067] Furthermore, the values of the substation distribution transformer load, distributed power supply online output and charging load in each time period in each sample sequence are divided by the corresponding substation distribution transformer rated capacity, distributed power supply rated capacity and charging pile charging capacity to obtain the substation distribution transformer load rate, distributed power supply utilization rate and charging pile utilization rate in each time period. These are used as new data together with the original data of the corresponding time period as the prediction sample of the corresponding time period, thereby obtaining a prediction sample sequence and putting it into the prediction data set.
[0068] It should be noted that the length of the predicted sample sequence can be adjusted according to actual conditions and does not need to be the same as the length of the historical sample sequence constructed in step S1.
[0069] Specifically, the data in the prediction samples of each time period in the prediction data set include: substation distribution transformer load, distributed power supply online output, charging load, user conventional load, substation distribution transformer load rate, distributed power supply utilization rate and charging pile utilization rate.
[0070] The time series neural network model in this embodiment is used to learn from the input prediction sample sequence and predict the line loss rate for the next time period. Therefore, the label value (i.e., the actual line loss rate) of each prediction sample sequence is calculated based on the distribution transformer load, distributed generation power output, charging load, and user conventional load in the next time period after the prediction sample sequence. This is used for comparison with the line loss rate prediction output by the model during the training process.
[0071] In this step, any time series neural network model, such as RNN, LSTM, GRU, or Transformer model, is used to train it using the prediction sample set. After the training, a line loss rate prediction model is obtained.
[0072] This embodiment converts the fixed rated capacity of the equipment into a load rate and utilization rate that varies over time, quantifies the current operating status of the equipment relative to its maximum capacity, and obtains more meaningful multiple time series features to supplement the original operating data. This enables the model to more accurately learn the changing pattern of the line loss rate and improves the accuracy of the model's prediction of the line loss rate.
[0073] S4. Collect the operating data and equipment rated capacity used for prediction, and use the line loss rate prediction model to obtain the line loss rate at the prediction time.
[0074] In actual application, the operating data used for prediction is collected: the substation distribution transformer load, distributed power supply online output and charging load in multiple time periods. The substation distribution transformer load rate, distributed power supply utilization rate and charging pile utilization rate in each time period are calculated based on the rated capacity of the equipment. The time series data used for prediction is input into the line loss rate prediction model, and the line loss rate of the next time period of the time series data is output.
[0075] Compared with the existing technology, this embodiment provides a low-voltage active substation line loss rate prediction method based on improved GAN. By improving the generator of the generative adversarial network, the generated samples are consistent with the historical samples in terms of the continuity of the time series. Moreover, after screening with data constraints, the quality of the generated samples in each time period is guaranteed, and the rationality of the generated samples is improved. Multiple generated sample sequences are obtained by using each historical sample sequence, covering more data features and pattern changes, avoiding data homogeneity caused by a single generation path, and increasing the diversity of samples, which facilitates the line loss rate prediction model to learn more comprehensive substation operation characteristics, so that the line loss rate prediction model can more accurately capture the law of line loss changes, thereby significantly improving the prediction accuracy of the line loss rate. The sample generation strategy is divided into two parts. The first generation strategy focuses on sample generation in the first period. Based on the data constraints and effective data quantity constraints of the initial period, it lays a good foundation for subsequent generation and prevents the quality of the initial sample from affecting the entire sequence. The second generation strategy considers the data constraints and change constraints of adjacent periods. Based on the high-quality samples of the first period, it generates samples for subsequent periods, ensuring the rationality and coherence of the generated sample sequence in the time dimension, making the generated time series data more realistic and reliable. The features of multiple indicators are extracted from the historical sample sequence and the generated sample sequence generated by the generator, allowing the discriminator to discriminate the comprehensive and implicit deep features in the sample. It has stronger robustness when facing data with different distributions or noise interference, and improves the stability and reliability of the network.
[0076] Example 2 Another embodiment of the present invention discloses a low-voltage active substation line loss rate prediction system based on an improved GAN, thereby implementing the low-voltage active substation line loss rate prediction method based on an improved GAN in Example 1. The specific implementation of each module refers to the corresponding description in Example 1. The system includes: The sample collection module is used to collect historical operation data of low-voltage active substations, construct historical samples for each period, and form multiple historical sample sequences through a sliding time window and put them into a historical sample set; The sample generation module is used to construct a generative adversarial network including a generator, a feature extractor, and a discriminator, and train it using a historical sample set. The generator is used to obtain multiple generated sample sequences based on each historical sample sequence and the sample generation strategy. The feature extractor is used to extract real features and multiple generated features based on each historical sample sequence and its corresponding multiple generated sample sequences. The discriminator is used to distinguish all real features from generated features. The model training module is used to obtain the generated sample sequence corresponding to the generated features judged as true by the discriminator after training, put it into the historical sample set, and construct a prediction data set together with the rated capacity of the equipment. The time series neural network model is trained to obtain the line loss rate prediction model; The line loss rate prediction module is used to collect operating data and equipment rated capacity for prediction, and use the line loss rate prediction model to obtain the line loss rate at the prediction time.
[0077] Since the present embodiment of a low-voltage active substation line loss rate prediction system based on an improved GAN and the aforementioned low-voltage active substation line loss rate prediction method based on an improved GAN can be mutually referenced in their related aspects, they are redundantly described and will not be repeated here. Since the present system embodiment and the aforementioned method embodiment share the same principles, the present system embodiment also has the corresponding technical effects of the aforementioned method embodiment.
[0078] Those skilled in the art will appreciate that all or part of the process steps of the above-described embodiments can be implemented by instructing related hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, such as a magnetic disk, an optical disk, a read-only memory, or a random access memory.
[0079] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.
Claims
1. A method for predicting line loss rate in low-voltage active area based on improved GAN, characterized in that: The following steps are involved: Collect historical operation data of low-voltage active substations, construct historical samples for each period, and form multiple historical sample sequences through sliding time windows, and put them into the historical sample set; A generative adversarial network comprising a generator, a feature extractor, and a discriminator is constructed and trained using the historical sample set; the generator is used to obtain multiple generated sample sequences based on each historical sample sequence and a sample generation strategy; the feature extractor is used to extract real features and multiple generated features based on each historical sample sequence and its corresponding multiple generated sample sequences; The discriminator is used to discriminate all real features and generated features; After training, the generated sample sequence corresponding to the generated features judged as true by the discriminator is obtained and put into the historical sample set. Together with the equipment rated capacity, a prediction data set is constructed. The time series neural network model is trained to obtain a line loss rate prediction model. The operating data and equipment rated capacity used for prediction are collected, and the line loss rate at the prediction time is obtained using the line loss rate prediction model.
2. The low-voltage active area line loss rate prediction method based on improved GAN according to claim 1 is characterized in that: The generator is used to obtain multiple generated sample sequences based on each historical sample sequence and the sample generation strategy. It uses the historical samples of the first time period in each historical sample sequence and uses the sample generation strategy to sequentially generate multiple initial samples for each time period, and obtains valid samples that meet the data constraints of each time period as multiple generated samples for the corresponding time period, thereby obtaining multiple generated sample sequences with the same time window length.
3. The low-voltage active area line loss rate prediction method based on improved GAN according to claim 2 is characterized in that: The data constraints include: initial time period data constraints, adjacent time period data constraints, adjacent time period change constraints and valid sample quantity constraints; the sample generation strategy includes: a first generation strategy and a second generation strategy in sequence; the first generation strategy is used to obtain multiple generated samples of the first time period that meet the initial time period data constraints and valid data quantity constraints based on the historical samples of the first time period in each historical sample sequence; the second generation strategy is used to iteratively obtain multiple generated samples of subsequent consecutive time periods that meet the adjacent time period data constraints, adjacent time period change constraints and valid data quantity constraints based on each generated sample of the first time period according to the length of the time window, to form multiple generated sample sequences.
4. The method for predicting line loss rate of low-voltage active substation area based on improved GAN according to claim 3 is characterized in that: The historical samples and generated samples for each period include multiple operating data: distribution and transformation load in the substation, distributed power generation output, charging load and user conventional load; The first generation strategy performs the following steps: According to each operating data, fluctuation value and random number in the historical sample of the first time period in each historical sample sequence, corresponding generated data are obtained multiple times as multiple initial samples of the first time period, and valid samples that meet the data constraints of the initial time period are obtained from the multiple initial samples of the first time period. The first generation strategy is repeatedly executed until the valid samples of the first time period meet the valid sample quantity constraint, thereby obtaining multiple generated samples of the first time period.
5. The low-voltage active area line loss rate prediction method based on improved GAN according to claim 4 is characterized in that: The second generation strategy performs the following steps: Take the first period as the current period, and obtain the generated data corresponding to the next period based on each generated data, fluctuation value, and random number in each generated sample of the current period. Obtain the initial sample of the next period corresponding to each generated sample of the current period, and obtain valid samples that meet the adjacent period data constraints and adjacent period change constraints. If the valid sample does not meet the valid sample quantity constraint, repeat the second generation strategy; Otherwise, the obtained valid samples are used as the generated samples for the next period, and the next period is used as the current period, and the second generation strategy is executed again until multiple generated samples for multiple consecutive periods of the time window length are obtained to form multiple generated sample sequences.
6. The method for predicting line loss rate of low-voltage active substation area based on improved GAN according to claim 5 is characterized in that: The corresponding generated data is obtained multiple times based on each operating data, fluctuation value and random number in the historical sample of the first period in each historical sample sequence, and is generated by adding a first random perturbation value to each operating data multiple times, wherein the first random perturbation value is obtained by multiplying the random perturbation ratio by the corresponding operating data; The generating data corresponding to the next period is obtained according to each generating data, the fluctuation value and the random number in each generating sample of the current period, and is generated by adding a second random perturbation value to each generating data in each generating sample, wherein the second random perturbation value is obtained by multiplying the random perturbation ratio by the generating data of the corresponding generating sample; The random disturbance ratio is obtained according to the fluctuation value and the random number.
7. The method for predicting line loss rate of low-voltage active substation area based on improved GAN according to claim 5 is characterized in that: The initial period data constraint is that the absolute value of the difference between the line loss rate of the initial sample of the first period and the line loss rate of the historical sample of the first period is not greater than a first threshold; The adjacent time period data constraint is that the absolute value of the difference in line loss rate of initial samples in adjacent time periods is not greater than a first threshold; The adjacent time period change constraint is that the change value of the distribution transformer load of the initial sample in the adjacent time period is greater than the change value of the line loss; The effective sample quantity constraint is that the number of effective samples is greater than or equal to a second threshold.
8. The method for predicting line loss rate of low-voltage active substation area based on improved GAN according to claim 5 is characterized in that: The extracting of real features and multiple generated features based on each historical sample sequence and its corresponding multiple generated sample sequences includes: According to multiple preset indicators, the features of each indicator of each operating data are extracted from each historical sample sequence and spliced together to obtain the real features corresponding to each historical sample sequence; the features of each indicator of each operating data are extracted from each generated sample sequence and spliced together to obtain the generated features corresponding to each generated sample sequence.
9. The method for predicting line loss rate of low-voltage active substation area based on improved GAN according to claim 6, characterized in that: The fluctuation value is a trainable parameter in the generator, and the fluctuation value is automatically updated and optimized through gradient descent during the generative adversarial network training process.
10. A low-voltage active area line loss rate prediction system based on improved GAN, characterized by: include: The sample collection module is used to collect historical operation data of low-voltage active substations, construct historical samples for each period, and form multiple historical sample sequences through a sliding time window and put them into a historical sample set; A sample generation module is used to construct a generative adversarial network including a generator, a feature extractor, and a discriminator, and train it using the historical sample set; the generator is used to obtain multiple generated sample sequences based on each historical sample sequence and the sample generation strategy; the feature extractor is used to extract real features and multiple generated features based on each historical sample sequence and its corresponding multiple generated sample sequences; The discriminator is used to discriminate all real features and generated features; The model training module is used to obtain the generated sample sequence corresponding to the generated features judged as true by the discriminator after training, put it into the historical sample set, and construct a prediction data set together with the rated capacity of the equipment. The time series neural network model is trained to obtain the line loss rate prediction model; The line loss rate prediction module is used to collect operating data and equipment rated capacity for prediction, and use the line loss rate prediction model to obtain the line loss rate at the prediction time.
Citation Information
Patent Citations
Method, device and system for calculating theoretical line loss of transformer area
CN116304526A
Conditional generative adversarial network-based short-term wind power prediction method
CN116488151A
Low-voltage transformer area line loss index differentiation prediction method and system
CN118485179A
Ultra-short-term photovoltaic power prediction method and system based on short-term weather forecast
CN119539215A
Runoff forecasting method and system based on improved generative adversarial network
CN120180858A
Cited By
Timegan-based transformer area extreme scenario operation data enhancement method and system
CN122712205A