Transformer area capacity configuration method and device of electric power system, terminal equipment and storage medium

By locally clustering, balancing processing and classified voting on the load data of the power system, the maximum load data vector is extracted, and the capacity value of the station area is predicted, the problem of low accuracy of capacity configuration in the middle station area in the existing technology is solved, and the operation efficiency and safety of the power system are improved.

CN119944673AInactive Publication Date: 2025-05-06GUANGDONG POWER GRID CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510421612.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The capacity configuration accuracy of the existing technology middle-end zone is low, resulting in unreasonable load distribution during power supply, which may cause light load or heavy overload of the distribution transformer, and even burnt of the distribution transformer.

Method used

By obtaining the initial load data of the power system to be configured, local data clustering and balancing processing are performed, load data classification and voting is used using multiple base classifiers, the maximum load data vector corresponding to the load type is extracted, and the power system to be configured is predicted based on these data to be configured, and the capacity value of the station area is obtained.

Benefits of technology

It improves the calculation accuracy of the capacity value of the station area, avoids light load or heavy overload of the distribution, and enhances the operational economy and safety of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119944673A_ABST
    Figure CN119944673A_ABST
Patent Text Reader

Abstract

The invention discloses a transformer area capacity configuration method and device of a power system, terminal equipment and a storage medium, and belongs to the technical field of power systems, and the method comprises the steps: carrying out the local data clustering of the obtained initial load data of a to-be-configured power system, obtaining a load clustering sample, and carrying out the balance processing of the load clustering sample, obtaining a load data set; classifying the current load data through a plurality of base classifiers, voting based on classification results, and taking the classification result with the most votes as the load type of the current load data; extracting a maximum load data vector corresponding to each load type; and based on the load type of each piece of load data in the load data set and the maximum load data vector corresponding to each load type, obtaining a transformer area capacity value of the to-be-configured power system, and configuring the to-be-configured power system. Therefore, by implementing the method and the device, the problem of low distribution room capacity configuration accuracy in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power systems, and in particular to a method, device, terminal equipment and storage medium for configuring the capacity of a power station area in a power system. Background Art

[0002] With the rapid development of my country's economy and the continuous improvement of material and cultural levels, various types of electricity demand are increasing, and the power load in the substation is showing a strong growth momentum. For a long time, the open capacity of the substation has been evaluated based on the user's own operating experience, and the evaluation accuracy is low, which may lead to unreasonable load distribution during power supply. If the load of the connected users is small, the distribution transformer (referred to as distribution transformer) will always be in a light-load state, and the economic operation level of the distribution transformer will be low; if the load of the connected users is too much, some distribution transformers in the substation will be overloaded, and even the distribution transformers in the substation will burn out.

[0003] Therefore, there is an urgent need for a power system substation capacity configuration strategy to solve the problem of low accuracy of substation capacity configuration. Summary of the invention

[0004] The embodiments of the present invention provide a method, an apparatus, a terminal device and a storage medium for configuring the capacity of an electric power system to solve the problem of low accuracy of the configuration of the capacity of the electric power system.

[0005] In order to solve the above problem, an embodiment of the present invention provides a method for configuring the capacity of a power system, including: Obtaining initial load data of the power system to be configured; Performing local data clustering on the initial load data to obtain load cluster samples, and performing balancing processing on the load cluster samples to obtain a load data set; For the load data in the load data set, the current load data is classified by several base classifiers, and voting is performed based on the classification results of the current load data output by each base classifier, and the classification result with the most votes is used as the load type of the current load data; By using the Spearman correlation coefficient calculation method, the maximum load data vector corresponding to each load type in the load data set is extracted; Based on the load type of each load data in the load data set and the maximum load data vector corresponding to each load type, the power system to be configured is predicted, the substation capacity value of the power system to be configured is obtained, and the power system to be configured is configured according to the substation capacity value.

[0006] As an improvement of the above solution, the local data clustering of the initial load data to obtain load cluster samples includes: determining a plurality of initial class centers in the initial load data; Taking each of the initial class centers as input, repeatedly performing the class center cyclic change operation, and stopping the class center cyclic change operation after the class center no longer changes, to obtain the final class center; Determine a number of final clusters according to the final cluster center, and output a load clustering sample based on each final cluster; The class center loop change operation is specifically as follows: According to the Euclidean distance between the characteristic vector of each initial load data and each input data, classification is performed to obtain several initial groups; In each initial cluster, the characteristic vector distance between each initial load data is calculated, the sum of the distances between each initial load data and other load data is determined, and the initial load data with the smallest sum of distances is selected as the change cluster center; Determine whether the changed class center of each initial class group is consistent with the initial class center; If yes, the class center is no longer changing, the class center cycle change operation is stopped, and the final class center is output; If not, the class center has changed, and the changed class center is used as input data to re-execute the class center cyclic change operation.

[0007] As an improvement of the above solution, the balancing process is performed on the load cluster samples to obtain the load data set, including: Obtain the ratio between the sample size of each small category in each load clustering sample and the sample size of the largest category; wherein, a category of load data in the load clustering sample whose data volume is less than the sample threshold range is a small category; each category corresponds to a final group; Determine the composite multiple of each small category based on the ratio of the sample size of each small category to the sample size of the largest category; In each small category, samples are synthesized based on the synthesis multiple to obtain the number of synthetic samples; In each small category, according to the number of synthetic samples, the SMOTE oversampling method is used to synthesize samples to obtain synthetic samples that meet the number of synthetic samples, and the synthetic samples are added to the current small category to obtain a synthetic category; Based on each synthetic category and the category corresponding to the maximum category sample size, a load data set is obtained.

[0008] As an improvement of the above scheme, for the load data in the load data set, the current load data is classified by a plurality of base classifiers, and voting is performed based on the classification results of the current load data output by each base classifier, and the classification result with the most votes is used as the load type of the current load data, including: Extract data from the load data set according to a preset probability to obtain sampling data; The sampled data are preprocessed by deleting missing values ​​and normalizing them, and the preprocessed sampled data are sampled by Bootstrapping with replacement to obtain the sampled data; Input the sampled data into each base classifier respectively to obtain the classification result of the current load data output by each base classifier; wherein, the training of each base classifier includes: inputting the pre-stored historical load data into a BP neural network in Spark for network training, obtaining the error value generated by the training, and generating a base classifier when the error value is less than the error threshold; the network parameters corresponding to each base classifier are different; the historical load data is the historical load with load type; Vote for each classification result, and take the classification result with the most votes as the load type of the current load data.

[0009] As an improvement of the above solution, the method of calculating the Spearman correlation coefficient to extract the maximum load data vector corresponding to each load type in the load data set includes: In the load data corresponding to each load type in the load data set, the similarity between a load data and other load data except itself is calculated by the Spearman correlation coefficient calculation formula, and the sum of the similarities of the load data is accumulated; Among the load data corresponding to each load type in the load data set, the load data with the highest total similarity is selected as the maximum load data vector corresponding to the current load type.

[0010] As an improvement of the above scheme, the method of predicting the power system to be configured based on the load type of each load data in the load data set and the maximum load data vector corresponding to each load type to obtain the substation capacity value of the power system to be configured includes: Taking each load data in the load data set as an input of a local weighted periodic trend decomposition algorithm, and adjusting each load data in the load data set by the local weighted periodic trend decomposition algorithm to obtain an adjusted load data set; Based on the load type of each load data in the adjusted load data set and the maximum load data vector corresponding to each load type, the load type is input into the load prediction model to obtain the model prediction value; Determining a maximum simultaneous rate of the power system to be configured based on a maximum system load of the power system to be configured and maximum loads of several subsystems; Substitute the model prediction value, the maximum simultaneous rate, and the capacity of the distribution facilities into the open capacity calculation formula to obtain the area capacity value; wherein the open capacity calculation formula is specifically: P= Where, P is the capacity value of the area; For the capacity of the power distribution facilities, refer to the transformer marking; is the maximum simultaneous rate; is the model's predicted value.

[0011] As an improvement of the above solution, the training of the load forecasting model includes: Obtain historical area load data of the power system to be configured; wherein the historical area load data includes: a number of historical load samples, a load type of each historical load sample, and a maximum historical load sample data vector corresponding to each load type; Using the historical area load data as input of a local weighted periodic trend decomposition algorithm, adjusting the historical area load data through the local weighted periodic trend decomposition algorithm to obtain adjusted historical area load data; Performing a stationarity test and a differential process on the adjusted historical area load data to obtain a stationary time series of the adjusted historical area load data; According to the stationary time series, the autocorrelation coefficient and the partial autocorrelation coefficient of the adjusted historical load data of the substation are calculated, and the parameters of the SARIMA model are determined in the autocorrelation coefficient and the partial autocorrelation coefficient of the adjusted historical load data of the substation through the Akaike information criterion and the Bayesian information criterion; The historical substation load data is input into the SARIMA model with determined parameters for prediction to obtain the prediction error; when the prediction error is greater than or equal to the error threshold, the coefficient corresponding to the current parameter is excluded from the autocorrelation coefficient and partial autocorrelation coefficient of the adjusted historical substation load data, and the parameters of the SARIMA model are determined again through the Akaike information criterion and the Bayesian information criterion in the autocorrelation coefficient and partial autocorrelation coefficient of the adjusted historical substation load data after excluding the coefficient; when the prediction error is less than the error threshold, the load prediction model is output.

[0012] Accordingly, an embodiment of the present invention further provides a device for configuring the capacity of a power system, comprising: a data acquisition module, a data clustering module, a data classification module, a data calculation module and a data prediction module; The data acquisition module is used to acquire initial load data of the power system to be configured; The data clustering module is used to perform local data clustering on the initial load data to obtain load clustering samples, and perform balancing processing on the load clustering samples to obtain a load data set; The data classification module is used to classify the current load data in the load data set through a plurality of base classifiers, vote based on the classification results of the current load data output by each base classifier, and use the classification result with the most votes as the load type of the current load data; The data calculation module is used to extract the maximum load data vector corresponding to each load type in the load data set by using the Spearman correlation coefficient calculation method; The data prediction module is used to predict the power system to be configured based on the load type of each load data in the load data set and the maximum load data vector corresponding to each load type, obtain the substation capacity value of the power system to be configured, and configure the power system to be configured according to the substation capacity value.

[0013] Correspondingly, an embodiment of the present invention also provides a computer terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements a method for configuring the substation capacity of an electric power system as described in the present invention.

[0014] Correspondingly, an embodiment of the present invention further provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a method for configuring the substation capacity of an electric power system as described in the present invention.

[0015] As can be seen from the above, the present invention has the following beneficial effects: The present invention provides a method for configuring the substation capacity of an electric power system, which comprises the following steps: obtaining initial load data of an electric power system to be configured; performing local data clustering on the initial load data to obtain load clustering samples, and performing balancing processing on the load clustering samples to obtain a load data set; for the load data in the load data set, classifying the current load data through a plurality of base classifiers, voting based on the classification results of the current load data output by each base classifier, and taking the classification result with the largest number of votes as the load type of the current load data; extracting the maximum load data vector corresponding to each load type in the load data set through a Spearman correlation coefficient calculation method; predicting the electric power system to be configured based on the load type of each load data in the load data set and the maximum load data vector corresponding to each load type, obtaining the substation capacity value of the electric power system to be configured, and configuring the electric power system to be configured according to the substation capacity value. The present invention performs local clustering on the load data of the power system, and after the local clustering is completed, balances the data volume through balancing processing, so that the load data set input to the base classifier is more stable, and classification voting of the load data is performed through multiple base classifiers to make the load type more accurate, so as to predict the power system to be configured under the load data of the determined load type, which is beneficial to improving the calculation accuracy of the substation capacity value of the power system to be configured. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flow chart of a method for configuring the capacity of a power system area provided by an embodiment of the present invention; Figure 2 It is a structural schematic diagram of a power system area capacity configuration device provided by an embodiment of the present invention; Figure 3 It is a schematic diagram of the structure of a terminal device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0018] Embodiment 1 See also Figure 1 , Figure 1 The present invention is a flowchart of a method for configuring the capacity of a power system area provided by an embodiment of the present invention, which belongs to the technical field of power systems. Figure 1 As shown, this embodiment includes steps 101 to 105, and each step is specifically as follows: Step 101: Acquire initial load data of the power system to be configured.

[0019] In this embodiment, the initial load data is loaded through the data set Determined, the data set consists of n feature vectors represented as ={ , ,… }.

[0020] Step 102: Perform local data clustering on the initial load data to obtain load cluster samples, and perform balancing processing on the load cluster samples to obtain a load data set.

[0021] In this embodiment, the local data clustering of the initial load data to obtain load cluster samples includes: determining a plurality of initial class centers in the initial load data; Taking each of the initial class centers as input, repeatedly performing the class center cyclic change operation, and stopping the class center cyclic change operation after the class center no longer changes, to obtain the final class center; Determine a number of final clusters according to the final cluster center, and output a load clustering sample based on each final cluster; The class center loop change operation is specifically as follows: According to the Euclidean distance between the characteristic vector of each initial load data and each input data, classification is performed to obtain several initial groups; In each initial cluster, the characteristic vector distance between each initial load data is calculated, the sum of the distances between each initial load data and other load data is determined, and the initial load data with the smallest sum of distances is selected as the change cluster center; Determine whether the changed class center of each initial class group is consistent with the initial class center; If yes, the class center is no longer changing, the class center cycle change operation is stopped, and the final class center is output; If not, the class center has changed, and the changed class center is used as input data to re-execute the class center cyclic change operation.

[0022] In a specific embodiment, there are multiple initial cluster centers, which are the same as the number of final cluster centers.

[0023] In a specific embodiment, the load types include: According to the generation and supply, it is divided into: power load, power supply load, power generation load, etc.; According to the occurrence time, it can be divided into: peak load, minimum load and average load; According to the interruption loss, it is divided into: first, second and third level load; According to the electricity consumption sectors, it is divided into: industry, agriculture, transportation and municipal life electricity consumption.

[0024] In a specific embodiment, the following example is given to illustrate local clustering: 1) Initialize class center: load data set It is composed of n eigenvectors and is represented as ={ , ,… },exist Randomly select K vectors as the initial cluster center C= { , ,… }.

[0025] 2) Classification: All load feature vectors are divided into the centers of each class according to the closest Euclidean distance. arrive The distance calculation formula is: In the formula: r is and T is the dimension of the time series load feature vector, that is, the number of load collection periods for a certain user in a day. T, n, .

[0026] 3) Update the class center: In each class, calculate the distance sum from each load feature vector to other data vectors of the current class, and select the load data vector with the smallest distance sum as the new class center, as follows: Where: is the feature vector arrive The distance; J is The number of feature vectors in the class; for The sum of distances to all vectors of the current class.

[0027] 4) Repeat steps 2) and 3) until the cluster center no longer changes, and the clustering results are used as load training samples.

[0028] 5) According to step 3), the sum of the distances between each vector in each type of load and other data vectors in its class is calculated, and the load data in the training sample that exceeds the set threshold is filtered out.

[0029] In this embodiment, the balancing process is performed on the load cluster samples to obtain a load data set, including: Obtain the ratio between the sample size of each small category in each load clustering sample and the sample size of the largest category; wherein, a category of load data in the load clustering sample whose data volume is less than the sample threshold range is a small category; each category corresponds to a final group; Determine the composite multiple of each small category based on the ratio of the sample size of each small category to the sample size of the largest category; In each small category, samples are synthesized based on the synthesis multiple to obtain the number of synthetic samples; In each small category, according to the number of synthetic samples, the SMOTE oversampling method is used to synthesize samples to obtain synthetic samples that meet the number of synthetic samples, and the synthetic samples are added to the current small category to obtain a synthetic category; Based on each synthetic category and the category corresponding to the maximum category sample size, a load data set is obtained.

[0030] In a specific embodiment, the following example is given to illustrate the balancing process: 1) First, to ensure that the amount of small sample data after synthesis is equivalent to that of large category samples, the following settings are made for their sample weights: ∈(0 ,0.1] , =10; ∈(0.1 ,0.2] , =6; ∈(0.2 ,0.3] , =3; ∈(0.3 ,0.4] , =2; ∈(0.4 ,0.5] , =1; In the formula is the kth class of small samples to be processed after clustering With the maximum class sample size The ratio of = / , synthetic multiple , the number of synthesized samples required for the kth class for: 2) For any feature vector in the k-th small sample , The cluster center after its clustering The distance is , then the sample The weight calculation formula is: The number of new samples that need to be synthesized The calculation formula is: = And the actual number of synthesized new samples is obtained by rounding up the number of synthesized samples with corresponding weights: Where INT is the rounding function.

[0031] 3) In each category, sample synthesis is performed according to the SMOTE oversampling method. , within the class to which it belongs, by distance to the sample The distance size is selected in sequence. The sample data constitutes a set , expressed as ={ , ,…, }, Generate a new sample as follows: = +ζ( ) Where: ξ∈(0, 1) is a randomly generated value; for Any sample vector in .

[0032] 4) Repeat step 3) according to the number of samples generated to obtain the required number of load training data sets .

[0033] Step 103: For the load data in the load data set, the current load data is classified by a number of base classifiers, and voting is performed based on the classification results of the current load data output by each base classifier, and the classification result with the most votes is used as the load type of the current load data.

[0034] In this embodiment, for the load data in the load data set, the current load data is classified by a plurality of base classifiers, and voting is performed based on the classification results of the current load data output by each base classifier, and the classification result with the most votes is used as the load type of the current load data, including: Extract data from the load data set according to a preset probability to obtain sampling data; The sampled data are preprocessed by deleting missing values ​​and normalizing them, and the preprocessed sampled data are sampled by Bootstrapping with replacement to obtain the sampled data; Input the sampled data into each base classifier respectively to obtain the classification result of the current load data output by each base classifier; wherein, the training of each base classifier includes: inputting the pre-stored historical load data into a BP neural network in Spark for network training, obtaining the error value generated by the training, and generating a base classifier when the error value is less than the error threshold; the network parameters corresponding to each base classifier are different; the historical load data is the historical load with load type; Vote for each classification result, and take the classification result with the most votes as the load type of the current load data.

[0035] In a specific embodiment, extracting data from the load data set according to a preset probability to obtain sample data is specifically as follows: Bootstrapping is a process of repeatedly sampling from the load data set to construct a new sample set that can represent the distribution of the original samples. This process involves uniform sampling with replacement from the existing training set. Each time a sample is drawn, it is likely to be selected again and added to the training set again. The process of extracting training blocks is as follows: 1) Assume that the load data set is = { , , …, }, where any eigenvector =( , , …, ), where 1≤r≤R.

[0036] 2) In Randomly select a load data record The probability of Z is = 1 / R, using random sampling with replacement, The probability of being drawn next time remains unchanged.

[0037] In a specific embodiment, the training of the base classifier in this embodiment involves a distributed BPNN load classification process based on Spark, specifically: 1) Load data preprocessing: Delete the load data records containing vacant values ​​and perform normalization processing on the load collection points throughout the day, as shown below: T Where: , , , They respectively represent the maximum and minimum load data in a time period within a day, the load value in any time period, and the normalized value in that time period.

[0038] 2) Sampling load training block: all load training data = { , ,…, } Obtain M sample training blocks through Bootstrapping with replacement sampling ={ , ,…, In order to avoid the situation where an even number of base categories in the binary classification problem have equal numbers of classification results for the two types, and thus it is impossible to vote to obtain the final classification result, the value of M is 3, 5, or 7.

[0039] 3) Data storage format: The remaining load data are added as samples to be classified to each training block file, and each file is saved in the Hadoop distributed file system (Hadoop distributed file system, HDFS). The format of the load training data in the file is <"train", data, class>, and the format of the data to be classified is <"classify", data>. The first column is the distinguishing label of the two types of data, data represents the load vector, and class is the category of the load training data represented by a binary number.

[0040] 4) Network learning process: Spark reads files from HDFS and starts the same number of Mappers as the number of load data blocks. The BPNN in each Mapper initializes the network parameters. The number of neurons in the network input layer is equal to the dimension of the processed vector and can be flexibly adjusted. The error value is limited to the threshold through the forward calculation of the BP neural network and the back propagation of the error. The network training is completed and a BPNN base classifier with different classification performance is formed.

[0041] 5) Load classification stage: The load data to be classified is input into all trained BPNN base classifiers. Each base classifier only performs forward calculation of the input signal to obtain its own classification results. , The same load type exists.

[0042] 6) Classification type voting: The classification results of all base classifiers for the same load data are voted by majority vote, as shown below: Where: A is the number of classifiers; H is the number of categories; a=1, 2,…,A; h=1, 2,…,H; is the result of base classifier a classifying a load data into the hth category, ∈{0, 1}, when the base classifier a classifies the load data into the hth class =1, otherwise = 0. The classification type with the most votes is determined as the load type to which it belongs. .

[0043] Step 104: extracting the maximum load data vector corresponding to each load type in the load data set by using the Spearman correlation coefficient calculation method.

[0044] In this embodiment, the method of calculating the Spearman correlation coefficient to extract the maximum load data vector corresponding to each load type in the load data set includes: In the load data corresponding to each load type in the load data set, the similarity between a load data and other load data except itself is calculated by the Spearman correlation coefficient calculation formula, and the sum of the similarities of the load data is accumulated; Among the load data corresponding to each load type in the load data set, the load data with the highest total similarity is selected as the maximum load data vector corresponding to the current load type.

[0045] In a specific embodiment, the following case is provided to illustrate the Spearman correlation coefficient calculation method: The Spearman correlation coefficient is a correlation index in statistics that uses a monotonic equation to evaluate two statistical variables. It indicates the correlation direction of two independent variables. The calculation formula of the Spearman correlation coefficient is: Where: C is the Spearman correlation coefficient between any two vectors; T is the vector dimension; d is the ranking difference set of elements in the two vectors. The specific steps for selecting the load shape model are as follows: 1) Among various load data, the similarity between two load vectors is calculated according to the Spearman correlation coefficient; 2) For a load data vector, the sum of its similarity with all the data in the same class is as follows: Where: is the sum of the similarities between a load vector and all the data in its class; G is the number of vectors in this class.

[0046] 3) Select the data with the highest similarity to all data in the class, that is The largest load data vector is taken as the center of the morphology.

[0047] Step 105: Based on the load type of each load data in the load data set and the maximum load data vector corresponding to each load type, predict the power system to be configured, obtain the substation capacity value of the power system to be configured, and configure the power system to be configured according to the substation capacity value.

[0048] In this embodiment, the prediction of the power system to be configured based on the load type of each load data in the load data set and the maximum load data vector corresponding to each load type to obtain the substation capacity value of the power system to be configured includes: Taking each load data in the load data set as an input of a local weighted periodic trend decomposition algorithm, and adjusting each load data in the load data set by the local weighted periodic trend decomposition algorithm to obtain an adjusted load data set; Based on the load type of each load data in the adjusted load data set and the maximum load data vector corresponding to each load type, the load type is input into the load prediction model to obtain the model prediction value; Determining a maximum simultaneous rate of the power system to be configured based on a maximum system load of the power system to be configured and maximum loads of several subsystems; Substitute the model prediction value, the maximum simultaneous rate, and the capacity of the distribution facilities into the open capacity calculation formula to obtain the area capacity value; wherein the open capacity calculation formula is specifically: P= Where, P is the capacity value of the area; For the capacity of the power distribution facilities, refer to the transformer marking; is the maximum simultaneous rate; is the model's predicted value.

[0049] In this embodiment, the training of the load forecasting model includes: Obtain historical area load data of the power system to be configured; wherein the historical area load data includes: a number of historical load samples, a load type of each historical load sample, and a maximum historical load sample data vector corresponding to each load type; Using the historical area load data as input of a local weighted periodic trend decomposition algorithm, adjusting the historical area load data through the local weighted periodic trend decomposition algorithm to obtain adjusted historical area load data; Performing a stationarity test and a differential process on the adjusted historical area load data to obtain a stationary time series of the adjusted historical area load data; According to the stationary time series, the autocorrelation coefficient and the partial autocorrelation coefficient of the adjusted historical load data of the substation are calculated, and the parameters of the SARIMA model are determined in the autocorrelation coefficient and the partial autocorrelation coefficient of the adjusted historical load data of the substation through the Akaike information criterion and the Bayesian information criterion; The historical substation load data is input into the SARIMA model with determined parameters for prediction to obtain the prediction error; when the prediction error is greater than or equal to the error threshold, the coefficient corresponding to the current parameter is excluded from the autocorrelation coefficient and partial autocorrelation coefficient of the adjusted historical substation load data, and the parameters of the SARIMA model are determined again through the Akaike information criterion and the Bayesian information criterion in the autocorrelation coefficient and partial autocorrelation coefficient of the adjusted historical substation load data after excluding the coefficient; when the prediction error is less than the error threshold, the load prediction model is output.

[0050] It should be noted that this embodiment first uses a local weighted periodic trend decomposition algorithm to decompose the historical substation load data into trend terms, seasonal terms and residual terms; secondly, the adjusted historical substation load data is input into the load forecasting model to predict future substation load changes and load peaks; at the same time, the substation DSR criterion is established based on the historical substation load data; finally, the SARIMA-DSR model is used to reasonably adjust the configuration coefficient in the capacity calculation method to achieve accurate calculation of the substation capacity.

[0051] In a specific embodiment, the following case is provided to illustrate the local weighted periodic trend decomposition algorithm: First, the collected historical load data of the substation area is subjected to STL seasonal adjustment, and its model expression is as follows: = + + Where: represents the original time series; Represents the trend component of the time series; Represents the seasonal component of the time series; Represents the remaining component of the time series, that is, the residual component.

[0052] The key to the STL algorithm lies in the iterative process of Loess, which is divided into an inner loop and an outer loop. The iterative process of the inner loop is as follows: 1) Initial parameter assignment: k=0, =0.

[0053] 2) Detrending component: - .

[0054] 3) Periodic subsequence smoothing: Use Loess to regress and extend each subsequence to form a temporary seasonal sequence .

[0055] 4) Periodic subsequence low-pass filtering: Do 3 sliding averages and 1 Loess regression, and we get .

[0056] 5) To smooth the trend of periodic subsequences: = - .

[0057] 6) De-cycle: - .

[0058] 7) Trend smoothing: Perform Loess regression on step 6) to obtain .

[0059] 8) Result verification: If the maximum number of iterations is reached or , then output the decomposition result. If it is not satisfied, repeat steps 2) to 8).

[0060] The larger value of the remainder obtained in the inner loop is regarded as an outlier, and the outer loop introduces robust weights during Loess smoothing to handle outliers and improve the robustness of the algorithm.

[0061] The time series composed of transformer historical load data is decomposed into trend component, seasonal component and residual component to find out the long-term trend and seasonal characteristics of transformer load data, and the historical load data of the substation area decomposed by STL time series is adjusted for season and trend. The formula is as follows: = + In the formula t represents the adjusted time series.

[0062] In a specific embodiment, the formula of the SARIMA model is as follows: Φ( )φ(B) ( =θ(B)θ( ) Where: φ(B)=1- B- B-⋯- θ(B)=1- B - B-⋯- q Φ(B)=1- - -⋯- Θ(B)=1- B- B-⋯- Where: is the time series of the load data of the substation; s is the seasonal cycle parameter; is the stationary time series after difference; B represents the lag operator; 1-B represents the difference operator; Φ( )φ(B) is a seasonal autoregressive model; φ(B) represents a p-order autoregressive polynomial; , , is the non-seasonal autoregressive parameter; Φ( ) represents the seasonal autoregressive polynomial, , , is the F-order seasonal autoregressive parameter; θ(B)Θ( ) represents the seasonal moving average model, where θ(B) represents the q-order moving average polynomial, , , is the non-seasonal moving average parameter, Θ( ) represents the seasonal moving average polynomial; , , is the Q-order seasonal moving average parameter; is Gaussian noise.

[0063] In the formula, θ(B) and Θ(B) reflect the seasonal periodic relationship in the sequence, φ(B) and Φ(B) reflect the quantitative relationship between adjacent moments of the sequence; when P, D, and Q are all 0, it means that the sequence does not contain seasonal factors, and the SARIMA model degenerates into the ARIMA model.

[0064] In a specific embodiment, the following case is provided to illustrate the load forecasting model: the prediction process of the load power forecasting model based on SARIMA is as follows: 1) Stationarity test: The Augmented Dickey-Fuller (ADF) unit root test method is used to test whether the time series composed of the historical load data of the substation is a stationary time series. If the time series is a stationary series, the test statistic is significantly less than the three confidence critical values ​​of 1%, 5%, and 10%, or the P-value will be extremely close to 0. If the time series is a non-stationary series, proceed to step 2); otherwise, proceed to step 3).

[0065] 2) Difference processing: Perform d-order difference processing on the original time series to make it a stationary series. The formula for difference operation is as follows: = - 3) Model identification and parameter determination: The autocorrelation coefficient (ACF) and partial autocorrelation coefficient (PACF) of the stationary time series are calculated successively, and the model parameters are preliminarily screened from them. The Akaike information criterion and Bayesian information criterion are used to screen and determine the model parameters.

[0066] The calculation formula of Akaike's information criterion is as follows: AIC=-2ln(L)+2k The calculation formula of Bayesian information criterion is as follows: BIC=-2ln(L)+ln(d)·k Where: L is the maximum likelihood under the model; d is the number of samples; k is the number of variables in the model.

[0067] 4) Model verification: The above three steps can be used to preliminarily determine the parameters of the SARIMA model. After the model is established, the SARIMA model of the data can be optimized through residual analysis, error evaluation and other methods. If the error between the prediction result of the established model and the actual data is large, the parameters of the SARIMA model of the data need to be adjusted.

[0068] See also Figure 2 , Figure 2 It is a structural schematic diagram of a power system area capacity configuration device provided by an embodiment of the present invention, comprising: a data acquisition module 201, a data clustering module 202, a data classification module 203, a data calculation module 204 and a data prediction module 205; The data acquisition module is used to acquire initial load data of the power system to be configured; The data clustering module is used to perform local data clustering on the initial load data to obtain load clustering samples, and perform balancing processing on the load clustering samples to obtain a load data set; The data classification module is used to classify the current load data in the load data set through a plurality of base classifiers, vote based on the classification results of the current load data output by each base classifier, and use the classification result with the most votes as the load type of the current load data; The data calculation module is used to extract the maximum load data vector corresponding to each load type in the load data set by using the Spearman correlation coefficient calculation method; The data prediction module is used to predict the power system to be configured based on the load type of each load data in the load data set and the maximum load data vector corresponding to each load type, obtain the substation capacity value of the power system to be configured, and configure the power system to be configured according to the substation capacity value.

[0069] It can be understood that the above-mentioned system item embodiments correspond to the method item embodiments of the present invention, which can implement the substation capacity configuration method of the power system provided by any one of the above-mentioned method item embodiments of the present invention.

[0070] This embodiment obtains initial load data of the power system to be configured; performs local data clustering on the initial load data to obtain load clustering samples, and performs balancing processing on the load clustering samples to obtain a load data set; for the load data in the load data set, classifies the current load data through a plurality of base classifiers, and votes based on the classification results of the current load data output by each base classifier, and takes the classification result with the largest number of votes as the load type of the current load data; extracts the maximum load data vector corresponding to each load type in the load data set through the Spearman correlation coefficient calculation method; predicts the power system to be configured based on the load type of each load data in the load data set and the maximum load data vector corresponding to each load type, obtains the substation capacity value of the power system to be configured, and configures the power system to be configured according to the substation capacity value. The present invention performs local clustering on the load data of the power system, and after the local clustering is completed, balances the data volume through balancing processing, so that the load data set input to the base classifier is more stable, and classification voting of the load data is performed through multiple base classifiers to make the load type more accurate, so as to predict the power system to be configured under the load data of the determined load type, which is beneficial to improving the calculation accuracy of the substation capacity value of the power system to be configured.

[0071] Embodiment 2 See also Figure 3 , Figure 3 It is a schematic diagram of the structure of a terminal device provided in one embodiment of the present invention.

[0072] A terminal device of this embodiment includes: a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program, the steps of the above-mentioned method for configuring the capacity of each power system in the embodiment are implemented, for example: Figure 1 Alternatively, when the processor executes the computer program, the functions of each module in the above-mentioned device embodiments are realized, for example: Figure 2All modules of the area capacity configuration device of the power system shown.

[0073] In addition, an embodiment of the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the substation capacity configuration method of the power system as described in any of the above embodiments.

[0074] Those skilled in the art will understand that the schematic diagram is merely an example of a terminal device and does not constitute a limitation on the terminal device. The terminal device may include more or fewer components than shown in the diagram, or a combination of certain components, or different components. For example, the terminal device may also include input and output devices, network access devices, buses, etc.

[0075] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor 301 is the control center of the terminal device, and uses various interfaces and lines to connect various parts of the entire terminal device.

[0076] The memory 302 can be used to store the computer program and / or module. The processor 301 implements various functions of the terminal device by running or executing the computer program and / or module stored in the memory and calling the data stored in the memory 302. The memory 302 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0077] Wherein, if the module / unit integrated in the terminal device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0078] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art may understand and implement it without paying any creative effort.

[0079] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for configuring the capacity of a power system, characterized in that: include: Obtaining initial load data of the power system to be configured; Performing local data clustering on the initial load data to obtain load cluster samples, and performing balancing processing on the load cluster samples to obtain a load data set; For the load data in the load data set, the current load data is classified by several base classifiers, and voting is performed based on the classification results of the current load data output by each base classifier, and the classification result with the most votes is used as the load type of the current load data; By using the Spearman correlation coefficient calculation method, the maximum load data vector corresponding to each load type in the load data set is extracted; Based on the load type of each load data in the load data set and the maximum load data vector corresponding to each load type, the power system to be configured is predicted, the substation capacity value of the power system to be configured is obtained, and the power system to be configured is configured according to the substation capacity value.

2. The method for configuring the area capacity of the power system according to claim 1, characterized in that: The performing local data clustering on the initial load data to obtain load cluster samples includes: determining a plurality of initial class centers in the initial load data; Taking each of the initial class centers as input, repeatedly performing the class center cyclic change operation, and stopping the class center cyclic change operation after the class center no longer changes, to obtain the final class center; Determine a number of final clusters according to the final cluster center, and output a load clustering sample based on each final cluster; The class center loop change operation is specifically as follows: According to the Euclidean distance between the characteristic vector of each initial load data and each input data, classification is performed to obtain several initial groups; In each initial cluster, the characteristic vector distance between each initial load data is calculated, the sum of the distances between each initial load data and other load data is determined, and the initial load data with the smallest sum of distances is selected as the change cluster center; Determine whether the changed class center of each initial class group is consistent with the initial class center; If yes, the class center is no longer changing, the class center cycle change operation is stopped, and the final class center is output; If not, the class center has changed, and the changed class center is used as input data to re-execute the class center cyclic change operation.

3. The method for configuring the area capacity of the power system according to claim 2, characterized in that: The balancing process is performed on the load cluster samples to obtain a load data set, including: Obtain the ratio between the sample size of each small category in each load clustering sample and the sample size of the largest category; wherein, a category of load data in the load clustering sample whose data volume is less than the sample threshold range is a small category; each category corresponds to a final group; Determine the composite multiple of each small category based on the ratio of the sample size of each small category to the sample size of the largest category; In each small category, samples are synthesized based on the synthesis multiple to obtain the number of synthetic samples; In each small category, according to the number of synthetic samples, the SMOTE oversampling method is used to synthesize samples to obtain synthetic samples that meet the number of synthetic samples, and the synthetic samples are added to the current small category to obtain a synthetic category; Based on each synthetic category and the category corresponding to the maximum category sample size, a load data set is obtained.

4. The method for configuring the area capacity of the power system according to claim 3, characterized in that: The load data in the load data set is classified by using a plurality of base classifiers, and voting is performed based on the classification results of the current load data output by each base classifier, and the classification result with the most votes is used as the load type of the current load data, including: Extract data from the load data set according to a preset probability to obtain sampling data; The sampled data are preprocessed by deleting missing values ​​and normalizing them, and the preprocessed sampled data are sampled by Bootstrapping with replacement to obtain the sampled data; Input the sampled data into each base classifier respectively to obtain the classification result of the current load data output by each base classifier; wherein, the training of each base classifier includes: inputting the pre-stored historical load data into a BP neural network in Spark for network training, obtaining the error value generated by the training, and generating a base classifier when the error value is less than the error threshold; the network parameters corresponding to each base classifier are different; the historical load data is the historical load with load type; Vote for each classification result, and take the classification result with the most votes as the load type of the current load data.

5. The method for configuring the area capacity of the power system according to claim 4, characterized in that: The method of calculating the Spearman correlation coefficient to extract the maximum load data vector corresponding to each load type in the load data set includes: In the load data corresponding to each load type in the load data set, the similarity between a load data and other load data except itself is calculated by the Spearman correlation coefficient calculation formula, and the sum of the similarities of the load data is accumulated; Among the load data corresponding to each load type in the load data set, the load data with the highest total similarity is selected as the maximum load data vector corresponding to the current load type.

6. The method for configuring the area capacity of the power system according to claim 5, characterized in that: The predicting of the power system to be configured based on the load type of each load data in the load data set and the maximum load data vector corresponding to each load type to obtain the substation capacity value of the power system to be configured includes: Taking each load data in the load data set as an input of a local weighted periodic trend decomposition algorithm, and adjusting each load data in the load data set by the local weighted periodic trend decomposition algorithm to obtain an adjusted load data set; Based on the load type of each load data in the adjusted load data set and the maximum load data vector corresponding to each load type, the load type is input into the load prediction model to obtain the model prediction value; Determining a maximum simultaneous rate of the power system to be configured based on a maximum system load of the power system to be configured and maximum loads of several subsystems; Substitute the model prediction value, the maximum simultaneous rate, and the capacity of the distribution facilities into the open capacity calculation formula to obtain the area capacity value; wherein the open capacity calculation formula is specifically: P= Where, P is the capacity value of the area; For the capacity of the power distribution facilities, refer to the transformer marking; is the maximum simultaneous rate; is the model's predicted value.

7. The method for configuring the area capacity of the power system according to claim 6, characterized in that: The training of the load forecasting model includes: Obtain historical area load data of the power system to be configured; wherein the historical area load data includes: a number of historical load samples, a load type of each historical load sample, and a maximum historical load sample data vector corresponding to each load type; Using the historical area load data as input of a local weighted periodic trend decomposition algorithm, adjusting the historical area load data through the local weighted periodic trend decomposition algorithm to obtain adjusted historical area load data; Performing a stationarity test and a differential process on the adjusted historical area load data to obtain a stationary time series of the adjusted historical area load data; According to the stationary time series, the autocorrelation coefficient and the partial autocorrelation coefficient of the adjusted historical load data of the substation are calculated, and the parameters of the SARIMA model are determined in the autocorrelation coefficient and the partial autocorrelation coefficient of the adjusted historical load data of the substation through the Akaike information criterion and the Bayesian information criterion; The historical substation load data is input into the SARIMA model with determined parameters for prediction to obtain the prediction error; when the prediction error is greater than or equal to the error threshold, the coefficient corresponding to the current parameter is excluded from the autocorrelation coefficient and partial autocorrelation coefficient of the adjusted historical substation load data, and the parameters of the SARIMA model are determined again through the Akaike information criterion and the Bayesian information criterion in the autocorrelation coefficient and partial autocorrelation coefficient of the adjusted historical substation load data after excluding the coefficient; when the prediction error is less than the error threshold, the load prediction model is output.

8. A device for configuring the capacity of a power system, characterized in that: include: Data acquisition module, data clustering module, data classification module, data calculation module and data prediction module; The data acquisition module is used to acquire initial load data of the power system to be configured; The data clustering module is used to perform local data clustering on the initial load data to obtain load clustering samples, and perform balancing processing on the load clustering samples to obtain a load data set; The data classification module is used to classify the current load data in the load data set through a plurality of base classifiers, vote based on the classification results of the current load data output by each base classifier, and use the classification result with the most votes as the load type of the current load data; The data calculation module is used to extract the maximum load data vector corresponding to each load type in the load data set by using the Spearman correlation coefficient calculation method; The data prediction module is used to predict the power system to be configured based on the load type of each load data in the load data set and the maximum load data vector corresponding to each load type, obtain the substation capacity value of the power system to be configured, and configure the power system to be configured according to the substation capacity value.

9. A computer terminal device, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a method for configuring the substation capacity of an electric power system as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute a method for configuring the substation capacity of a power system as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Non-invasive load identification algorithm based on hybrid neural network and ensemble learning

    CN107122790A

  • Daily load curve clustering method based on ant colony algorithm and C-K algorithm

    CN113392877A

  • SARIMA and coincidence rate-based area openable capacity calculation method

    CN115575703A

  • Air conditioner load prediction method and system, electronic equipment and computer storage medium

    CN116468138A

  • Typical load mode extraction method and system

    CN116701901A