A distribution network abnormal data prediction method and system based on data analysis
By clustering and average calculation of distribution network data, combining locust optimization algorithm and BP neural network model, the accuracy problem of prediction of distribution network abnormal data is solved, and efficient identification and prediction of future abnormal data is achieved to ensure the stability of the power system.
Patent Information
- Application Number
- CN202411444515.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-11
- Filing Date
- 2024-10-16
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2044-10-16
AI Technical Summary
The prior art cannot effectively predict abnormal data of distribution networks, resulting in reduced stability and reliability of the power system, increased maintenance costs, and may even cause large-scale power outages.
By collecting distribution network data at multiple historical time points, clustering and average calculations are performed, combining locust optimization algorithm and BP neural network model, distribution network data is classified and predicted, non-failed and faulty abnormal time periods are identified, and existing periodic data is used for comparison and analysis to improve prediction accuracy.
It improves the accuracy and comprehensiveness of abnormal data prediction in distribution networks, reduces the data contingency, enhances the predictive ability of future abnormal data, and ensures the stability and reliability of the power system.
Smart Images

Figure CN119322774B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of abnormal data prediction and analysis, and in particular, relates to a distribution network abnormal data prediction method and system based on data analysis. Background Art
[0002] Chinese patent CN117472898B discloses a fusion-based distribution network abnormal data correction method and system, which collects distribution network data and determines whether there is abnormal data therein; if there is second power abnormal data, the location of the abnormal data is obtained, and the power data sequence is divided into a normal power data subsequence and an abnormal power data subsequence; the normal power data subsequence is input into a preset LSTM neural network for training to obtain a first target LSTM model; at least one normal power data associated with a certain abnormal power data is input into the first target LSTM model to obtain predicted power data, and the predicted power data replaces the abnormal power data.
[0003] Abnormal distribution network data may reduce the stability and reliability of the power system, affect power quality, increase maintenance costs, and even cause large-scale power outages. Therefore, if abnormal distribution network data cannot be predicted, measures cannot be taken in advance to prevent the occurrence of the above situations, resulting in economic losses to enterprises and even affecting people's daily lives. Summary of the Invention
[0004] In response to the problems in the related art, the present invention proposes a distribution network abnormal data prediction method and system based on data analysis to overcome the above technical problems existing in the existing related art.
[0005] To solve the above technical problems, the present invention is achieved through the following technical solutions:
[0006] The present invention is a method for predicting abnormal data of distribution network based on data analysis, comprising the following steps:
[0007] S1. Collect distribution network data of various types and multiple historical time points to obtain a distribution network data matrix set;
[0008] S2. Perform a clustering operation on the distribution network data matrix set to obtain a distribution network data classification matrix; then classify the distribution network data classification matrix to obtain an initial non-fault abnormal time period matrix and an initial fault abnormal time period matrix; calculate the average values of the initial non-fault abnormal time period matrix and the initial fault abnormal time period matrix to obtain an average matrix of the non-fault abnormal time period set and an average matrix of the fault abnormal time period set;
[0009] S3. Collecting distribution network data of multiple types and multiple historical time points to be predicted again to obtain a historical distribution network data matrix; predicting and classifying the distribution network data at future moments based on the historical distribution network data matrix to obtain a first analysis result;
[0010] S4. Classify the distribution network abnormality data in the historical distribution network data matrix to obtain a historical non-fault abnormality time period matrix and a historical fault abnormality time period matrix;
[0011] S5. Perform a prediction analysis on the historical non-fault abnormal time period matrix and the historical fault abnormal time period matrix based on the average matrix of the non-fault abnormal time period set and the average matrix of the fault abnormal time period set to obtain a second analysis result;
[0012] S6. Compare the first analysis result and the second analysis result to obtain a final non-fault abnormal time period matrix and a final fault abnormal time period matrix;
[0013] Different distribution network data have different changing trends over time and need to be analyzed separately. Therefore, the collection of distribution network data of various types and multiple historical time points provides data support for the subsequent analysis of the time point distribution of abnormal data in the distribution network data; by clustering the distribution network data matrix set, the distribution network data with similar data are aggregated, which facilitates the subsequent calculation of the average value of the initial non-fault type abnormal time period matrix and the initial fault type abnormal time period matrix; by calculating the average value of the initial non-fault type abnormal time period matrix and the initial fault type abnormal time period matrix, the sporadic nature of the data is reduced, thereby improving the accuracy of subsequent analysis; by collecting and predicting distribution network data of various types and multiple historical time points, a data basis is provided for subsequent prediction operations; this scheme predicts the historical distribution network data matrix from two perspectives. One is to predict the distribution network data in the future based on the changing trend of the distribution network data of multiple historical time points over time, and to perform an analysis on the predicted unseen data. The first step is to classify and analyze abnormal data in the distribution network data at the previous time point; since electricity consumption and production are affected by factors such as daily life patterns, industrial production cycles, and seasonal changes; for example, household and commercial electricity consumption will have different patterns on weekdays and weekends, daytime and nighttime, while industrial electricity consumption may have more obvious periodicity, such as changes by shift or production cycle; in addition, seasonal changes will also affect electricity demand, such as the increase in power load caused by increased use of air conditioners in summer, so distribution network data usually has periodicity; therefore, the second step is to find the periodicity of abnormal data in the distribution network, compare the distribution network data collected at multiple historical time points with the existing distribution network abnormality data with known cycles, and use the most similar existing distribution network abnormality data as the distribution network data at the future time point, and then classify and analyze the distribution network data at the future time point; obtain two analysis results from these two perspectives, and finally compare and analyze the two analysis results comprehensively, so that the final analysis result is more comprehensive and accurate.
[0014] Preferably, the S1 comprises the following steps:
[0015] S11, set a variety of distribution network data types and corresponding data statistical periods to obtain a distribution network data type set a={a1, a2, ..., a i ,...,a a′} and the data statistical period set b′={b1′,b2′,...,b i ′,...,b a ″},a i Indicates the type of the set i-th network distribution data, a′ indicates the total number of set network distribution types; b i ' represents the data statistics period set for the i-th type of distribution network data;
[0016] Set the same number of historical time points for the data statistical period corresponding to each distribution network data type to obtain the historical time point matrix as follows,
[0017]
[0018] in, Indicates the jth historical time point set for the data statistics period corresponding to the i-th distribution network data type. Indicates the total number of historical time points set for the data statistics period corresponding to each distribution network data type;
[0019] According to the historical time point matrix And the data type set of the distribution network is collected a={a1,a2,...,a i ,...,a a′}, and obtain a distribution network data matrix set;
[0020] S12: Set the corresponding abnormal data threshold for each distribution network data type to obtain the abnormal data threshold set. represents the abnormal data threshold set for the i-th type of distribution network data;
[0021] The distribution network data types include electrical parameters such as voltage, current, active power, reactive power, frequency, as well as circuit breaker and isolation status information. The change cycles of these different types of distribution network data are different. Therefore, a targeted data statistical cycle is set for each type of distribution network data. This makes the collected distribution network data of each type more complete and reflects its periodicity. Secondly, the abnormal data thresholds of different types of distribution network data are also different. Therefore, setting corresponding abnormal data thresholds for different types of distribution network data improves the accuracy of subsequent analysis.
[0022] Preferably, said S2 comprises the following steps:
[0023] S21, performing a clustering operation on each distribution network data matrix in the distribution network data matrix set to obtain a distribution network data classification matrix c; as follows,
[0024]
[0025] Among them, c ij represents the j-th classification of the distribution network data matrix corresponding to the i-th distribution network data type; c ijkk′ Indicates c ij The distribution network data at the k′th historical time point in the kth data statistical period; Indicates c ijThe total number of data statistical cycles corresponding to the distribution network data;
[0026] S22. Setting the threshold value set for the number of consecutive time points Indicates the threshold value of the number of consecutive time points corresponding to the i-th type of distribution network data type; in conjunction with the threshold value set of the number of consecutive time points and abnormal data threshold set When the time point data of the distribution network abnormal data in each cycle of each distribution network data matrix corresponding to each classification of each distribution network data type in the distribution network data classification matrix c is less than the corresponding continuous time point number threshold, the time period corresponding to the distribution network abnormal data is used as the non-fault abnormal time period; otherwise, the time period corresponding to the distribution network abnormal data is used as the fault abnormal time period; the initial non-fault abnormal time period matrix d1 and the initial fault abnormal time period matrix d2 are obtained; as follows,
[0027]
[0028] Among them, d 1ij and d 2ij They represent the non-fault abnormal time period matrix and fault abnormal time period matrix of the jth classification corresponding to the i-th distribution network data type respectively; they are as follows,
[0029]
[0030] in, Respectively represent d 1ij The starting time point and ending time point of the k′th non-fault abnormal time period in the kth data statistical period; Respectively represent d 2ij The starting and ending time points of the k′th fault-type abnormal time period in the kth data statistical period; d1′ ij d2′ ij Respectively represent d 1ij The total number of non-fault abnormal time periods in each data statistical period and d 2ij The total number of fault-type abnormal time periods in each data statistical period;
[0031] S23, respectively calculating the average value of the starting time point and the ending time point of each non-fault abnormal time period in each category and the average value of the starting time point and the ending time point of each fault abnormal time period in the initial non-fault abnormal time period matrix d1 and the initial fault abnormal time period matrix d2, to obtain the average matrix d3 of the non-fault abnormal time period set and the average matrix d4 of the fault abnormal time period set; respectively as follows,
[0032]
[0033] Among them, d 3ij d 4ij They represent the non-fault type abnormal average time period set and the fault type abnormal average time period set of the jth classification corresponding to the i-th distribution network data type; they are as follows,
[0034]
[0035] in, Respectively represent d 3ij The average value of the starting time point and the ending time point of the k′th non-fault type abnormal average time period in ; Respectively represent d 4ij The average value of the starting time point and the ending time point of the k′th fault type abnormal average time period in the calculation formula are as follows:
[0036]
[0037] Since there are many possibilities for the occurrence of abnormal data, such as sudden centralized power consumption or failure of each electrical component; for the former, it is temporary, so the abnormal data caused by it is short-lived, and when the centralized power consumption ends, the abnormal data will return to normal; for the latter, it is permanent and irreversible, requiring manual repair; therefore, it lasts longer; based on the above, in this scheme, the abnormal data appearing in the distribution network data is divided into non-fault-type abnormal time periods according to the number of continuous time points, that is, the time period where the abnormal data is automatically recovered, and fault-type abnormal time periods, that is, the time period where the abnormal data is irreversible.
[0038] Preferably, the S21 includes the following steps:
[0039] S211. Building a locust population represents the i-th locust in the locust population, Indicates the size of the locust population; Set the maximum number of iterations of the locust population to e1, the current number of iterations to e2, and the dimension of the search space to
[0040] S212, randomly selecting a number of distribution network data sets from the distribution network data matrix set as initial classification center data multiple times to obtain an initial classification center data matrix set; using the initial classification center data matrix set as the initial position matrix set of the locust population e′ k′ represents the initial position matrix of the i-th locust in the locust population; as follows,
[0041]
[0042] Among them, e′ k′ij represents e′ k′ The initial classification center data set of the jth classification corresponding to the i-th distribution network data type, e′ k′ijk represents e′ k′ij The initial classification center data at the k-th historical time point;
[0043] S213, calculate the distribution network data matrix set b2={b 21 ,b 22 ,...,b 2i ,...,b 2a′ The sum of the Euclidean distances between each distribution network dataset and the corresponding classification center dataset in} is used to obtain the Euclidean distance data matrix as follows,
[0044]
[0045] in, represents e′ k′ij Each distribution network data in the corresponding category is sent to e' k′ij The sum of the Euclidean distances between them;
[0046] According to the Euclidean distance data matrix Set the fitness function of the locust population as follows,
[0047]
[0048] S214, start iteration, before the iteration, set the current iteration number e2 to 1; in the first round of iteration, the fitness function of the locust population is used. Calculate the initial position matrix set The fitness value of the initial position matrix of each locust in the first fitness value set is obtained to obtain a first fitness value set; the maximum fitness value and the initial position of the corresponding locust are selected from the first fitness value set as the first global optimal fitness value and the first global optimal position respectively; the initial position matrix set is adjusted according to the first global optimal fitness value and the first global optimal position. The initial position matrix of each locust is updated, and after the update is completed, the current iteration number e2 is increased by 1, and the next round of iteration is entered;
[0049] Except for the first round of iteration, the fitness function of the locust population is used in each other round of iteration. Calculating the fitness value of the position matrix of each locust obtained in the previous iterative process to obtain a second fitness value set; selecting the maximum fitness value and the initial position of the corresponding locust from the second fitness value set as the second global optimal fitness and the second global optimal position, respectively; updating the position matrix of each locust obtained in the previous iterative process according to the second global optimal fitness and the second global optimal position; after the update is completed, increasing the current iteration number e2 by 1, and entering the next iteration;
[0050] S215, when e2≥e1, stop iteration and obtain the final global optimal position; use the final global optimal position as the final classification center data matrix; and calculate the distribution network data matrix set b2={b 21 ,b 22 ,...,b 2i ,...,b 2a′} to classify and obtain the distribution network data classification matrix c;
[0051] The locust optimization algorithm has the advantages of balancing global search and local search, having strong global search capabilities, and being able to find better solutions in complex search spaces. Therefore, in this scheme, the locust optimization algorithm is used to perform multiple iterations to optimize each classification center data in the distribution network data matrix set, and the overall discrete degree of the distribution network data matrix set is used as its fitness function. As the iteration proceeds, the overall discrete degree of the distribution network data matrix set will become lower and lower, that is, the classification center data found will be more accurate.
[0052] Preferably, the step S3 includes the following steps:
[0053] S31. Setting a set of time points to be predicted and future time points f 1i 、f 2i They represent the i-th time point to be predicted and the future time point respectively, and f1′ and f2′ represent the total number of time points to be predicted and the future time points respectively;
[0054] According to the distribution network data type set a={a1, a2, ..., a i ,...,a a′} and the set of time points to be predicted Collect distribution network data of various distribution network data types to obtain historical distribution network data matrix as follows,
[0055]
[0056] in, The network distribution data representing the i-th network distribution data type corresponding to the j-th time point to be predicted;
[0057] S32, constructing an initial BP neural network prediction model and setting the training data ratio and the test data ratio; and calculating the historical distribution network data matrix according to the training data ratio and the test data ratio. Perform data division to obtain the historical distribution network training data matrix, the historical distribution network test data matrix, and the historical distribution network to-be-predicted data matrix;
[0058] The historical distribution network training data matrix and the historical distribution network test data matrix are respectively used to train and test the initial BP neural network prediction model; after the training and testing are completed, the final BP neural network prediction model is obtained;
[0059] S33, cooperate with the future time point set The final BP neural network prediction model is used to perform prediction operations on the historical distribution network to be predicted data matrix to obtain the future distribution network data matrix as follows,
[0060]
[0061] in, Represents the predicted distribution network data of the i-th distribution network data type corresponding to the j-th future time point;
[0062] S34, the future distribution network data matrix The distribution network data and abnormal data threshold set at each future time point in Compare the corresponding abnormal data thresholds in to obtain the future distribution network abnormal time period matrix;
[0063] S35, cooperate with the threshold value set of the number of consecutive time points Classify each future distribution network abnormal time period in the future distribution network abnormal time period matrix to obtain the first non-fault abnormal time period matrix to be analyzed and the first fault type abnormal time period matrix h to be analyzed; respectively as follows,
[0064]
[0065] in, Represent the future distribution network data matrix The starting time point and ending time point of the jth non-fault abnormal time period corresponding to the i-th distribution network data type; Represent the future distribution network data matrix The starting time point and ending time point of the jth fault type abnormal time period corresponding to the i-th distribution network data type; hi ′ represents the future distribution network data matrix The total number of non-fault abnormal time periods and fault abnormal time periods corresponding to the i-th distribution network data type;
[0066] By constructing the final BP neural network prediction model, a prediction model is provided for the subsequent prediction operation of the historical distribution network data matrix to be predicted.
[0067] Preferably, in S32, the historical distribution network training data matrix and the historical distribution network test data matrix are respectively used to train and test the initial BP neural network prediction model; after the training and testing are completed, obtaining the final BP neural network prediction model includes the following steps:
[0068] S321. Setting a training error threshold; inputting the historical distribution network training data matrix into the initial BP neural network prediction model for training; during the training process, when the training error is less than the training error threshold, stopping the training operation to obtain a trained BP neural network prediction model; otherwise, continuing the training until the training error is less than the training error threshold;
[0069] S322, setting a test accuracy threshold; inputting the historical distribution network test data matrix into the trained BP neural network prediction model for testing, and obtaining the test accuracy after the test is completed; when the test accuracy is greater than or equal to the test accuracy threshold, using the trained BP neural network prediction model as the final BP neural network prediction model; otherwise, returning to S321 to continue training the trained BP neural network prediction model until the test accuracy is greater than or equal to the test accuracy threshold;
[0070] By using the historical distribution network training data matrix and the historical distribution network test data matrix to train and test the initial BP neural network prediction model, the final BP neural network prediction model obtained has good prediction ability for each type of distribution network time series data.
[0071] Preferably, the S4 comprises the following steps:
[0072] S41, the historical distribution network data matrix The distribution network data and abnormal data threshold set for each time point to be predicted Compare the corresponding abnormal data thresholds in the , and obtain the historical distribution network abnormal time period matrix;
[0073] S42, cooperate with the threshold value set of the number of consecutive time points Classify each historical distribution network abnormal time period in the historical distribution network abnormal time period matrix to obtain the historical non-fault abnormal time period matrix g and the historical fault abnormal time period matrix g′; they are as follows:
[0074]
[0075] in, Represents the historical distribution network data matrix The starting time point and ending time point of the jth non-fault abnormal time period corresponding to the i-th distribution network data type; Represents the historical distribution network data matrix The starting time point and ending time point of the jth fault type abnormal time period corresponding to the i-th distribution network data type; g i ′ represents the historical distribution network data matrix The total number of non-fault abnormal time periods and fault abnormal time periods corresponding to the i-th distribution network data type;
[0076] The distribution network data collected at multiple historical time points are compared with the existing distribution network anomaly data with known cycles, and the most similar existing distribution network anomaly data is used as the distribution network data at a future time point, and the distribution network data at the future time point is classified and analyzed.
[0077] Preferably, the S5 comprises the following steps:
[0078] S51, calculate the Euclidean distance between the historical non-fault abnormal time period matrix g and each non-fault abnormal time period set in the corresponding row in the average matrix d3 of the non-fault abnormal time period set, and the Euclidean distance between the historical fault abnormal time period matrix g′ and each fault abnormal time period set in the corresponding row in the average matrix d4 of the fault abnormal time period set, to obtain a first Euclidean distance matrix l1 and a second Euclidean distance matrix l2; respectively,
[0079]
[0080] Among them, l 1ij 、l 2ij Respectively represent the data of row i and row d in the matrix g of the historical non-fault abnormal time period 3ij The Euclidean distance between the data in row i and the data in row d in the historical fault type abnormal time period matrix g′ 4ij The Euclidean distance between
[0081] S52, select the smallest Euclidean distance in each row of the first Euclidean distance matrix l1 and the second Euclidean distance matrix l2, and use the corresponding non-fault abnormal time period set and fault abnormal time period set as the future time data of the historical non-fault abnormal time period set of the corresponding row in the historical non-fault abnormal time period matrix g and the future time data of the historical fault abnormal time period set of the corresponding row in the historical fault abnormal time period matrix g′, respectively, to obtain the second non-fault abnormal time period matrix to be analyzed and the second fault type abnormal time period matrix to be analyzed They are as follows:
[0082]
[0083]
[0084] in, as well as Represents the second non-fault abnormal time period matrix to be analyzed and the second fault type abnormal time period matrix to be analyzed The starting time point and ending time point of the jth non-fault abnormal time period and the fault abnormal time period corresponding to the i-th distribution network data type;
[0085] The distribution network data collected at multiple historical time points are compared with the average matrix of the non-fault type abnormal time period set and the average matrix of the fault type abnormal time period set. Since the average matrix of the non-fault type abnormal time period and the average matrix of the fault type abnormal time period set are periodic abnormal time periods, the most similar existing distribution network abnormal data can directly reflect the distribution network abnormal data in the future.
[0086] Preferably, the S6 comprises the following steps:
[0087] S61: Calculate the first non-fault abnormal time period matrix to be analyzed And the second non-fault abnormal time period matrix to be analyzed The union of the fault-type abnormal time periods at the corresponding positions in is used to obtain the final non-fault-type abnormal time period matrix;
[0088] S62: Calculate the first to-be-analyzed fault type abnormal time period matrix h and the second to-be-analyzed fault type abnormal time period matrix h. The union of the fault type abnormal time periods at the corresponding positions in is used to obtain the final fault type abnormal time period matrix;
[0089] By obtaining the first non-fault type abnormal time period matrix to be analyzed and the second non-fault type abnormal time period matrix to be analyzed, as well as the first fault type abnormal time period matrix to be analyzed and the second fault type abnormal time period matrix to be analyzed, and finally performing a union operation on the first non-fault type abnormal time period matrix to be analyzed and the second non-fault type abnormal time period matrix to be analyzed, as well as the first fault type abnormal time period matrix to be analyzed and the second fault type abnormal time period matrix to be analyzed, the final analysis result is more comprehensive and accurate.
[0090] A distribution network abnormal data prediction system based on data analysis includes a first distribution network data acquisition module, a clustering module, a first distribution network abnormal data classification module, an abnormal data time point mean value acquisition module, a second distribution network data acquisition module, a first prediction analysis module, a second distribution network abnormal data classification module, a second prediction analysis module and a comparative analysis module.
[0091] The present invention has the following beneficial effects:
[0092] 1. The present invention reduces the sporadic nature of the data by calculating the average values of the initial non-fault-type abnormal time period matrix and the initial fault-type abnormal time period matrix, thereby improving the accuracy of subsequent analysis; predicts the historical distribution network data matrix from two perspectives. One is to predict the distribution network data for the future time based on the changing trend of the distribution network data at multiple historical time points over time; the other is to find the periodicity of the abnormal data in the distribution network, compare it with the existing distribution network abnormal data with a known period, and use the most similar existing distribution network abnormal data as the distribution network data for the future time point; two analysis results are obtained from these two perspectives, and finally the two analysis results are comprehensively compared and analyzed, making the final analysis result more comprehensive and accurate.
[0093] 2. In the present invention, the locust optimization algorithm is used to perform multiple iterations to optimize each classification center data in the distribution network data matrix set, and the overall discrete degree of the distribution network data matrix set is used as its fitness function, so that as the iteration proceeds, the classification center data found becomes more accurate.
[0094] 3. In the present invention, the initial BP neural network prediction model is trained and tested by using the historical distribution network training data matrix and the historical distribution network test data matrix respectively, so that the final BP neural network prediction model obtained has good prediction ability for each type of distribution network time series data.
[0095] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] In order to more clearly illustrate the technical solutions of the embodiments of the invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the invention. For ordinary technicians in this field, they can also obtain drawings based on these drawings without paying any creative work.
[0097] Figure 1 The present invention is a flow chart of a distribution network abnormal data prediction system based on data analysis for predicting distribution network abnormal data. DETAILED DESCRIPTION
[0098] The following will clearly and completely describe the technical solutions in the embodiments of the invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0099] In the description of the present invention, it should be understood that the terms "opening", "upper", "lower", "top", "middle", "inside" and the like indicating orientation or positional relationship are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the components or elements referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the invention.
[0100] Example 1
[0101] This embodiment is a method for predicting abnormal distribution network data based on data analysis, comprising the following steps:
[0102] S1. Collect distribution network data of various types and multiple historical time points to obtain a distribution network data matrix set;
[0103] Said S1 comprises the following steps:
[0104] S11, set a variety of distribution network data types and corresponding data statistical periods to obtain a distribution network data type set a={a1, a2, ..., a i ,...,a a′} and the data statistical period set b′={b1′,b2′,...,b i ′,...,b a ″},a i Indicates the type of the set i-th network distribution data, a′ indicates the total number of set network distribution types; b i ' represents the data statistics period set for the i-th type of distribution network data;
[0105] Set the same number of historical time points for the data statistical period corresponding to each distribution network data type to obtain the historical time point matrix as follows,
[0106]
[0107] in, represents the jth historical time point set for the data statistical period corresponding to the i-th distribution network data type, and a~ represents the total number of historical time points set for the data statistical period corresponding to each distribution network data type;
[0108] According to the historical time point matrix And the data type set of the distribution network is collected a={a1,a2,...,a i ,...,a a′}, and obtain a distribution network data matrix set;
[0109] S12: Set the corresponding abnormal data threshold for each distribution network data type to obtain the abnormal data threshold set. represents the abnormal data threshold set for the i-th type of distribution network data;
[0110] S2. Perform a clustering operation on the distribution network data matrix set to obtain a distribution network data classification matrix; then classify the distribution network data classification matrix to obtain an initial non-fault abnormal time period matrix and an initial fault abnormal time period matrix; calculate the average values of the initial non-fault abnormal time period matrix and the initial fault abnormal time period matrix to obtain an average matrix of the non-fault abnormal time period set and an average matrix of the fault abnormal time period set;
[0111] The S2 comprises the following steps:
[0112] S21, performing a clustering operation on each distribution network data matrix in the distribution network data matrix set to obtain a distribution network data classification matrix c; as follows,
[0113]
[0114] Among them, c ij represents the j-th classification of the distribution network data matrix corresponding to the i-th distribution network data type; c ijkk′ Indicates c ij The distribution network data at the k′th historical time point in the kth data statistical period; Indicates c ij The total number of data statistical cycles corresponding to the distribution network data;
[0115] The S21 includes the following steps:
[0116] S211. Building a locust population represents the i-th locust in the locust population, Indicates the size of the locust population; Set the maximum number of iterations of the locust population to e1, the current number of iterations to e2, and the dimension of the search space to
[0117] S212, randomly selecting a number of distribution network data sets from the distribution network data matrix set as initial classification center data multiple times to obtain an initial classification center data matrix set; using the initial classification center data matrix set as the initial position matrix set of the locust population e′ k′ represents the initial position matrix of the i-th locust in the locust population; as follows,
[0118]
[0119] Among them, e′ k′ij represents e′ k′ The initial classification center data set of the jth classification corresponding to the i-th distribution network data type, e′ k′ijk represents e′ k′ij The initial classification center data at the k-th historical time point;
[0120] S213, calculate the distribution network data matrix set b2={b 21 ,b 22 ,...,b 2i ,...,b 2a′ The sum of the Euclidean distances between each distribution network dataset and the corresponding classification center dataset in} is used to obtain the Euclidean distance data matrix as follows,
[0121]
[0122] in, represents e′ k′ij Each distribution network data in the corresponding category is sent to e' k′ij The sum of the Euclidean distances between them;
[0123] According to the Euclidean distance data matrix Set the fitness function of the locust population as follows,
[0124]
[0125] S214, start iteration, before the iteration, set the current iteration number e2 to 1; in the first round of iteration, the fitness function of the locust population is used. Calculate the initial position matrix set The fitness value of the initial position matrix of each locust in the first fitness value set is obtained to obtain a first fitness value set; the maximum fitness value and the initial position of the corresponding locust are selected from the first fitness value set as the first global optimal fitness value and the first global optimal position respectively; the initial position matrix set is adjusted according to the first global optimal fitness value and the first global optimal position. The initial position matrix of each locust is updated, and after the update is completed, the current iteration number e2 is increased by 1, and the next round of iteration is entered;
[0126] Except for the first round of iteration, the fitness function of the locust population is used in each other round of iteration. Calculating the fitness value of the position matrix of each locust obtained in the previous iterative process to obtain a second fitness value set; selecting the maximum fitness value and the initial position of the corresponding locust from the second fitness value set as the second global optimal fitness and the second global optimal position, respectively; updating the position matrix of each locust obtained in the previous iterative process according to the second global optimal fitness and the second global optimal position; after the update is completed, increasing the current iteration number e2 by 1, and entering the next iteration;
[0127] S215, when e2≥e1, stop iteration and obtain the final global optimal position; use the final global optimal position as the final classification center data matrix; and calculate the distribution network data matrix set b2={b 21 ,b 22 ,...,b 2i ,...,b 2a′} to classify and obtain the distribution network data classification matrix c;
[0128] S22. Setting the threshold value set for the number of consecutive time points Indicates the threshold value of the number of consecutive time points corresponding to the i-th type of distribution network data type; in conjunction with the threshold value set of the number of consecutive time points and abnormal data threshold set When the time point data of the distribution network abnormal data in each cycle of each distribution network data matrix corresponding to each classification of each distribution network data type in the distribution network data classification matrix c is less than the corresponding continuous time point number threshold, the time period corresponding to the distribution network abnormal data is used as the non-fault abnormal time period; otherwise, the time period corresponding to the distribution network abnormal data is used as the fault abnormal time period; the initial non-fault abnormal time period matrix d1 and the initial fault abnormal time period matrix d2 are obtained; as follows,
[0129]
[0130] Among them, d 1ij and d 2ij They represent the non-fault abnormal time period matrix and fault abnormal time period matrix of the jth classification corresponding to the i-th distribution network data type respectively; they are as follows,
[0131]
[0132] in, Respectively represent d 1ij The starting time point and ending time point of the k′th non-fault abnormal time period in the kth data statistical period; Respectively represent d 2ij The starting and ending time points of the k′th fault-type abnormal time period in the kth data statistical period; d1′ ij d2′ ij Respectively represent d 1ij The total number of non-fault abnormal time periods in each data statistical period and d 2ij The total number of fault-type abnormal time periods in each data statistical period;
[0133] S23, respectively calculating the average value of the starting time point and the ending time point of each non-fault abnormal time period in each category and the average value of the starting time point and the ending time point of each fault abnormal time period in the initial non-fault abnormal time period matrix d1 and the initial fault abnormal time period matrix d2, to obtain the average matrix d3 of the non-fault abnormal time period set and the average matrix d4 of the fault abnormal time period set; respectively as follows,
[0134]
[0135] Among them, d 3ij d 4ij They represent the non-fault type abnormal average time period set and the fault type abnormal average time period set of the jth classification corresponding to the i-th distribution network data type; they are as follows,
[0136]
[0137] in, Respectively represent d 3ij The average value of the starting time point and the ending time point of the k′th non-fault type abnormal average time period in ; Respectively represent d 4ij The average value of the starting time point and the ending time point of the k′th fault type abnormal average time period in the calculation formula are as follows:
[0138]
[0139] S3. Collecting distribution network data of multiple types and multiple historical time points to be predicted again to obtain a historical distribution network data matrix; predicting and classifying the distribution network data at future moments based on the historical distribution network data matrix to obtain a first analysis result;
[0140] The S3 includes the following steps:
[0141] S31. Setting a set of time points to be predicted and future time points f 1i 、f 2i They represent the i-th time point to be predicted and the future time point respectively, and f1′ and f2′ represent the total number of time points to be predicted and the future time points respectively;
[0142] According to the distribution network data type set a={a1, a2, ..., a i ,...,a a′} and the set of time points to be predicted Collect distribution network data of various distribution network data types to obtain historical distribution network data matrix as follows,
[0143]
[0144] in, The network distribution data representing the i-th network distribution data type corresponding to the j-th time point to be predicted;
[0145] S32, constructing an initial BP neural network prediction model and setting the training data ratio and the test data ratio; and calculating the historical distribution network data matrix according to the training data ratio and the test data ratio. Perform data division to obtain the historical distribution network training data matrix, the historical distribution network test data matrix, and the historical distribution network to-be-predicted data matrix;
[0146] The historical distribution network training data matrix and the historical distribution network test data matrix are respectively used to train and test the initial BP neural network prediction model; after the training and testing are completed, the final BP neural network prediction model is obtained;
[0147] In S32, the historical distribution network training data matrix and the historical distribution network test data matrix are used to train and test the initial BP neural network prediction model respectively; after the training and testing are completed, the final BP neural network prediction model is obtained, which includes the following steps:
[0148] S321. Setting a training error threshold; inputting the historical distribution network training data matrix into the initial BP neural network prediction model for training; during the training process, when the training error is less than the training error threshold, stopping the training operation to obtain a trained BP neural network prediction model; otherwise, continuing the training until the training error is less than the training error threshold;
[0149] S322, setting a test accuracy threshold; inputting the historical distribution network test data matrix into the trained BP neural network prediction model for testing, and obtaining the test accuracy after the test is completed; when the test accuracy is greater than or equal to the test accuracy threshold, using the trained BP neural network prediction model as the final BP neural network prediction model; otherwise, returning to S321 to continue training the trained BP neural network prediction model until the test accuracy is greater than or equal to the test accuracy threshold;
[0150] S33, cooperate with the future time point set The final BP neural network prediction model is used to perform prediction operations on the historical distribution network to be predicted data matrix to obtain the future distribution network data matrix as follows,
[0151]
[0152] in, Represents the predicted distribution network data of the i-th distribution network data type corresponding to the j-th future time point;
[0153] S34, the future distribution network data matrix The distribution network data and abnormal data threshold set at each future time point in Compare the corresponding abnormal data thresholds in to obtain the future distribution network abnormal time period matrix;
[0154] S35, cooperate with the threshold value set of the number of consecutive time points Classify each future distribution network abnormal time period in the future distribution network abnormal time period matrix to obtain the first non-fault abnormal time period matrix to be analyzed and the first fault type abnormal time period matrix h to be analyzed; respectively as follows,
[0155]
[0156] in, Represent the future distribution network data matrix The starting time point and ending time point of the jth non-fault abnormal time period corresponding to the i-th distribution network data type; Represent the future distribution network data matrix The starting time point and ending time point of the jth fault type abnormal time period corresponding to the i-th distribution network data type; h i ′ represents the future distribution network data matrix The total number of non-fault abnormal time periods and fault abnormal time periods corresponding to the i-th distribution network data type;
[0157] S4. Classify the distribution network abnormality data in the historical distribution network data matrix to obtain a historical non-fault abnormality time period matrix and a historical fault abnormality time period matrix;
[0158] The S4 comprises the following steps:
[0159] S41, the historical distribution network data matrix The distribution network data and abnormal data threshold set for each time point to be predicted Compare the corresponding abnormal data thresholds in the , and obtain the historical distribution network abnormal time period matrix;
[0160] S42, cooperate with the threshold value set of the number of consecutive time points Classify each historical distribution network abnormal time period in the historical distribution network abnormal time period matrix to obtain the historical non-fault abnormal time period matrix g and the historical fault abnormal time period matrix g′; they are as follows:
[0161]
[0162] in, Represents the historical distribution network data matrix The starting time point and ending time point of the jth non-fault abnormal time period corresponding to the i-th distribution network data type; Represents the historical distribution network data matrix The starting time point and ending time point of the jth fault type abnormal time period corresponding to the i-th distribution network data type; g i ′ represents the historical distribution network data matrix The total number of non-fault abnormal time periods and fault abnormal time periods corresponding to the i-th distribution network data type;
[0163] S5. Perform a prediction analysis on the historical non-fault abnormal time period matrix and the historical fault abnormal time period matrix based on the average matrix of the non-fault abnormal time period set and the average matrix of the fault abnormal time period set to obtain a second analysis result;
[0164] The S5 comprises the following steps:
[0165] S51, calculating the Euclidean distance between the historical non-fault abnormal time period matrix g and each non-fault abnormal time period set in the corresponding row in the average matrix d3 of the non-fault abnormal time period set, and the Euclidean distance between the historical fault abnormal time period matrix g′ and each fault abnormal time period set in the corresponding row in the average matrix d4 of the fault abnormal time period set, to obtain a first Euclidean distance matrix and a second Euclidean distance matrix;
[0166] S52. Select the smallest Euclidean distance in each row of the first Euclidean distance matrix and the second Euclidean distance matrix, and use the corresponding non-fault abnormal time period set and fault abnormal time period set as the future time data of the historical non-fault abnormal time period set of the corresponding row in the historical non-fault abnormal time period matrix and the future time data of the historical fault abnormal time period set of the corresponding row in the historical fault abnormal time period matrix, respectively, to obtain a second non-fault abnormal time period matrix to be analyzed and a second fault abnormal time period matrix to be analyzed;
[0167] S6. Compare the first analysis result and the second analysis result to obtain a final non-fault abnormal time period matrix and a final fault abnormal time period matrix;
[0168] The S6 comprises the following steps:
[0169] S61, calculating the union of the fault-type abnormal time periods at corresponding positions in the first non-fault-type abnormal time period matrix to be analyzed and the second non-fault-type abnormal time period matrix to be analyzed, to obtain a final non-fault-type abnormal time period matrix;
[0170] S62: Calculate the union of the fault type abnormal time periods at corresponding positions in the first fault type abnormal time period matrix to be analyzed and the second fault type abnormal time period matrix to be analyzed to obtain a final fault type abnormal time period matrix.
[0171] Example 2
[0172] This embodiment discloses a distribution network abnormal data prediction system based on data analysis. The system can implement the method of the above embodiment, including a first distribution network data acquisition module, a clustering module, a first distribution network abnormal data classification module, an abnormal data time point mean acquisition module, a second distribution network data acquisition module, a first prediction analysis module, a second distribution network abnormal data classification module, a second prediction analysis module and a comparative analysis module;
[0173] The first distribution network data acquisition module is used to collect distribution network data of multiple types and multiple historical time points to obtain a distribution network data matrix set;
[0174] The clustering module is used to perform clustering operations on the distribution network data matrix set to obtain a distribution network data classification matrix;
[0175] The first distribution network abnormal data classification module is used to classify the distribution network data classification matrix to obtain an initial non-fault type abnormal time period matrix and an initial fault type abnormal time period matrix;
[0176] The abnormal data time point mean value acquisition module is used to calculate the average value of the initial non-fault abnormal time period matrix and the initial fault abnormal time period matrix to obtain the non-fault abnormal time period set average matrix and the fault abnormal time period set average matrix;
[0177] The second distribution network data collection module is used to collect distribution network data of multiple types and multiple historical time points to be predicted again to obtain a historical distribution network data matrix;
[0178] The first prediction and analysis module is used to predict and classify the distribution network data at a future time according to the historical distribution network data matrix to obtain a first analysis result;
[0179] The second distribution network abnormality data classification module is used to classify the distribution network abnormality data in the historical distribution network data matrix to obtain a historical non-fault abnormality time period matrix and a historical fault abnormality time period matrix;
[0180] The second prediction and analysis module is used to perform prediction analysis on the historical non-fault abnormal time period matrix and the historical fault abnormal time period matrix according to the average matrix of the non-fault abnormal time period set and the average matrix of the fault abnormal time period set to obtain a second analysis result;
[0181] The comparison and analysis module is used to compare the first analysis result and the second analysis result to obtain a final non-fault type abnormal time period matrix and a final fault type abnormal time period matrix.
[0182] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0183] The preferred embodiments of the invention disclosed above are intended only to help illustrate the invention. These preferred embodiments do not exhaust all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.
Claims
1. A distribution network abnormal data prediction method based on data analysis, characterized in that: The following steps are involved: S1. Collect distribution network data of various types and multiple historical time points to obtain a distribution network data matrix set; S2. Performing a clustering operation on the distribution network data matrix set to obtain a distribution network data classification matrix; then classifying the distribution network data classification matrix to obtain an initial non-fault type abnormal time period matrix and an initial fault type abnormal time period matrix; Calculating the average values of the initial non-fault abnormal time period matrix and the initial fault abnormal time period matrix to obtain a non-fault abnormal time period set average matrix and a fault abnormal time period set average matrix; S3. Collect distribution network data of multiple types and multiple historical time points to be predicted again to obtain a historical distribution network data matrix; Predicting and classifying the distribution network data at a future time according to the historical distribution network data matrix to obtain a first analysis result; S4. Classify the distribution network abnormality data in the historical distribution network data matrix to obtain a historical non-fault abnormality time period matrix and a historical fault abnormality time period matrix; S5. Perform a prediction analysis on the historical non-fault abnormal time period matrix and the historical fault abnormal time period matrix based on the average matrix of the non-fault abnormal time period set and the average matrix of the fault abnormal time period set to obtain a second analysis result; S6. Compare the first analysis result and the second analysis result to obtain a final non-fault abnormal time period matrix and a final fault abnormal time period matrix; The S5 comprises the following steps: S51, calculating the Euclidean distance between each non-fault abnormal time period set in the corresponding row of the historical non-fault abnormal time period matrix and the average matrix of the non-fault abnormal time period set, and the Euclidean distance between each fault abnormal time period set in the corresponding row of the historical fault abnormal time period matrix and the average matrix of the fault abnormal time period set, to obtain a first Euclidean distance matrix and a second Euclidean distance matrix; S52. Select the smallest Euclidean distance in each row of the first Euclidean distance matrix and the second Euclidean distance matrix, and use the corresponding non-fault type abnormal time period set and fault type abnormal time period set as the data of the historical non-fault type abnormal time period set of the corresponding row in the historical non-fault type abnormal time period matrix at a future time and the data of the historical fault type abnormal time period set of the corresponding row in the historical fault type abnormal time period matrix at a future time, respectively, to obtain a second non-fault type abnormal time period matrix to be analyzed and a second fault type abnormal time period matrix to be analyzed.
2. A distribution network abnormal data prediction method based on data analysis according to claim 1, characterized in that: Said S1 comprises the following steps: S11. Setting multiple distribution network data types and corresponding data statistical periods to obtain a distribution network data type set and a data statistical period set; setting the same number of historical time points for the data statistical period corresponding to each distribution network data type to obtain a historical time point matrix; and obtaining a distribution network data matrix set based on the historical time point matrix and the distribution network data of multiple data statistical periods with abnormal distribution network data corresponding to each distribution network data type in the distribution network data type set; S12. Setting a corresponding abnormal data threshold for each type of distribution network data to obtain an abnormal data threshold set.
3. A distribution network abnormal data prediction method based on data analysis according to claim 2, characterized in that: The S2 comprises the following steps: S21. Perform a clustering operation on each distribution network data matrix in the distribution network data matrix set to obtain a distribution network data classification matrix; S22. Set a threshold set for the number of consecutive time points; in conjunction with the threshold set for the number of consecutive time points and the threshold set for abnormal data, when the number of time point data of the distribution network abnormal data in each cycle of each distribution network data matrix corresponding to each classification of each distribution network data type in the distribution network data classification matrix is less than the corresponding threshold set for the number of consecutive time points, then the time period corresponding to the distribution network abnormal data is regarded as a non-fault abnormal time period; otherwise, the time period corresponding to the distribution network abnormal data is regarded as a fault abnormal time period; and obtain an initial non-fault abnormal time period matrix and an initial fault abnormal time period matrix; S23. Calculate the average value of the start time point and the end time point of each non-fault type abnormal time period in each category in the initial non-fault type abnormal time period matrix and the initial fault type abnormal time period matrix, as well as the average value of the start time point and the end time point of each fault type abnormal time period, to obtain the average matrix of the non-fault type abnormal time period set and the average matrix of the fault type abnormal time period set.
4. A distribution network abnormal data prediction method based on data analysis according to claim 3, characterized in that: The S21 includes the following steps: S211, constructing a locust population; setting the maximum number of iterations of the locust population to , the current number of iterations is ; S212, randomly selecting a plurality of distribution network data sets from the distribution network data matrix set as initial classification center data multiple times to obtain an initial classification center data matrix set; and using the initial classification center data matrix set as an initial position matrix set of the locust population; S213, calculating the sum of the Euclidean distances between each distribution network data set in the distribution network data matrix set and the corresponding classification center data set based on the initial position matrix set to obtain a Euclidean distance data matrix; and setting a fitness function of the locust population based on the Euclidean distance data matrix; S214, starting iteration; in each round of iteration, using the fitness function of the locust population to calculate the fitness value of the position matrix of each locust obtained in the previous round of iteration and updating the position matrix of each locust obtained in the previous round of iteration; S215, when When , the iteration is stopped to obtain the final global optimal position; the final global optimal position is used as the final classification center data matrix; the distribution network data matrix set is classified according to the final classification center data matrix to obtain the distribution network data classification matrix.
5. A distribution network abnormal data prediction method based on data analysis according to claim 4, characterized in that: The S3 includes the following steps: S31, setting a set of time points to be predicted and a set of future time points; collecting distribution network data of multiple distribution network data types according to the distribution network data type set and the set of time points to be predicted, to obtain a historical distribution network data matrix; S32, constructing an initial BP neural network prediction model; performing data partitioning on the historical distribution network data matrix to obtain a historical distribution network training data matrix, a historical distribution network test data matrix, and a historical distribution network to-be-predicted data matrix; respectively using the historical distribution network training data matrix and the historical distribution network test data matrix to train and test the initial BP neural network prediction model; after the training and testing are completed, obtaining the final BP neural network prediction model; S33, using the final BP neural network prediction model to perform a prediction operation on the historical distribution network data matrix to be predicted in conjunction with the future time point set to obtain a future distribution network data matrix; S34, comparing the distribution network data at each future time point in the future distribution network data matrix with the abnormal data threshold corresponding to the abnormal data threshold set to obtain a future distribution network abnormal time period matrix; S35. Classify each future distribution network abnormal time period in the future distribution network abnormal time period matrix according to the continuous time point number threshold set to obtain a first non-fault abnormal time period matrix to be analyzed and a first fault abnormal time period matrix to be analyzed.
6. A distribution network abnormal data prediction method based on data analysis according to claim 5, characterized in that: The S4 comprises the following steps: S41, comparing the distribution network data of each time point to be predicted in the historical distribution network data matrix with the abnormal data threshold corresponding to the abnormal data threshold set to obtain a historical distribution network abnormal time period matrix; S42. Classify each historical distribution network abnormal time period in the historical distribution network abnormal time period matrix in accordance with the continuous time point number threshold set to obtain a historical non-fault abnormal time period matrix and a historical fault abnormal time period matrix.
7. A distribution network abnormal data prediction method based on data analysis according to claim 6, characterized in that: The S6 comprises the following steps: S61, calculating the union of the fault-type abnormal time periods at corresponding positions in the first non-fault-type abnormal time period matrix to be analyzed and the second non-fault-type abnormal time period matrix to be analyzed, to obtain a final non-fault-type abnormal time period matrix; S62: Calculate the union of the fault type abnormal time periods at corresponding positions in the first fault type abnormal time period matrix to be analyzed and the second fault type abnormal time period matrix to be analyzed to obtain a final fault type abnormal time period matrix.
8. A system for implementing the distribution network abnormal data prediction method based on data analysis as described in any one of claims 1 to 7.
Citation Information
Patent Citations
A distribution network abnormal data error correction method and system based on fusion
CN117472898B
Distribution network voltage abnormal data detection method and device
CN112819373A
Identification and correction method for power distribution network line loss abnormal data
CN115834424A