A method, system and medium for cleaning operation data of power grid equipment
Through backpropagation neural network and improved DBSCAN algorithm, the operation data of power grid equipment is processed, and noise, missing values and duplicate data are identified and processed, and the problem of inefficient data cleaning in the prior art is solved, achieving efficient and accurate data cleaning effect.
Patent Information
- Application Number
- CN202211023992.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-08-24
AI Technical Summary
现有技术难以有效处理电网设备运行数据中的错误数据和重复数据,导致数据清洗效率低下,影响数据分析的准确性和可靠性。
The data prediction model is constructed using backpropagation feedforward neural network, combined with the improved DBSCAN algorithm, it identifies and processes noise data, missing values and duplicate data, and cleanses data through clustering algorithms.
It improves the efficiency and accuracy of data cleaning, reduces the consumption of resources, and ensures the reliability and accuracy of data analysis.
Smart Images

Figure CN115423008B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power grid equipment data processing, and specifically relates to a method, system and medium for cleaning operation data of power grid equipment. Background Art
[0002] With the diversification of energy types and the continuous optimization of the energy structure in China, various types of units from different systems will generate a large amount of data streams during operation, and the information that can be mined based on these data is very rich. Since it includes error data generated by various factors such as sensor failures and external interferences, as well as data conflicts caused by different business systems, these data not only cannot be utilized, but even hinder information mining. Therefore, this type of data is called "dirty" data. In view of this, it is of great significance to clean the "dirty" data, improve the quality of the "dirty" data, and ensure the accuracy of subsequent data analysis results.
[0003] Considering the complexity and unpredictability of random factors in the actual operation process, the existing research on the cleaning of operation data of unit equipment can be summarized into two categories: the cleaning of attribute value error data and the cleaning of duplicate data. For data with abnormal attribute values, the first thing to consider is the detection of outliers. A common, simple and easy-to-implement method is the statistical method. Based on Chebyshev's theorem, outliers are distinguished according to the number of standard deviations n from the average of the data set; however, traditional methods are difficult to apply to the complex distribution of outliers with a high data dimension. For the problem of missing values in operation data, if filling methods such as median, mean, and mode are used, it will often affect the reliability of the data, and even mislead the analysis results and subsequent decisions. At the same time, there may be duplicate records in multi-source data from different systems. If the most direct method of pairwise comparison in the database is used, the steps are too cumbersome and resource-consuming, and the time complexity is also high, which is not conducive to the development of data cleaning work. Summary of the Invention
[0004] The main purpose of the present invention is to overcome the deficiencies of the prior art, and provide a method, system and medium for cleaning operation data of power grid equipment. The present invention comprehensively considers two types of error data and duplicate data, identifies noise data and sets null values to obtain missing values, then constructs a data prediction model for prediction filling, and finally realizes the cleaning of operation data of power grid equipment through a clustering algorithm. The method is simple, has a low complexity, and has a good cleaning effect.
[0005] To achieve the above object, on the one hand, the present invention provides a method for cleaning operation data of power grid equipment, which is characterized by including the following steps:
[0006] Obtain the original operation data of the power grid equipment from the database and perform preprocessing;
[0007] Identify the noisy data with abnormal attributes in the preprocessed operation data and set them to null values to obtain the operation data containing missing values;
[0008] Construct and train a data prediction model based on the backpropagation feedforward neural network, predict the operation data containing missing values, obtain the data values to be filled and perform data filling to obtain the filled operation data;
[0009] For the filled operation data, use the improved DBSCAN algorithm for clustering, select and delete the duplicate data records with high matching degree to obtain the cleaned operation data.
[0010] As a preferred technical solution, the step of identifying the noisy data with abnormal attributes in the preprocessed operation data is as follows:
[0011] Randomly select a data point p from the preprocessed operation data and calculate the neighborhood N ε of the data point p. The neighborhood N ε has a radius of ε, and the calculation formula is:
[0012] N ε (p) = {q ∈ D|dist(p, q) ≤ ε}
[0013] Among them, N ε (p) represents the neighborhood of the data point p, dist(p, q) represents the distance from other data points q to the data point p in the operation data, and D represents the number of data points;
[0014] When dist(p, q) ≤ ε, the number of data points in the neighborhood of the data point p is incremented by 1, and the loop calculation is performed until the distance values of all data points to the data point p are found;
[0015] Suppose that the neighborhood N ε of the data point p contains at least Minpts data points. If q ∈ N ε (p) and |N ε (p)| ≥ Minpts, then the data point p is marked as a core point, otherwise the data point p is marked as noisy data and the data point p is set to a null value as a missing value;
[0016] Repeat the above steps until all data points in the preprocessed operation data are marked, and obtain the operation data containing missing values.
[0017] As a preferred technical solution, the backpropagation feedforward neural network, i.e., the BP neural network, includes an input layer, a hidden layer, and an output layer, which are sequentially connected in order; the input layer, the hidden layer, and the output layer all contain a number of neuron nodes, and the neurons are connected by weights; the transformation function of the neurons adopts the Sigmoid function; the input layer inputs the attribute values of the operation data containing missing values; the output layer outputs the predicted data values to be filled;
[0018] The weights and thresholds of the BP neural network are optimized and trained using a genetic algorithm to obtain a data prediction model.
[0019] As a preferred technical solution, the steps of optimizing and training using a genetic algorithm to obtain a data prediction model are as follows:
[0020] Determine the network topology of the BP neural network, and divide the pre-prepared data sample set into training samples and test samples;
[0021] Encode the weights and thresholds of the BP neural network to obtain an initial population;
[0022] Use the training samples to train the BP neural network, use the test samples to test the BP neural network, and calculate the error between the output value and the expected value;
[0023] Calculate the fitness of the chromosomes in the BP neural network, select the chromosomes with high fitness for replication, and perform crossover and mutation operations to generate a new population;
[0024] Judge whether the BP neural network reaches the performance index or the maximum number of iterations. If so, decode the chromosomes obtained by encoding the weights and thresholds of the BP neural network to obtain the weights and thresholds of the best neural network as the initial weights and thresholds of the data prediction model;
[0025] Otherwise, obtain a new initial population and continue to execute.
[0026] As a preferred technical solution, the network topology of the BP neural network is represented as: m-l-n, where m, l, and n respectively represent the number of neuron nodes in the input layer, the hidden layer, and the output layer;
[0027] The pre-prepared data sample set consists of N power grid data samples and is randomly divided into training samples and test samples which are respectively represented as:
[0028]
[0029]
[0030] Among them, (x(), y()) represents a sample in the sample space, x() represents the sample value associated with y(), y() represents the actual sample value to be predicted by the BP neural network, t is the time parameter, R m and R n are real numbers, N1 represents the number in the training samples, and N represents the number of samples in the data sample set;
[0031] The training samples are used to establish the input-output mapping relationship of the BP neural network; the test samples are used to verify the correctness of the input-output mapping relationship;
[0032] The input-output mapping relationship of the BP neural network is expressed as:
[0033]
[0034] Among them, represents the output of the BP neural network, l is the number of neurons in the hidden layer, v jk represents the weight from the j-th neuron in the hidden layer to the k-th neuron in the output layer, k = 1, 2,..., n, w ij is the weight from the i-th neuron in the input layer to the j-th neuron, x i (t) represents the input of the BP neural network, θ j is the threshold at the j-th neuron in the hidden layer, r k is the threshold at the k-th neuron in the output layer, f[] is the activation function, expressed as:
[0035]
[0036] If the total error E1 of the BP neural network is less than or equal to the total error target threshold e1, then there is:
[0037]
[0038] Among them, y k (t) represents the actual sample value;
[0039] If the average detection error E2 of the BP neural network is less than or equal to the average detection sample error target threshold e2, then there is:
[0040]
[0041] Among them, n represents the number of neurons in the output layer.
[0042] As a preferred technical solution, encoding the weights and thresholds of the BP neural network to obtain the initial population is specifically:
[0043] The real - number coding method is used to form a chromosome with the network weights of each node in the BP neural network, arranged in sequence into a string, and the connection weights between neurons are optimized to initialize the population P(t); the transfer function of the hidden layer in the BP neural network adopts the sigmoid cross - entropy loss function, introducing the non - linear mapping ability.
[0044] As a preferred technical solution, calculating the fitness of the chromosome in the BP neural network is specifically as follows:
[0045] Take out a chromosome i from the initialized population P(t), input the network weights of each node in it into the BP neural network in sequence, calculate the total error E of the BP neural network, and define the fitness f of this chromosome i i , expressed as:
[0046]
[0047] Select chromosomes with high fitness for replication, set the crossover probability P c and the mutation probability P m and the selection probability p of chromosome i in the initialized population i ; The selection probability is defined as:
[0048] p i = α*(1 - α)
[0049] where α is a random number in [0, 1], i = 1, 2, …, t;
[0050] Perform crossover and mutation operations. Among them, the crossover operation is specifically as follows:
[0051] Sort each chromosome according to the fitness from high to low, and calculate the cumulative probability q of each chromosome i i :
[0052]
[0053] Use the roulette wheel algorithm to generate a random number r ∈ [0, 1] in each round. If q i-1 < r ≤ q i , then take chromosome i as the parent chromosome, perform the roulette wheel algorithm t times to obtain t parent chromosomes, and form chromosome pairs with adjacent parent chromosomes; for each chromosome pair, determine the crossover position k according to the crossover probability P c and swap the genes between 1 and k numbers of the two parent chromosomes in the chromosome pair, thus obtaining two new chromosomes;
[0054] The mutation operation is specifically as follows:
[0055] According to the mutation probability P mDetermine u mutation positions on two new chromosomes respectively, and perform mutation operations on the genes at the mutation positions, that is, add a random decimal number uniformly distributed in the range of [-1, 1] to the genes on the two new chromosomes to obtain two new sub-chromosomes;
[0056] Insert the new individuals into the population P(t) to generate a new population P(t + 1).
[0057] As a preferred technical solution, the improved DBSCAN algorithm is used to select and delete highly matching duplicate data records. Specifically:
[0058] Establish a word-document matrix in the form of an inverted index, quickly obtain the list of documents containing a certain word according to the word, and divide the filled operation data into several subsets according to the same type of equipment;
[0059] Use the improved DBSCAN algorithm to cluster several subsets so that the duplicate data records form a cluster, calculate the similarity of the records in the cluster and judge whether they are duplicate data records. Specifically:
[0060] Determine the parameters of the improved DBSCAN algorithm, and the parameters include the radius ε′, the minimum number of points Minpts′ within the radius, and the initial similarity threshold R;
[0061] Randomly select a data point A from a certain subset as the core point, and use the similarity distance function approxDist() function to calculate the similarity distance between the remaining data points and the data point A;
[0062] Cluster the filled operation data in the subset through the set distance thresholds N1 and N2:
[0063] Calculate the similarity of any two filled operation data in combination with the attribute weights. The calculation formula is:
[0064]
[0065] Among them, n is the total number of attributes of the filled operation data, S Ai (x, y) represents the attribute similarity between the filled operation data x and the filled operation data y, which is expressed as:
[0066]
[0067] Among them, d represents the distance between the points where x and y are mapped to the two-dimensional space;
[0068] Judge whether to update and iterate the radius ε′ according to the initial similarity threshold R. The iteration formula is:
[0069]
[0070] When the iteration reaches the acceptable range where the output result meets the similarity threshold, the clustering is completed to obtain the set of duplicate data records;
[0071] Based on the set of duplicate data records, calculate the similarity of each data record, retain the data record with the maximum similarity, and delete the rest to obtain the cleaned operation data.
[0072] On the other hand, the present invention provides a cleaning system for operation data of power grid equipment, which is applied to the above-mentioned cleaning method for operation data of power grid equipment, and includes a data acquisition module, a noise identification module, a data filling module, and a data deletion module;
[0073] The data acquisition module is used to acquire the original operation data of the power grid equipment from the database and perform preprocessing;
[0074] The noise identification module is used to identify the noise data with abnormal attributes in the preprocessed operation data and set null values to obtain the operation data containing missing values;
[0075] The data filling module constructs and trains a data prediction model based on a backpropagation feedforward neural network, predicts the operation data containing missing values, obtains the data values to be filled, and performs data filling to obtain the filled operation data;
[0076] The data deletion module uses the improved DBSCAN algorithm for the filled operation data, selects the highly matching duplicate data records and deletes them to obtain the cleaned operation data.
[0077] On the other hand, the present invention also provides a computer-readable storage medium storing a program, which when executed by a processor, implements the above-mentioned cleaning method for operation data of power grid equipment.
[0078] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0079] 1. In the present invention, the identification of data outliers and the cleaning of duplicate data are both centered around the density-based clustering algorithm. Its advantages are that it does not require inputting the number of clusters, can cluster dense data sets of any shape; can discover outliers while clustering and is not sensitive to outliers in the data set; the clustering results have no bias. In contrast, clustering algorithms such as K-Means are greatly affected by the initial values on the clustering results.
[0080] 2. The method for predicting and filling missing values in the present invention is to construct a data prediction model based on a genetic neural network, which fully utilizes the global search ability of the genetic algorithm and the non-linear mapping ability of the neural network, greatly improving the prediction accuracy of the data, and the data prediction accuracy is controllable.
[0081] 3. In view of the problem of low accuracy in detecting duplicate records in the density clustering algorithm, the improved DBSCAN clustering algorithm proposed by the present invention improves the accuracy of detecting duplicate data records to a certain extent, ensuring the effectiveness of data cleaning. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0083] Figure 1 It is a schematic flowchart of a method for cleaning operation data of power grid equipment according to an embodiment of the present invention;
[0084] Figure 2 It is a schematic flowchart of training a data prediction model according to an embodiment of the present invention;
[0085] Figure 3 It is a flowchart of an improved DBSCAN algorithm according to an embodiment of the present invention;
[0086] Figure 4 It is a schematic structural diagram of a system for cleaning operation data of power grid equipment according to an embodiment of the present invention;
[0087] Figure 5 It is a schematic structural diagram of a computer-readable storage medium according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0088] In order to enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0089] References to "embodiments" in this application mean that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described in this application can be combined with other embodiments.
[0090] This embodiment is based on the data cleaning problem in the BPA database of China Southern Power Grid NF2022DX. Combining Figures 1 to 3 , a method for cleaning operation data of power grid equipment provided by the present invention is elaborated in detail, including the following steps:
[0091] S1. Obtain the original operation data of power grid equipment from the database and perform preprocessing;
[0092] S2. Identify the noisy data with abnormal attributes in the preprocessed operation data and set them to null values to obtain the operation data containing missing values;
[0093] S3. Construct and train a data prediction model based on a backpropagation feedforward neural network, predict the operation data containing missing values, obtain the data values to be filled and perform data filling to obtain the filled operation data;
[0094] S4. For the filled operation data, use the improved DBSCAN algorithm for clustering, select and delete the duplicate data records with high matching degree to obtain the cleaned operation data.
[0095] More specifically, in step S1, since the original operation data comes from different power grid equipment, it is necessary to perform preprocessing on the data to facilitate the effective cleaning of the operation data; in this embodiment, the data attributes are initially selected according to experience to perform preprocessing on the data.
[0096] More specifically, in step S2, the present invention adopts the guiding ideology of the density clustering algorithm to cluster the preprocessed operation data and identify the noisy data with abnormal attributes. The steps are as follows:
[0097] S21. Randomly select a data point p from the preprocessed operation data and calculate the neighborhood N ε of the data point p. For the data point p, its neighborhood N ε has a radius of ε, and the calculation formula for the neighborhood N ε is:
[0098] N ε (p) = {q ∈ D|dist(p, q) ≤ ε}
[0099] where N ε(p) represents the neighborhood of data point p, dist(p,q) represents the distance from another data point q in the running data to data point p, and represents the number of data points;
[0100] When dist(p,q) ≤ ε, the number of data points in the neighborhood of data point p is incremented by 1, and the calculation is looped until the distance values from all data points to data point p are found;
[0101] Define the core object. Let the neighborhood N of data point p ε contain at least Minpts data points. In the case of determining the neighborhood N ε and Minpts, if q ∈ N ε (p) and |N ε (p)| ≥ Minpts, then mark data point p as a core point; otherwise, mark data point p as noise data and set data point p to a null value as a missing value to facilitate the subsequent prediction and filling steps for missing values;
[0102] Repeat the above steps until all data points in the preprocessed running data are marked, and the running data containing missing values is obtained.
[0103] More specifically, in step S3, the backpropagation feedforward neural network, i.e., the BP neural network, includes an input layer, a hidden layer, and an output layer. Each layer contains several neuron nodes (depending on the specific data attribute dimensions). The neurons are connected by weights, and the transformation function of the neurons uses the Sigmoid function; the input layer is the attribute value of the running data containing missing values, and the output layer is the predicted data value to be filled;
[0104] As Figure 2 shown, the weights and thresholds of the BP neural network are optimized and trained using a genetic algorithm to obtain a data prediction model and improve the prediction performance of the model. Specifically:
[0105] S31. First, determine the network topology of the BP neural network, and divide the pre-prepared data sample set into training samples and test samples;
[0106] Among them, the network topology of the BP neural network is expressed as: m - l - n, where m, l, and n represent the number of neuron nodes in the input layer, hidden layer, and output layer respectively; the pre-prepared data sample set consists of N power grid data samples and is randomly divided into training samples and test samples which are respectively expressed as:
[0107]
[0108]
[0109] Among them, (x(), y()) represents a certain sample in the sample space, x() represents the sample value associated with y(), y() represents the actual sample value to be predicted by the BP neural network, t is the time parameter, R m and R n are real numbers, N1 represents the number in the training samples, and N represents the number of samples in the data sample set;
[0110] In this embodiment, the data sample set prepared in advance comes from the grid data of the 100KV bus AB in Conghua. Taking the voltage of line AB as the prediction target of the BP neural network, then x() represents the active power associated with the voltage, and y() is the actual voltage value to be predicted by the BP neural network; after training, the BP neural network obtains the output prediction value by inputting x(). When it is almost equal to the actual voltage value y(), that is, the prediction performance of the BP neural network meets the standard; thereafter, the BP neural network can be used to fill in the missing data of the voltage of line AB, and the corresponding voltage of line AB can be predicted by inputting the recorded active power.
[0111] In the present invention, the training samples are used to establish the input-output mapping relationship of the BP neural network; the test samples are used to verify the correctness of the input-output mapping relationship; the input-output mapping relationship of the BP neural network is expressed as:
[0112]
[0113] Among them, represents the output of the BP neural network, l is the number of neurons in the hidden layer, v jk represents the weight from the j-th neuron in the hidden layer to the k-th neuron in the output layer, k = 1, 2,..., n, w ij is the weight from the i-th neuron in the input layer to the j-th neuron, x i (t) represents the input of the BP neural network, θ j is the threshold at the j-th neuron in the hidden layer, r k is the threshold at the k-th neuron in the output layer, and f[] is the activation function, expressed as:
[0114]
[0115] Suppose the total error E1 of the BP neural network is less than or equal to the network total error target threshold e1, then there is:
[0116]
[0117] Among them, y k (t) represents the actual sample value;
[0118] If the average detection error E2 of the BP neural network is less than or equal to the target threshold e2 of the average detection sample error, then:
[0119]
[0120] Among them, n represents the number of neuron nodes in the output layer.
[0121] S32. Encode the weights and thresholds of the BP neural network to obtain the initial population;
[0122] Adopt the real number encoding method to avoid the change of weight supplementation. Form a chromosome from the network weights of each node in the BP neural network, arrange them in sequence, and optimize the connection weights between neurons to initialize the population P(t); The transfer function of the hidden layer in the BP neural network adopts the sigmoid cross-entropy loss function to introduce the ability of nonlinear mapping; In this implementation, the population is initialized with random decimals uniformly distributed on [-3, 3] to reduce the problem that the algorithm converges too slowly due to too small weight adjustment.
[0123] S33. Use the training samples to train the BP neural network, use the test samples to test the BP neural network and calculate the error between the output value and the expected value;
[0124] S34. Calculate the fitness of the chromosomes in the BP neural network, select the chromosomes with high fitness for replication, and perform crossover and mutation operations to generate a new population;
[0125] Take out a chromosome i from the initialized population P(t), input the network weights of each node in it into the BP neural network in sequence, calculate the total error E of the BP neural network, and define the fitness f of the chromosome i i , which is expressed as:
[0126]
[0127] Select the chromosomes with high fitness for replication, set the crossover probability P c , mutation probability P m and the selection probability p of chromosome i in the initialized population i , which is defined as:
[0128] p i =α*(1 - α)
[0129] where α is a random number in [0, 1], i = 1, 2,..., t;
[0130] Perform crossover and mutation operations. Among them, the crossover operation is specifically:
[0131] Sort each chromosome from high to low according to fitness, and calculate the cumulative probability q of each chromosome ii :
[0132]
[0133] Use the roulette wheel algorithm to generate a random number r ∈ [0, 1] in each round. If q i-1 <r ≤ q i , then take chromosome i as the parent chromosome, perform the roulette wheel algorithm t times to obtain t parent chromosomes, and form chromosome pairs with adjacent parent chromosomes; for each chromosome pair, determine the crossover position k according to the crossover probability P c , and swap the genes between the two parent chromosomes in the chromosome pair numbered between [1, k], so as to obtain two new chromosomes;
[0134] According to the mutation probability P m , determine u mutation positions on the two new chromosomes respectively, and perform mutation operations on the genes at the mutation positions, that is, add a random decimal uniformly distributed between [-1, 1] to the genes on the two new chromosomes to obtain two new sub-chromosomes;
[0135] Insert the new individuals into the population P(t) to generate a new population P(t + 1).
[0136] S35. Determine whether the BP neural network reaches the performance index or the maximum number of iterations. If so, decode the chromosome obtained by encoding the weights and thresholds of the BP neural network to obtain the weights and thresholds of the best neural network as the initial weights and thresholds of the data prediction model; if not, return to step S32 to obtain the initial population again and continue to execute.
[0137] More specifically, in step S4, as Figure 3 shown, use the improved DBSCAN algorithm to iteratively classify approximately duplicate data records into the same class for multiple times. The steps are as follows:
[0138] In the first stage, establish a word-document matrix in the way of inverted index, quickly obtain the document list containing a certain word according to the word, and divide the filled operation data into several subsets according to the same type of equipment. The improved DBSCAN algorithm in the second stage can perform clustering and screening of duplicate records for each subset of data, reducing the computing power resources and time complexity in the second stage;
[0139] In the second stage, use the improved DBSCAN algorithm to cluster several subsets so that duplicate data records form a cluster, calculate the similarity of the records in the cluster and judge whether they are duplicate data records. Specifically:
[0140] S41. Determine the parameters of the improved DBSCAN algorithm, including the radius ε′, the minimum number of points Minpts′ within the radius, and the initial similarity threshold R;
[0141] S42. Randomly select a data point A from a certain subset as the core point, and use the similarity distance function approxDist() to calculate the similarity distances between the remaining data points and the data point A;
[0142] S43. Cluster the filled operation data in the subset through the set distance thresholds N1 and N2:
[0143] Calculate the similarity between any two filled operation data in combination with the attribute weights. The calculation formula is:
[0144]
[0145] where n is the total number of attributes of the filled operation data, S Ai (x, y) represents the attribute similarity between the filled operation data x and the filled operation data y, expressed as:
[0146]
[0147] where d represents the distance between the points obtained by mapping x and y to the two-dimensional space respectively;
[0148] Judge whether to update and iterate the radius ε′ according to the similarity initial threshold R. The iteration formula is:
[0149]
[0150] When the iteration reaches the acceptable range where the output result meets the similarity threshold, complete the clustering to obtain the set of duplicate data records;
[0151] S44. After the clustering is completed, obtain the set of duplicate data records, calculate the similarity of each data record, retain the data record with the maximum similarity, and delete the rest to obtain the cleaned operation data.
[0152] To verify the cleaning performance of the present invention for power grid operation data, in the embodiments of the present invention, based on the power grid operation data of Conghua 110KV bus AB, Conghua 110KV bus BC, and Conghua 110KV bus CA, the mean method (K-Means) and the BP neural network proposed by the present invention are used to predict the missing values, and the data shown in Table 1 below is obtained:
[0153] Conghua 110KV Busbar AB Conghua 110KV Busbar BC Conghua 110KV Busbar CA Mean method 112.560 112.950 114.220 BP neural network 112.739 113.533 113.148 True value 112.741 113.531 113.150
[0154] Table 1
[0155] It can be seen that the missing values predicted by the BP neural network proposed by the present invention are closer to the true values and are more accurate than the prediction results of the mean method.
[0156] Meanwhile, in order to verify the recognition accuracy of the improved DBSCAN algorithm of the present invention for duplicate records, the present invention randomly selected 3 groups of operation data with different record numbers (500, 5000, 50000) from the BPA database in the South China Power Grid NF2022DX. After predicting missing values through a BP neural network, the basic record matching algorithm and the improved DBSCAN algorithm in the present invention were respectively used for recognition, and the recognition results are shown in Table 2 below:
[0157] Number of data records Basic record matching algorithm Improved DBSCAN algorithm 500 90.33% 90.78% 5000 80.54% 86.66% 50000 69.47% 79.68%
[0158] Table 2
[0159] As can be seen from Table 2, as the number of data records increases, the accuracy rate of the improved DBSCAN algorithm has always been higher than that of the basic record matching algorithm, and the rate of decrease in the accuracy rate is lower than that of the basic record matching algorithm. Therefore, the improved DBSCAN algorithm of the present invention has a high accuracy rate and better performance in duplicate data screening.
[0160] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.
[0161] Based on the same idea as a method for cleaning operation data of power grid equipment in the above embodiment, the present invention also provides a system for cleaning operation data of power grid equipment, and this system can be used to execute the above method for cleaning operation data of power grid equipment. For the sake of convenience of description, in the structural schematic diagram of an embodiment of a system for cleaning operation data of power grid equipment, only the part related to the embodiment of the present invention is shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and it may include more or fewer components than those illustrated, or combine certain components, or have different component arrangements.
[0162] Please refer to Figure 4 , in another embodiment of the present application, a system for cleaning operation data of power grid equipment is provided. This system includes a data acquisition module, a noise point recognition module, a data filling module, and a data deletion module;
[0163] Among them, the data acquisition module is used to acquire the original operation data of power grid equipment from the database and perform preprocessing;
[0164] The noise point recognition module is used to identify noise point data with abnormal attributes in the preprocessed operation data and set null values to obtain operation data containing missing values;
[0165] The data filling module constructs and trains a data prediction model based on a backpropagation feedforward neural network, predicts the operation data containing missing values, obtains the data values to be filled and fills the data to obtain the filled operation data;
[0166] The data deletion module uses the improved DBSCAN algorithm for the filled operation data, selects and deletes the highly matching duplicate data records, and obtains the cleaned operation data.
[0167] It should be noted that a cleaning system for grid equipment operation data of the present invention corresponds one-to-one with a cleaning method for grid equipment operation data of the present invention. The technical features and their beneficial effects described in the embodiments of the above-mentioned cleaning method for grid equipment operation data are all applicable to the embodiments of a cleaning system for grid equipment operation data. For specific content, reference can be made to the description in the method embodiments of the present invention, which will not be repeated here. This is hereby declared.
[0168] In addition, in the implementation manner of the cleaning system for grid equipment operation data in the above embodiments, the logical division of each program module is only an example. In actual applications, according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the cleaning system for grid equipment operation data is divided into different program modules to complete all or part of the functions described above.
[0169] Please refer to Figure 5 , the embodiments of the present invention also provide a computer-readable storage medium, storing a program in the memory. When the program is executed by a processor, a cleaning method for grid equipment operation data is implemented, specifically:
[0170] Obtain the original operation data of grid equipment from the database and perform preprocessing;
[0171] Identify the noisy data with abnormal attributes in the preprocessed operation data and set them to null values to obtain the operation data containing missing values;
[0172] Construct and train a data prediction model based on a backpropagation feedforward neural network, predict the operation data containing missing values, obtain the data values to be filled and fill the data to obtain the filled operation data;
[0173] For the filled operation data, use the improved DBSCAN algorithm for clustering, select and delete the highly matching duplicate data records, and obtain the cleaned operation data.
[0174] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0175] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered that the scope described in this specification is covered.
[0176] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention should be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for cleaning operation data of power grid equipment, characterized in that, It includes the following steps: Obtain the original operation data of power grid equipment from the database and perform preprocessing; Identify the noisy data with abnormal attributes in the preprocessed operation data and set them to null values to obtain the operation data containing missing values; The step of identifying the noisy data with abnormal attributes in the preprocessed operation data is as follows: Randomly select a data point p from the preprocessed running data, and calculate the neighborhood N of the data point p ε , the neighborhood N ε has a radius of ε, and the calculation formula is: N ε (p) = {q ∈ D | dist(p, q) ≤ ε} Among them, N ε (p) represents the neighborhood of data point p, dist(p, q) represents the distance from other data point q in the running data to data point p, and D represents the number of data points; When dist(p,q) ≤ ε, the number of data points in the neighborhood of data point p is incremented by 1, and the loop calculation is performed until the distance values from all data points to data point p are found; Let the neighborhood N of the data point p ε contain at least Minpts data points. If q ∈ n ε (p) and |n ε (p)| ≥ Minpts, then mark the data point p as a core point; otherwise, mark the data point p as noise data and set the data point p to a null value as a missing value. Repeat the above steps until all data points in the preprocessed operation data are marked, and the operation data containing missing values is obtained; Construct and train a data prediction model based on a backpropagation feedforward neural network, predict the operation data containing missing values, obtain the data values to be filled and perform data filling to obtain the filled operation data; The backpropagation feedforward neural network, i.e., the BP neural network, includes an input layer, a hidden layer, and an output layer, which are sequentially connected; The input layer, hidden layer, and output layer all contain several neuron nodes, and the neurons are connected by weights; The transformation function of the neuron uses the Sigmoid function; The input layer inputs the attribute values of the operation data containing missing values; The output layer outputs the predicted data values to be filled; The weights and thresholds of the BP neural network are optimized and trained using a genetic algorithm to obtain a data prediction model; For the filled operation data, use the improved DBSCAN algorithm for clustering, select and delete the duplicate data records with high matching degree, and obtain the cleaned operation data.
2. The cleaning method for the operation data of power grid equipment according to claim 1, characterized in that The step of using a genetic algorithm for optimization training to obtain a data prediction model is as follows: Determine the network topology of the BP neural network, and divide the pre-prepared data sample set into training samples and test samples; Encode the weights and thresholds of the BP neural network to obtain an initial population; Use the training samples to train the BP neural network, use the test samples to test the BP neural network and calculate the error between the output value and the expected value; Calculate the fitness of the chromosomes in the BP neural network, select the chromosomes with high fitness for replication, and perform crossover and mutation operations to generate a new population; Judge whether the BP neural network reaches the performance index or the maximum number of iterations. If so, decode the chromosomes obtained by encoding the weights and thresholds of the BP neural network to obtain the weights and thresholds of the best neural network as the initial weights and thresholds of the data prediction model; Otherwise, obtain a new initial population and continue to execute.
3. The cleaning method of power grid equipment operation data according to claim 2, characterized in that The network topology of the BP neural network is represented as: m-l-n, where m, l, and n represent the number of neuron nodes in the input layer, hidden layer, and output layer respectively; The pre-prepared data sample set consists of N power grid data samples and is randomly divided into training samples and test samples which are respectively expressed as: Among them, (x(), y()) represents a certain sample in the sample space, x() represents the sample value associated with y(), y() represents the actual value of the sample to be predicted by the BP neural network, t is the time parameter, R m and R n are real numbers, N1 represents the number in the training samples, and N represents the number of samples in the data sample set; The training samples are used to establish the input-output mapping relationship of the BP neural network; The test samples are used to verify the correctness of the input-output mapping relationship; The input-output mapping relationship of the BP neural network is represented as: Among them, represents the output of the BP neural network, l is the number of neurons in the hidden layer, v jk represents the weight from the j-th neuron in the hidden layer to the k-th neuron in the output layer, k = 1, 2, …, n, w ij is the weight from the i-th neuron in the input layer to the j-th neuron, x i (t) represents the input of the BP neural network, θ j is the threshold at the j-th neuron in the hidden layer, r k is the threshold at the k-th neuron in the output layer, f[] is the activation function, expressed as: Suppose the total error E1 of the BP neural network is less than or equal to the network total error target threshold e1, then there is: where y k (t) represents the actual value of the sample; Suppose the average detection error E2 of the BP neural network is less than or equal to the detection sample average error target threshold e2, then there is: Where, n represents the number of neuron nodes in the output layer.
4. The cleaning method of the operation data of the power grid equipment according to claim 3, characterized in that, Encoding the weights and thresholds of the BP neural network to obtain the initial population, specifically: Using the real number encoding method to form a chromosome for the network weights of each node in the BP neural network, arranging them in sequence into a string, and optimizing the connection weights between neurons to initialize the population P(t); the transfer function of the hidden layer in the BP neural network adopts the sigmoid cross-entropy loss function to introduce the non-linear mapping ability.
5. The cleaning method for operation data of power grid equipment according to claim 4, characterized in that Calculating the fitness of the chromosomes in the BP neural network, specifically: Take out a chromosome i from the initial population P(t), input the network weights of each node in it into the BP neural network in sequence, calculate the total error E of the BP neural network, and define the fitness f of this chromosome i i , which is expressed as: ; Select chromosomes with high fitness for replication, and set the crossover probability P c , the mutation probability P m , and the selection probability p of chromosome i in the initial population i ; The selection probability is defined as: p i =α*(1-α) where α is a random number in [0, 1], and i = 1, 2, …, t; Performing crossover and mutation operations. Among them, the crossover operation is specifically: Sort each chromosome according to the fitness from high to low, and calculate the cumulative probability q of each chromosome i i : Generate a random number \(r\in[0,1]\) in each round using the roulette wheel algorithm. If \(q\) i-1 \(< r\leq q\) i , then take chromosome \(i\) as the parent chromosome, and perform the roulette wheel algorithm \(t\) times to obtain \(t\) parent chromosomes, and form chromosome pairs by adjacent parent chromosomes; for each chromosome pair, determine the crossover position \(k\) according to the crossover probability \(P\) c . Swap the genes between the two parent chromosomes in the chromosome pair numbered in \([1,k]\) to obtain two new chromosomes. The mutation operation is specifically: According to the mutation probability P m Determine u mutation positions on the two new chromosomes respectively, and perform mutation operations on the genes at the mutation positions, that is, add a random decimal uniformly distributed in [-1, 1] to the genes on the two new chromosomes to obtain two new daughter chromosomes; Inserting the new individuals into the population P(t) to generate a new population P(t + 1).
6. The cleaning method for operation data of power grid equipment according to claim 5, characterized in that, Using the improved DBSCAN algorithm to select and delete highly matching duplicate data records, specifically: Establishing a word-document matrix in the form of an inverted index, quickly obtaining the document list containing a certain word according to the word, and dividing the filled operation data into several subsets according to the same type of equipment; Using the improved DBSCAN algorithm to cluster several subsets so that duplicate data records form a cluster, calculating the similarity of the records in the cluster and judging whether they are duplicate data records, specifically: Determining the parameters of the improved DBSCAN algorithm, where the parameters include the radius ε′, the minimum number of points Minpts′ within the radius, and the initial similarity threshold R; Randomly selecting a data point A from a certain subset as the core point, and using the approximate distance function approxDist() to calculate the approximate distance between the remaining data points and the data point A; Clustering the filled operation data in the subset through the set distance thresholds N1 and N2: Combining the attribute weights to calculate the similarity of any two filled operation data, and the calculation formula is: where n is the total number of attributes of the filled running data, S Ai (x, y) represents the attribute similarity between the filled running data x and the filled running data y, expressed as: where d represents the distance between the points obtained by mapping x and y to the two-dimensional space respectively; Judging whether to update and iterate the radius ε′ according to the initial similarity threshold R, and the iteration formula is: When the iteration reaches the acceptable range of the similarity threshold for the output result, the clustering is completed to obtain the set of duplicate data records; Based on the set of duplicate data records, calculating the similarity of each data record, retaining the data record with the largest similarity, and deleting the rest to obtain the cleaned operation data.
7. A cleaning system for operation data of power grid equipment, characterized in that Applied to a method for cleaning operation data of power grid equipment described in any one of claims 1 - 6, including a data acquisition module, a noise recognition module, a data filling module, and a data deletion module; The data acquisition module is used to acquire the original operation data of power grid equipment from the database and perform preprocessing; The noise recognition module is used to identify the noise data with abnormal attributes in the preprocessed operation data and set them to null values to obtain the operation data containing missing values; The data filling module constructs and trains a data prediction model based on the backpropagation feedforward neural network, predicts the operation data containing missing values, obtains the data values to be filled, and performs data filling to obtain the filled operation data; The data deletion module uses the improved DBSCAN algorithm for the filled operation data to select and delete highly matching duplicate data records, and obtains the cleaned operation data.
8. A computer-readable storage medium storing a program, characterized in that, When the program is executed by a processor, it implements a method for cleaning operation data of grid equipment according to any one of claims 1-6.
Citation Information
Patent Citations
Air quality prediction method and device, equipment and storage medium
CN113610297A
Point cloud denoising method, image processing device and apparatus having storage function
WO2020114321A1