Low-voltage multi-source data checking and cleaning method

Through low-voltage multi-source data verification and cleaning methods, including data sorting, dimensionality reduction, cleaning, clustering and intelligent verification, bad data in the distribution network is identified and corrected, and the problem of bad data identification and correction in the existing technology is solved, and efficient analysis and auxiliary decision-making of power supply reliability of medium and low voltage users is achieved.

CN119939121APending Publication Date: 2025-05-06YUXI POWER SUPPLY BUREAU OF YUNNAN POWER GRID
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510013395.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to accurately and quickly identify and correct bad data in the distribution network, affecting the evaluation of the power supply reliability of medium and low voltage users.

Method used

Low-voltage multi-source data verification and cleaning methods are adopted, including collecting and sorting power user big data from different sources, using principal component analysis method to reduce dimensionality, processing missing data and inconsistent data, preliminary clustering and high-dimensional intelligent correlation verification are carried out through the M-BIRCH algorithm and LSTM neural network, combining the combined statistical model to identify bad data, and using curve similarity to correct bad data.

Benefits of technology

It realizes accurate identification and correction of low-voltage user data, improves data accuracy and reliability, supports efficient analysis of power supply reliability of medium- and low-voltage users, and provides support for assisting decision-making and risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939121A_ABST
    Figure CN119939121A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a low-voltage multi-source data checking and cleaning method, which comprises the following steps: collecting and sorting power consumer big data from different sources, and establishing an automatic data communication acquisition and acquisition mechanism; performing dimension reduction on the data by using a principal component analysis method, eliminating noise data, and simplifying a calculation process; the data is cleaned by processing missing data, repeated data and inconsistent data; standardization processing is carried out on data from different sources, and differences of units and dimensions are eliminated; the M-BIRCH algorithm is improved, and initial clustering and low-latitude anomaly checking processing are carried out on summarized data; carrying out high-dimensional intelligent association check on the data by adopting an LSTM improved neural network based on artificial intelligence; and identifying bad data by using a method based on a combined statistical model, and providing a bad data correction method based on curve similarity to correct the bad data. The method has the beneficial effect that bad data in the power distribution network can be accurately and quickly identified and corrected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of data processing, and in particular to a low-voltage multi-source data verification and cleaning method. Background Art

[0002] With the development of smart grids, the data of medium and low voltage users collected by distribution networks may contain various types of bad load data, and the existing bad data may affect the evaluation of the reliability of power supply to low voltage users. Data preprocessing is an important step in data mining, which usually includes operations such as deleting and supplementing data, standardizing data, and correcting data. Therefore, accurately and quickly identifying and correcting bad data in the distribution network is one of the important tasks faced by establishing a stable and reliable method for evaluating the reliability of power supply to low voltage users. In order to ensure the validity and reliability of medium and low voltage user data, the bad data correction method based on similarity is often used to repair the detected bad data. There are many characteristic factors that affect user behavior in a complex power grid environment. Directly modeling and analyzing user data will result in complex results, making it difficult to discover the laws and patterns of user behavior. Therefore, preprocessing is often required to simplify the number of user behavior features. Summary of the invention

[0003] In view of the above problems or problems existing in the prior art, the present invention is proposed.

[0004] Therefore, the object of the present invention is to provide a low voltage multi-source data verification and cleaning method, which can accurately and quickly identify and correct bad data in the distribution network to ensure the validity and reliability of medium and low voltage user data.

[0005] To solve the above technical problems, the present invention provides the following technical solutions: a low voltage multi-source data verification and cleaning method, which includes collecting and collating big data of power users from different sources, and establishing an automated data communication collection and acquisition mechanism;

[0006] Use principal component analysis to reduce the dimension of data, remove noise data, and simplify the calculation process;

[0007] Clean the data by dealing with missing data, duplicate data, and inconsistent data;

[0008] By standardizing data from different sources, differences in units and dimensions can be eliminated;

[0009] By improving the M-BIRCH algorithm, the summary data is preliminarily clustered and low-latitude anomaly verification is performed;

[0010] Adopt LSTM improved neural network based on artificial intelligence to perform high-dimensional intelligent correlation verification on data;

[0011] A method based on combined statistical models is used to identify bad data, and a bad data correction method based on curve similarity is proposed to correct the bad data.

[0012] As a preferred solution of the low-voltage multi-source data verification and cleaning method of the present invention, the big data of power users includes data of the dispatching automation system, the electric energy metering system, and the distribution automation terminal and the marketing business domain.

[0013] As a preferred solution of the low-voltage multi-source data verification and cleaning method of the present invention, data dimensionality reduction includes completing dimensionality reduction by building a data covariance matrix, calculating eigenvalues ​​and eigenvariables, selecting main components and converting data to a new space.

[0014] As a preferred solution of the low-voltage multi-source data verification and cleaning method of the present invention, data cleaning includes filling in missing data, merging or clearing duplicate data, and detecting and correcting deviations of inconsistent data.

[0015] As a preferred solution of the low-voltage multi-source data verification and cleaning method of the present invention, data standardization is to compare and analyze data from different sources and remove different units and dimensions of various types of data. The formula is as follows:

[0016]

[0017] Among them, max and min are the maximum and minimum values ​​of the samples respectively.

[0018] As a preferred solution of the low-voltage multi-source data verification and cleaning method of the present invention, the M-BIRCH algorithm improves the quality of clustering by introducing additional parameters and steps on the basis of the BIRCH algorithm.

[0019] As a preferred solution of the low-voltage multi-source data verification and cleaning method of the present invention, the BIRCH algorithm describes the clustering characteristics of data points by constructing a clustering feature tree, and the M-BIRCH algorithm adds a secondary analysis of the cluster blocks on this basis.

[0020] As a preferred solution of the low-voltage multi-source data verification and cleaning method of the present invention, the intelligent verification includes using the "gate" structure of the LSTM network to record the user's power consumption status, mine the user's short-term power consumption characteristics, and memorize the power consumption information at each moment.

[0021] As a preferred solution of the low-voltage multi-source data verification and cleaning method of the present invention, bad data identification includes processing missing data, judging mutation type, and continuous bad data.

[0022] As a preferred solution of the low-voltage multi-source data verification and cleaning method of the present invention, the bad data correction method measures the degree of correlation between different curves by calculating the grey correlation coefficient, and uses this correlation to correct the bad data.

[0023] Beneficial effects of the present invention: The present invention improves the accuracy of low-voltage user data by automatically identifying and correcting abnormal data and filling in missing data. It provides auxiliary decision-making for distribution network operation by supporting efficient and accurate analysis of power supply reliability for medium and low voltage users. It provides support for automatic analysis of power grid operation risk assessment and power outage events through precise data cleaning, thereby improving power supply reliability. It introduces a bad data correction method based on curve similarity and uses the grey correlation coefficient to measure the degree of correlation between curves, providing a new technical means for cleaning power big data. It can adapt to operating data from different sources, different equipment and different times, and has wide applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:

[0025] Figure 1 Schematic diagram of the missing data filling process for the low-voltage multi-source data verification and cleaning method.

[0026] Figure 2 Schematic diagram of the LSTM network recurrent unit structure for the low-voltage multi-source data verification and cleaning method.

[0027] Figure 3 Schematic diagram of a bad data identification method based on a combined statistical model for low-voltage multi-source data verification and cleaning methods.

[0028] Figure 4 Schematic diagram of the bad data correction method based on curve similarity for the low-pressure multi-source data verification and cleaning method. DETAILED DESCRIPTION

[0029] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0030] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0031] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0032] Example 1

[0033] Reference Figure 1 to Figure 4 , is an embodiment of the present invention, which provides a low-voltage multi-source data verification and cleaning method, which can accurately and quickly identify and correct bad data in the distribution network to ensure the validity and reliability of medium and low voltage user data.

[0034] Specifically, collect and organize big data of power users from different sources and establish an automated data communication collection and acquisition mechanism;

[0035] Use principal component analysis to reduce the dimension of data, remove noise data, and simplify the calculation process;

[0036] Clean the data by dealing with missing data, duplicate data, and inconsistent data;

[0037] By standardizing data from different sources, differences in units and dimensions can be eliminated;

[0038] By improving the M-BIRCH algorithm, the summary data is preliminarily clustered and low-latitude anomaly verification is performed;

[0039] Adopt LSTM improved neural network based on artificial intelligence to perform high-dimensional intelligent correlation verification on data;

[0040] A method based on combined statistical models is used to identify bad data, and a bad data correction method based on curve similarity is proposed to correct the bad data.

[0041] Furthermore, the big data of electricity users includes data from dispatching automation systems, electric energy metering systems, and distribution automation terminals and marketing business domains.

[0042] It should be noted that electricity consumption data can be obtained through TTUs, branch monitoring units, and terminal sensing terminals (or smart meters with HPLC functions) installed in low-voltage substations. The data covers a wide range and has a large transmission volume, including power outage-related data collected and calculated by distribution network terminals, power outage events identified by distribution network automation systems, switch signals of the main network, and power outage and restoration information in the management system. Specific data include distribution network line topology data, circuit breakers and reclosing devices, section switches, fault indicators, distribution transformers, smart meters and other information. A unified distribution information model is constructed using the principle of the Common Information Model (CIM), thereby forming low-voltage multi-source data based on household meters and distribution automation.

[0043] Furthermore, data dimensionality reduction includes building a data covariance matrix, calculating eigenvalues ​​and eigenvariables, selecting principal components and transforming data into a new space to complete dimensionality reduction.

[0044] It should be noted that the main steps of data dimensionality reduction include:

[0045] Build the data covariance matrix;

[0046] Calculate the eigenvalues ​​and eigenvariates of the covariance matrix respectively;

[0047] Arrange the eigenvalues ​​according to their contribution;

[0048] After selecting the first K eigenvalues ​​as the main components, the data is converted into a new data space, and dimensionality reduction is performed on it. The data space is constructed using new eigenvectors.

[0049] Furthermore, data cleaning includes filling in missing data, merging or removing duplicate data, and detecting and correcting deviations in inconsistent data.

[0050] It should be noted that data cleaning specifically includes:

[0051] Detect and correct the inconsistent data according to data rules;

[0052] For missing data, you can fill in the data, clear the tuples, or not process them. In order to preserve the integrity of the original data to the greatest extent, this time we use the method of filling in missing data. Figure 1 The figure shows the flowchart for filling missing data.

[0053] Calculate the similarity and use it to determine whether there is overlap. If there is duplication, merge or remove it. Calculate the distance to obtain the similarity, that is, the actual distance between two points in N-dimensional space. The distance calculation formula for N-dimensional space is as follows:

[0054]

[0055] Where d is the distance between two points; (x2-x1) 2 Represents the square of the difference between the coordinates of two points in the first dimension, where x2 and x1 are the coordinates of the two points in the first dimension; (y2-y1) 2 Represents the square of the difference between the coordinates of two points in the second dimension; (z n -z n-1 ) 2 Represents the square of the difference between the coordinates of two points in the nth dimension.

[0056] Furthermore, data standardization is to conduct comparative analysis of data from different sources and eliminate different units and dimensions of various types of data. The formula is as follows:

[0057]

[0058] Among them, max and min are the maximum and minimum values ​​of the samples respectively.

[0059] Furthermore, the M-BIRCH algorithm improves the quality of clustering by introducing additional parameters and steps based on the BIRCH algorithm.

[0060] Furthermore, the BIRCH algorithm describes the clustering characteristics of data points by constructing a clustering feature tree, and the M-BIRCH algorithm adds a secondary analysis of cluster blocks on this basis.

[0061] It should be noted that the BIRCH algorithm is a hierarchical clustering algorithm. Its clustering idea is summarized by clustering features and feature trees (CF trees). BIRCH defines: i}(i=1,2,3,…,N), which has N d-dimensional data points, and the feature vector definition is described by the following formula:

[0062] CF=(N,LS,SS)

[0063] Among them, N describes the set of all points in the cluster, and the linear sum of N points is It is described by LS and used to interpret the center of the cluster. The sum of squares of the data points is described by SS and is used to measure the length of the cluster diameter.

[0064] Among them, the clustering feature theorem is: use CF1=(N1,LS1,SS1), CF2=(N2,LS2,SS2) and CF1+CF2=(N1+N2,LS1+LS2,SS1+SS2) to describe the clustering features of the two classes and the new class features obtained by fusion respectively.

[0065] The algorithm calculates the center, radius, and distance between clusters through clustering features. The features of hierarchical clustering are located in the CF tree, which is composed of a highly balanced tree with two parameters: branching factor B and threshold T. Among them, the maximum number of non-leaf nodes depends on the branching factor, and the longest diameter of the sub-cluster located in the leaf node of the tree is determined by the threshold value. The CF tree can read all data into memory, or read data items separately into external memory.

[0066] It should be noted that the M-BIRCH algorithm performs secondary analysis based on the initial results obtained by the BIRCH clustering algorithm to obtain more accurate results. P is used to describe the abnormal probability of power user behavior, and the percentage, the average distance in the current class, and the average distance between the point and the rest of the points in the class are respectively expressed as d avg d new The threshold is described by T. The newly started data points need to be counted before processing. When the data point is included in the original cluster block, the BIRCH clustering algorithm pre-calculates and corrects the cluster feature values ​​according to the set threshold T, and integrates the processing results into the cluster block; otherwise, the average distance d of all data points in the current cluster block of the data point is collected. new , the average distance d between it and the current cluster block avg For comparison. avg Multiply the value of the initial percentage P by more than d new After completing the correction operation of the clustering feature values ​​in the clustering block, the correction results are integrated into the clustering block. On the contrary, the subsequent clustering blocks are operated, and if they do not match, a new clustering block is built.

[0067] The specific process is:

[0068] Initialize the CF tree, set the branching factor B and threshold T;

[0069] Perform sliding window processing on the data stream and use the BIRCH algorithm for preliminary clustering;

[0070] For each new data point, calculate its fitness with the existing clustering blocks;

[0071] According to the fitness and threshold T, decide whether the new data point should be integrated into the existing cluster block or create a new cluster block;

[0072] The cluster blocks were subjected to secondary analysis and the M-BIRCH algorithm was used to make more refined cluster adjustments.

[0073] The big data clustering algorithm process based on the M-BIRCH algorithm is as follows:

[0074] M-BIRCH-Cluster(T,d new ,d avg ,P),

[0075] {First, accumulate data flows in the sliding window and use the BRICH algorithm to cluster the data volume, and each clustering block is divided according to its output result}

[0076] For (not reaching the end of the data stream) {

[0077] Select a new data point and read it in;

[0078] For (calculate the existing cluster blocks one by one) {

[0079] If (T threshold >= maximum diameter) {the data point is absorbed into the cluster block and the cluster feature value is corrected}

[0080] Else{The average distance between the data point and all data points in the current clustering block}

[0081] Collect the average distance d of the cluster block avg ,

[0082] If(d avg Multiply by the initial percentage P value greater than d new ){The data point is absorbed into the cluster block and the cluster feature value is corrected}

[0083] Else{This data point is calculated synchronously with the next clustering block}

[0084] If the injury condition is not met, a new cluster block is built

[0085] }}

[0086] Furthermore, intelligent verification includes using the "gate" structure of the LSTM network to record the user's electricity usage status, explore the user's short-term electricity usage characteristics, and memorize the electricity usage information at each moment.

[0087] It should be noted that LSTM, as an improved recurrent neural network, solves the gradient vanishing and explosion problems of recurrent neural networks (RNN) in long-term series training, and is widely used in text processing, speech recognition, time series prediction and other fields. Figure 2 The figure shows the LSTM recurrent network structure. The LSTM network uses a "gate" structure to achieve memory and control of cell states.

[0088] Among them, the forget gate (forgetgate, denoted as f t ), decides whether to remember the state information of the previous unit.

[0089] The calculation formula is as follows: t =σ(W f ·[h t-1 , xt ]+b f )

[0090] Among them, W f 、b f is the forget gate weight and bias term; σ is the sigmoid activation function; [h t-1 ,x t ] is the output h of the previous unit t-1 and the current input x t The matrix composed of.

[0091] Input gate (inputgate, denoted as i t ), which is used to update the cell state and memorize important information about the previous and current cell states. The calculation formula is as follows:

[0092] i t =σ(W i ×[h t-1 ,x t ]+b i )

[0093]

[0094]

[0095] Among them, W i 、b i is the input gate weight and bias term; W C 、b c are weights and bias items for the updated values; is the candidate value in the state; tanh is the activation function of the candidate cell information; C t is the updated cell state.

[0096] Output gate (output gate, denoted as O t ), determines which information in the cell state is stored as the output state h t The calculation formula is as follows:

[0097] O t =σ(W o ×[h t-1 ,x t ]+b o )

[0098] h t =O t ×tanh(C t )

[0099] Among them, W o 、b o are the weight and bias of the output gate.

[0100] Cell state (denoted as C t ), only a small amount of linear transformation is performed to memorize and transmit the state information of each unit. The calculation formula is as follows:

[0101]

[0102] Among them, ■ is the product of vector elements, C t-1 It is the memory unit of the previous moment.

[0103] LSTM records the user’s electricity consumption status C at time t t The gate structure can not only help to mine the short-term power consumption characteristics of users, but the new cell state can also memorize the power consumption information at each moment. t , the electricity consumption information of each moment together with the current moment is used as the input of the next moment, so as to effectively mine the characteristics of users' short-term electricity consumption data.

[0104] Bad data identification method based on combined statistical model. Firstly, missing data is processed, then mutation bad data is judged by 3sigma criterion, then continuous load is identified by comparing the change rate of load at each moment relative to the previous moment, and finally bad data is accurately identified. Suppose there is an m×n dimensional load data set L m×n The steps of the bad data identification method based on the combined statistical model are as follows:

[0105] First, the load data is preprocessed to form a dimensional load data set:

[0106]

[0107] Identify missing values ​​in load data and mark them as bad data;

[0108] Calculate the mean μ for each row of the load data set m×1 and standard deviation σ m×1 ;

[0109] According to the 3sigma criterion, the upper and lower limits of normal load data are calculated. The upper limit of the load data at the i-th moment on the load curve is The lower limit is

[0110] To judge mutation type bad data, if If the load data l ij is the normal load data; otherwise, the load data l ij It is mutation type bad data;

[0111] Calculate the change rate C of the load data sample at each moment relative to the previous moment m×n ;

[0112] Determine the continuity of bad data, if C i,j =0 or C i,j-1 =C i,j or C i,j+1 =C i,j , then the load data l ij This is continuous bad data.

[0113] Furthermore, bad data identification includes processing missing data, judging mutation type, and continuous bad data.

[0114] Furthermore, the bad data correction method measures the correlation between different curves by calculating the grey correlation coefficient and uses this correlation to correct the bad data.

[0115] It should be noted that if Figure 4 As shown in the figure, it is a flow chart of the bad data correction method based on curve similarity. The algorithm first determines the user category and seasonal attributes of the data to be corrected; according to the seasonal periodicity analysis of the load, the daily loads in the same season have similar changes, so the seasonal typical daily load curve of the corresponding season is selected, and the gray correlation coefficient between the remaining time of the day where the data to be corrected is located and the seasonal typical daily load curve is calculated as the similarity, and the average value of the similarity at the remaining time is used as the weight; the weight is multiplied by the load at the time when the data to be corrected is located on the typical daily load curve, and the load correction value calculated by the formula is obtained.

[0116] Among them, the grey correlation coefficient is used to calculate the similarity, which can be used to measure the degree of correlation between curves and judge whether their connection is close. The specific design steps of the bad data correction method based on curve similarity are given:

[0117] Determine the user type to which the load to be corrected belongs;

[0118] Determine the season to which the load to be corrected belongs;

[0119] Calculate the seasonal typical daily load curve Y=y(i), i=1,2…n of the season to which the corresponding user type belongs;

[0120] The remaining normal time X=x(j), j=1,2…m of the day where the data to be corrected is located and the load value corresponding to the time Y=y(j), j=1,2…m of the seasonal typical daily load curve is processed without quantization to obtain X'=x'(j), j=1,2…mY=y′(j), j=1,2…m;

[0121] The correlation coefficient between the two is calculated as follows:

[0122]

[0123] Among them, ε(j)=|y'(j)-x'(j)|, Generally, 0.5, μj It represents the correlation coefficient between the jth moment of the day to be corrected and the jth moment of the seasonal typical day load curve;

[0124] 6. The correlation coefficient is the closeness value at each moment in the curve. There is more than one value, so the average value of the correlation coefficient at each moment is used as the curve similarity. The calculation formula is as follows:

[0125]

[0126] To calculate the load correction value, use the curve similarity to weight the load value of the typical daily load curve of the season at the time of correction. The calculation formula is as follows:

[0127]

[0128] Among them, k is the time when bad data appears;

[0129] Replace the bad data with the corrected value x(k) and end.

[0130] In summary, the present invention improves the accuracy of low-voltage user data by automatically identifying and correcting abnormal data and completing missing data. It provides auxiliary decision-making for distribution network operation by supporting efficient and accurate analysis of power supply reliability for medium and low voltage users. It provides support for automatic analysis of power grid operation risk assessment and power outage events through precise data cleaning, thereby improving power supply reliability. It introduces a bad data correction method based on curve similarity and uses the grey correlation coefficient to measure the degree of correlation between curves. It provides a new technical means for cleaning power big data, which can adapt to operating data from different sources, different equipment and different times, and has wide applicability.

[0131] Importantly, it should be noted that the construction and arrangement of the present application shown in a plurality of different exemplary embodiments are only exemplary. Although only a few embodiments are described in detail in this disclosure, it should be readily understood by those who refer to this disclosure that many modifications are possible without substantially departing from the novel teachings and advantages of the subject matter described in the application (e.g., mounting arrangements, use of materials, color, changes in orientation, etc.). For example, the element shown as integrally formed may be composed of a plurality of parts or elements, the position of the element may be inverted or otherwise changed, and the nature or number or position of the discrete element may be altered or changed. Therefore, all such modifications are intended to be included within the scope of the present invention. The order or sequence of any process or method steps may be changed or reordered according to an alternative embodiment. In the claims, any "bracket plus function" clause is intended to cover the structure of the execution function described herein, and is not only structurally equivalent but also equivalent structure. Without departing from the scope of the present invention, other replacements, modifications, changes and omissions may be made in the design, operating conditions and arrangement of the exemplary embodiments. Therefore, the present invention is not limited to a specific embodiment, but extends to a variety of modifications that still fall within the scope of the appended claims.

[0132] Additionally, in an effort to provide a concise description of example embodiments, all features of an actual implementation may not be described.

[0133] It should be understood that in the development of any actual implementation, as in any engineering or design project, numerous implementation-specific decisions may be made. Such a development effort may be complex and time-consuming, but for those of ordinary skill having the benefit of this disclosure, the development effort will be a routine task of design, fabrication, and production without undue experimentation.

[0134] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A low-voltage multi-source data verification and cleaning method, characterized in that: include, Collect and organize big data of power users from different sources and establish an automated data communication collection and acquisition mechanism; Use principal component analysis to reduce the dimension of data, remove noise data, and simplify the calculation process; Clean the data by dealing with missing data, duplicate data, and inconsistent data; Eliminate differences in units and dimensions by standardizing data from different sources; By improving the M-BIRCH algorithm, the summary data is preliminarily clustered and low-latitude anomaly verification is performed; Adopt LSTM improved neural network based on artificial intelligence to perform high-dimensional intelligent correlation verification on data; A method based on combined statistical models is used to identify bad data, and a bad data correction method based on curve similarity is proposed to correct the bad data.

2. The low-voltage multi-source data verification and cleaning method according to claim 1, characterized in that: The power user big data includes data from the dispatching automation system, the electric energy metering system, and the distribution automation terminal and marketing business domain.

3. The low-voltage multi-source data cleaning method as described in claim 2 is verified and characterized in that: the data dimensionality reduction includes completing dimensionality reduction by building a data covariance matrix, calculating eigenvalues ​​and eigenvariates, selecting main components and converting data to a new space.

4. The low-voltage multi-source data verification and cleaning method according to claim 3, characterized in that: The data cleaning includes filling in missing data, merging or clearing duplicate data, and detecting and correcting deviations in inconsistent data.

5. The low-voltage multi-source data verification and cleaning method according to claim 4, characterized in that: The data standardization is to compare and analyze data from different sources and eliminate different units and dimensions of various types of data. The formula is as follows: Among them, max and min are the maximum and minimum values ​​of the samples respectively.

6. The low-voltage multi-source data verification and cleaning method according to claim 5, characterized in that: The M-BIRCH algorithm improves the quality of clustering by introducing additional parameters and steps based on the BIRCH algorithm.

7. The low-voltage multi-source data verification and cleaning method according to claim 6, characterized in that: The BIRCH algorithm describes the clustering features of data points by constructing a clustering feature tree, and the M-BIRCH algorithm adds a secondary analysis of cluster blocks on this basis.

8. The low-voltage multi-source data verification and cleaning method according to any one of claims 6 or 7, characterized in that: The intelligent verification includes using the "gate" structure of the LSTM network to record the user's electricity usage status, explore the user's short-term electricity usage characteristics, and memorize the electricity usage information at each moment.

9. The low-voltage multi-source data verification and cleaning method according to claim 8, characterized in that: The bad data identification includes processing missing data, judging mutation type, and continuous bad data.

10. The low-voltage multi-source data verification and cleaning method according to claim 9, characterized in that: The bad data correction method measures the correlation between different curves by calculating the grey correlation coefficient, and uses this correlation to correct the bad data.