A data governance method and system for power equipment archive data

Through the Spark parallel computing model and regular expression standard library, the problem of low data quality in power equipment archives is solved, an efficient data governance process is realized, and data quality and governance efficiency are improved.

CN113919427BActive Publication Date: 2025-07-04STATE GRID NINGXIA ELECTRIC POWER CO +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111201888.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-15
Publication Date
2025-07-04
Estimated Expiration
2041-10-15

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively manage the archive data of power equipment, resulting in low data quality and affecting the internal business of the power grid and the construction of upper-level applications.

Method used

The Spark parallel computing model and regular expression standard library are adopted to establish an abnormal data filtering algorithm model through data understanding, filtering and correction processes to improve data quality.

Benefits of technology

It improves the quality of power equipment archive data, reduces the difficulty of data governance, and improves the work efficiency of data governance and the scalability of system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113919427B_ABST
    Figure CN113919427B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for data governance of power equipment archives. By steps such as business and data understanding, data extraction, model establishment, and data correction, the data governance method of power equipment archives is improved into a set of processes, reducing the work difficulty of data governance. Moreover, the Spark parallel computing model is used for memory computing with high reuse rate, and the modeling method combined with regular expressions is used for algorithm modeling. The present invention can improve the work efficiency of abnormal identification of archive data, and has high scalability and good performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power equipment data management, and particularly to a method and system for governing power equipment archive data. Background Art

[0002] The archive data of power equipment is the basis for the development of power grid production work. Grassroots team members are responsible for inputting, updating, and maintaining the archive data of power equipment. All production work such as on-site power equipment operation and maintenance, inspection, testing, and experiment needs to be based on the archive data of power equipment. Only when the archive data is accurate can it be ensured that the relevant operation and maintenance records can be accurately registered in its production management system. However, due to problems such as different data sources, different statistical calibers for the same data, data input by front-line personnel, abnormal behaviors, etc., and the lack of a corresponding data quality control system, abnormal data often occurs. These data problems affect the development of internal business in the power grid and also affect the construction of upper-layer applications based on these archive data. Therefore, it is necessary to design a power equipment archive data governance platform based on distributed computing to improve the efficiency of data quality control, enhance the quality of archive data, and thus realize the in-depth application of these data resources.

[0003] Most of the existing data governance methods are for the measurement data during the operation process of power equipment, that is, some artificial intelligence algorithms are used to identify anomalies in the measurement data, and then the data is corrected and filled. However, since most of the archive data is text data, it is difficult to govern the data using machine learning algorithms and it is more difficult to implement. Summary of the Invention

[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a method and system for governing power equipment archive data, which can improve the quality of power archive data.

[0005] To achieve the above purpose, the present invention provides the following solutions:

[0006] A method for governing power equipment archive data includes:

[0007] Understanding the data of power equipment operations and the archive data of the power equipment operations to obtain a data input specification library;

[0008] According to the data input specification library, extracting power equipment archive data from the power equipment data warehouse into a Spark parallel computing model to obtain a resilient distributed dataset;

[0009] Based on the data input specification library, calling the regular expression standard library in the Spark parallel computing model to establish an archive data anomaly screening algorithm model;

[0010] Filter and count the elastic distributed dataset according to the above-mentioned algorithm model for screening abnormal data of archive category to obtain the dataset of archive category to be corrected;

[0011] Based on the API call function in the Spark parallel computing model, correct the dataset of archive category to be corrected to obtain the dataset of archive category with data governance completed.

[0012] Preferably, the data understanding of the power equipment business and the archive data of the power equipment business to obtain a data entry specification library includes:

[0013] Analyze the power equipment business to obtain business knowledge data; the business knowledge data includes: power equipment business structure data, data governance requirement data, and target completion data;

[0014] Judge the abnormal data set and the reasons for abnormal data in the archive data;

[0015] Build the data entry specification library based on the archive data, the business knowledge data, the abnormal data set, and the reasons for abnormal data.

[0016] Preferably, according to the data entry specification library, extracting power equipment archive category data from the power equipment data warehouse into the Spark parallel computing model to obtain an elastic distributed dataset includes:

[0017] According to the data entry specification library, import the power equipment archive category data from the power equipment data warehouse into the Spark parallel computing model to obtain an initial dataset;

[0018] Conduct an integrity check on the initial dataset to obtain the inspected elastic distributed dataset.

[0019] Preferably, based on the data entry specification library, calling the regular expression standard library in the Spark parallel computing model to establish an algorithm model for screening abnormal data of archive category includes:

[0020] With reference to the data entry specification library, establish abnormal data recognition rules for power equipment archive category data according to preset business requirements;

[0021] According to the abnormal data recognition rules for power equipment archive category data, call the regular expression standard library provided by the Spark parallel computing model to establish the algorithm model for screening abnormal data of archive category.

[0022] Preferably, after establishing the screening algorithm model for abnormal archive data by invoking the regular expression standard library provided by the Spark parallel computing model according to the abnormal identification rule for power equipment archive data, the following steps are further included:

[0023] Use the cluster driver in the Spark parallel computing model to read the input tasks of the screening algorithm model for abnormal archive data, and distribute the input tasks to multiple executors for processing to obtain the running results;

[0024] Based on the running results, count the abnormal data in the archive data metrics;

[0025] Improve the abnormal identification rule for power equipment archive data according to the running results and the abnormal data to obtain the improved abnormal identification rule for power equipment archive data;

[0026] Optimize the screening algorithm model for abnormal archive data according to the improved abnormal identification rule for power equipment archive data to obtain the optimized screening algorithm model for abnormal archive data.

[0027] Preferably, after optimizing the screening algorithm model for abnormal archive data according to the improved abnormal identification rule for power equipment archive data to obtain the optimized screening algorithm model for abnormal archive data, the following steps are further included:

[0028] Establish a statistical information error table according to the output results of the optimized screening algorithm model for abnormal archive data.

[0029] Preferably, the correction of the archive dataset to be corrected based on the API call function in the Spark parallel computing model to obtain the archive dataset with data governance completed includes:

[0030] Based on the structured API call function provided by Spark, perform data deletion, data filling, and data correction on the archive dataset to be corrected to obtain the archive dataset with data governance completed;

[0031] Among them, the steps of data deletion include:

[0032] Deduplicate and save the duplicate data in the archive dataset to be corrected;

[0033] Delete the missing data in the archive dataset to be corrected according to the data understanding and the abnormal identification rule for power equipment archive data;

[0034] The steps of data filling include:

[0035] Fill in the missing value data in the file - type dataset to be corrected according to the missing value filling method;

[0036] The steps of the data correction include:

[0037] Batch - correct the format - error data in the file - type dataset to be corrected.

[0038] A data governance system for power equipment file - type data includes:

[0039] A data understanding module, which is used to understand the power equipment business and the file data of the power equipment business to obtain a data entry specification library;

[0040] A data extraction module, which is used to extract the power equipment file - type data from the power equipment data warehouse into the Spark parallel computing model according to the data entry specification library to obtain a resilient distributed dataset;

[0041] A model establishment module, which is used to call the regular expression standard library in the Spark parallel computing model to establish a screening algorithm model for file - type abnormal data based on the data entry specification library;

[0042] A screening module, which is used to screen and count the resilient distributed dataset according to the screening algorithm model for file - type abnormal data to obtain a file - type dataset to be corrected;

[0043] A data correction module, which is used to correct the file - type dataset to be corrected based on the API call function in the Spark parallel computing model to obtain a file - type dataset with data governance completed.

[0044] Preferably, the data understanding module specifically includes:

[0045] An analysis unit, which is used to analyze the power equipment business to obtain business knowledge data; the business knowledge data includes: power equipment business structure data, data governance requirement data, and target completion data;

[0046] A judgment unit, which is used to judge the abnormal data set and the reason for abnormal data in the file data;

[0047] A specification library construction unit, which is used to construct the data entry specification library based on the file data, the business knowledge data, the abnormal data set, and the reason for abnormal data.

[0048] Preferably, the data extraction module specifically includes:

[0049] An import unit for importing the power equipment archive data from the power equipment data warehouse into the Spark parallel computing model according to the data entry specification library to obtain an initial data set;

[0050] An inspection unit for performing integrity inspection on the initial data set to obtain the inspected resilient distributed data set.

[0051] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0052] The present invention provides a method and system for governing power equipment archive data. By steps such as business and data understanding, data extraction, model establishment, and data correction, the method for governing power equipment archive data is improved into a set of processes, reducing the work difficulty of data governance. And the Spark parallel computing model is used for high-reusability in-memory computing, and an algorithm model is built by combining a modeling method with regular expressions. The present invention can improve the work efficiency of abnormal identification of archive data and has high scalability and good performance. Description of the Drawings

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0054] Figure 1 It is a flowchart of the method for governing power equipment archive data in the embodiments provided by the present invention;

[0055] Figure 2 It is a schematic diagram of the implementation steps of the method for governing power equipment archive data in the embodiments provided by the present invention;

[0056] Figure 3 It is a module connection diagram of the power equipment archive data governance system in the embodiments provided by the present invention. Detailed Embodiments

[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0058] Reference to "embodiment" in this document means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears at various positions in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0059] The terms "first", "second", "third", "fourth", etc. in the specification, claims, and drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a series of steps, processes, methods, etc. included do not limit to the listed steps, but optionally further include steps not listed, or optionally further include other step elements inherent to these processes, methods, products, or devices.

[0060] The present invention addresses the problems in the prior art: the lack of key parameters in the power equipment archives, and the non-standard filling of the power equipment archive names, resulting in the inability to identify and incorrect filling of the power equipment archive parameters or inconsistency with the actual situation on site of the power equipment. The present invention provides a method and system for data governance of power equipment archive data, which can improve the quality of power archive data.

[0061] To make the above objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0062] Figure 1 This is the flowchart of the method for data governance of power equipment archive data in the embodiments provided by the present invention, as Figure 1 shown, the present invention provides a method for data governance of power equipment archive data, including:

[0063] Step 100: Understand the power equipment business and the archive data of the power equipment business to obtain a data entry specification library;

[0064] Step 200: According to the data entry specification library, extract the power equipment archive data from the power equipment data warehouse into the Spark parallel computing model to obtain a resilient distributed dataset;

[0065] Step 300: Based on the data entry specification library, call the regular expression standard library in the Spark parallel computing model to establish an algorithm model for screening abnormal archive data;

[0066] Step 400: Screen and count the elastic distributed dataset according to the algorithm model for abnormal file data to obtain the file dataset to be corrected;

[0067] Step 500: Correct the file dataset to be corrected based on the API call function in the Spark parallel computing model to obtain the file dataset with data governance completed.

[0068] Specifically, in this embodiment, the understanding of power equipment business and its file data can standardize data entry and ensure the effectiveness of the model.

[0069] Preferably, the step 100 includes:

[0070] Analyze the power equipment business to obtain business knowledge data; the business knowledge data includes: power equipment business structure data, data governance requirement data, and target completion data;

[0071] Judge the abnormal data set and the reasons for abnormal data in the file data;

[0072] Build the data entry specification library based on the file data, the business knowledge data, the abnormal data set, and the reasons for abnormal data.

[0073] Optionally, the understanding of power equipment business and its file data in the step 100 should start from the application scenarios of power data and include the following four points:

[0074] a) Understand the power equipment business structure;

[0075] b) Convert business knowledge into the requirements for data governance problems and the preliminary plan for achieving the goals;

[0076] c) Initially judge the possible abnormal and invalid data in the file data and their causes;

[0077] d) Establish a data entry specification library to standardize various index attributes of the modeling data.

[0078] Preferably, the step 200 includes:

[0079] Import the power equipment file data from the power equipment data warehouse into the Spark parallel computing model according to the data entry specification library to obtain an initial dataset;

[0080] Check the integrity of the initial dataset to obtain the inspected elastic distributed dataset.

[0081] Specifically, the steps of extracting power equipment archive data from the power equipment data warehouse to the Spark parallel computing platform (Spark parallel computing model) are as follows:

[0082] a) According to the data entry specification, import the required modeling data from the power equipment data warehouse into Spark to form a resilient distributed dataset for parallel computing;

[0083] b) Check the power equipment archive data, including the uniqueness screening of the data's geospatial attributes, power equipment station lines, business member attributes, and power equipment component / business unit attributes. The modeling data needs to maintain the unique identifier of its main equipment parameters.

[0084] Preferably, the step 300 includes:

[0085] Referring to the data entry specification library, establish an abnormal data recognition rule for power equipment archive data according to the preset business requirements;

[0086] According to the abnormal data recognition rule for power equipment archive data, call the regular expression standard library provided by the Spark parallel computing model to establish the screening algorithm model for abnormal archive data.

[0087] Preferably, after establishing the screening algorithm model for abnormal archive data by calling the regular expression standard library provided by the Spark parallel computing model according to the abnormal data recognition rule for power equipment archive data, it further includes:

[0088] Use the cluster driver in the Spark parallel computing model to read the input tasks of the screening algorithm model for abnormal archive data and distribute the input tasks to multiple executors for processing to obtain the running results;

[0089] Based on the running results, count the abnormal data in the archive data metrics;

[0090] Improve the abnormal data recognition rule for power equipment archive data according to the running results and the abnormal data to obtain the improved abnormal data recognition rule for power equipment archive data;

[0091] Optimize the screening algorithm model for abnormal archive data according to the improved abnormal data recognition rule for power equipment archive data to obtain the optimized screening algorithm model for abnormal archive data.

[0092] Preferably, after optimizing the screening algorithm model for abnormal archive data according to the improved abnormal data recognition rule for power equipment archive data to obtain the optimized screening algorithm model for abnormal archive data, it further includes:

[0093] Establish a statistical information error table based on the output result of the optimized algorithm model for screening abnormal data of archives.

[0094] Specifically, the construction steps of the abnormal archive data extraction model are as follows:

[0095] a) Refer to the data entry specification library and initially establish the abnormal data recognition rules for power equipment archive data according to business requirements.

[0096] b) According to the abnormal data recognition rules, call the regular expression standard library provided in Spark to establish an algorithm model for screening abnormal archive data.

[0097] c) After the Spark cluster driver reads the input of the algorithm model, it distributes the tasks to several executors for processing. Each executor is responsible for implementing different abnormal data recognition algorithms and counting the abnormalities in the archive data metrics based on the running results of the algorithm model.

[0098] d) According to the results in c), improve the abnormal data recognition rules to meet the requirements of uniqueness, integrity, consistency, validity, and accuracy of data quality. The Spark platform runs the model again according to the improved abnormal data recognition rules and establishes a statistical information error table.

[0099] Preferably, the step 500 includes:

[0100] Based on the structured API call function provided by Spark, perform data deletion, data filling, and data correction on the archive data set to be corrected, and obtain the archive data set with data governance completed.

[0101] Among them, the step of data deletion includes:

[0102] Deduplicate and save the duplicate data in the archive data set to be corrected.

[0103] Delete the missing data in the archive data set to be corrected according to the data understanding and the abnormal data recognition rules for power equipment archive data.

[0104] The step of data filling includes:

[0105] Fill in the missing value data in the archive data set to be corrected according to the missing value filling method.

[0106] The step of data correction includes:

[0107] Batch correct the format error data in the archive data set to be corrected.

[0108] Specifically, for the abnormal data screened by the model, based on the functions called by the structured API provided by Spark, the steps for correcting the power equipment archive dataset are as follows:

[0109] a) Data deletion: For duplicate data, one record needs to be directly de-duplicated and retained; for missing data, it is judged whether to delete according to the understanding of power archive data and the abnormal identification rules.

[0110] b) Data filling: Applied to missing value data, different missing value filling methods are adopted according to different attributes of the indicators.

[0111] c) Data correction: For some data with format errors (such as non-standard date format, classification error, feature matching error), in this embodiment, the data can be corrected in batches, and the incorrect data or non-standard format data can be corrected.

[0112] d) Obtain a power equipment archive dataset with high data quality.

[0113] Figure 2 It is a schematic diagram of the implementation steps of the power equipment archive data governance method in the embodiment provided by the present invention. As Figure 2 shown, in this embodiment, specific steps for applying the above method to the actual data governance field are also provided. Taking the wire data of the Power Production Management System (PMS) as an example, the specific implementation steps of its data governance platform are as follows:

[0114] The specific process of Step 1 is as follows:

[0115] (1) Understand the PMS wire business and requirements: The business process of PMS manufacturing wires (cables) includes basic ledger data management and defect management. The line equipment is organized by one pole and one card. Each cable has a corresponding equipment ledger, called a ledger card. The database composed of the ledger cards of the cables records all cable information and all cable changes. The query of the database needs to ensure the high quality and accuracy of the archive data.

[0116] (2) Preliminary judgment on PMS wire data anomalies and their causes: Due to reasons such as human misoperation and recording errors, there may be a large amount of abnormal data in the archive data. To modify the data targeted, it is necessary to record the types of archive data ledgers in the database, the id codes in each ledger, the names of error labels, the parameter names corresponding to the error labels, the incorrect data, the error types, the detection time, the error information, the description, etc.

[0117] (3) Establish a basic data entry specification library for PMS wires, including parameter names, measurement units, requirements for entry methods, and filling instructions for each parameter in the database.

[0118] 2. The specific process of Step 2 is as follows:

[0119] (1) According to the basic data entry specification library of PMS conductors, import the required modeling data from the PMS data warehouse into the Spark platform and establish a resilient distributed dataset, with a total of 100,000 data records.

[0120] (2) Conduct an integrity check on the PMS conductor database, including the following 39 parameters: affiliated line, affiliated city, starting tower, ending tower, length (m), commissioning date, power supply area, affiliated main feeder, equipment status, model, whether it is maintained by others, number of conductor strands and specifications, manufacturer, rotation direction, whether it is a rural grid, conductor type, equipment code, registration time, conductor material type, equipment owner, PM code, voltage level name, affiliated sectional line, professional classification, affiliated main feeder id, maximum feeder branch id, number of splits, conductor cross-section (mm 2 ), maximum allowable current of the conductor (A), breaking tensile force (N), maximum design stress (MPa), rated current-carrying capacity (A), safety factor, remarks, voltage level code, equipment type code, and equipment id.

[0121] 3. The specific process of Step 3 is as follows:

[0122] (1) Based on the data entry specification library established in Step 1, item 3, preliminarily determine the rules for identifying abnormal PMS conductor data:

[0123] a) Uniqueness rule: All conductor parameters are unique and without repetition.

[0124] b) Integrity rule: All evaluated fields cannot be empty.

[0125] c) Accuracy rule: 1. 10 ≤ conductor cross-section (mm 2 ) ≤ 400, 2. 50 ≤ rated current-carrying capacity (A) ≤ 800.

[0126] d) Consistency rule:

[0127] ① The factory date is earlier than the commissioning date.

[0128] ② If the erection method is "mixed", then none of the overhead line length, cable line length, and total line length can be 0.

[0129] ③ If the model contains "YJ", then the overhead type is "insulated conductor".

[0130] ④ If the model starts with "LGY", then the overhead type is "bare conductor".

[0131] ⑤ Length (m) = (ending tower number - starting tower number) × 80.

[0132] ⑥ Whether the rural power grid matches the regional characteristics ('city center area', 'county urban area', 'urban area' do not match the rural power grid, 'rural area', 'town', 'township' match the rural power grid).

[0133] (2) According to the PMS data anomaly recognition rules in (1), use the API in the Spark platform to call the regular expression standard library to implement the corresponding regular expression method, so as to match the corresponding text and character data of PMS, and establish a PMS wire anomaly data screening algorithm model.

[0134] (3) Based on the model operation results obtained after executing the model in (2), count the anomalies in the archive data indicators (listing some: model judgment, rural power grid classification) as shown in Table 1 and Table 2; Table 1 is a schematic table of model judgment (partial), and Table 2 is a schematic table of rural power grid judgment (partial).

[0135] Table 1

[0136]

[0137] Table 2

[0138]

[0139] (4) According to the results in (3), revise the PMS data anomaly recognition rules to

[0140] a) Uniqueness rule: All wire parameters are unique and non-repetitive;

[0141] b) Integrity rule: All assessment fields must not be empty;

[0142] c) Accuracy rule: 1.10 ≤ wire cross-section (mm2) ≤ 400, 2.50 ≤ rated current-carrying capacity (A) ≤ 800;

[0143] d) Consistency rule:

[0144] ① The commissioning date is earlier than the registration date;

[0145] ② If the erection method is "mixed", then any one of the overhead line length, cable line length, and total line length shall not be 0;

[0146] ③ Based on most samples being correct and a small number of samples being wrong, find out the incorrect classification of wire types;

[0147] ④ Based on most samples being correct and a small number of samples being wrong, find out the incorrect classification of rural power grids;

[0148] ⑤ Length (m) ≤ (ending pole number - starting pole number) × 80;

[0149] ⑥The text of the starting tower pole does not match that of the ending tower pole, or there is no tower pole number, and the index value is extracted.

[0150] (5) Establish a statistical information error table as shown in Appendix 3.

[0151] Table 3

[0152]

[0153] 4. The specific process of Step 4 is as follows:

[0154] Based on the running results of the model in Step 3, call functions such as map, reduce, and drop using the structured API provided by Spark to correct the PMS wire file data:

[0155] (1) For duplicate PMS wire data, directly delete one and keep one record;

[0156] (2) For missing data, for categorical and numerical data, compare according to other attribute classifications. If there are data with the same type of attributes, they are grouped into one category. If not, the missing values of categorical data can be filled with "Unknown", and the numerical data can be filled with "0" or 'empty'.

[0157] (3) For misclassified data, such as incorrect wire type classification or incorrect rural power grid matching, batch correct the data according to the PMS wire data anomaly identification rules;

[0158] (4) For data that does not conform to the anomaly identification rules, correct it according to the standard format. If it cannot be corrected, add a mark to it;

[0159] Finally, a PMS wire data set with data governance completed is obtained.

[0160] Figure 3 This is the module connection diagram of the power equipment file data governance system in the embodiments provided by the present invention. As Figure 3 shown, this embodiment also provides a power equipment file data governance system, including:

[0161] A data understanding module for understanding the power equipment business and the file data of the power equipment business to obtain a data entry specification library;

[0162] A data extraction module for extracting power equipment file data from the power equipment data warehouse into the Spark parallel computing model according to the data entry specification library to obtain a resilient distributed dataset;

[0163] A model building module, which is used to call the regular expression standard library in the Spark parallel computing model to build an algorithm model for screening abnormal data of archives based on the data entry specification library;

[0164] A screening module, which is used to screen and count the elastic distributed dataset according to the algorithm model for screening abnormal data of archives, and obtain a dataset of archives to be corrected;

[0165] A data correction module, which is used to correct the dataset of archives to be corrected based on the API call function in the Spark parallel computing model, and obtain a dataset of archives with data governance completed.

[0166] Preferably, the data understanding module specifically includes:

[0167] An analysis unit, which is used to analyze the power equipment business to obtain business knowledge data; the business knowledge data includes: power equipment business structure data, data governance requirement data, and target completion data;

[0168] A judgment unit, which is used to judge the abnormal data set and the reasons for abnormal data in the archive data;

[0169] A specification library construction unit, which is used to build the data entry specification library based on the archive data, the business knowledge data, the abnormal data set, and the reasons for abnormal data.

[0170] Preferably, the data extraction module specifically includes:

[0171] An import unit, which is used to import the power equipment archive data from the power equipment data warehouse into the Spark parallel computing model according to the data entry specification library, and obtain an initial data set;

[0172] An inspection unit, which is used to perform integrity inspection on the initial data set to obtain the inspected elastic distributed data set.

[0173] The beneficial effects of the present invention are as follows:

[0174] (1) The present invention provides a complete process for governing power archive data, which can improve the quality of power archive data.

[0175] (2) Due to the high efficiency of data reading and model execution of the Spark computing platform in the present invention, the working efficiency of archive data governance can be improved.

[0176] (3) The present invention reduces the working difficulty of governing power equipment archive data, and provides an effective means for data quality screening and data rectification in the early stage of the application of power grid equipment archive data.

[0177] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0178] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A method for managing data of power equipment archives, characterized in that, Including: Conduct data understanding on the power equipment business and the archive data of the power equipment business to obtain a data entry specification library; According to the data entry specification library, extract power equipment archive data from the power equipment data warehouse into the Spark parallel computing model to obtain a resilient distributed dataset; Based on the data entry specification library, call the regular expression standard library in the Spark parallel computing model to establish an archive data anomaly screening algorithm model; According to the archive data anomaly screening algorithm model, screen and count the resilient distributed dataset to obtain an archive dataset to be corrected; Based on the API call function in the Spark parallel computing model, correct the archive dataset to be corrected to obtain an archive dataset with completed data governance; The step of, based on the data entry specification library, calling the regular expression standard library in the Spark parallel computing model to establish an archive data anomaly screening algorithm model includes: Referring to the data entry specification library, establish an anomaly recognition rule for power equipment archive data according to preset business requirements; According to the anomaly recognition rule for power equipment archive data, call the regular expression standard library provided by the Spark parallel computing model to establish the archive data anomaly screening algorithm model; After establishing the archive data anomaly screening algorithm model according to the anomaly recognition rule for power equipment archive data by calling the regular expression standard library provided by the Spark parallel computing model, it further includes: Use the cluster driver in the Spark parallel computing model to read the input tasks of the archive data anomaly screening algorithm model and distribute the input tasks to multiple executors for processing to obtain a running result; Based on the running result, count the anomaly data in the archive data metrics; Improve the anomaly recognition rule for power equipment archive data according to the running result and the anomaly data to obtain an improved anomaly recognition rule for power equipment archive data; Optimize the archive data anomaly screening algorithm model according to the improved anomaly recognition rule for power equipment archive data to obtain an optimized archive data anomaly screening algorithm model.

2. The method for governing power equipment file - type data according to claim 1, characterized in that, The step of conducting data understanding on the power equipment business and the archive data of the power equipment business to obtain a data entry specification library includes: Analyze the power equipment business to obtain business knowledge data; the business knowledge data includes: power equipment business structure data, data governance requirement data, and target completion data; Judge the anomaly data set and the reasons for the anomaly data in the archive data; Construct the data entry specification library based on the archive data, the business knowledge data, the anomaly data set, and the reasons for the anomaly data.

3. The method for governing power equipment archive data according to claim 1, wherein The step of, according to the data entry specification library, extracting power equipment archive data from the power equipment data warehouse into the Spark parallel computing model to obtain a resilient distributed dataset includes: According to the data entry specification library, import the power equipment archive data from the power equipment data warehouse into the Spark parallel computing model to obtain an initial data set; Conduct an integrity check on the initial data set to obtain the resilient distributed data set after the check.

4. The method for governing power equipment file data according to claim 1, wherein After optimizing the archive anomaly data screening algorithm model according to the improved power equipment archive data anomaly recognition rule to obtain an optimized archive anomaly data screening algorithm model, it further includes: Establish a statistical information error table based on the output result of the optimized archive anomaly data screening algorithm model.

5. The method for governing power equipment file data according to claim 1, characterized in that, The correction of the archive data set to be corrected based on the API call function in the Spark parallel computing model to obtain a completely data-governed archive data set includes: Based on the structured API call function provided by Spark, perform data deletion, data filling, and data correction on the archive data set to be corrected to obtain the completely data-governed archive data set; Among them, the steps of the data deletion include: Deduplicate and save the duplicate data in the archive data set to be corrected; Delete the missing data in the archive data set to be corrected according to the data understanding and the power equipment archive data anomaly recognition rule; The steps of the data filling include: Fill the missing value data in the archive data set to be corrected according to the missing value filling method; The steps of the data correction include: Batch correct the format error data in the archive data set to be corrected.

6. A data governance system for power equipment archive data, characterized in that, It includes: A data understanding module for understanding the power equipment business and the archive data of the power equipment business to obtain a data entry specification library; A data extraction module for extracting power equipment archive data from the power equipment data warehouse into the Spark parallel computing model according to the data entry specification library to obtain a resilient distributed data set; A model establishment module for establishing an archive anomaly data screening algorithm model by calling the regular expression standard library in the Spark parallel computing model based on the data entry specification library; A screening module for screening and counting the resilient distributed data set according to the archive anomaly data screening algorithm model to obtain an archive data set to be corrected; A data correction module for correcting the archive data set to be corrected based on the API call function in the Spark parallel computing model to obtain a completely data-governed archive data set; The establishment of an archive anomaly data screening algorithm model by calling the regular expression standard library in the Spark parallel computing model based on the data entry specification library includes: Refer to the data entry specification library and establish a power equipment archive data anomaly recognition rule according to the preset business requirements; According to the power equipment archive data anomaly recognition rule, call the regular expression standard library provided by the Spark parallel computing model to establish the archive anomaly data screening algorithm model; After establishing the screening algorithm model for the abnormal archive data by invoking the regular expression standard library provided by the Spark parallel computing model according to the abnormal identification rule for the power equipment archive data, the following steps are further included: Use the cluster driver in the Spark parallel computing model to read the input tasks of the screening algorithm model for the abnormal archive data, and distribute the input tasks to multiple executors for processing to obtain the running results; Based on the running results, count the abnormal data in the archive data metrics; Improve the abnormal identification rule for the power equipment archive data according to the running results and the abnormal data to obtain the improved abnormal identification rule for the power equipment archive data; Optimize the screening algorithm model for the abnormal archive data according to the improved abnormal identification rule for the power equipment archive data to obtain the optimized screening algorithm model for the abnormal archive data.

7. The power equipment file data governance system according to claim 6, characterized in that The data understanding module specifically includes: An analysis unit for analyzing the power equipment business to obtain business knowledge data; the business knowledge data includes: power equipment business structure data, data governance requirement data, and target completion data; A judgment unit for judging the abnormal data set and the reasons for the abnormal data in the archive data; A specification library construction unit for constructing the data entry specification library based on the archive data, the business knowledge data, the abnormal data set, and the reasons for the abnormal data.

8. The power equipment file data governance system according to claim 6, characterized in that The data extraction module specifically includes: An import unit for importing the power equipment archive data from the power equipment data warehouse into the Spark parallel computing model according to the data entry specification library to obtain an initial data set; An inspection unit for performing integrity inspection on the initial data set to obtain the inspected resilient distributed data set.

Citation Information

Patent Citations

  • An error correction method of power network equipment archives data based on machine learning

    CN109472293A