A power grid system data asset management system

Through the power grid system data asset management system, the data asset professional definition module, data blood model building module and AI optimization engine module are used to solve the problem of unified processing and model description of data of different professionals, and improve data governance efficiency.

CN113901028BActive Publication Date: 2025-08-05GUANGDONG POWER GRID CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111187691.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-12
Publication Date
2025-08-05
Estimated Expiration
2041-10-12

AI Technical Summary

Technical Problem

The existing data governance technology lacks unified processing of data from different specialties, and there is no unified modeled description, resulting in low data governance efficiency.

Method used

It provides a power grid system data asset management system, including data asset professional definition module, data blood model construction module and AI optimization engine module. Through professional definition, data correlation analysis and AI optimization model, the processing process of data assets is optimized.

Benefits of technology

It improves the accuracy and management quality of data professional definition, realizes unified processing and modeled description of data of different professions, and improves data governance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113901028B_ABST
    Figure CN113901028B_ABST
Patent Text Reader

Abstract

The present application discloses a power grid system data asset management system, including a data asset professional definition module for professionally defining the accessed target data assets according to preset professional categories to obtain a professional definition result; a data lineage model construction module for performing preset analysis operations on the target data assets to obtain data association information, and constructing a data lineage model based on the data association information. The preset analysis operations include data regression analysis, preset feature analysis, span analysis, etc.; an AI optimization engine module for training a preset AI optimization model according to a preset sample library and the data lineage model to obtain a target AI optimization model, and the target AI optimization model is used to screen the target data assets through preset screening rules to optimize the result of the professional definition. The present application can solve the technical problems that the existing data governance technologies lack unified processing for different professional data and do not have a unified model-based description, resulting in low data governance efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a data asset management system for a power grid system. Background Art

[0002] The rapid development of China's economy and the continuous improvement of people's living standards have led to an increasing demand for electricity and a higher requirement for data reliability. As a core part of the power grid, substations have continuously developed and improved data governance content, explored an efficient data governance framework, built a full-process intelligent one-stop data governance platform, gradually solved problems such as the comprehensiveness, accuracy, integrity, consistency, and timeliness of data, improved the level and quality of data asset management, enhanced data service capabilities, and provided data support for the whole bank's data management, product innovation, digital transformation, etc.

[0003] However, existing data governance solutions still require different patrol and maintenance centers to control various types of operation tasks for the same device according to different specialties, which is easy to miss tasks and cause repeated power outages; moreover, work tickets at different stages after the execution of operation tasks are filled in by different personnel, resulting in a too large difference in the quality of work tickets; and there is no unified model description for the records of operation tasks, leading to low data governance efficiency. Summary of the Invention

[0004] This application provides a data asset management system for a power grid system, which is used to solve the technical problem that existing data governance technologies lack unified processing for data of different specialties and have no unified model description, resulting in low data governance efficiency.

[0005] In view of this, in the first aspect of this application, a data asset management system for a power grid system is provided, including: a data asset professional definition module, a data lineage model construction module, and an AI optimization engine module;

[0006] The data asset professional definition module is used to perform professional definition on the accessed target data assets according to preset professional categories to obtain a professional definition result. The preset professional categories include operation protection specialty, high voltage specialty, maintenance specialty, instrument specialty, and automation specialty;

[0007] The data lineage model construction module is used to perform preset analysis operations on the target data assets to obtain data association information, and construct a data lineage model according to the data association information. The preset analysis operations include data regression analysis, preset feature analysis, span analysis, and qualitative retrieval information management;

[0008] The AI optimization engine module is used to train a preset AI optimization model based on a preset sample library and the data lineage model to obtain a target AI optimization model, and the target AI optimization model is used to screen the target data assets through preset screening rules to optimize the results of professional definitions.

[0009] Preferably, it further includes: a data acquisition module;

[0010] The data acquisition module is used to obtain power grid data information in the inner data framework of the power grid system and organize the power grid data information into target data assets, and the power grid data information includes fault information, maintenance information, and overhaul information.

[0011] Preferably, it further includes: a preprocessing module;

[0012] The preprocessing module is used to perform preliminary data cleaning, data parsing, data filling, and data annotation operations on the target data assets.

[0013] Preferably, the data lineage model construction module is specifically used for:

[0014] Performing regression analysis on the mapping relationship between the classified target data assets and substation data and attributes to obtain attribute-associated data assets;

[0015] Extracting data feature formulas from the target data assets to obtain a preset feature set, and the data feature formulas are used to describe power peaks and power fluctuation amplitudes;

[0016] Marking the data information exceeding the preset span in the target data assets after clustering difference analysis to obtain large-span data assets;

[0017] Extracting multiple key search terms from the target data assets and setting search specification information based on preset search permissions and the key search terms;

[0018] Constructing a data lineage model according to the data association information, where the data association information includes attribute-associated data assets, the preset feature set, the large-span data assets, and the search specification information.

[0019] Preferably, the preset screening rules include: preset primary rules, preset secondary rules, and preset tertiary rules;

[0020] The preset primary rules are preset cleaning conditions and preset business conditions set for operation data;

[0021] The preset secondary rules are primary screening mechanisms set according to the matching rate of the first key characters within the primary permission range;

[0022] The preset three-level rule is an advanced screening mechanism set according to the matching rate of the second key characters within the highest authority range.

[0023] From the above technical solutions, it can be seen that the embodiments of this application have the following advantages:

[0024] In this application, a power grid system data asset management system is provided, including: a data asset professional definition module, a data lineage model construction module, and an AI optimization engine module; the data asset professional definition module is used to perform professional definition on the accessed target data assets according to preset professional categories to obtain a professional definition result, and the preset professional categories include operation protection, high voltage, maintenance, instrument, and automation; the data lineage model construction module is used to perform preset analysis operations on the target data assets to obtain data association information, and construct a data lineage model according to the data association information. The preset analysis operations include data regression analysis, preset feature analysis, span analysis, and qualitative retrieval information management; the AI optimization engine module is used to train a preset AI optimization model according to a preset sample library and the data lineage model to obtain a target AI optimization model, and the target AI optimization model is used to screen the target data assets through preset screening rules to optimize the result of the professional definition.

[0025] The power grid system data asset management system provided by this application performs professional definition on all target data assets, which can be regarded as a classification of a data processing type; then, through specific analysis operations on the target data assets, the association relationships between different data are extracted, and a data lineage model that can uniformly describe the data relationships is constructed. According to this model and the sample library, an AI optimization model for optimizing data assets can be trained, which can not only improve the accuracy of data professional definition but also ensure the quality of data management. Therefore, this application solves the technical problem that the existing data governance technology lacks unified processing for different professional data and has no unified model description, resulting in low data governance efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a schematic structural diagram of a power grid system data asset management system provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0028] For ease of understanding, please refer to Figure 1 An embodiment of a power grid system data asset management system provided by this application includes: a data asset professional definition module 101, a data lineage model construction module 102, and an AI optimization engine module 103.

[0029] The data asset professional definition module 101 is used to perform professional definition on the accessed target data assets according to preset professional categories to obtain a professional definition result. The preset professional categories include operation protection, high voltage, maintenance, instrumentation, and automation.

[0030] The data assets specifically include design drawings, contract orders, and various services carried on any documents related to usage; they also include technical content of each department, equipment data, and department operation and maintenance data; they further include company management platform data, pictures, videos, and platform management accounts.

[0031] The essence of data asset professional definition is to establish the owners of data assets, usually based on the substation asset management platform, which is used to authorize users to upload and download relevant professional reserve knowledge and update the operation and maintenance situation after current operations in a timely manner; then, according to the operation departments of the substation, a secondary connection database of the platform is set up, that is, a database of preset professional types; that is, the definition of the owners of data assets is carried out according to the substation profession.

[0032] The data lineage model construction module 102 is used to perform preset analysis operations on the target data assets to obtain data association information, and construct a data lineage model according to the data association information. The preset analysis operations include data regression analysis, preset feature analysis, span analysis, and qualitative retrieval information management.

[0033] Further, the data lineage model construction module 102 is specifically used for:

[0034] Performing regression analysis on the mapping relationship between substation data and attributes of the classified target data assets to obtain attribute-associated data assets;

[0035] Extracting data feature formulas from the target data assets to obtain a preset feature set. The data feature formulas are used to describe the power peak and power fluctuation amplitude;

[0036] Performing marking processing on the data information exceeding the preset span in the target data assets after clustering difference analysis to obtain large-span data assets;

[0037] Extracting multiple key retrieval words from the target data assets and setting retrieval specification information based on the preset retrieval permissions and key retrieval words;

[0038] Construct a data lineage model based on data association information, which includes attribute-associated data assets, preset feature sets, large-span data assets, and retrieval specification information.

[0039] The data lineage model is used to systematically describe the association relationships between data at multiple levels. To ensure the accuracy of the description of the association relationships between data, various different data analyses need to be performed on the data, such as regression analysis, preset feature analysis, span analysis, and qualitative retrieval information management, etc.

[0040] Among them, regression analysis can generate a mapping relationship between the substation data in the substation database of the target data asset and the attribute values in terms of time. This mapping relationship is a function of real-valued predictive variables and is used to describe the dependency relationship between the substation data and the attributes. The attribute-associated data assets include substation data, attribute values, and the description of the mapping relationship between the two.

[0041] Preset feature analysis is the process of extracting data feature expressions. The data relationships are described using feature expressions, mainly referring to some data with strong regular correlations, such as power peaks and power fluctuation amplitudes, etc.; these features represent the overall characteristics of the data assets.

[0042] Before performing span analysis, it is also necessary to perform clustering analysis on the target data asset, that is, to classify the target data asset according to similarity and difference. The purpose is to make the similarity between data belonging to the same category as large as possible, and the similarity between different categories as small as possible; the difference can be adaptively configured according to requirements to reduce null value fields. For example, if the similarity reaches more than 90%, and the similarity of key data reaches more than 60%, it can be determined as associated data or data with a lineage relationship, which is used to describe the proximity of the relationships between data items in the data asset. Then, the information span between data can be marked and processed, and large-span data assets can be obtained by specifically marking the data assets with a larger information span screened according to the preset span.

[0043] In qualitative retrieval management analysis, the target information can be retrieved in the model according to the set key retrieval terms. The specific key retrieval terms can be set according to needs and are not limited here; the key retrieval terms can mark relevant content; and retrieval records and user information that cannot be deleted can be set according to security needs, or retrieval permissions can be set, etc.

[0044] In addition to performing the above data analyses, other forms of in-depth mining can also be performed according to needs, mainly to describe more of the association information between data, so that the data lineage model constructed based on the obtained data association information is more accurate and reliable.

[0045] The AI optimization engine module 103 is used to train a preset AI optimization model based on a preset sample library and a data lineage model to obtain a target AI optimization model. The target AI optimization model is used to screen target data assets through preset screening rules to optimize professionally defined results.

[0046] The basic model framework of the preset AI optimization model is a neural network model, composed of multiple feature extraction network layers. Data input is processed through different neurons to perform function calculations, resulting in different expression feature vectors. Target data prediction is then achieved based on these feature vectors. The AI optimization model primarily optimizes data assets, screening the data within the target data assets to optimize the professionally defined results of the target data assets. The preset sample library is also a data asset acquired based on preset screening rules, obtained by integrating multiple sub-sample libraries on the data asset platform. The process of training the preset AI optimization model based on the preset sample library and data lineage model is a process of continuously screening data assets. The model achieves data standardization through continuous data input and output, allowing for human-readable data. It is built on the sample library and assists the data lineage model in data association. The associated data can still be compared with the sample library for association training, thereby improving data uniformity.

[0047] Model training uses the sample library and the panoramic lineage model as datasets for exploratory data analysis. The model can perform operations such as pivoting, grouping, and filtering on the data. Data processing involves various checks and scrutiny of the data, including correcting missing values, spelling errors, normalizing / standardizing values for comparability, and transforming data for operational purposes. Data splitting can also be performed if necessary. For example, to simulate new types of data, the original data can be split into two parts (sometimes called a training / test split). Typically, the first part is a larger subset of the data, serving as the training set (e.g., 80% of the original data), while the second part is typically a smaller subset, serving as the testing set (the remaining 20%). Next, a predictive model is built using the training set, and this trained model is then tested on the test set. The best model is selected based on its performance on the test set. To obtain the best model, hyperparameter optimization can also be performed. Another approach is to perform a training-validation-test split. A common data splitting approach is to divide the data into three parts: training, validation, and testing. The training set is used to build a predictive model, while the validation set is evaluated. Predictions are made based on this data, allowing for model parameter tuning and selecting the best-performing model based on the validation set results. The test set is not involved in any model building or preparation. Therefore, the test set can truly serve as new, unknown data. Cross-validation is then performed, with existing data reused for training. Finally, a mature training model is established, which can be used to repeatedly train and cleanse the data in the sample library and panoramic lineage model, i.e., screen the data.

[0048] Regarding the fitting of the model, if the learning of data assets is too precise, the model will also learn the associated data features in non-keywords, which will affect the subsequent test results and lead to poor generalization ability of the model. If the feature capture of data assets is shallow, the model does not read the associated information between keywords and cannot efficiently fit the data to achieve optimized screening of the data, which is the under-fitting state of the model. When the model is over-fitted, manual intervention needs to be carried out through methods such as data augmentation or network pruning; the method of data augmentation includes adding random noise, and data augmentation can also be processed through some existing other methods. At the same time, the retrieval surface and repetition rate of keywords in the sample library can also be improved. The network pruning method generally adjusts the training efficiency of the model by setting a pruning coefficient and cutting the network operators participating in each training according to the pruning coefficient. The specific network pruning scheme can be obtained according to the existing technology and will not be elaborated here.

[0049] Furthermore, it further includes: a data acquisition module 104;

[0050] The data acquisition module is used to acquire power grid data information in the inner data framework of the power grid system and organize the power grid data information into target data assets. The power grid data information includes fault information, maintenance information, and overhaul information.

[0051] The data is transmitted to the processing center through the near and far transmission networks and enters the inner data processing framework server for comprehensive processing of the data. The framework server has independent processing devices such as processors and is established in the substation area management center for direct access to the data. In addition to including information such as fault information, other types of data information can also be acquired for the power grid data information, which is not limited here. The power grid data information can be either static data information or dynamic data information.

[0052] Furthermore, it further includes: a preprocessing module 105;

[0053] The preprocessing module is used to perform preliminary data cleaning, data parsing, data filling, and data annotation operations on the target data assets.

[0054] Preliminary data cleaning can screen out invalid data with format errors and inconsistent data types in the target data assets through preset cleaning rules to reduce the data volume. Data parsing also proposes data that does not meet business or technical requirements through preset business rules. Data filling can fill in the relevant tables, null fields, or key values abandoned in the target data assets through RPA and AI intelligent operations. The data annotation operation annotates the data according to different situations so that each data has a specific label for subsequent data analysis.

[0055] Further, the preset screening rules include: preset primary rules, preset secondary rules, and preset tertiary rules;

[0056] The preset primary rules are preset cleaning conditions and preset business conditions set for the operation data;

[0057] The preset secondary rules are primary screening mechanisms set according to the matching rate of the first key characters within the scope of primary permissions;

[0058] The preset tertiary rules are advanced screening mechanisms set according to the matching rate of the second key characters within the scope of the highest permissions.

[0059] In addition to the preset cleaning conditions and preset business conditions, the main part of the preset screening rules is the key character screening mechanism. The key characters include keywords and key symbols. The preset tertiary rules need to be generated according to the superior review and are summarized from the personal experience, methods, and means of the authorized users; the preset secondary rules need to conform to the enterprise rule settings and can be generated after the peer review; in addition, when conflicts occur among the preset primary, secondary, and tertiary rules, the higher one is selected to ensure the stability of the data screening process according to the preset screening rules.

[0060] The power grid system data asset management system provided by the embodiments of this application can professionally define all target data assets, which can be regarded as a type of data processing classification; then, by performing specific analysis operations on the target data assets, the association relationships between different data can be extracted, and a data lineage model that can uniformly describe the data relationships can be constructed. According to this model and the sample library, an AI optimization model for optimizing data assets can be trained, which can not only improve the accuracy of data professional definition but also ensure the quality of data management. Therefore, the embodiments of this application solve the technical problems of the existing data governance technology, which lacks unified processing for different professional data and has no unified model-based description, resulting in low data governance efficiency.

[0061] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other forms.

[0062] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0063] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0064] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to execute all or part of the steps of the methods described in various embodiments of this application through a computer device (which can be a personal computer, a server, or a network device, etc.). The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (English full name: Read-Only Memory, English abbreviation: ROM), random access memories (English full name: Random Access Memory, English abbreviation: RAM), magnetic disks or optical discs and other various media that can store program codes.

[0065] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. A power grid system data asset management system, characterized in that: include: Data asset professional definition module, data lineage model construction module and AI optimization engine module; The data asset professional definition module is used to professionally define the accessed target data assets according to preset professional categories to obtain professional definition results. The preset professional categories include operation protection, high voltage, maintenance, instrumentation, and automation. The data lineage model construction module is used to perform preset analysis operations on the target data assets to obtain data association information and construct a data lineage model based on the data association information. The preset analysis operations include data regression analysis, preset feature analysis, span analysis, and qualitative retrieval information management. The data lineage model construction module is specifically used to: Performing a regression analysis on the mapping relationship between substation data and attributes on the classified target data assets to obtain attribute-associated data assets; Extracting a data feature formula from the target data asset to obtain a preset feature set, wherein the data feature formula is used to describe power peak value and power fluctuation amplitude; Marking the data information exceeding a preset span in the target data assets after cluster difference analysis to obtain large-span data assets; Extracting multiple key search terms according to the target data asset, and setting search specification information based on preset search permissions and the key search terms; Constructing a data lineage model based on the data association information, wherein the data association information includes attribute-related data assets, the preset feature set, the large-span data assets, and the retrieval specification information; The AI optimization engine module is used to train a preset AI optimization model based on a preset sample library and the data lineage model to obtain a target AI optimization model. The target AI optimization model is used to screen the target data assets through preset screening rules to optimize professionally defined results.

2. The power grid system data asset management system according to claim 1, characterized in that: Also includes: Data acquisition module; The data acquisition module is used to acquire power grid data information in the inner data framework of the power grid system and organize the power grid data information into target data assets. The power grid data information includes fault information, maintenance information and repair information.

3. The power grid system data asset management system according to claim 1, characterized in that: Also includes: Preprocessing module; The pre-processing module is used to perform preliminary data cleaning, data analysis, data filling and data labeling operations on the target data assets.

4. The power grid system data asset management system according to claim 1, characterized in that: The preset screening rules include: preset first-level rules, preset second-level rules and preset third-level rules; The preset first-level rules are preset cleaning conditions and preset business conditions set for the operation data; The preset secondary rule is a primary screening mechanism set according to the matching rate of the first key character within the primary authority range; The preset third-level rule is an advanced screening mechanism set according to the matching rate of the second key character within the highest authority range.

Citation Information

Patent Citations

  • Integrated power grid resource model and construction and maintenance method

    CN104616101A

  • Power grid operation data analysis evaluation and report system

    CN108764683A