Electric power asset intelligent classification and grading method, system and equipment based on large model, and medium

Through the intelligent classification and grading method of power data assets based on large models, the problems of low efficiency and poor accuracy in the power industry are solved, automated data detection and grading are realized, data processing capabilities and security are improved, and adaptability to dynamic business scenarios is supported.

CN120429692APending Publication Date: 2025-08-05HUBEI CENT CHINA TECH DEV OF ELECTRIC POWER
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510553227.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing technology has problems in the power industry with low efficiency, poor accuracy, difficulty in adapting to dynamic business scenarios, and insufficient unstructured data processing capabilities, and lack of a unified management framework for multi-source heterogeneous data, resulting in high data security compliance risks and low classification and grading efficiency.

Method used

The intelligent classification and grading method of power data assets based on large models is adopted, including data collection, preprocessing, sensitive feature marking, sensitive data template construction and recording, and automated identification and grading are combined with machine learning algorithms to establish a unified data classification and grading system.

Benefits of technology

It realizes automatic data detection, sensitive data discovery, classification and grading template management and automated tagging, improves data processing capabilities and dynamic adaptability, supports batch or manual grading, and improves the efficiency and security of data asset management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429692A_ABST
    Figure CN120429692A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent classification and grading method, system and equipment for electric power assets based on a large model and a medium, and the method comprises the steps: obtaining data in a data storage device, and extracting field attributes and field contents; cleaning, de-noising and standardizing the acquired data, and then constructing structured and unstructured data sets; performing classification and grading marking on sensitive data in fields, library tables and descriptions in the data set according to the sensitivity degree through a preset grading algorithm; constructing the marked sensitive data to obtain a sensitive data template; inputting the sensitive data template into a data query grading system, and detecting, grading and recording data in a data storage device through the data query grading system; and collecting the recorded sensitive data to obtain a sensitive data log. The method has the sensitive data discovery capability, the data asset grading capability and the classification and grading template management capability, and can improve the data processing capability and the dynamic adaptability in combination with a large model technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method, system, device and medium for intelligent classification and grading of power assets based on a large model. Background Art

[0002] As the power industry accelerates its digital transformation, data volumes and types are surging and becoming increasingly complex. Traditional classification and grading methods, which rely on manual rules, suffer from low efficiency, poor accuracy, and difficulty adapting to dynamic business scenarios. Furthermore, existing technologies are insufficiently capable of processing unstructured data and lack a unified management framework for multi-source, heterogeneous data. This leads to high data security and compliance risks and inefficient classification and grading.

[0003] Existing solutions often rely on static rule bases for sensitive data identification, which are unable to adapt to semantically complex business scenarios. Classification and grading template updates rely on manual labor and lack automated feedback mechanisms. Therefore, there is an urgent need for an efficient and intelligent power data asset classification and grading method that combines large-scale modeling technology to enhance data processing capabilities and dynamic adaptability. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to overcome the defects of the existing technology and provide a method, system, equipment and medium for intelligent classification and grading of power data assets based on a large model.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: A method for intelligent classification and grading of power data assets based on a large model includes the following steps: Data collection, obtaining data from data storage devices and extracting field attributes and field contents; Preprocessing: Cleaning, denoising, and standardizing the acquired data to construct structured and unstructured data sets; Sensitive feature marking: using a preset classification algorithm to classify and grade sensitive data in the fields, tables, and descriptions of the dataset according to their sensitivity; Sensitive data template construction: construct the sensitive data template from the marked sensitive data; Sensitive data recording, inputting the sensitive data template into the data query and classification system, and detecting, classifying and recording the data in the data storage device through the data query and classification system; Sensitive data collection: collect the recorded sensitive data to obtain sensitive data logs.

[0006] Furthermore, the data collection includes the following sub-steps: Load the tags and rules of the power industry template; Connect to the Hive database and obtain information about all databases and tables in it; Traverse all the above databases and data table information, and extract all field attributes and field contents therein, wherein the field attributes include field name, table name, and description; Match the extracted field attributes and field contents with the tags and rules in the power industry template; If the field attributes and field content match the rules of the corresponding tag, the classification and grading information of the field is considered to be the classification and grading information corresponding to the tag and is recorded.

[0007] Furthermore, the sensitive feature marking includes the following sub-steps: Field classification: Based on the meaning of the fields, all fields contained in the graded database table are classified into different levels to determine the security level N of each field; Classify the database tables according to their business application scenarios and provision methods, and determine their security level M. Comprehensive grading: the maximum value of the field grading Nmax=MAX{N1,N2…Nn} and the library table grading M are taken as the final grading L=MAX{M,Nmax} of the resource to be graded.

[0008] Furthermore, the method further includes the following steps: Learning the steps of labeling sensitive features through machine learning and generating built-in rules; The machine learning algorithm identifies and labels sensitive features of new data in the data storage device through the built-in rules.

[0009] Furthermore, a large-scale model-based intelligent classification and grading system for power data assets is characterized by including the following modules: The data acquisition module is used to obtain data from the data storage device and extract field attributes and field contents; The preprocessing module is used to clean, denoise and standardize the acquired data to construct structured and unstructured data sets; A sensitive feature marking module is used to classify and grade the sensitive data in the fields, database tables, and descriptions in the data set according to their sensitivity using a preset grading algorithm; A sensitive data template construction module is used to construct a sensitive data template from the marked sensitive data; A sensitive data recording module, configured to input the sensitive data template into a data query and grading system, and detect, grade and record the data in the data storage device through the data query and grading system; The sensitive data collection module is used to collect the recorded sensitive data and obtain sensitive data logs.

[0010] Furthermore, it also includes a machine learning module, which includes a data acquisition layer, a basic component layer and a result fusion layer.

[0011] An electronic device comprises a memory, a processor and a computer program stored in the memory and operable on the processor, wherein the steps of the method are implemented when the processor executes the program.

[0012] A computer-readable storage medium stores a computer program, which implements the steps of the method when executed by a processor.

[0013] Compared with the prior art, the present invention has the following beneficial effects: 1. Automatic data detection capability: Identify data assets in traditional and big data environments, including storage location, size, attributes, and other information; 2. Sensitive data discovery capability: Based on the sensitive data feature library, it can identify sensitive information in database tables and support customized sensitive data features; 3. Data asset classification capability: Classify and label sensitive data, supporting batch or manual classification. Statuses include confirmed, pending, and unknown. 4. Classification and grading template management capabilities: Establish a unified data classification and grading labeling system to support template revision and adjustment; 5. Automatic classification and grading capabilities: Combine automation technology to automatically mark the business category and security level of fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 This is a flow chart of a method for intelligent classification and grading of power data assets based on a large model according to an embodiment of the present invention; Figure 2 is a flow chart of the sub-steps of data collection according to an embodiment of the present invention; Figure 3 is a classification and grading flow chart of an embodiment of the present invention; Figure 4 is a flowchart of the sub-steps of sensitive feature marking according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a large-scale model-based intelligent classification and grading system for power data assets according to an embodiment of the present invention; Figure 6 This is a system module view of an embodiment of the present invention; Figure 7 is a machine learning framework diagram of an embodiment of the present invention; Figure 8This is a diagram of an intelligent classification algorithm based on field names, table names, and descriptions according to an embodiment of the present invention; Figure 9 is a structural block diagram of a similarity calculation module according to an embodiment of the present invention; Figure 10 This is a diagram of a neural network classification algorithm integrating multi-dimensional information coding according to an embodiment of the present invention. DETAILED DESCRIPTION

[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0016] like Figure 1 As shown, the present invention provides a method for intelligent classification and grading of power data assets based on a large model, comprising the following steps: S1, data collection, obtains data from the data storage device and extracts field attributes and field contents.

[0017] S2, preprocessing, cleans, denoises and standardizes the acquired data to construct structured and unstructured datasets; The data preprocessing module is mainly responsible for processing the data collected in the previous step, including: removing special tag symbols, data word segmentation, character segmentation, camel case naming resolution, removing stop words, etc., to retain as much useful information as possible in the data and eliminate noise data.

[0018] S3, sensitive feature marking, classifies and grades the sensitive data in the fields, libraries, tables, and descriptions in the dataset according to their sensitivity through a preset grading algorithm.

[0019] S4, sensitive data template construction, constructs the sensitive data template from the marked sensitive data.

[0020] S5, sensitive data recording, inputting the sensitive data template into the data query and classification system, and detecting, classifying and recording the data in the data storage device through the data query and classification system.

[0021] S6, sensitive data collection, collects the recorded sensitive data to obtain sensitive data logs.

[0022] In this embodiment, the preset grading algorithm is to comprehensively grade the database tables according to a certain algorithm. It makes specific requirements for the structured data grading method based on the "Electric Power Industry Data Classification and Grading Specification", and comprehensively grades the data from both the database table and field aspects. The structured data level is divided into a database table bit unit, and the classification is implemented by comprehensive judgment based on the meaning of the field and the business application scenario of the database table. The classification of power data is generally based on the following dimensions: Business subject area: Power generation side: four categories: production, safety, consumption, and economy; Power grid enterprises: production domain (such as equipment operating conditions and system logs), marketing domain (customer information and electricity consumption data), management domain (supply chain and financial data), etc.

[0023] Power generation type: Data for different power generation types such as thermal power, hydropower, wind power, photovoltaic power, and nuclear power need to be classified separately. Data source: Internal data: data generated by the company's own systems; External data: government shared data, user-side data, etc.

[0024] Power data is usually divided into three levels, based on the degree of potential harm: Three-level classification (common in the industrial field) Level 3: May cause a particularly serious safety incident, affecting national security or the national economy; Level 2: Causing significant economic losses or industry-wide impacts; Level 1: Only affects local business or causes minor economic losses.

[0025] The data query classification system adopts existing technology.

[0026] like Figure 2-3 As shown, further, the data collection includes the following sub-steps: S11, load the tags and rules of the power industry template; S12: Connect to the Hive database and obtain information about all databases and tables in it; S13, traverse all the above databases and data table information, and extract all the field attributes and field contents therein; S14, matching the retrieved field attributes and field contents with the tags and rules in the power industry template; S15: If the field attributes and field content match the rules of the corresponding tag, the classification and grading information of the field is considered to be the classification and grading information corresponding to the tag and is recorded.

[0027] In this embodiment, this step is to scan the sensitive data on the pages of the internal website to discover the distribution of the sensitive data.

[0028] Take Hive classification and grading scanning as an example: 1) Load the power industry standard template, including the classification and grading information in the industry standard, the correspondence between the classification and grading and the labels, and the rule information corresponding to the labels.

[0029] 2) Connect to Hive data assets and connect to the Hive database through Java JDBC.

[0030] 3) Get all Hive database and data table information.

[0031] 4) Traverse all databases and tables and extract all field attributes and field contents from the data tables. When extracting field contents, use select column from tablename where limit N, where N is the configured number of samples.

[0032] 5) The extracted field attributes and field contents are matched against the tags and rules in the industry standard. If a match is found, the classification and grading information for the field is considered to be the classification and grading information corresponding to the tag. This information is recorded in the result table and summary table of the database.

[0033] 6) Then continue to loop to retrieve fields and repeat steps 4 and 5 until all databases, tables, and field information are retrieved.

[0034] The SQL commands corresponding to the database, table, and field are as follows: -Get all libraries: show databases; -Get all tables under the library: show tables; -Get table structure: desc formatted tablename; The scanning methods of other databases are similar to those of Hive.

[0035] like Figure 4 As shown, further, the sensitive feature marking includes the following sub-steps: S31, field classification, based on the meaning of the field, classify all fields contained in the classification table into different levels and determine the security level N of each field; S32, library table classification: classify the library tables to be classified according to their business application scenarios and provision methods, and determine the security level M of the library tables; S33, comprehensive classification, the field classification maximum value Nmax=MAX{N1,N2…Nn} and the library table classification M, the maximum value of the two is taken as the final classification L=MAX{M,Nmax} of the resource to be classified.

[0036] Furthermore, the method further includes the following steps: Learning the steps of labeling sensitive features through machine learning and generating built-in rules; The machine learning algorithm identifies and labels sensitive features of new data in the data storage device through its built-in rules.

[0037] In this embodiment, it is difficult to identify data only through keywords and regular expressions, so it is necessary to use machine learning algorithms to perform data identification.

[0038] Detect field values, database table names, data types, and sensitive categories.

[0039] The data query and classification system retrieves table and field attributes and data information from the database and matches them with industry standard tags and rules. If a tag in the industry standard is matched, the classification and classification information of the field is output as the information corresponding to the tag.

[0040] If a match cannot be found in the data query and classification system, a prediction is made using the table name and field name model in machine learning. If the predicted similarity exceeds the threshold, the classification and classification of the field is considered to be the classification and classification information predicted by the model. If it is less than or equal to the threshold, the classification is passed to the neural network classification module in machine learning. Neural network classification models are primarily learned from field data and primarily predict field data. If the neural network can predict the classification and classification of the field, the classification and classification of the field is considered to be the classification and classification predicted by the neural network model. Otherwise, the classification and classification of the field cannot be confirmed.

[0041] like Figure 5 As shown in the figure, a large-scale model-based intelligent classification and grading system for power data assets includes: The data acquisition module 100 is used to obtain data from the data storage device and extract field attributes and field contents; The pre-processing module 200 is used to clean, denoise and standardize the acquired data to construct structured and unstructured data sets; Sensitive feature marking module 300, used to classify and grade sensitive data in fields, database tables, and descriptions in the data set according to their sensitivity using a preset grading algorithm; A sensitive data template construction module 400 is used to construct a sensitive data template from the marked sensitive data; The sensitive data recording module 500 is used to input the sensitive data template into the data query and classification system, and detect, classify and record the data in the data storage device through the data query and classification system; The sensitive data collection module 600 is used to collect the recorded sensitive data to obtain a sensitive data log.

[0042] like Figure 7 As shown, further, a machine learning module is also included, which includes a data acquisition layer, a basic component layer and a result fusion layer.

[0043] In this embodiment, in order to build and update the intelligent algorithm model we need, it is necessary to continuously update samples, train models, and adjust parameters. At the same time, machine learning and deep learning also involve many Python third-party libraries and specific Python versions. Therefore, a more convenient way is to use Anaconda to manage all required open source Python libraries and then package the corresponding Anaconda environment and the corresponding code and deploy them to the corresponding environment.

[0044] The machine learning module architecture consists of three layers: the data collection layer, the basic component layer, and the result fusion layer. The data collection layer collects and processes data for intelligent algorithms. The basic component layer provides algorithmic support based on the Anaconda environment and separate algorithm modules. It mainly includes two categories of algorithms, each of which includes some common process modules for machine learning. The result fusion layer is responsible for fusing the results of these two categories of algorithms using different strategies, ultimately connecting them with the results of the data query and grading system rules.

[0045] The intelligent algorithm solution for data classification and grading consists of two modules: machine learning model construction and algorithm self-learning. The algorithm model construction is primarily divided into two categories: an intelligent classification algorithm based on field names, table names, and descriptions; and a neural network classification module that incorporates multidimensional information coding (primarily based on field content information). These two algorithms comprehensively classify and grade data based on multiple dimensions, including field names, table names, descriptions, and field content. The intelligent classification algorithm based on field names, table names, and descriptions consists of six submodules: sample collection module, data preprocessing module, feature selection module, model training module, model evaluation module, and model update module. It is responsible for providing accurate and reliable algorithm models based on field names, table names, and descriptions for preliminary screening. The neural network classification algorithm that incorporates multidimensional information coding consists of five submodules: sample collection module, data preprocessing module, model training module, model evaluation module, and model update module. It is responsible for providing accurate and reliable algorithm models. After the algorithm is deployed in the product, it supports self-learning. Manual corrections and new data and labels can be made based on the algorithm's recognition results. Efficient model self-learning is achieved through similarity analysis, enabling continuous iteration and improvement.

[0046] like Figure 3 、 8As shown in Figure 9, intelligent algorithms are further divided into ① intelligent classification algorithms based on field names, table names, and descriptions. These include similarity-based algorithms, NB, LR, KNN, DT, MLP, RF, GBDT, and other intelligent algorithms that use feature information such as field names, table names, and descriptions to classify sensitive labels, solving the problem of data that cannot be classified by field content. Similarity calculation algorithm based on field names, table names, and descriptions: To perform similarity calculation on field names, first define standard sensitive names and their corresponding labels. During testing, calculate the semantic and character-level similarity of the test data field names, table names, descriptions, and standard sensitive names, and finally integrate and sort them. The advantage is that labels can be added by directly adding data without retraining the algorithm.

[0047] like Figure 10 As shown, the neural network classification module integrates multi-dimensional information encoding (currently, data is primarily based on field value information). Based on pre-trained language models such as BERT, it integrates multi-dimensional information encodings such as table names, field names, field values, and field types to perform sensitive label classification. Specifically, the data is split into multi-dimensional information such as field values, field names, data types, and database table names. After encoding each using a neural network, the data is input into a self-attention model to obtain the final encoding, and finally input into a fully connected layer for classification. The loss function uses Focal Loss to mitigate uneven sample distribution. Classification training is used during training, and similarity measurement is used during testing. The advantage is that it integrates multi-dimensional information for joint decision-making and deep semantic analysis, solving some classification problems that require deep semantic analysis and joint analysis of multi-dimensional information.

[0048] After the above two types of algorithms, after the data query and grading system performs rule recognition, the unrecognized results are input into the intelligent classification algorithm based on field name, table name, and description to obtain similarity sorting. After threshold screening, the remaining data is input into the neural network model for classification. Finally, the three parts of the results are integrated to obtain the final classification and grading results.

[0049] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The electronic device may include a processor (Center Processing Unit, CPU), a memory, an input device, and an output device. The input device may include a keyboard, a mouse, a touch screen, and the like. The output device may include a display device such as a liquid crystal display (LCD), a cathode ray tube (CRT), and the like. The memory may include a read-only memory (ROM) and a random access memory (RAM), and provide the processor with program instructions and data stored in the memory. In an embodiment of the present application, the memory may be used to store a program of any of the methods for intelligent classification and grading of power assets based on a large model in the embodiments of the present application. The processor calls the program instructions stored in the memory, and the processor is used to execute any of the methods for intelligent classification and grading of power assets based on a large model in the embodiments of the present application according to the obtained program instructions.

[0050] A computer-readable storage medium stores a computer program, which, when executed, implements the intelligent classification and grading of power assets in any of the above method embodiments.

[0051] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A method for intelligent classification and grading of power data assets based on a large model, characterized in that: The following steps are involved: Data collection, obtaining data from data storage devices and extracting field attributes and field contents; Preprocessing: Cleaning, denoising, and standardizing the acquired data to construct structured and unstructured data sets; Sensitive feature marking: using a preset classification algorithm to classify and grade sensitive data in fields, database tables, and descriptions in the dataset according to their sensitivity; Sensitive data template construction: construct the sensitive data template from the marked sensitive data; Sensitive data recording, inputting the sensitive data template into the data query and classification system, and detecting, classifying and recording the data in the data storage device through the data query and classification system; Sensitive data collection: collect the recorded sensitive data to obtain sensitive data logs.

2. The method for intelligent classification and grading of power data assets based on a large model according to claim 1 is characterized in that: The data collection includes the following sub-steps: Load the tags and rules of the power industry template; Connect to the Hive database and obtain information about all databases and tables in it; Traverse all the above databases and data table information, and extract all field attributes and field contents therein, wherein the field attributes include field name, table name, and description; Match the extracted field attributes and field contents with the tags and rules in the power industry template; If the field attributes and field content match the rules of the corresponding tag, the classification and grading information of the field is considered to be the classification and grading information corresponding to the tag and is recorded.

3. The method for intelligent classification and grading of power data assets based on a large model according to claim 1 is characterized in that: The sensitive feature marking includes the following sub-steps: Field classification: Based on the meaning of the fields, all fields contained in the classified database table are classified into different levels to determine the security level N of each field; Classify the database tables according to their business application scenarios and provision methods, and determine their security level M. Comprehensive grading: the maximum value of the field grading Nmax=MAX{N1,N2…Nn} and the library table grading M, the maximum value of the two is taken as the final grading L=MAX{M,Nmax} of the resource to be graded.

4. The method for intelligent classification and grading of power data assets based on a large model according to claim 1, characterized in that: The following steps are also included: Learning the steps of labeling sensitive features through machine learning and generating built-in rules; The machine learning algorithm identifies and labels sensitive features of new data in the data storage device through the built-in rules.

5. A large-scale model-based intelligent classification and grading system for power data assets, characterized by: Includes the following modules: The data acquisition module is used to obtain data from the data storage device and extract field attributes and field contents; The preprocessing module is used to clean, denoise and standardize the acquired data to construct structured and unstructured data sets; A sensitive feature marking module is used to classify and grade the sensitive data in the fields, database tables, and descriptions in the data set according to their sensitivity using a preset grading algorithm; A sensitive data template construction module is used to construct a sensitive data template from the marked sensitive data; A sensitive data recording module, configured to input the sensitive data template into a data query and grading system, and detect, grade and record the data in the data storage device through the data query and grading system; The sensitive data collection module is used to collect the recorded sensitive data and obtain sensitive data logs.

6. The intelligent classification and grading system for power data assets based on a large model according to claim 5 is characterized in that: It also includes a machine learning module, which includes a data acquisition layer, a basic component layer, and a result fusion layer.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 4 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Sensitive data classification and grading identification method and system

    CN114511019A

  • Power sensitive data classification and grading method and device, storage medium and electronic equipment

    CN116108393A

  • Power sensitive data processing method and device, electronic equipment and storage medium

    CN116628584A

  • Database table classification and grading method and system based on training model

    CN117271679A

  • Data classification and grading method and device based on feature recognition

    CN117609897A