Data classification and grading method and device and data governance system
By building a knowledge base and using natural language processing and machine learning technology to classify and classify enterprise data, the problem of low data governance efficiency is solved, automated data management and security protection is realized, and data utilization efficiency and security are improved.
Patent Information
- Application Number
- CN202510338788.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-18
AI Technical Summary
Data governance in the prior art is low efficiency and cannot effectively distinguish the importance of different types of data, resulting in waste of resources and management difficulties.
By building a knowledge base, using natural language processing and machine learning technology, enterprise data is classified and hierarchical, matching the data content, field names and table names, determining its classification and security levels, and realizing automated data governance.
It improves the efficiency of data management, reduces manual intervention, ensures data security and compliance, optimizes data storage and access control, and promotes the rational use of data and value mining.
Smart Images

Figure CN120336536A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data development and application. Specifically, it relates to a method, device, computer program product, and data governance system for classifying and grading data. Background Art
[0002] In modern enterprises, "data-driven business" has become a consensus. During the digital transformation process of a company, a large amount of data will be generated, such as employee information, contracts, financial statements, corporate strategic plans, etc. For different departments, the importance of different types of data is not the same, and a one-size-fits-all management method cannot be adopted. If the same protection is applied to all data, it will cause huge waste of resources for the enterprise. Therefore, the current data governance solutions have low efficiency. Summary of the Invention
[0003] The main objective of this application is to provide a method, device, computer program product, and data governance system for classifying and grading data, so as to at least solve the problem of low efficiency in data governance in the prior art.
[0004] To achieve the above objective, according to one aspect of this application, a method for classifying and grading data is provided, including: obtaining initial data of an enterprise, where the initial data includes at least one or more of human resource data, contract data, financial data, and planning data; performing data governance on the initial data at least according to a knowledge base to obtain target data, where the data governance methods include classification and / or grading, classification means dividing the initial data into different categories, and grading means dividing the initial data into different security levels; storing the target data in the knowledge base, and the data in the knowledge base is at least used for business analysis.
[0005] Optionally, performing data governance on the initial data at least according to a knowledge base to obtain target data includes: matching the content of the initial data with the content of the data in the knowledge base to obtain a first matching rate; matching the field names of the initial data with the field names of the data in the knowledge base to obtain a second matching rate; matching the table names of the initial data with the table names of the data in the knowledge base to obtain a third matching rate; performing weighted averaging according to the first matching rate, the second matching rate, and the third matching rate to obtain a comprehensive matching rate; determining the classification and grading of the data in the knowledge base with the highest comprehensive matching rate with the initial data as the classification and grading of the initial data to obtain the target data.
[0006] Optionally, performing data governance on the initial data at least according to a knowledge base to obtain target data includes: performing classification processing on the initial data to obtain intermediate data; performing grading processing on the intermediate data to obtain the target data.
[0007] Optionally, classify the initial data to obtain intermediate data, including: constructing a classification model, where the classification model is trained using multiple sets of training data, and each set of training data in the multiple sets of training data includes historical initial data obtained within a historical time period and historical intermediate data corresponding to the historical initial data; input the initial data into the classification model to obtain the intermediate data corresponding to the initial data.
[0008] Optionally, classify the initial data to obtain intermediate data, including: obtaining the intersection - union ratio of the initial data and the data in the knowledge base; determining the similarity according to the intersection - union ratio, where the intersection - union ratio and the similarity are in a positive - correlation relationship; determining the category of the data in the knowledge base with the highest similarity to the initial data as the category of the initial data to obtain the intermediate data.
[0009] Optionally, perform grading processing on the intermediate data to obtain the target data, including: constructing a grading model, where the grading model is trained using multiple sets of training data through the PaddleNLP algorithm, and each set of training data in the multiple sets of training data includes historical intermediate data obtained within a historical time period and historical target data corresponding to the historical intermediate data; input the intermediate data into the grading model to obtain the target data corresponding to the intermediate data.
[0010] Optionally, perform grading processing on the intermediate data to obtain the target data, including: obtaining the cosine similarity between the intermediate data and the data in the target data; determining the level of the data in the knowledge base with the highest cosine similarity to the intermediate data as the level of the intermediate data to obtain the target data.
[0011] According to another aspect of the present application, there is provided a device for classifying and grading data, including: an acquisition unit for acquiring the initial data of an enterprise, where the initial data includes at least one or more of human - resource data, contract data, financial data, and planning data; a classification and grading unit for performing data governance on the initial data at least according to a knowledge base to obtain target data, where the data governance methods include classification and / or grading, classification means dividing the initial data into different categories, and grading means dividing the initial data into different security levels; a storage unit for storing the target data into the knowledge base, and the data in the knowledge base is at least used for business analysis.
[0012] According to another aspect of the present application, there is provided a computer program product, including a computer program, which when executed by a processor implements the steps of any one of the methods for classifying and grading the data.
[0013] According to yet another aspect of the present application, there is provided a data governance system, including: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of the methods for classifying and grading the data.
[0014] Applying the technical solution of the present application, a knowledge base is pre-constructed. As the standard for classification and grading, the data of the enterprise can be matched with the data in the knowledge base, and the matching result is used as the result of classifying and grading the data, greatly simplifying the data management method of the enterprise, reducing manual intervention, and thus improving the data governance efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings forming a part of this application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0016] Figure 1 The hardware structure block diagram of a mobile terminal showing a method for classifying and grading data provided in an embodiment of the present application is shown;
[0017] Figure 2 The schematic flow diagram of a method for classifying and grading data provided in an embodiment of the present application is shown;
[0018] Figure 3 The schematic architecture diagram of the present solution is shown;
[0019] Figure 4 The schematic diagram showing the business classification and grading standard is shown;
[0020] Figure 5 The schematic diagram showing the classification and security level of the constructed knowledge base is shown;
[0021] Figure 6 The schematic diagram showing the creation of a data source is shown;
[0022] Figure 7 The schematic flow diagram of another method for classifying and grading data is shown;
[0023] Figure 8 The schematic flow diagram of classification and grading is shown;
[0024] Figure 9Shows a schematic diagram of the intersection over union;
[0025] Figure 10 Shows a schematic diagram of the classification criteria of the knowledge base;
[0026] Figure 11 Shows a schematic diagram of the data to be classified;
[0027] Figure 12 Shows a schematic diagram of the results obtained by word segmentation and intersection over union calculation;
[0028] Figure 13 Shows a schematic diagram of the classification matching results;
[0029] Figure 14 Shows a schematic diagram of the hierarchical criteria of the knowledge base;
[0030] Figure 15 Shows a schematic diagram of the data to be hierarchically classified;
[0031] Figure 16 Shows a schematic diagram of calculating similarity;
[0032] Figure 17 Shows a schematic diagram of the similarity between multiple data;
[0033] Figure 18 Shows a schematic diagram of the hierarchical matching results;
[0034] Figure 19 Shows a structural block diagram of a device for classifying and hierarchically classifying a kind of data provided according to an embodiment of the present application.
[0035] Among them, the above-mentioned drawings include the following reference numerals:
[0036] 102, processor; 104, memory; 106, transmission device; 108, input / output device. Detailed implementation manners
[0037] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0038] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts shall fall within the protection scope of the present application.
[0039] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so as to implement the embodiments of this application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0040] As introduced in the background art, the efficiency of data governance in the prior art is relatively low. To solve the above problems, embodiments of this application provide a method, apparatus, computer program product and data governance system for classifying and grading data.
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0042] The method embodiments provided in the embodiments of this application can be executed on a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a method of classifying and grading data according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0043] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the display method of device information in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include the wireless network provided by the communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0044] In this embodiment, a method for classifying and grading data running on a mobile terminal, computer terminal, or similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0045] Figure 2 It is a schematic flowchart of a method for classifying and grading data according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:
[0046] Step S201, obtain the initial data of the enterprise, where the above initial data includes at least one or more of human resources data, contract data, financial data, and planning data;
[0047] Specifically, in the initial stage of data governance, "initial data" needs to be collected from various data sources within the enterprise. These data sources can include the enterprise's information systems, databases, file systems, etc., and the scope covered should at least include human resources data, contract data, financial data, and planning data. Human resources data involves employee information such as name, position, work experience, etc.; contract data covers various agreements and contracts signed by the enterprise; financial data includes the enterprise's financial statements, transaction records, etc.; planning data involves information such as the enterprise's development strategy and project plans. Collecting this initial data is the first step in subsequent data governance, aiming to comprehensively understand the current situation of the enterprise's data and lay a foundation for subsequent classification and grading work.
[0048] Step S202, perform data governance on the above initial data at least according to the knowledge base to obtain target data. Among them, the methods of data governance include classification and / or grading. Classification means dividing the above initial data into different categories, and grading means dividing the above initial data into different security levels.
[0049] Specifically, in this step, the "knowledge base" will be used to govern the collected "initial data" to generate "target data". The knowledge base is a database integrating industry knowledge, enterprise rules, laws and regulations, etc., and is used to guide the standards for data classification and grading. Data governance includes two key steps: classification and grading, which can be used alone or in combination.
[0050] Classification: It refers to dividing the initial data into different categories according to its business attributes, types and other characteristics. For example, all data related to human resources will be classified as human resources data, and financial data will be classified as financial data, etc. Classification helps the enterprise build a clear data asset structure, facilitating data management, search and application.
[0051] Grading: It refers to dividing data into different security levels according to the sensitivity, importance or potential risk of the data. The purpose of grading is to formulate corresponding protection measures for different levels of data, such as encryption, access control, backup strategies, etc., to ensure the security and compliance of the data. For example, financial data involving personal privacy may be classified as a high level and thus requires more stringent protection.
[0052] Step S203, store the above target data into the above knowledge base, and the data in the above knowledge base is at least used for business analysis.
[0053] Specifically, as the basis for data classification and grading, the knowledge base stores standard classification and grading information. By adding the newly identified classification and grading numbers to the knowledge base, the standardization and consistency of enterprise data management can be ensured, avoiding different departments or systems having different understandings and processing methods for the same data type.
[0054] The classification and grading information in the knowledge base can be used to build an intelligent decision-making system. For example, when the system receives a new dataset, it can automatically query the knowledge base to quickly determine the classification and grading of the data, and then automatically select appropriate data processing and protection strategies, such as encryption, access control, backup frequency, etc.
[0055] Storing classified and graded data in the knowledge base helps optimize the enterprise data architecture. Enterprises can build data lakes, data warehouses based on data classification, or design more efficient data indexing and storage strategies. The grading information can then guide the setting of data access permissions to ensure that access to sensitive data is strictly controlled.
[0056] The classification and grading information stored in the knowledge base is crucial for enterprises to comply with data security laws and regulations. For example, enterprises are required to classify and grade data for protection. The knowledge base can be used as a basis for compliance checks to help enterprises conduct internal audits and ensure that data processing activities meet the requirements of laws and regulations.
[0057] Through classification and grading, enterprises can better understand their data assets, and thus more effectively tap the value of the data. For example, generally less classified data can be used more freely for data analysis, machine learning and other applications, while highly classified sensitive data needs to be used under more strict conditions to ensure the reasonable utilization and maximization of data value.
[0058] The classification and grading information in the knowledge base can be used as a reference for employees' data security training, enabling employees to understand the importance of different data types and grades, and improving their data security awareness and management capabilities.
[0059] According to data grading, enterprises can formulate more reasonable disaster recovery plans and business continuity strategies, prioritize the protection and recovery of highly graded data, and ensure the normal operation of critical businesses.
[0060] The system in the direct application solution of the enterprise internal data management platform realizes automated data classification and grading, which is used to improve data management efficiency and security.
[0061] Regulatory agencies can use similar knowledge bases and classification and grading mechanisms to supervise and manage industry data to ensure compliance with data security regulations.
[0062] Data security service providers can use this solution to provide consulting services and audit services for data classification and grading to customers.
[0063] Industry organizations can formulate industry data security standards and guidelines based on extensive knowledge bases and classification and grading practices.
[0064] Big data analysis and mining projects ensure the reasonable use of data resources when dealing with big data, and avoid business risks or legal risks caused by improper data classification.
[0065] In the last step of data governance, the "target data" after classification and grading is stored back into the "knowledge base". This step may seem like simply storing the data back into the knowledge base, but in fact, the knowledge base here is no longer a simple data repository. Instead, it is an intelligent library containing the data information after governance, used to store data meta-information, classification and grading results, governance strategies, etc.
[0066] Storing the target data in the knowledge base can have the following functions:
[0067] Feedback mechanism: Feeding back the data information after governance to the knowledge base can continuously enrich and optimize the knowledge base, making it a more powerful basis for classification and grading.
[0068] Business analysis: The data classification and grading results stored in the knowledge base can be used for subsequent business analysis. For example, by analyzing the data distribution and importance, enterprises can optimize business processes, enhance decision-making support, and ensure the compliance and security of data use.
[0069] Through this embodiment, a knowledge base is pre-constructed. As the standard for classification and grading, the data of the enterprise can be matched with the data in the knowledge base, and the matching results are used as the classification and grading results of the data, greatly simplifying the data management method of the enterprise, reducing manual intervention, and thus improving the data governance efficiency.
[0070] Specifically, the solution of this application aims to perform intelligent classification and grading on the data of the platform, and then perform security management and application on the classified and graded data.
[0071] The purpose of the present invention is to achieve intelligent classification and grading processing of the data of the platform. After achieving classification and grading, corresponding management and application can be carried out according to data classification and security levels. It can be reflected in the following 4 aspects:
[0072] 1. Protecting important data is necessary to protect national security, public interests, and citizens' rights. The ultimate and primary purpose of data classification and grading is to protect important data.
[0073] 2. If classification and grading are not carried out and high-standard protection is provided for all data, first, the cost is huge; second, it is difficult to implement; third, there are many loopholes; fourth, it limits the great value that can be exerted by the reasonable utilization of data.
[0074] 3. Classification and grading can facilitate the utilization of general data and are necessary for promoting the high-quality development of the digital economy. After the important data is screened out and protected with the data classification and grading protection system, it gives other data a "reassuring pill", allowing them to be safely transmitted, shared, developed, and utilized.
[0075] 4. Data classification and grading are beneficial for each department to clarify its powers and responsibilities, avoiding both the situation where the data that should be managed cannot be managed and the overlapping management of the same type of data.
[0076] The classification and grading of data serve the management and application of data.
[0077] Specifically, as Figure 3 shown, the specific content of this solution includes the following:
[0078] First, system deployment: The invention can be built based on an independent platform or deployed on a single server and connected to multiple platforms to achieve the classification and grading recognition of data from multiple platforms simultaneously and output the results.
[0079] Second, system architecture: This system mainly includes a data dictionary module, a platform knowledge base module, an AI intelligent recognition, classification, and grading module, a classification and grading result auditing and adjustment module, and a classification and grading result display module. There are a total of 5 modules, with a simple structure and clear function division.
[0080] Third, introduction of AI technology: The relatively popular AI artificial intelligence technology is introduced, and open-source AI semantic recognition resources are used. This makes the processing of classification and grading in this invention more intelligent, user-friendly, and automated.
[0081] Fourth, knowledge base construction: This invention supports importing the sorted industry knowledge information into the knowledge base. The accuracy of the classification and grading depends on the comprehensiveness and accuracy of the content in the knowledge base. It is necessary to gradually enrich the content of the knowledge base to improve the system's capabilities.
[0082] Fifth, multi-database support: This invention supports connecting to multiple database platforms simultaneously, and the database types can be different. It can support mainstream databases such as Mysql, Oracle, Hive Telepg, Teledb, Hbase, and Mongodb.
[0083] Sixth, security: This invention needs to process a large amount of data, but its processing logic is based on a sample data obtained from the data dictionary module and does not query a large number of detailed data in the platform table. By cutting off the direct access of users to batch detailed data, the leakage of detailed-level data is avoided, ensuring data security.
[0084] Seventh, loop optimization improvement: The present invention can set to extract result items with abnormal or unsatisfactory scores, screen impact factors, refine relevant factors into new corpora and add them to the platform knowledge base. The more powerful and comprehensive the knowledge base is built, the higher and more comprehensive the accuracy of recognition will be.
[0085] In the specific implementation process, at least data governance is performed on the above initial data according to the knowledge base to obtain target data, which can be achieved through the following steps: matching the content of the above initial data with the content of the data in the above knowledge base to obtain a first matching rate; matching the field names of the above initial data with the field names of the data in the above knowledge base to obtain a second matching rate; matching the table names of the above initial data with the table names of the data in the above knowledge base to obtain a third matching rate; performing weighted averaging according to the above first matching rate, the above second matching rate and the above third matching rate to obtain a comprehensive matching rate; determining the classification and grading of the data in the above knowledge base with the highest comprehensive matching rate with the above initial data as the classification and grading of the above initial data to obtain the above target data.
[0086] In this solution, by combining the matching of the content, field names and table names of the data, the meaning of the data can be understood more comprehensively, thereby improving the accuracy of classification and grading.
[0087] Specifically, in the first step of data governance, semantic matching is performed on the actual content (sample data) of the initial data and the data content stored in the knowledge base. This usually involves natural language processing (NLP) technologies, such as word segmentation, semantic understanding, text similarity calculation, etc., to compare the field sample text in the initial data with the text of the corresponding classification criteria in the knowledge base. The first matching rate reflects the similarity between the data content and the knowledge base content and is one of the important bases for determining data classification and grading.
[0088] Specifically, the field names in the initial data are textually matched with the field names in the knowledge base to calculate the similarity. The second matching rate reflects the matching degree between the field names and the field name standards defined in the knowledge base. Field names are usually more directly related to the business attributes of the data. Therefore, the second matching rate can be used as another important indicator for classification and grading.
[0089] Specifically, the table names of the initial data are also matched with the table names stored in the knowledge base to calculate the similarity. The third matching rate reflects the proximity of the table names to the table standards in the knowledge base. Table names usually can more comprehensively reflect the business meaning of the data set. Therefore, it also plays an important role in data classification and grading.
[0090] Specifically, after obtaining the three matching rates, they will be averaged according to the preset weights to obtain a comprehensive matching rate. The key to this step lies in determining the weights of each matching rate. The setting of the weights should be based on the understanding of the enterprise's data governance requirements. Usually, the content matching rate (the first matching rate) may be given a higher weight because the semantic understanding of the data content is crucial for determining the classification and grading of the data.
[0091] Specifically, finally, according to the obtained comprehensive matching rate, the data classification and grading information with the highest matching degree will be found in the knowledge base and used as the classification and grading result of the initial data, and finally the target data will be generated. The target data contains the classification and grading information after the governance of the initial data, which provides a basis for subsequent data security management and compliance use.
[0092] Specifically, platform knowledge base construction: While building the platform, collect information such as platform business knowledge, data scope, information content, and security specifications, determine the classification categories and security levels of the platform (see Figure 4 ), refine the platform data into corpus and map it to the classification and grading, and create the knowledge base content of the platform (see Figure 5 ).
[0093] Data dictionary building and connecting to the target platform: Once the system is built, the database of the target platform to be classified and graded can be accessed. After connecting the network and providing the configured address and user information, the access can be completed. The data dictionary module only obtains the metadata and a sample data of the target library and does not view other information of the target library. The following is an example for different platform accesses. Taking the platform to be classified and graded as a mysql library as an example, the operation example is as follows:
[0094] (1) In the left navigation bar, click "Data Source" under "Data Recognition".
[0095] (2) Click "New".
[0096] (3) Under the root directory on the left side of the pop-up new data source page, add or select the data source classification (you can also select libraries of other types such as oracle, sql server mongo, etc.).
[0097] (4) On the new data source page, set the data source information, as Figure 6 shown.
[0098] The configuration information is shown in Table 1.
[0099] Table 1: Adding MySQL Data Source
[0100]
[0101] After the configuration is completed, click "Test Connection" to check whether the configuration information is correct and whether the database is successfully connected.
[0102] (6) After passing the test, click "Next".
[0103] (7) On the popped-up metadata page, click "Register All", or select multiple metadata items and then click "Register".
[0104] Click "Save", and the data dictionary module will synchronize the table structure and extract a sample data from each table as the benchmark data for the following classification and grading process.
[0105] AI intelligent classification and grading process: After the above two steps, there is a classification and grading knowledge base and the information of the target table to be classified and graded. Obtain the two parts of data, call the AI semantic recognition function to parse the semantics of each field in the target table, match the semantics with the corpus in the knowledge base, generate score values and classification and grading results, and thus complete the first round of classification and grading marking (see the AI recognition classification and grading part in the Figure 7 method flow schematic diagram).
[0106] The platform starts to be built (collecting platform business knowledge, data scope, information content, security specifications, etc.), constructing a knowledge base module (sorting out classification and grading standards and corpus according to business and regulations), constructing each database to be classified and graded on the platform, managing data fields, obtaining the metadata to be classified and graded on the platform, and obtaining the latest classification and grading knowledge base information.
[0107] The AI semantic recognition module conducts recognition, mainly including the following three contents: recognizing the sample data of metadata to match the classification and grading corpus and giving the matching rate, recognizing the Chinese field names of metadata to match the classification and grading corpus and giving the matching rate, and recognizing the Chinese table names of metadata to match the classification and grading corpus and giving the matching rate.
[0108] Summarize the results of the above three steps, take the optimal value according to the matching rate and priority to determine the classification and grading of each field, extract the matching rate. If the matching rates of classification and grading are all greater than 0.4, then output a satisfactory result, extract the corpus of field meanings to enrich the knowledge base. If it is less than 0.4, then re-obtain the sample data and source data information corresponding to the fields with a matching rate lower than 0.4.
[0109] Now take an example Figure 8 Implementation process of the classification and grading system:
[0110] The data sources include databases and Excel in the knowledge base. Through scanning the source data for data recognition, through algorithm matching for resource confirmation, and then manually confirming whether to release resources, checking the released information and then establishing a data directory, and finally exporting the classification and grading results.
[0111] In some embodiments, at least based on the knowledge base, the above initial data is governed to obtain target data, which can be specifically achieved through the following steps: classifying the above initial data to obtain intermediate data; and grading the above intermediate data to obtain the above target data.
[0112] In this solution, classification processing, as a prerequisite step for grading processing, can first quickly divide the data according to business attributes, reducing the complexity of grading processing and improving the efficiency and speed of overall data governance.
[0113] Specifically, the above initial data is classified to obtain intermediate data. Here, the "initial data" refers to the original data collected by the enterprise, covering multiple business areas such as human resources, contracts, finance, and planning. Data classification processing is to divide this data into different categories according to their business attributes, data types, subject domains, and other characteristics. This process can be based on the analysis of data content, field name matching, or intelligent classification combining enterprise knowledge and business rules. The "intermediate data" obtained after classification has clear business labels, facilitating subsequent grading processing and management.
[0114] The above intermediate data is graded to obtain the above target data. On the basis that the data has been clearly classified, grading processing is to evaluate the sensitivity and importance of each category of data and divide it into different security levels. Grading is usually based on factors such as regulatory requirements, enterprise internal policies, the risk of data leakage, and the impact of data on business. The "target data" after grading not only contains the classification information of the data but also attaches security level labels, facilitating the enterprise to formulate corresponding data protection strategies and access control measures.
[0115] Specifically, data classification refers to classifying data into different categories according to the attributes, content, and usage scenarios of the data. Classification is usually based on the business attributes of the data. For example, personal basic information, financial data, contract documents, etc. Each type of data can correspond to different business areas of the enterprise. Classification helps the enterprise understand the composition of its data assets and facilitates the adoption of rational and professional strategies in the use, storage, and management of data.
[0116] Data grading refers to evaluating the importance or sensitivity of data and dividing it into different security levels according to the evaluation results. Grading takes into account the impact of data on aspects such as personal privacy, enterprise interests, and legal compliance. Different levels of data will be protected to different degrees. The purpose of grading is to ensure that sensitive data is subject to more stringent security controls, while general data can be used more openly to achieve the purpose of reasonable allocation and utilization of resources.
[0117] In the specific implementation process, the above-mentioned initial data is classified to obtain intermediate data, which can be achieved through the following steps: construct a classification model, where the above-mentioned classification model is trained using multiple sets of training data, and each set of training data in the above-mentioned multiple sets of training data includes historical initial data obtained within a historical time period and historical intermediate data corresponding to the above-mentioned historical initial data; input the above-mentioned initial data into the above-mentioned classification model to obtain the above-mentioned intermediate data corresponding to the above-mentioned initial data.
[0118] In this solution, the machine learning model can automatically classify a large amount of data, greatly reducing the resources and time required for manual classification and improving the efficiency of data governance. By using a large amount of historical classification data for training, the model can learn the complex rules of data classification, and its accuracy and consistency are higher compared with simple rule matching or manual classification.
[0119] Specifically, the "classification model" mentioned here is a model based on machine learning technology for automatically classifying data. Constructing this model requires "multiple sets of training data", which are "historical initial data" obtained within a historical time period and their corresponding "historical intermediate data". Historical initial data refers to the original data collected in the past, while historical intermediate data is the result of classifying these initial data. By training the model with a large amount of historical data, the model can learn the rules and patterns of data classification.
[0120] After the model is constructed, these multiple sets of training data are used to train the model. The training process involves adjusting the model parameters to maximize the classification accuracy of the model on the training data. Common classification models include decision trees, support vector machines, neural networks, random forests, etc. Through the training of historical data, the model can establish a mapping relationship from the initial data to the intermediate data (classification result).
[0121] Input the current "initial data" into the trained classification model, and the model will output the "intermediate data corresponding to the initial data" according to the learned patterns and rules, that is, the result of data classification. The intermediate data contains classification labels for the initial data, such as the data belongs to categories such as "financial data" and "contract data".
[0122] In some embodiments, the above-mentioned initial data is classified to obtain intermediate data, which can be specifically achieved through the following steps: obtain the intersection and union ratio of the above-mentioned initial data and the data in the above-mentioned knowledge base; determine the similarity according to the above-mentioned intersection and union ratio, where the above-mentioned intersection and union ratio and the above-mentioned similarity are in a positive correlation relationship; determine the category of the data in the above-mentioned knowledge base with the highest similarity to the above-mentioned initial data as the category of the above-mentioned initial data to obtain the above-mentioned intermediate data.
[0123] In this solution, by quantifying the Intersection over Union (IoU) to calculate the similarity, the business attributes of the initial data can be identified more accurately, improving the accuracy of data classification. Using IoU to calculate similarity and classification can achieve a high degree of automation, reduce manual intervention, and improve the efficiency of data governance.
[0124] Specifically, Intersection over Union (IoU) is a metric that measures the overlap between two sets and is commonly used in fields such as object detection and image segmentation. However, in the context of data classification, it can also be used to evaluate the similarity between data sets. Specifically, the IoU between the initial data and each category of data in the knowledge base will be calculated. IoU measures the similarity between two by comparing the common part (intersection) of the fields in the data and the union of all fields. For example, if the "financial data" category in the knowledge base contains fields such as "income", "expenses", and "bills", and the initial data also contains information related to "income" and "bills", then the IoU between the two will be relatively high, indicating a high similarity.
[0125] The similarity is positively correlated with the IoU, meaning that the larger the value of the IoU, the higher the similarity between the data. The similarity degree between data sets will be quantified based on the calculated IoU, and this step is crucial for subsequent determination of data categories.
[0126] The data category in the knowledge base with the highest similarity to the initial data will be found, and the category label will be assigned to the initial data to complete the data classification process. The resulting "intermediate data" is the classified initial data with a clear business category label, facilitating subsequent hierarchical processing and data management.
[0127] Specifically, the classification algorithm calculates the similarity between the metadata and the knowledge base classification criteria based on the jieba Chinese word segmentation component library and the IoU. Jieba is an excellent Chinese word segmentation library, which is efficient and accurate and is widely used in fields such as natural language processing and information retrieval. IoU is one of the most commonly used evaluation metrics for tasks such as object detection and semantic segmentation. IoU refers to the ratio of the intersection to the union of two regions. When the two regions completely overlap, the IoU is maximized at 1, and when the two regions do not overlap at all, the IoU is minimized at 0. The calculation of this ratio can reflect the similarity between the two. The IoU calculation formula used in the classification algorithm is as Figure 9 shown.
[0128] For example, a knowledge base classification standard is as Figure 10 shown. The metadata to be classified after field tagging through AI semantic recognition is as Figure 11 shown. The results obtained through word segmentation and IoU calculation are as Figure 12As shown. Then, integrate the calculation results of the metadata in units of tables, and further obtain the data classification and matching results based on similarity, such as Figure 13 As shown.
[0129] In the specific implementation process, the above intermediate data is processed hierarchically to obtain the above target data, which can be achieved through the following steps: construct a hierarchical model, where the above hierarchical model is trained using multiple sets of training data through the PaddleNLP algorithm, and each set of training data in the above multiple sets of training data includes historical intermediate data obtained within a historical time period and the historical target data corresponding to the above historical intermediate data; input the above intermediate data into the above hierarchical model to obtain the above target data corresponding to the above intermediate data.
[0130] In this solution, the hierarchical model can automatically complete data classification, reducing a large amount of manual work and improving the efficiency of data processing. At the same time, since the model classifies based on the semantic features of the data, this provides an intelligent decision-making basis for data classification.
[0131] Specifically, the construction of the hierarchical model depends on the PaddleNLP algorithm, which is an open-source natural language processing toolkit. The training of the model requires based on "multiple sets of training data", and these training data come from the "historical intermediate data" obtained within a historical time period and the corresponding "historical target data". The historical intermediate data is data after classification, and the historical target data is data with security level labels for these classified data. Through these data, the model can learn the relationship between the semantic features of the data and the security level.
[0132] Use the above multiple sets of training data to train the hierarchical model through the PaddleNLP algorithm. During the training process, the model will try to identify patterns related to the security level from the semantic features of the data to achieve the purpose of accurately predicting the data security level. PaddleNLP provides rich models, such as pre-trained models like BERT and ERNIE, which have high accuracy in NLP tasks and are suitable for extracting features from the text description of the data and classifying them.
[0133] Input the currently processed "intermediate data" into the trained hierarchical model. The model will output the security level corresponding to the intermediate data according to the learned mapping relationship between the features and the security level, thereby obtaining the "target data". The target data refers to the data with a security level label after the intermediate data is processed hierarchically.
[0134] In some embodiments, the above intermediate data is classified to obtain the above target data, which can be specifically implemented through the following steps: Obtain the cosine similarity between the data in the above intermediate data and the above target data; Determine the level of the data in the above knowledge base with the highest cosine similarity to the above intermediate data as the level of the above intermediate data, so as to obtain the above target data.
[0135] In this solution, the cosine similarity can accurately quantify the semantic similarity between data, which provides a scientific quantitative basis for data classification. By calculating the cosine similarity and selecting the security level corresponding to the highest value, the accuracy and consistency of classification can be ensured.
[0136] Specifically, the cosine similarity is an index that measures the angular difference between vectors and is often used in the calculation of text similarity. In this scenario, the "intermediate data" refers to the data that has been classified, and the "target data" is the historical data with security level labels. The cosine similarity between the intermediate data and each piece of target data in the knowledge base will be calculated. This calculation process involves converting the text features of the data (such as table names, field names, and field contents) into vector representations, and then calculating the cosine similarity between these vectors. The cosine similarity value ranges from -1 to 1, but in this scenario, we usually focus on the range from 0 to 1, where 1 indicates complete similarity and 0 indicates no similarity.
[0137] The target data in the knowledge base with the highest cosine similarity to the intermediate data will be found, and then the security level of this target data will be assigned to the intermediate data. The obtained "target data" is the intermediate data, but now it already has a security level label, completing the classification process.
[0138] Specifically, the classification algorithm is based on a pre-trained model to sequentially calculate the semantic matching similarity between the metadata and the classification criteria of the knowledge base. A pre-trained model refers to a model that has been trained on a large dataset. These models have good performance in some general tasks and can be used as a starting point for specific tasks. In NLP, pre-trained models can learn the grammatical structure of the language and the relationships between vocabulary. This system uses the open-source pre-trained word vector model of PaddleNLP and combines jieba segmentation. First, the Chinese names of the tables and fields of the metadata are segmented, then the words are converted into word vectors, and then the word vectors are merged into text vectors. Finally, the cosine similarity is calculated based on the text vectors.
[0139] For example, a classification standard of a knowledge base is as Figure 14 shown. The data to be classified is as Figure 15 shown. After calculating the semantic similarity through the pre-trained model to convert the text vectors, as Figure 16As shown. Since a table usually contains multiple fields, and the number of fields contained in different tables may not be the same, the maximum similarity of the fields of the table is taken, and an average value of these maximum field similarities is calculated. This average value is used to represent the similarity between tables, as Figure 17 shown. Finally, the above calculation results are weighted with equal weights to obtain the data classification and matching results based on similarity, as Figure 18 shown.
[0140] Result verification and optimization: After the results are generated, the results can be screened and viewed. Extract the ones with low scoring values, and enrich the knowledge base by sorting out the corpus. If the data of the target table is unavailable, go to the data dictionary module to re-obtain the sample data and metadata. After the knowledge base and data dictionary are updated, the classification and grading recognition can be redone, and iterate repeatedly until all data has satisfactory results. Confirm the satisfactory results for subsequent use.
[0141] Final result display: A result viewing interface is provided, and it also supports exporting the results or feedbacking them to the target platform through an interface. The subsequent platform can perform security management based on classification and grading.
[0142] In summary, this solution has scalability. The knowledge base of the present invention can be infinitely expanded as the business knowledge accumulates and the classification and grading specifications and standards increase. The data dictionary module also supports connecting to multiple databases, facilitating expansion according to different requirements.
[0143] This solution has simplified manageability. The modules of the present invention are simple in function, the interface is easy to operate, the installation and deployment requirements are low, and there are almost no usage barriers, reducing the complexity of use.
[0144] This solution has data security. The present invention only obtains metadata information and single sample data, cutting off the way for users to access batch detailed data, preventing risks, and ensuring the safe use of data.
[0145] This solution supports docking multiple platforms simultaneously and can be integrated across industries and businesses. The present invention supports docking multiple database platforms of multiple databases simultaneously. Even if different industries and different businesses are integrated at the same time, it is no problem as long as the resources of the knowledge base are rich. In theory, it supports wireless expansion.
[0146] This solution introduces AI technology support and incorporates the popular AI intelligent semantic recognition function, making the platform intelligent. The traditional method is to map to hierarchical classification based on the Chinese or English names of fields, which involves a huge amount of work in sorting out the mapping relationships. Now, it is to identify the semantics of problems or data from the real sample data in the table and classify and grade according to the semantics, improving the accuracy. It reduces the cost of maintaining the Chinese and English names of the target table fields in the target library and improves the efficiency. Through deep learning, artificial intelligence can also improve the recognition accuracy and scope, with the advantage of self-upgrading function.
[0147] The embodiments of the present application also provide a device for classifying and grading data. It should be noted that the device for classifying and grading data in the embodiments of the present application can be used to execute the method for classifying and grading data provided in the embodiments of the present application. The device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0148] The following introduces the device for classifying and grading data provided in the embodiments of the present application.
[0149] Figure 19 It is a structural block diagram of a device for classifying and grading data according to the embodiments of the present application. As Figure 19 shown, the device includes:
[0150] An acquisition unit 10, configured to acquire the initial data of the enterprise, where the initial data includes at least one or more of human resource data, contract data, financial data, and planning data;
[0151] A classification and grading unit 20, configured to perform data governance on the initial data at least according to the knowledge base to obtain target data, where the data governance methods include classification and / or grading. Classification means dividing the initial data into different categories, and grading means dividing the initial data into different security levels;
[0152] A storage unit 30, configured to store the target data into the knowledge base, and the data in the knowledge base is at least used for business analysis.
[0153] Through this embodiment, a knowledge base is pre-constructed. As the standard for classification and grading, the data of the enterprise can be matched with the data in the knowledge base, and the matching result is used as the result of classifying and grading the data, greatly simplifying the data management method of the enterprise, reducing manual intervention, and thus improving the data governance efficiency.
[0154] In the specific implementation process, the classification and grading unit includes a first matching module, a second matching module, a third matching module, a weighted average module, and a determination module. The first matching module is used to match the content of the above initial data with the content of the data in the above knowledge base to obtain a first matching rate; the second matching module is used to match the field names of the above initial data with the field names of the data in the above knowledge base to obtain a second matching rate; the third matching module is used to match the table names of the above initial data with the table names of the data in the above knowledge base to obtain a third matching rate; the weighted average module is used to perform a weighted average based on the above first matching rate, the above second matching rate, and the above third matching rate to obtain a comprehensive matching rate; the determination module is used to determine the classification and grading of the data in the above knowledge base with the highest comprehensive matching rate with the above initial data as the classification and grading of the above initial data to obtain the above target data.
[0155] In this solution, by combining the matching of the content, field names, and table names of the data, the meaning of the data can be more comprehensively understood, thereby improving the accuracy of classification and grading.
[0156] In some embodiments, the classification and grading unit includes a classification module and a grading module. The classification module is used to perform classification processing on the above initial data to obtain intermediate data; the grading module is used to perform grading processing on the above intermediate data to obtain the above target data.
[0157] In this solution, classification processing, as a prerequisite step for grading processing, can quickly divide the data according to business attributes first, reducing the complexity of grading processing and improving the efficiency and speed of overall data governance.
[0158] In the specific implementation process, the classification module includes a first construction module and a first processing module. The first construction module is used to construct a classification model. Among them, the above classification model is trained using multiple sets of training data, and each set of the above multiple sets of training data includes historical initial data obtained within a historical time period and the historical intermediate data corresponding to the above historical initial data; the first processing module is used to input the above initial data into the above classification model to obtain the above intermediate data corresponding to the above initial data.
[0159] In this solution, the machine learning model can automatically classify a large amount of data, greatly reducing the resources and time required for manual classification and improving the efficiency of data governance. By using a large amount of historical classification data for training, the model can learn the complex rules of data classification, and its accuracy and consistency are higher compared to simple rule matching or manual classification.
[0160] In some embodiments, the classification module includes a first acquisition sub-module, a first determination sub-module, and a second determination sub-module. The first acquisition sub-module is configured to acquire the intersection-union ratio between the above-mentioned initial data and the data in the above-mentioned knowledge base. The first determination sub-module is configured to determine the similarity according to the above-mentioned intersection-union ratio, wherein the above-mentioned intersection-union ratio and the above-mentioned similarity are in a positive correlation relationship. The second determination sub-module is configured to determine that the category of the data in the above-mentioned knowledge base with the highest similarity to the above-mentioned initial data is the category of the above-mentioned initial data, so as to obtain the above-mentioned intermediate data.
[0161] In this solution, by quantifying the intersection-union ratio to calculate the similarity, the business attributes of the initial data can be identified more accurately, and the accuracy of data classification can be improved. Using the intersection-union ratio to calculate the similarity and classification can achieve a high degree of automation, reduce manual intervention, and improve the efficiency of data governance.
[0162] In the specific implementation process, the grading module includes a second construction sub-module and a second processing sub-module. The second construction sub-module is configured to construct a grading model, wherein the above-mentioned grading model is trained by using a multi-group of training data through the PaddleNLP algorithm. Each group of training data in the above-mentioned multi-group of training data includes historical intermediate data obtained within a historical time period and the historical target data corresponding to the above-mentioned historical intermediate data. The second processing sub-module is configured to input the above-mentioned intermediate data into the above-mentioned grading model to obtain the above-mentioned target data corresponding to the above-mentioned intermediate data.
[0163] In this solution, the grading model can automatically complete data grading, reducing a large amount of manual work and improving the efficiency of data processing. At the same time, since the model grades based on the semantic features of the data, this provides an intelligent decision-making basis for data grading.
[0164] In some embodiments, the grading module includes a second acquisition sub-module and a third determination sub-module. The second acquisition sub-module is configured to acquire the cosine similarity between the above-mentioned intermediate data and the data in the above-mentioned target data. The third determination sub-module is configured to determine that the level of the data in the above-mentioned knowledge base with the highest cosine similarity to the above-mentioned intermediate data is the level of the above-mentioned intermediate data, so as to obtain the above-mentioned target data.
[0165] In this solution, the cosine similarity can accurately quantify the semantic similarity between data, which provides a scientific quantitative basis for data grading. By calculating the cosine similarity and selecting the security level corresponding to the highest value, the accuracy and consistency of grading can be ensured.
[0166] The device for classifying and grading the above data includes a processor and a memory. The above acquisition unit, classification and grading unit, storage unit, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions. The above modules are all located in the same processor; alternatively, the above modules are separately located in different processors in any combination form.
[0167] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem of low efficiency in data governance in the prior art can be solved.
[0168] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one storage chip.
[0169] An embodiment of the present invention provides a computer-readable storage medium. The above computer-readable storage medium includes a stored program, wherein when the above program runs, it controls the device where the above computer-readable storage medium is located to execute the method for classifying and grading the above data.
[0170] An embodiment of the present invention provides a processor. The above processor is used to run a program, wherein when the above program runs, it executes the method for classifying and grading the above data.
[0171] An embodiment of the present invention provides a device. The device includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the method steps for classifying and grading data. The device herein can be a server, a PC, a PAD, a mobile phone, etc.
[0172] A computer program product includes a non-volatile computer-readable storage medium. The above non-volatile computer-readable storage medium stores a computer program, and when the above computer program is executed by a processor, it implements the steps of the method for classifying and grading the above data in each embodiment of the present application.
[0173] The present application also provides a data governance system, including one or more processors, a memory, and one or more programs. Among them, the above one or more programs are stored in the above memory and are configured to be executed by the above one or more processors. The above one or more programs include those for executing any one of the above methods for classifying and grading data.
[0174] Obviously, those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described herein can be executed in a different order, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.
[0175] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0176] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or a combination of multiple flows and / or blocks
[0177] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more of the flows Figure 1 or a combination of multiple flows and / or blocks
[0178] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the functions in the flowFigure 1 one or more processes and / or blocks Figure 1 steps of functions specified in one or more blocks
[0179] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0180] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0181] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassettes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0182] It should also be noted that the term "comprises," "comprising," or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0183] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, or improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for classifying and grading data, characterized in that, Including: Obtain the initial data of the enterprise, where the initial data includes at least one or more of human resource data, contract data, financial data, and planning data; Perform data governance on the initial data at least according to the knowledge base to obtain target data, where the data governance methods include classification and / or grading. Classification means dividing the initial data into different categories, and grading means dividing the initial data into different security levels; Store the target data in the knowledge base, and the data in the knowledge base is at least used for business analysis.
2. The method according to claim 1, wherein Perform data governance on the initial data at least according to the knowledge base to obtain target data, including: Match the content of the initial data with the content of the data in the knowledge base to obtain a first matching rate; Match the field names of the initial data with the field names of the data in the knowledge base to obtain a second matching rate; Match the table names of the initial data with the table names of the data in the knowledge base to obtain a third matching rate; Perform weighted averaging according to the first matching rate, the second matching rate, and the third matching rate to obtain a comprehensive matching rate; Determine the classification and grading of the data in the knowledge base with the highest comprehensive matching rate with the initial data as the classification and grading of the initial data to obtain the target data.
3. The method according to claim 1, wherein Perform data governance on the initial data at least according to the knowledge base to obtain target data, including: Perform classification processing on the initial data to obtain intermediate data; Perform grading processing on the intermediate data to obtain the target data.
4. The method according to claim 3, wherein Perform classification processing on the initial data to obtain intermediate data, including: Construct a classification model, where the classification model is trained using multiple sets of training data, and each set of training data in the multiple sets of training data includes historical initial data obtained within a historical time period and the historical intermediate data corresponding to the historical initial data; Input the initial data into the classification model to obtain the intermediate data corresponding to the initial data.
5. The method according to claim 3, characterized in that, Perform classification processing on the initial data to obtain intermediate data, including: Obtain the intersection-union ratio of the initial data and the data in the knowledge base; Determine the similarity according to the intersection-union ratio, where the intersection-union ratio and the similarity are in a positive correlation relationship; Determine the category of the data in the knowledge base with the highest similarity to the initial data as the category of the initial data to obtain the intermediate data.
6. The method according to claim 3, characterized in that, Perform grading processing on the intermediate data to obtain the target data, including: Construct a grading model, where the grading model is trained using multiple sets of training data through the PaddleNLP algorithm, and each set of training data in the multiple sets of training data includes historical intermediate data obtained within a historical time period and the historical target data corresponding to the historical intermediate data; Input the intermediate data into the grading model to obtain the target data corresponding to the intermediate data.
7. The method according to claim 3, wherein Perform grading processing on the intermediate data to obtain the target data, including: Obtain the cosine similarity between the intermediate data and the data in the target data; Determine the level of the data in the knowledge base that has the highest cosine similarity with the intermediate data as the level of the intermediate data to obtain the target data.
8. An apparatus for classifying and grading data, characterized in that It includes: An acquisition unit for acquiring the initial data of an enterprise, where the initial data includes at least one or more of human resource data, contract data, financial data, and planning data; A classification and grading unit for performing data governance on the initial data at least according to the knowledge base to obtain target data, where the data governance methods include classification and / or grading. Classification means dividing the initial data into different categories, and grading means dividing the initial data into different security levels; A storage unit for storing the target data into the knowledge base, and the data in the knowledge base is at least used for business analysis.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for classifying and grading the data described in any one of claims 1 to 7.
10. A data governance system, characterized in that, It includes: One or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include the method for classifying and grading the data described in any one of claims 1 to 7.
Citation Information
Cited By
Hierarchical navigation method and system for dynamically generated text content
CN121071135A