Data classification and grading method and device, equipment and medium
By building a data classification and grading catalog and quantitative assessment, and automating data classification and grading, we solve the problems of low efficiency and lack of accuracy in existing technologies, and achieve efficient and accurate data management and security.
Patent Information
- Application Number
- CN202510802380.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-26
AI Technical Summary
The existing data classification and grading methods rely on manual operations, resulting in low efficiency and difficulty in ensuring accuracy, and cannot meet data security and compliance requirements.
By analyzing data classification and building a catalog based on files, the database's value assessment indicators are collected for quantitative evaluation, non-critical databases are eliminated, and critical database data is matched with the catalog for batch labeling processing.
It improves the efficiency and accuracy of data classification and grading, and ensures the efficiency and security of data management.
Smart Images

Figure CN120705227A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular to a data classification and grading method, device, equipment and medium. Background Art
[0002] With the advent of the information age, data has become an indispensable and important resource in project management and decision-making. In the big data era, data is multi-source and heterogeneous, with different values. It is particularly important to classify and grade data to facilitate the adoption of different data protection measures and prevent data leakage. Therefore, data classification and grading management in different application fields is one of the important links in data security protection.
[0003] For example, in the healthcare sector, medical institutions handle a large amount of sensitive patient information, such as personal health information, medical records, diagnosis results, treatment plans, etc. According to relevant laws, regulations, and industry standards, medical institutions need to classify and manage data in a classified and graded manner to ensure patient privacy and data security.
[0004] For example, in the financial sector, insurance and financial institutions handle large amounts of sensitive information, such as personal customer information, financial data, and transaction records. According to financial regulatory requirements, insurance and financial institutions are also required to classify and grade data and strengthen data security level management to ensure data security and compliance.
[0005] The existing data classification and grading method usually involves organizing all departments to manually label each field one by one, which consumes huge manpower and time and makes it difficult to ensure the accuracy and timeliness of data classification and grading, reducing the efficiency and reliability of data classification and grading. Summary of the Invention
[0006] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a data classification and grading method, device, equipment and medium that can be applied to the medical field, financial technology or other related fields. Its main purpose is to ensure the efficiency and accuracy of data classification and grading, and improve data management efficiency and security.
[0007] The technical solutions of the present invention are as follows:
[0008] A first aspect of the present invention provides a data classification and grading method, comprising:
[0009] Obtaining a data classification basis file, parsing the data classification basis file, and constructing a corresponding data classification and grading directory;
[0010] Determine all databases to be processed, collect value assessment indicators for each database to be processed, perform value quantitative assessment on each database to be processed based on the value assessment indicators, and obtain corresponding value quantitative assessment results;
[0011] Determine non-critical databases in the database to be processed according to the value quantification evaluation result, and remove the non-critical databases from the database to be processed to obtain remaining critical databases to be processed;
[0012] The data in the key database is matched with the data classification and grading directory, and the corresponding data is batch-labeled according to the matched data categories and data levels.
[0013] A second aspect of the present invention provides a data classification and grading device, comprising:
[0014] A directory construction module is used to obtain a data classification basis file, parse the data classification basis file, and construct a corresponding data classification and grading directory;
[0015] A value assessment module is used to determine all databases to be processed, collect value assessment indicators of each database to be processed, perform value quantitative assessment on each database to be processed according to the value assessment indicators, and obtain corresponding value quantitative assessment results;
[0016] A database elimination module is used to determine non-critical databases in the database to be processed according to the value quantification evaluation result, and eliminate the non-critical databases from the database to be processed to obtain the remaining critical databases to be processed;
[0017] The marking processing module is used to match the data in the key database with the data classification and grading directory, and perform batch marking processing on the corresponding data according to the matched data categories and data levels.
[0018] A third aspect of the present invention provides a computer device comprising at least one processor; and
[0019] a memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned data classification and grading method.
[0021] A fourth aspect of the present invention provides a non-volatile computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, the one or more processors can execute the above-mentioned data classification and grading method.
[0022] Beneficial effects: The present invention discloses a data classification and grading method, device, equipment and medium. Compared with the prior art, the embodiment of the present invention obtains a data classification basis file, parses the data classification basis file, and constructs a corresponding data classification and grading directory; determines all databases to be processed, collects value assessment indicators of each database to be processed, and performs value quantitative assessment on each database to be processed according to the value assessment indicators to obtain corresponding value quantitative assessment results; determines non-critical databases in the database to be processed according to the value quantitative assessment results, and eliminates the non-critical databases from the database to be processed to obtain the remaining critical databases to be processed; matches the data in the critical database with the data classification and grading directory, and batch tags the corresponding data according to the data category and data level obtained by matching. By parsing and constructing a data classification and grading directory and automatically matching after confirming the critical database through value assessment, and batch automatically tagging the data based on the matching results, the efficiency and accuracy of data classification and grading are improved, and data management efficiency and security are ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the solutions in the present invention, a brief introduction is given below to the drawings required for use in describing the embodiments of the present invention. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] Figure 1 A schematic diagram of an application environment of the data classification and grading method provided by an embodiment of the present invention;
[0025] Figure 2 A flow chart of a data classification and grading method provided in an embodiment of the present invention;
[0026] Figure 3 A flowchart of step S201 in the data classification and grading method provided in an embodiment of the present invention;
[0027] Figure 4 A flow chart of step S203 in the data classification and grading method provided in an embodiment of the present invention;
[0028] Figure 5 A flow chart of step S204 in the data classification and grading method provided in an embodiment of the present invention;
[0029] Figure 6 A schematic diagram of the functional modules of the data classification and grading device provided in an embodiment of the present invention;
[0030] Figure 7A schematic diagram of the hardware structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0031] To make the objectives, technical solutions, and effects of the present invention more clear and distinct, the present invention is further described in detail below. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. The embodiments of the present invention are described below with reference to the accompanying drawings.
[0032] The data classification and grading method provided by the embodiment of the present invention can be applied in the following situations: Figure 1 In an application environment, the system includes a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0033] The user may use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).
[0034] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0035] The server 105 may be a server that provides various services, such as a backend server that provides support for the content browsed by the user using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (for example only). The backend server may analyze and process the received user requests and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to the user request) to the terminal device. The server 105 may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server 105 may also be a server for a distributed system, or a server combined with a blockchain.
[0036] It should be noted that the data classification and grading method provided in the embodiments of the present application can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the data classification and grading apparatus provided in the embodiments of the present invention can also be set in the first terminal device 101, the second terminal device 102, or the third terminal device 103. Alternatively, the data classification and grading method provided in the embodiments of the present invention can generally be executed by the server 105. Accordingly, the data classification and grading apparatus provided in the embodiments of the present invention can generally be set in the server 105.
[0037] It should be understood that the numbers of the above terminal devices, networks and servers are merely illustrative and any number of terminal devices, networks and servers may be provided as required.
[0038] like Figure 2 As shown, the data classification and grading method provided by the embodiment of the present invention specifically includes the following steps:
[0039] S201. Obtain a data classification basis file, parse the data classification basis file, and construct a corresponding data classification and grading directory.
[0040] In this embodiment, relevant data classification basis documents are obtained from regulatory agencies (such as the State Financial Supervision and Administration Bureau, the National Health Commission), industry standard organizations (such as the China Financial Standardization Technical Committee, the China Hospital Association), or internal enterprises. These documents may include national or industry regulatory requirements (such as the "Data Security Management Measures for Banking and Insurance Institutions" and the "Medical Information Security Management Measures"), international standards (such as ISO 27001), and internal company management regulations (such as data security policies and data governance manuals). They can be in PDF, Word, Excel, HTML and other formats, and the content includes text descriptions, tables, charts, etc.
[0041] Data classification is performed by parsing the file to extract key information, such as classification dimensions (e.g., customer information, transaction data, financial data), and grading criteria (e.g., high sensitivity, medium sensitivity, low sensitivity, etc.). For example, NLP tools can be used to identify keywords, phrases, and sentences within the file to extract data classification and grading rules. The parsed data classification and grading rules are stored in a structured format, such as JSON, XML, or a database table, to construct a corresponding data classification and grading directory. The specific data classification and grading directory is presented in a hierarchical structure, for example: a first-level directory for business areas (e.g., finance, healthcare); a second-level directory for business modules (e.g., customer management, transaction processing, patient diagnosis, treatment records); and a third-level directory for specific data categories (e.g., name, ID number, transaction amount, diagnosis results). The grading criteria for each data category are clearly defined within the directory. For example, data involving personal privacy and financial information is classified as high sensitivity; data involving business operations but not privacy information is classified as medium sensitivity; and public information or auxiliary information is classified as low sensitivity. By constructing a corresponding directory based on data classification file parsing, a unified standard and framework is provided for accurate and efficient data classification and grading, ensuring accuracy and consistency in classification and grading.
[0042] For example, in the field of medical health, according to the "Medical Information Security Management Measures" and other regulations, medical institutions need to classify and grade patient personal information, medical records, medical images, etc., build a data classification and grading catalog, and clearly define "personal identity information" (such as ID number, name) as a high-sensitivity level, "medical record information" (such as diagnosis results, treatment plans) as a high-sensitivity level, "public information" (such as hospital announcements) as a low-sensitivity level, and so on.
[0043] In the financial field, in accordance with regulatory requirements such as the "Data Security Management Measures for Banking and Insurance Institutions", financial institutions need to classify and grade customer identity information, financial data, transaction records, etc., build a data classification and grading catalog, and clearly define "personal identity information" (such as ID number, name) as a high-sensitivity level, "financial data" (such as account balance, transaction details) as a high-sensitivity level, "public information" (such as market conditions) as a low-sensitivity level, and so on.
[0044] S202: Determine all databases to be processed, collect value assessment indicators of each database to be processed, perform value quantitative assessment on each database to be processed according to the value assessment indicators, and obtain corresponding value quantitative assessment results.
[0045] In this embodiment, a data asset inventory is conducted to identify all databases to be processed, including production databases, backup databases, log databases, and test databases. Based on the inventory results, a list of all databases is generated, recording basic information for each database, such as the database name, system, storage location, and data volume. Subsequently, a valuation assessment is performed on each database by collecting valuation indicators for each database to be processed. These indicators include business relevance, data usage frequency, data update frequency, data integrity, and data sensitivity. Business relevance refers to the degree to which a database supports core business operations. For example, a bank's core transaction database is extremely important to the business, while a log database is relatively less important. Data usage frequency refers to the frequency of usage of a statistical database, including the frequency of operations such as data queries, updates, and deletions. For example, a customer information database is frequently used, while a backup database is less frequently used. Data update frequency refers to the frequency of data updates in a database. For example, a real-time transaction database requires frequent updates, while an annual financial report database requires less frequent updates. Data integrity refers to checking the integrity of the database to ensure the integrity of the data records. For example, a customer relationship management (CRM) database typically contains complete customer information, while temporarily stored user log data may be incomplete. Data sensitivity refers to assessing whether a database contains sensitive data. For example, a database containing personal customer information (such as ID card numbers and bank card numbers) is highly sensitive, while a public information database (such as company news announcements) is less sensitive.
[0046] Based on the collected value assessment indicators, a quantitative assessment of the database is performed. For example, based on business needs and experience, each assessment indicator is assigned a corresponding weight. The weighted sum of the value assessment indicators and the corresponding weights is then calculated to obtain the corresponding quantitative value assessment results. This quantitative assessment can accurately screen valuable databases, avoid unnecessary processing of low-value databases, and improve resource utilization efficiency.
[0047] For example, in the healthcare sector, a quantitative assessment of the value of a hospital's patient information database and diagnostic database is conducted to ensure that critical medical data is prioritized for protection. For example, the patient information database has a value score of 90, the diagnostic database has a value score of 85, and the backup database has a value score of 40.
[0048] In the financial sector, banks' core transaction databases and customer information databases are quantitatively assessed for value, with high-value databases prioritized. For example, a core transaction database might have a value score of 90, a customer information database might have a value score of 85, and a backup database might have a value score of 40.
[0049] S203 : Determine non-critical databases from the database to be processed according to the value quantification evaluation result, and remove the non-critical databases from the database to be processed to obtain remaining critical databases to be processed.
[0050] In this embodiment, a threshold is set based on business needs and resource availability to distinguish between critical and non-critical databases. For example, databases with a value score below 60 are classified as non-critical. Based on the quantitative value assessment results, databases below the threshold are screened and marked as non-critical. Non-critical databases are removed from the list of databases to be processed, and only critical databases are retained for subsequent processing. By eliminating non-critical databases and concentrating resources on critical databases, the efficiency of data classification and grading is improved, ensuring that critical databases receive priority processing and protection, and enhancing data security management effectiveness.
[0051] Furthermore, the decision threshold for non-critical databases is dynamically adjusted based on business needs and resource availability. For example, if resources are sufficient, the threshold can be appropriately lowered to process more databases.
[0052] For example, in the healthcare field, non-critical databases such as backup databases and test databases are eliminated, and resources are concentrated on processing patient information databases and diagnostic databases. For example, the test database has a value score of 30 and is determined to be a non-critical database and eliminated.
[0053] In the financial sector, non-critical databases such as temporary and log databases are eliminated, and resources are focused on core transaction databases and customer information databases. For example, a temporary database with a value score of 40 is considered non-critical and is eliminated.
[0054] S204: Match the data in the key database with the data classification and grading directory, and perform batch tagging on the corresponding data according to the matched data categories and data levels.
[0055] In this embodiment, the data in the key database is matched with the data classification and grading directory. Specifically, the classification and grading of the data field can be identified through keyword matching. For example, if the field name contains keywords such as "ID number" and "bank card number", it will be marked as "highly sensitive" according to the data classification and grading directory; or the classification and grading of the data field can be identified through data format matching. For example, if the field is in the format of a mobile phone number (such as 11 digits), it will be marked as "medium sensitive" according to the definition in the directory; or complex fields can be classified and graded in combination with business rules and contextual information, etc. The data fields are batch labeled according to the matching results. Specifically, the matched data categories and data levels can be applied to all relevant fields through an automated script tool, and the labeling results are recorded in the data management platform or database to facilitate subsequent query and management. Batch labeling can improve the efficiency and consistency of data classification and grading, and reduce the workload of manual operations.
[0056] In the above embodiment, the present invention discloses a data classification and grading method, which obtains a data classification basis file, parses the data classification basis file, and constructs a corresponding data classification and grading directory; determines all databases to be processed, collects value assessment indicators of each database to be processed, and performs value quantitative assessment on each database to be processed according to the value assessment indicator to obtain a corresponding value quantitative assessment result; determines non-critical databases in the database to be processed according to the value quantitative assessment result, and eliminates the non-critical databases from the database to be processed to obtain the remaining critical databases to be processed; matches the data in the critical database with the data classification and grading directory, and performs batch tagging on the corresponding data according to the data category and data level obtained by matching. By parsing and constructing a data classification and grading directory and automatically matching after confirming the critical database through value assessment, and automatically tagging the data in batches based on the matching results, the efficiency and accuracy of data classification and grading are improved, and data management efficiency and security are ensured.
[0057] In one embodiment, Figure 3 As shown, step S201 includes:
[0058] S301, obtaining a data classification basis file, parsing and obtaining classification and grading elements in the data classification basis file;
[0059] S302: Screen the classification and grading elements according to a preset business scenario to obtain a classification and grading standard that is compatible with the preset business scenario;
[0060] S303: Construct an initial data classification and grading directory according to the classification and grading standard, and iteratively update the data classification and grading directory according to a preset update strategy.
[0061] In this embodiment, a data classification basis file is obtained through channels such as the company's internal document management system, the official website of the regulatory agency, and the website of the industry standards organization. The classification and grading elements in the data classification basis file are parsed and obtained, including the classification dimensions and grading standards for different data. The classification and grading elements are filtered based on preset business scenarios, where the preset business scenarios refer to data application scenarios that require focus in actual business. For example, in the financial sector, core business scenarios may include customer identity verification, transaction processing, and risk assessment; in the healthcare sector, core business scenarios may include patient diagnosis, treatment records, and medical insurance settlement. Filtering rules are formulated based on the preset business scenarios, and standards related to the business scenarios are filtered out from the parsed classification and grading elements. For example, for the customer identity verification scenario in the financial sector, classification and grading standards related to the customer name, ID number, mobile phone number, etc. are filtered out. The classification and grading standards obtained after screening should be highly compatible with the preset business scenarios and can be directly applied to data classification and grading work in actual business, thereby improving the accuracy and efficiency of data classification and grading.
[0062] An initial data classification and grading catalog is constructed based on the selected classification and grading standards that are compatible with the preset business scenarios. Each data category is clearly defined in the catalog using a hierarchical structure, and the grading standards within each data category are given. At the same time, the generated data classification and grading catalog is iteratively updated according to the preset update strategy. The specific update strategy can be regular updates, such as a comprehensive evaluation and update of the catalog every quarter or every six months, or dynamic updates, such as timely updates to the catalog when regulatory policies change, business needs adjust, or new technologies are applied. Through continuous updates to optimize the catalog structure and content, the effectiveness and efficiency of data management can be improved.
[0063] For example, in the healthcare sector, an initial data classification and grading directory is constructed in accordance with regulations such as the "Medical Information Security Management Measures." For example, patient names and ID numbers are marked as "highly sensitive." Based on the development of medical technology and changes in business needs, the directory is dynamically updated, with new data categories (such as genetic testing information) and adjustments to the grading standards.
[0064] In the financial sector, an initial data classification and grading catalog has been established in accordance with regulatory requirements such as the "Measures for the Administration of Data Security in Banking and Insurance Institutions." For example, customer names and ID numbers are marked as "highly sensitive." Based on regulatory policy changes and business needs, the catalog is regularly updated, with new data categories (such as biometric information) and adjustments to grading standards.
[0065] In one embodiment, Figure 4 As shown, step S203 includes:
[0066] S401, ranking the databases to be processed by value according to the value quantification evaluation results;
[0067] S402: According to the value ranking result, a specified number of databases ranked lower are identified as non-critical databases;
[0068] S403: Eliminate the non-critical databases from the databases to be processed to obtain remaining critical databases to be processed.
[0069] In this embodiment, a quantitative value assessment was performed on each database to be processed, resulting in a value score for each database. This score reflects the database's value across multiple dimensions, including business importance, usage frequency, update frequency, data integrity, and data sensitivity. Based on the value assessment results, all databases to be processed were sorted in descending order, with databases with higher value scores placed first and those with lower value scores placed last. This clearly prioritizes each database and provides a basis for subsequent resource allocation and processing.
[0070] A specified number of databases ranked low are identified as non-critical databases, and the identified non-critical databases are removed from the list of databases to be processed. The remaining databases are critical databases, and these databases will enter the subsequent data classification and grading process. Specifically, a ratio threshold can be set based on business needs and resource conditions to determine the number of databases to be eliminated, such as eliminating the 20% of databases with the lowest value scores; or a fixed number of databases can be directly set for elimination. For example, the last 10 databases after sorting can be eliminated. Based on the set ratio threshold or fixed number, the databases ranked low are filtered out from the sorting results and marked as non-critical databases. By identifying non-critical databases, resources can be concentrated on processing critical databases, thereby improving resource utilization efficiency.
[0071] Preferably, when a non-critical database is removed, it is not necessary to directly delete it. Instead, it can be marked as "non-critical" and the data can be retained for possible future use. Alternatively, the non-critical database can be backed up and archived before being removed to ensure data security and recoverability.
[0072] In one embodiment, Figure 5 As shown, step S204 includes:
[0073] S501, performing keyword recognition on the data in the key database to obtain corresponding field keywords;
[0074] S502: Perform field matching on the field keyword and the data classification and grading directory to obtain the data category and data level of the matching field;
[0075] S503: batch labeling the corresponding data according to the data category and data level.
[0076] In this embodiment, during batch tagging, all data fields requiring classification and grading, including text fields, numeric fields, and date fields, are extracted from the key database. These extracted fields are then preprocessed, including removing spaces, standardizing case, and removing special characters, to improve the accuracy of keyword recognition. Keyword recognition is performed on the data using natural language processing tools or regular expressions to obtain corresponding field keywords. These identified field keywords can then be recorded in a structured file or database table to facilitate subsequent matching and tagging operations.
[0077] Match field keywords with the classification and grading standards in the catalog to determine the data category and data level of the matching fields. For example, the field keyword "customer name" is matched to the "customer information" category and the data level is "highly sensitive," etc. Field matching ensures the consistency of data classification and grading results, accurately classifies fields into the correct category and level, and avoids errors caused by manual operations. Based on the field matching results, use automated tools for batch labeling. For example, you can use SQL update statements to batch label fields in the database, marking fields in the "high sensitivity" category as "highly sensitive," and fields in the "medium sensitivity" category as "medium sensitive," etc. Batch labeling improves the efficiency of data classification and grading and reduces the workload of manual operations.
[0078] Optimally, automated labeling can be supplemented with manual review to correct possible mislabeling. For example, some fields may be mistakenly labeled "highly sensitive" and need to be manually corrected to "medium sensitive." Manual corrections should be recorded, including the time, person responsible, and reason for the correction, for subsequent audit and traceability. Manual review and correction based on automated labeling ensures the accuracy of the labeling results.
[0079] In one embodiment, before step S501, the method further includes:
[0080] Identifying repeated fields in the key database, and establishing a corresponding repeated field mapping table according to the storage information of the repeated fields;
[0081] The repeated fields in the key database are removed to obtain the remaining data to be marked.
[0082] In this embodiment, before starting batch tagging, duplicate fields in the key database are first deduplicated. Since there are a large number of duplicate items in the data fields in a large-scale data environment, these duplicate fields may come from different databases, data tables, or even backups or redundant storage of the same field in different systems. If deduplication is not performed and all fields are directly classified and tagged, the same field will be classified and tagged multiple times, wasting a lot of manpower and time. Therefore, duplicate fields in the key database are first identified. Specifically, direct matching can be performed by field name, for example, fields with exactly the same field names are considered duplicate fields; matching can also be performed by field content, for example, using a hash algorithm (such as MD5, SHA-256) to generate a hash value for the field content, and then comparing the hash value to detect duplicate fields; similarity detection can also be performed, that is, for fields with different field names or contents but similar, string similarity algorithms (such as cosine distance, Levenshtein distance) can be used for detection, etc.
[0083] Based on the identified repeated fields, a repeated field mapping table is created. The storage information of each repeated field is recorded in the mapping table, as shown in Table 1, including the field name, database, table name, field location, etc.:
[0084] Table 1
[0085] Duplicate field names Database Table name Field location Customer Name Database A Table 1 Column 1 Customer Name Database B Table 2 Column 2
[0086] Remove duplicate fields in key databases to obtain the remaining data to be labeled. Through deduplication processing, ensure that each field is classified and labeled only once, avoiding the classification and grading inconsistency caused by duplicate fields. At the same time, it can also greatly reduce the field data that needs to be labeled, significantly improving the efficiency of data classification and grading.
[0087] In one embodiment, after step S204, the method further includes:
[0088] Applying the data marking result to the corresponding repeated fields according to the repeated field mapping table;
[0089] The repeated fields with the marking results added thereto are restored to the corresponding original storage locations according to the storage information in the repeated field mapping table, and data verification is performed on the restored key database.
[0090] In this embodiment, during the deduplication stage, duplicate fields are identified and only one copy is retained for classification and labeling, while the storage mapping relationship of these fields is recorded. After the deduplication fields are classified and graded and labeled, and are assigned corresponding classification and grading labels according to their business value, sensitivity and other standards, the labeling results need to be restored, that is, the classification and labeling results are applied to all corresponding duplicate fields according to the mapping relationship recorded in the previously recorded duplicate field mapping table. That is, through deduplication, only one copy of the duplicate field needs to be labeled during the labeling stage, but during the restoration process, these classification and grading labels are copied to all duplicate fields. For example, if the field "customer name" is marked as "highly sensitive", this label is applied to all duplicate "customer name" fields. The classification and grading labels of all duplicate fields are completely consistent, ensuring the accuracy and consistency of data management.
[0091] According to the storage information in the repeated field mapping table, the repeated fields with the marking results added are restored to their original storage locations. For example, the "Customer Name" field marked as "Highly Sensitive" is restored to Table 1 of Database A and Table 2 of Database B. And the restored key database is verified to ensure that the marking results of all fields are correct. Specific data verification can include marking consistency, that is, checking whether the markings of all repeated fields are consistent; data integrity, that is, checking whether the restored fields are complete and whether there are omissions or errors; and data correlation, that is, checking whether the correlation between fields is normal, etc. By applying and restoring the marking results of repeated fields and verifying the data of the restored key database, it is ensured that the marking results of all repeated fields are consistent, avoiding data management confusion caused by repeated fields, and ensuring that the restored fields are complete and there are no omissions or errors, thereby ensuring data integrity.
[0092] It should be noted that there is not necessarily a certain order between the above steps. A person skilled in the art can understand, based on the description of the embodiments of the present invention, that in different embodiments, the above steps may have different execution orders, that is, they may be executed in parallel, or may be executed interchangeably, etc.
[0093] Further references Figure 6 , as a response to the above Figure 2 The present invention provides an embodiment of a data classification and grading device. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0094] like Figure 6 As shown, the data classification and grading device 60 described in this embodiment includes:
[0095] A directory construction module 601 is used to obtain a data classification basis file, parse the data classification basis file, and construct a corresponding data classification and grading directory;
[0096] The value assessment module 602 is used to determine all databases to be processed, collect value assessment indicators of each database to be processed, perform a quantitative value assessment on each database to be processed based on the value assessment indicators, and obtain corresponding quantitative value assessment results;
[0097] A database elimination module 603 is configured to determine non-critical databases from the database to be processed based on the value quantification evaluation result, and eliminate the non-critical databases from the database to be processed to obtain remaining critical databases to be processed;
[0098] The labeling processing module 604 is used to match the data in the key database with the data classification and grading directory, and perform batch labeling processing on the corresponding data according to the matched data categories and data levels.
[0099] The module referred to in the present invention refers to a series of computer program instruction segments that can perform specific functions. It is more suitable for describing the data classification and grading execution process than a program. For the specific implementation of each module, please refer to the corresponding method embodiment above, which will not be repeated here.
[0100] In one embodiment, the directory construction module 601 includes:
[0101] A file acquisition unit, configured to acquire a data classification basis file, and parse and acquire classification and grading elements in the data classification basis file;
[0102] An element screening unit, configured to screen the classification and grading elements according to a preset business scenario to obtain a classification and grading standard that is compatible with the preset business scenario;
[0103] The directory construction unit is used to construct an initial data classification and grading directory according to the classification and grading standard, and iteratively update the data classification and grading directory according to a preset update strategy.
[0104] In one embodiment, the database culling module 603 includes:
[0105] A value ranking unit, configured to rank the values of the databases to be processed according to the value quantification evaluation results;
[0106] A database determination unit is used to determine a specified number of databases ranked lower as non-critical databases based on the value ranking result;
[0107] The database elimination unit is used to eliminate the non-critical databases from the databases to be processed to obtain the remaining critical databases to be processed.
[0108] In one embodiment, the marking processing module 604 includes:
[0109] A keyword recognition unit, configured to perform keyword recognition on the data in the key database to obtain corresponding field keywords;
[0110] A field matching unit, configured to perform field matching between the field keyword and the data classification and grading directory, and obtain the data category and data level of the matching field;
[0111] The batch marking unit is used to perform batch marking processing on the corresponding data according to the data category and data level.
[0112] In one embodiment, the marking processing module 604 further includes:
[0113] a repeat identification unit, configured to identify repeated fields in the key database and establish a corresponding repeated field mapping table according to the storage information of the repeated fields;
[0114] The deduplication unit is used to remove repeated fields in the key database to obtain the remaining data to be marked.
[0115] In one embodiment, the marking processing module 604 further includes:
[0116] A repeated field marking unit, configured to apply the marking result of the data to the corresponding repeated field according to the repeated field mapping table;
[0117] The restoration and verification unit is used to restore the repeated fields with the marking results added to the corresponding original storage locations according to the storage information in the repeated field mapping table, and perform data verification on the restored key database.
[0118] In one embodiment, the value assessment indicators include business relevance, data usage frequency, data update frequency, data integrity, and data sensitivity.
[0119] In the above embodiment, the present invention discloses a data classification and grading device, which obtains a data classification basis file, parses the data classification basis file, and constructs a corresponding data classification and grading directory; determines all databases to be processed, collects value assessment indicators of each database to be processed, and performs value quantitative assessment on each database to be processed according to the value assessment indicator to obtain a corresponding value quantitative assessment result; determines non-critical databases in the database to be processed according to the value quantitative assessment result, and eliminates the non-critical databases from the database to be processed to obtain the remaining critical databases to be processed; matches the data in the critical database with the data classification and grading directory, and performs batch tagging on the corresponding data according to the data category and data level obtained by matching. By parsing and constructing a data classification and grading directory and automatically matching after confirming the critical database through value assessment, and automatically tagging the data in batches based on the matching results, the efficiency and accuracy of data classification and grading are improved, and data management efficiency and security are ensured.
[0120] Another embodiment of the present invention provides a computer device, such as Figure 7 As shown, the computer device 70 includes:
[0121] One or more processors 701 and memory 702, Figure 7 In the description, a processor 701 is used as an example. The processor 701 and the memory 702 can be connected via a bus or other means. Figure 7 The bus connection is taken as an example.
[0122] The processor 701 is used to complete various control logics of the computer device 70. It can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components or any combination of these components. In addition, the processor 701 can also be any traditional processor, microprocessor or state machine. The processor 701 can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP and / or any other such configuration.
[0123] Memory 702, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as program instructions corresponding to the data classification and grading method in the embodiments of the present invention. Processor 701 executes the non-volatile software programs, instructions, and modules stored in memory 702 to execute various functional applications and data processing of computer device 70, thereby implementing the data classification and grading method in the above-mentioned method embodiments.
[0124] The memory 702 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device 70, etc. In addition, the memory 702 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 702 may optionally include a memory remotely located relative to the processor 701, and these remote memories may be connected to the computer device 70 via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. One or more units are stored in the memory 702, and when executed by one or more processors 701, the steps of the data classification and grading method in any of the above-mentioned method embodiments are executed.
[0125] In the above embodiment, the present invention discloses a computer device, which obtains a data classification basis file, parses the data classification basis file, and constructs a corresponding data classification and grading directory; determines all databases to be processed, collects value assessment indicators of each database to be processed, and performs value quantitative assessment on each database to be processed according to the value assessment indicator to obtain a corresponding value quantitative assessment result; determines non-critical databases in the database to be processed according to the value quantitative assessment result, and eliminates the non-critical databases from the database to be processed to obtain the remaining critical databases to be processed; matches the data in the critical database with the data classification and grading directory, and performs batch tagging processing on the corresponding data according to the data category and data level obtained by matching. By parsing and constructing the data classification and grading directory and automatically matching after confirming the critical database through value assessment, and automatically tagging the data in batches based on the matching results, the efficiency and accuracy of data classification and grading are improved, and the efficiency and security of data management are ensured.
[0126] An embodiment of the present invention provides a non-volatile computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, the steps of the data classification and grading method in any of the above method embodiments are executed.
[0127] In the above embodiment, the present invention discloses a non-volatile computer-readable storage medium, which obtains a data classification basis file, parses the data classification basis file, and constructs a corresponding data classification and grading directory; determines all databases to be processed, collects value assessment indicators of each database to be processed, and performs value quantitative assessment on each database to be processed according to the value assessment indicator to obtain a corresponding value quantitative assessment result; determines non-critical databases in the database to be processed according to the value quantitative assessment result, and eliminates the non-critical databases from the database to be processed to obtain the remaining critical databases to be processed; matches the data in the critical database with the data classification and grading directory, and performs batch tagging processing on the corresponding data according to the data category and data level obtained by matching. By parsing and constructing the data classification and grading directory and automatically matching after confirming the critical database through value assessment, and automatically tagging the data in batches based on the matching results, the efficiency and accuracy of data classification and grading are improved, and the efficiency and security of data management are ensured.
[0128] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0129] The present invention can be used in a wide variety of general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0130] In summary, the present invention discloses a data classification and grading method, device, equipment and medium, the method includes: obtaining a data classification basis file, parsing the data classification basis file, and constructing a corresponding data classification and grading directory; determining all databases to be processed, collecting value assessment indicators of each database to be processed, and conducting a value quantitative assessment on each database to be processed according to the value assessment indicator to obtain a corresponding value quantitative assessment result; determining a non-critical database in the database to be processed according to the value quantitative assessment result, and eliminating the non-critical database from the database to be processed to obtain the remaining critical database to be processed; matching the data in the critical database with the data classification and grading directory, and batch tagging the corresponding data according to the data category and data level obtained by matching. By parsing and constructing a data classification and grading directory and automatically matching after confirming the critical database through value assessment, and batch automatically tagging the data based on the matching results, the efficiency and accuracy of data classification and grading are improved, and data management efficiency and security are ensured.
[0131] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes in the above-described method embodiments. The storage medium can be a memory, a magnetic disk, a floppy disk, a flash memory, an optical storage device, etc.
[0132] It should be noted that if any software tools or components not developed by our company appear in the examples of this application, they are for illustration purposes only and do not represent actual use. It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications shall fall within the scope of protection of the appended claims.
Claims
1. A data classification and grading method, characterized in that: include: Obtaining a data classification basis file, parsing the data classification basis file, and constructing a corresponding data classification and grading directory; Determine all databases to be processed, collect value assessment indicators for each database to be processed, perform value quantitative assessment on each database to be processed based on the value assessment indicators, and obtain corresponding value quantitative assessment results; Determine non-critical databases in the database to be processed according to the value quantification evaluation result, and remove the non-critical databases from the database to be processed to obtain remaining critical databases to be processed; The data in the key database is matched with the data classification and grading directory, and the corresponding data is batch-labeled according to the matched data categories and data levels.
2. The data classification and grading method according to claim 1, characterized in that: The obtaining of the data classification basis file, parsing the data classification basis file, and constructing a corresponding data classification and grading directory includes: Obtaining a data classification basis file, parsing and obtaining classification and grading elements in the data classification basis file; Filtering the classification and grading elements according to a preset business scenario to obtain a classification and grading standard that is compatible with the preset business scenario; An initial data classification and grading directory is constructed according to the classification and grading standards, and the data classification and grading directory is iteratively updated according to a preset update strategy.
3. The data classification and grading method according to claim 1, characterized in that: The determining of non-critical databases in the database to be processed according to the value quantification assessment, and removing the non-critical databases from the database to be processed to obtain the remaining critical databases to be processed includes: Rank the values of the databases to be processed according to the quantitative value assessment results; According to the value ranking results, a specified number of databases ranked lower are identified as non-critical databases; The non-critical databases are removed from the databases to be processed to obtain the remaining critical databases to be processed.
4. The data classification and grading method according to claim 1, characterized in that: The step of matching the data in the key database with the data classification and grading directory, and batch tagging the corresponding data according to the matched data categories and data levels, includes: Perform keyword recognition on the data in the key database to obtain corresponding field keywords; Perform field matching on the field keyword and the data classification and grading directory to obtain the data category and data level of the matching field; The corresponding data is batch labeled according to the data category and data level.
5. The data classification and grading method according to claim 4, characterized in that: Before performing keyword recognition on the data in the key database to obtain corresponding field keywords, the method further includes: Identifying repeated fields in the key database, and establishing a corresponding repeated field mapping table according to the storage information of the repeated fields; The repeated fields in the key database are removed to obtain the remaining data to be marked.
6. The data classification and grading method according to claim 5, characterized in that: After batch labeling the corresponding data according to the data category and data level, the method further includes: Applying the data marking result to the corresponding repeated fields according to the repeated field mapping table; The repeated fields with the marking results added thereto are restored to the corresponding original storage locations according to the storage information in the repeated field mapping table, and data verification is performed on the restored key database.
7. The data classification and grading method according to any one of claims 1 to 6, characterized in that: The value assessment indicators include business relevance, data usage frequency, data update frequency, data integrity and data sensitivity.
8. A data classification and grading device, characterized in that: include: A directory construction module is used to obtain a data classification basis file, parse the data classification basis file, and construct a corresponding data classification and grading directory; A value assessment module is used to determine all databases to be processed, collect value assessment indicators of each database to be processed, perform value quantitative assessment on each database to be processed according to the value assessment indicators, and obtain corresponding value quantitative assessment results; A database elimination module is used to determine non-critical databases in the database to be processed according to the value quantification evaluation result, and eliminate the non-critical databases from the database to be processed to obtain the remaining critical databases to be processed; The marking processing module is used to match the data in the key database with the data classification and grading directory, and perform batch marking processing on the corresponding data according to the matched data categories and data levels.
9. A computer device, characterized in that: comprising at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the data classification and grading method according to any one of claims 1 to 7.
10. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores computer-executable instructions, which, when executed by one or more processors, enable the one or more processors to execute the data classification and grading method according to any one of claims 1 to 7.