Data classification method and device, electronic equipment and storage medium

By obtaining the database table type and multi-level cache results, and using fingerprint or lineage analysis methods, the problems of low classification efficiency and resource waste caused by the large number of data tables in the database are solved, and efficient and accurate data classification is achieved.

CN120631992APending Publication Date: 2025-09-12BEIJING YUANYUAN SHUAN TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510762961.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In the prior art, the large number of data tables in the database leads to low efficiency of traversal scanning and classification, high resource consumption, and the need to scan the same or similar tables repeatedly, which increases the classification time and wastes resources.

Method used

By obtaining the table type of the data table, selecting the corresponding data analysis method (such as fingerprint analysis or lineage analysis), combining multi-level cache results, priority levels and cached results for analysis, reducing repeated scans, using fingerprint information or dependent table and column mapping information for queries, and only scanning and caching the results of tables that have not been queried.

Benefits of technology

It improves the accuracy and efficiency of data classification, reduces resource consumption, lowers CPU and I/O load, and ensures that data classification can be performed without affecting other businesses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631992A_ABST
    Figure CN120631992A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a data classification method and device, electronic equipment and a storage medium. The method comprises the steps that a table type corresponding to a first data table in a database is obtained, and the first data table is any one of at least one data table which is not classified in the database; acquiring a data analysis mode corresponding to the table type; a multi-level cache result corresponding to the first data table is obtained, the multi-level cache result comprises a classification result of at least one second data table, and the at least one second data table is a data table of which the classification result is determined before the first data table; according to the priority corresponding to the data analysis mode and the multi-level cache result, the data analysis mode is adopted to analyze the first data table, and a classification result corresponding to the first data table is obtained. According to the invention, the data classification efficiency can be improved while the data classification accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a data classification method, device, electronic device, and storage medium. Background Art

[0002] As data security requirements continue to increase, the need for data classification and grading in databases is growing. For example, a scanning program can perform a traversal scan of the database, applying predefined classification rules to the information collected from a data table. The corresponding classification information for the data table can be determined based on whether the rules match. However, due to the large number of data tables in the database, traversal scanning increases classification time, resulting in low classification efficiency and increased resource consumption. Summary of the Invention

[0003] The present disclosure provides a data classification method, device, electronic device, and storage medium that can improve data classification accuracy while also improving data classification efficiency. The technical solutions of the present disclosure are as follows:

[0004] According to a first aspect of an embodiment of the present disclosure, a data classification method is provided, comprising:

[0005] Obtaining a table type corresponding to a first data table in a database, wherein the first data table is any data table in at least one unclassified data table in the database;

[0006] Obtaining a data analysis method corresponding to the table type;

[0007] Acquire a multi-level cache result corresponding to the first data table, wherein the multi-level cache result includes a classification result of at least one second data table, and the at least one second data table is a data table for which a classification result has been determined before the first data table;

[0008] According to the priority corresponding to the data analysis method and the multi-level cache result, the first data table is analyzed using the data analysis method to obtain a classification result corresponding to the first data table.

[0009] According to some embodiments, analyzing the first data table using the data analysis method to obtain a classification result corresponding to the first data table includes:

[0010] In a case where the table type of the first data table is the first data table type, determining to adopt a fingerprint analysis method to analyze the first data table to obtain fingerprint information of the first data table;

[0011] According to the similarity threshold requirement, query the fingerprint information in the multi-level cache results to obtain a first query result;

[0012] If the first query result indicates that a classification result set corresponding to the fingerprint information is found, determining a classification result of the first data table according to the classification result set, wherein the classification result set includes at least one classification result;

[0013] or,

[0014] If the first query result indicates that no classification result corresponding to the fingerprint information is found, the first data table is scanned to obtain the classification result corresponding to the first data table, and the classification result corresponding to the first data table is cached in multiple levels.

[0015] According to some embodiments, obtaining fingerprint information of the first data table includes:

[0016] Extracting key features from metadata of the first data table;

[0017] Processing the key features to obtain processed key features;

[0018] A hash value operation is performed on the processed key features to obtain a feature mapping value of a preset length, and the feature mapping value of the preset length is used as the fingerprint information of the first data table.

[0019] According to some embodiments, analyzing the first data table using the data analysis method to obtain a classification result corresponding to the first data table includes:

[0020] In a case where the table type of the first data table is the second data table type, determining to analyze the first data table in a lineage analysis manner to obtain dependent tables and column mapping information of the first data table;

[0021] According to the similarity threshold requirement, query the dependent table and column mapping information in the multi-level cache result to obtain a second query result;

[0022] When the second query result indicates that a dependent table result corresponding to the dependent table is found, obtaining a classification result of the first data table according to the column mapping information;

[0023] or,

[0024] When the second query result indicates that no dependency table result corresponding to the dependency table is found, the first data table is scanned to obtain the classification result corresponding to the first data table, and the classification result corresponding to the first data table is cached in multiple levels.

[0025] According to some embodiments, analyzing the first data table by using a lineage analysis method to obtain dependent tables and column mapping information of the first data table includes:

[0026] Obtaining a definition statement corresponding to the first data table;

[0027] Parsing the definition statement corresponding to the first data table to obtain structured query language SQL information corresponding to the first data table;

[0028] Acquire column mapping information according to the structured query language SQL information and reference information of the SELECT statement, wherein the column mapping information is used to indicate mapping information between view columns and source table columns;

[0029] The structured query language SQL information is analyzed and processed to obtain a dependency table.

[0030] According to some embodiments, obtaining a classification result of the first data table according to the column mapping information includes:

[0031] When the second query result indicates that the dependent table result corresponding to the dependent table is queried, and multiple historical classification results are obtained according to the column mapping information, the multiple historical classification results are analyzed according to the column information corresponding to each historical classification result to obtain the classification result of the first data table.

[0032] According to some embodiments, the multi-level caching of the classification results corresponding to the first data table includes:

[0033] Obtaining cache level information corresponding to the classification result of the first data table;

[0034] When the cache level information indicates that the classification result of the first data table is to be cached at the table level, storing the classification result of the first data table;

[0035] When the cache level information indicates that column-level caching is performed on the classification results of the first data table, the classification results of each column in the first data table are stored.

[0036] According to a second aspect of an embodiment of the present disclosure, there is provided a data classification device, comprising:

[0037] A table type acquiring unit, configured to acquire a table type corresponding to a first data table in a database, wherein the first data table is any data table in at least one unclassified data table in the database;

[0038] A method acquisition unit, configured to acquire a data analysis method corresponding to the table type;

[0039] a result acquisition unit, configured to acquire a multi-level cache result corresponding to the first data table, wherein the multi-level cache result includes a classification result of at least one second data table, and the at least one second data table is a data table for which a classification result has been determined before for the first data table;

[0040] A data classification unit is configured to analyze the first data table using the data analysis method according to the priority corresponding to the data analysis method and the multi-level cache result, and obtain a classification result corresponding to the first data table.

[0041] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, including:

[0042] processor;

[0043] a memory for storing instructions executable by the processor;

[0044] The processor is configured to execute the instructions to implement the data classification method described in any one of the aforementioned aspects.

[0045] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is provided. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the data classification method described in any one of the aforementioned aspects.

[0046] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which implements the method described in any one of the aforementioned aspects when executed by a processor.

[0047] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:

[0048] In some or related embodiments, a method is provided for determining a table type corresponding to a first data table in a database, wherein the first data table is any data table in at least one data table that has not been classified in the database; obtaining a data analysis method corresponding to the table type; obtaining a multi-level cache result corresponding to the first data table, wherein the multi-level cache result includes a classification result of at least one second data table, wherein the at least one second data table is a data table for which a classification result has been previously determined for the first data table; and analyzing the first data table using the data analysis method according to a priority corresponding to the data analysis method and the multi-level cache result to obtain a classification result corresponding to the first data table. Therefore, the data processing method can be determined by the table type of the data table, thereby improving the accuracy of data processing. The data table can be classified using the multi-level cache result without repeatedly scanning the same or similar tables or traversing the data tables in the database, thereby reducing the time required for data table classification and the resources required for traversing all data tables. Furthermore, the multi-level cache result and the data analysis method corresponding to the data table can reduce the situation where the same classification rules are used for all types of tables, resulting in inaccurate data table classification, thereby improving data classification accuracy and data classification efficiency.

[0049] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0051] Figure 1 is a flow chart of a first data classification method provided by an embodiment of the present disclosure;

[0052] Figure 2 is a flow chart of a second data classification method provided by an embodiment of the present disclosure;

[0053] Figure 3 is a flow chart of a third data classification method provided by an embodiment of the present disclosure;

[0054] Figure 4 This is an example diagram of a data table type provided by an embodiment of the present disclosure;

[0055] Figure 5 This is a schematic diagram of an example of a fingerprint information acquisition method provided by an embodiment of the present disclosure;

[0056] Figure 6 This is a schematic diagram of an example of a fingerprint matching method provided by an embodiment of the present disclosure;

[0057] Figure 7 is an example schematic diagram of an association algorithm provided by an embodiment of the present disclosure;

[0058] Figure 8 This is a schematic diagram illustrating an example of a blood relationship analysis method provided by an embodiment of the present disclosure;

[0059] Figure 9 This is an example schematic diagram of a multi-level caching method provided by an embodiment of the present disclosure;

[0060] Figure 10 is a flowchart of a fourth data classification method provided by an embodiment of the present disclosure;

[0061] Figure 11 is a block diagram of a data classification device provided by an embodiment of the present disclosure;

[0062] Figure 12 The figure is a schematic diagram showing an example of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0063] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0064] The present disclosure provides a data classification method, apparatus, electronic device, and storage medium. In some embodiments, the terms "data classification method" and "information processing method" and "communication method" are interchangeable; the terms "data classification apparatus" and "information processing apparatus" and "communication apparatus" are interchangeable; and the terms "information processing system" and "communication system" are interchangeable.

[0065] The embodiments of the present disclosure are not exhaustive and are merely illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined. For example, some or all steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.

[0066] In each embodiment of the present disclosure, unless otherwise specified or provided for by logic, the terms and / or descriptions between the embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form a new embodiment based on their inherent logical relationships.

[0067] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.

[0068] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular, such as "a", "an", "the", "above", "said", "the", "the", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun following the article may be understood as a singular expression or a plural expression.

[0069] In the embodiments of the present disclosure, “plurality” refers to two or more.

[0070] In some embodiments, the terms "at least one," "one or more," "a plurality of," "multiple," etc. may be used interchangeably.

[0071] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects and do not constitute any restriction on the position, order, priority, quantity or content of the description objects. For the statement of the description object, please refer to the description in the context of the claims or embodiments, and no unnecessary restriction should be constituted due to the use of prefixes. For example, if the description object is a "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields". "First" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the description object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of description objects is not limited by the ordinal number and can be one or more. Taking "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes can be the same or different. For example, if the description object is "device", then the "first device" and the "second device" can be the same device or different devices, and their types can be the same or different; for another example, if the description object is "information", then the "first information" and the "second information" can be the same information or different information, and their contents can be the same or different.

[0072] In some embodiments, "terminal" or "terminal device" may be referred to as "user equipment (UE)", "user terminal" "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc.

[0073] In some embodiments, data, information, etc. may be obtained with the user's consent.

[0074] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0075] As data security requirements become increasingly stringent, users are increasingly demanding the classification and grading of data in databases. For example, there is a need to identify the names of people or account information in a data table. Data classification and grading can provide the primary foundation for subsequent data management. When performing data classification and grading, a scanning program can be used to traverse and scan the target database, collecting metadata information for each database table. Some data may also be sampled. The collected information is then matched against pre-defined classification rules. If a rule is successfully matched, it is considered to be classified as the business tag corresponding to the rule. When all classification rules are matched sequentially for all tables in the database, the classification and grading task is confirmed to be complete.

[0076] However, when the database has a large amount of data, the amount of data to be scanned is enormous, and a complete scan can take a long time. Furthermore, the sampling and rule matching process must be repeated for identical or similar tables, resulting in low classification efficiency and wasted resources. Directly scanning views can incur additional computational overhead and classification steps, increasing the database burden.

[0077] Figure 1 This is a flow chart of the first data classification method provided by the embodiment of the present disclosure. Figure 1 As shown, the data classification method can be used in a scenario where a large number of data tables are classified in a database, and includes the following steps:

[0078] In step S11, a table type corresponding to a first data table in a database is obtained, wherein the first data table is any data table in at least one unclassified data table in the database;

[0079] In some embodiments, the execution subject of the embodiments of the present disclosure may be, for example, an electronic device. The electronic device does not specifically refer to a fixed electronic device. For example, when the device identification changes, the electronic device may also change accordingly. For example, when the structure of the electronic device changes, the electronic device may also change accordingly. Among them, the execution subject of the embodiments of the present disclosure may also be, for example, a server. The server may be, for example, a single server or a server cluster, and the embodiments of the present disclosure are not limited to this.

[0080] In some embodiments, the database may be, for example, a library storing at least one data table. The database is not specifically a fixed database. For example, when the number of data tables included in the database changes, the database may also change accordingly. For example, when a data table in the database changes, the database may also change accordingly.

[0081] In some embodiments, the first data table is any one of at least one unclassified data table in the database. The first data table does not specifically refer to a fixed data table. For example, when the content of the first data table changes, the first data table may also change accordingly. For example, when the table structure of the first data table changes, the first data table may also change accordingly.

[0082] In some embodiments, the table type may be used to indicate the type of the first data table. The table type may include, for example, a common table type and a view type, but the embodiments of the present disclosure are not limited thereto. When the first data table changes, the table type may also change accordingly. For example, when the method for determining the table type changes, the first data table may also change accordingly.

[0083] In some embodiments, a table type corresponding to a first data table in a database is obtained, wherein the first data table is any data table in at least one unclassified data table in the database.

[0084] In step S12, a data analysis method corresponding to the table type is obtained;

[0085] In some embodiments, the data analysis method may be used to indicate the analysis method used when obtaining the classification result of the first data table. The data analysis method may include, for example, a blood relationship analysis method and a fingerprint analysis method, or a combination of the two, which is not limited in the present embodiment.

[0086] In some embodiments, the data analysis method may correspond to a table type, for example, that is, different table types may correspond to different data analysis methods.

[0087] In some embodiments, a data analysis method corresponding to the table type may be obtained.

[0088] In step S13, a multi-level cache result corresponding to the first data table is obtained, wherein the multi-level cache result includes a classification result of at least one second data table, and the at least one second data table is a data table for which a classification result has been determined before the first data table;

[0089] According to some embodiments, the multi-level cache result may, for example, correspond to the first data table, and more specifically, may correspond to the table type of the first data table. The multi-level cache result does not specifically refer to a fixed result. For example, when the first data table changes, the multi-level cache result may also change accordingly. For example, when the storage method of the cache result changes, the multi-level cache result may also change accordingly.

[0090] In some embodiments, the multi-level cache results include classification results from at least one second data table, where the at least one second data table is a data table for which classification results have been determined before the first data table. For example, the second data table is a data table for which classification results have been determined before the first data table. The second in the second data table is used to distinguish it from the first data table and does not specifically refer to a fixed data table. For example, when the data corresponding to the second data table changes, the second data table may also change accordingly.

[0091] In some embodiments, a multi-level cache result corresponding to the first data table is obtained, wherein the multi-level cache result includes classification results of at least one second data table, and the at least one second data table is a data table for which classification results have been determined before the first data table.

[0092] In step S14 , the first data table is analyzed using the data analysis method according to the priority corresponding to the data analysis method and the multi-level cache result, to obtain the classification result corresponding to the first data table.

[0093] In some embodiments, the priority may be, for example, the priority corresponding to the data analysis method, i.e., different data analysis methods may correspond to different priorities. For example, the priority of the fingerprint analysis method may be higher than the priority of the bloodline analysis method. That is, the fingerprint analysis method may be used to process the corresponding data table first, and then the bloodline analysis method may be used to process the corresponding data table. Therefore, the optimal analysis path can be selected based on the first data table, which can reduce the cost of obtaining the classification results. The priority can be determined, for example, based on the table type, table size, and available information of the data table.

[0094] According to some embodiments, the classification result may be, for example, the result obtained after classifying the first data table. The classification result is not specifically a fixed result. For example, when the first data table changes, the classification result may also change accordingly. For example, when the multi-level cache results change, the classification result may also change accordingly.

[0095] In some or related embodiments, a method is provided for obtaining a table type corresponding to a first data table in a database, wherein the first data table is any data table in at least one data table that has not been classified in the database; obtaining a data analysis method corresponding to the table type; obtaining a multi-level cache result corresponding to the first data table, wherein the multi-level cache result includes a classification result of at least one second data table, and the at least one second data table is a data table for which a classification result has been determined before for the first data table; and analyzing the first data table using the data analysis method according to a priority corresponding to the data analysis method and the multi-level cache result to obtain a classification result corresponding to the first data table. Therefore, the data processing method can be determined by the table type of the data table, thereby improving the accuracy of data processing. By classifying the data table using the multi-level cache result, it is not necessary to traverse the data tables in the database, thereby reducing the time required for data table classification and the resources required for traversing all data tables. Furthermore, by using the multi-level cache result and the data analysis method corresponding to the data table, it is possible to reduce the situation where the same classification rules are used for all types of tables, resulting in inaccurate data table classification, thereby improving the accuracy of data classification and the efficiency of data classification. Furthermore, intelligent analysis and result reuse mechanisms can reduce database query times and sampling volume, lowering the CPU and I / O load, allowing data classification to proceed without impacting other business operations. Furthermore, data processing methods can be determined based on the data table, dynamically adjusting analysis strategies and sampling methods to provide optimal analysis paths for different table types.

[0096] Figure 2 is a flow chart of a second data classification method provided by an embodiment of the present disclosure, such as Figure 2 As shown, this method can be used in classification scenarios of complex data in large enterprises, including the following steps:

[0097] In step S21, a table type corresponding to a first data table in a database is obtained, wherein the first data table is any data table in at least one unclassified data table in the database;

[0098] Among them, the relevant descriptions have been mentioned above and will not be repeated here.

[0099] In some embodiments, the execution subject of the embodiments of the present disclosure may be, for example, an electronic device. The electronic device does not specifically refer to a fixed electronic device. For example, when the device identification changes, the electronic device may also change accordingly. For example, when the structure of the electronic device changes, the electronic device may also change accordingly. Among them, the execution subject of the embodiments of the present disclosure may also be, for example, a server. The server may be, for example, a single server or a server cluster, and the embodiments of the present disclosure are not limited to this.

[0100] According to some embodiments, Figure 3 is a flow chart of the third data classification method provided by the embodiment of the present disclosure, such as Figure 3 As shown, for example, a scan instruction for a database may be received, and a database scan task may be executed in response to the scan instruction. In the initialization phase, tables in the database may be traversed to obtain a table type of a first data table. Figure 4 This is an example diagram of a data table type provided by an embodiment of the present disclosure, such as Figure 4 As shown, the table type may include, for example, common tables, views, materialized views, temporary tables and other table types.

[0101] In step S22, a data analysis method corresponding to the table type is obtained;

[0102] Among them, the relevant descriptions have been mentioned above and will not be repeated here.

[0103] Among them, Figure 4 As shown, the data analysis method corresponding to the ordinary table type can be, for example, a fingerprint analysis method, the data analysis method corresponding to the view table type can be, for example, a lineage analysis method, the data analysis method corresponding to the materialized view can be, for example, a hybrid analysis method, and the data analysis method corresponding to the temporary table can be, for example, a simplified analysis method.

[0104] In some embodiments, such as Figure 3 As shown, the table structure fingerprint analysis can be performed on the common table, and the classification result can be determined based on the analysis result.

[0105] In step S23, a multi-level cache result corresponding to the first data table is obtained, wherein the multi-level cache result includes a classification result of at least one second data table, and the at least one second data table is a data table for which a classification result has been determined before the first data table;

[0106] Among them, the relevant descriptions have been mentioned above and will not be repeated here.

[0107] In step S24, the first data table is analyzed using the data analysis method according to the priority corresponding to the data analysis method and the multi-level cache result to obtain the classification result corresponding to the first data table;

[0108] Among them, the relevant descriptions have been mentioned above and will not be repeated here.

[0109] According to some embodiments, obtaining a classification result corresponding to the first data table includes:

[0110] When the table type of the first data table is the first data table type, determining to analyze the first data table in a fingerprint analysis manner to obtain fingerprint information of the first data table;

[0111] According to the similarity threshold requirement, the fingerprint information is queried in the multi-level cache results to obtain the first query result;

[0112] In a case where the first query result indicates that a classification result set corresponding to the fingerprint information is queried, determining a classification result of the first data table according to the classification result set, wherein the classification result set includes at least one classification result;

[0113] or,

[0114] If the first query result indicates that no classification result corresponding to the fingerprint information is found, the first data table is scanned to obtain the classification result corresponding to the first data table, and the classification result corresponding to the first data table is cached in multiple levels. Therefore, the classification result can be obtained based on the fingerprint information, which can improve the accuracy of the classification result.

[0115] According to some embodiments, the first data table type may be, for example, a common table type. The fingerprint information may be used to uniquely identify the first data table.

[0116] In some embodiments, obtaining the first query result may be, for example, obtaining the query result based on similarity information, and specifically, for example, determining the query result based on similarities between different data tables.

[0117] According to some embodiments, obtaining fingerprint information of the first data table includes:

[0118] extracting key features from metadata of the first data table;

[0119] Processing the key features to obtain the processed key features;

[0120] A hash value operation is performed on the processed key features to obtain a feature mapping value of a preset length, and the feature mapping value of the preset length is used as the fingerprint information of the first data table.

[0121] In some embodiments, table structure feature mapping can be a technique for identifying and matching similar data structures. Therefore, since a large number of data tables share the same or similar table structures, the table structures can be identified, reducing repeated scans of data tables and unnecessary duplication of work. By using table structure fingerprinting and classification result reuse, classification can be propagated along data flow paths, significantly improving classification accuracy and reducing the necessary scan scope. This can reduce the number of tables requiring a complete evaluation, shorten classification time, and improve classification efficiency.

[0122] In some embodiments, key features extracted from metadata may include, for example, column names, data types, field lengths, and constraint information. Processing the key features may include, for example, standardizing the key features to reduce differences in substantive structures.

[0123] When the complexity of table or column names is less than a complexity threshold, feature mapping can be used to improve the accuracy of obtaining structurally similar tables. For example, even if the table or column names are different, structurally similar tables can still be identified.

[0124] In some embodiments, Figure 5 This is an example diagram of a fingerprint information acquisition method provided by an embodiment of the present disclosure. Figure 5 As shown, metadata information for all columns of the first data table can be obtained, and the columns can be sorted by name to improve the consistency of the generated feature values. For each column, features such as the column name, data type, and length can be extracted. The obtained features can be combined and converted into a feature string in a fixed format. The feature string can be hashed to obtain a feature mapping value, and the feature mapping value can be used as fingerprint information.

[0125] In some embodiments, Figure 6 This is an example diagram of a fingerprint matching method provided by an embodiment of the present disclosure. Figure 6 As shown, the electronic device may include, for example, a scanner, a fingerprint analyzer, and a result cache module. The scanner may be controlled to request analysis of a data table structure. The fingerprint analyzer may obtain fingerprint information from the data table and use the result cache module to query whether the fingerprint information exists. If so, the cached classification result may be returned to the scanner. If no query result is obtained, an empty result may be returned to the scanner, and a regular scan of the data table may be performed to obtain a new scan result. The fingerprint information and scan result are then stored. The scan result is the classification result.

[0126] According to some embodiments, obtaining a classification result corresponding to the first data table includes:

[0127] In a case where the table type of the first data table is the second data table type, determining to analyze the first data table in a lineage analysis manner to obtain dependent tables and column mapping information of the first data table;

[0128] According to the similarity threshold requirement, query the dependent table and column mapping information in the multi-level cache results to obtain the second query result;

[0129] When the second query result indicates that a dependent table result corresponding to the dependent table is found, obtaining a classification result of the first data table according to the column mapping information;

[0130] or,

[0131] If the second query result indicates that no dependent table result corresponding to the dependent table has been found, the first data table is scanned and processed to obtain the classification results corresponding to the first data table, and the classification results corresponding to the first data table are cached at multiple levels. Therefore, classification results can be obtained based on the data source of each view column, and the dependency relationship and column mapping relationship between the view and the base table can be automatically identified. This can achieve intelligent automatic propagation without the additional computational overhead of directly scanning the view data, reducing the calculation of complex classifications and the burden on the database. Intelligently propagating the tags of the base table to the view column avoids direct querying of the view data and ensures the consistency of the classification results.

[0132] According to some embodiments, the second data table type may be, for example, a view table type. Lineage analysis may also be referred to as data association analysis, which refers to techniques for identifying and tracking dependencies between data objects and is particularly applicable to analyzing mapping relationships between views and base tables. Views are defined based on base tables. If the sensitive data classification results of the base tables are known, the sensitive classification results of the data in the view can be directly derived without actually scanning the view.

[0133] According to some embodiments, analyzing the first data table using a lineage analysis method to obtain dependent tables and column mapping information of the first data table includes:

[0134] Obtaining a definition statement corresponding to the first data table;

[0135] Parsing the definition statement corresponding to the first data table to obtain structured query language SQL information corresponding to the first data table;

[0136] Obtaining column mapping information according to structured query language SQL information and reference information of a SELECT statement, wherein the column mapping information is used to indicate mapping information between view columns and source table columns;

[0137] Analyze and process the structured query language SQL information to obtain dependent tables.

[0138] According to some embodiments, obtaining a classification result of the first data table according to the column mapping information includes:

[0139] When the second query result indicates that the dependent table result corresponding to the dependent table is queried, and multiple historical classification results are obtained according to the column mapping information, the multiple historical classification results are analyzed according to the column information corresponding to each historical classification result to obtain the classification result of the first data table.

[0140] In one embodiment of the present disclosure, data association analysis may include, for example:

[0141] a. Relationship extraction: Analyze view definition statements and identify the mapping relationship between view columns and source table columns;

[0142] b. Relationship graph construction: Build a complete data association graph, including direct and indirect dependencies;

[0143] c. Sensitivity propagation: Based on association relationships, known sensitive classifications are propagated from the source table to the view;

[0144] d. Sensitivity assessment: When multiple source columns affect a view column, the final sensitivity classification is comprehensively evaluated.

[0145] The comprehensive analysis may include, for example, a comprehensive evaluation based on the identified types, and may also include a comprehensive evaluation based on the name similarity.

[0146] in, Figure 7 is an example schematic diagram of an association algorithm provided by an embodiment of the present disclosure, such as Figure 7 As shown, this process can include: 1. Obtaining the view definition statement; 2. Parsing the definition statement using syntax analysis techniques to generate a structured representation; 3. Identifying column references in the SELECT statement and mapping view columns to source table columns; 4. Processing various SQL expressions (such as functions, computed columns, and aliases); and 5. Building a complete column-level dependency graph. Therefore, parsing the data alignment definition can improve the accuracy of obtaining column-level mapping relationships.

[0147] According to some embodiments, Figure 8 This is an example diagram of a blood relationship analysis method provided by an embodiment of the present disclosure. Figure 8 As shown, the electronic device may include, for example, a scanner, a fingerprint analyzer, and a result cache module. The scanner may be controlled to request an analysis view, parse the view SQL definition, extract dependent tables and column mappings, and query each dependent table. When a dependent table result is obtained, the dependent table result may be returned and the column mapping information may be applied. When no dependent table result is obtained, a message indicating that the dependent table result was not obtained may be returned, and a propagation failure message may be marked.

[0148] The electronic device can control the scanner to analyze all dependent table results to obtain query results. When partial query results or empty results are obtained, the data table can be scanned normally and the view scan results can be cached. The view scan results are the classification results.

[0149] In step S25, when it is determined that the classification result corresponding to the first data table is cached in multiple levels, cache level information corresponding to the classification result of the first data table is obtained;

[0150] Among them, the relevant descriptions have been mentioned above and will not be repeated here.

[0151] Multi-level caching can, for example, achieve efficient result sharing by storing and reusing classification results at different granularities. This mechanism includes three levels of caching, each targeting different reuse scenarios.

[0152] In step S26, when the cache level information indicates that the classification result of the first data table is to be cached at the table level, the classification result of the first data table is stored;

[0153] Among them, the relevant descriptions have been mentioned above and will not be repeated here.

[0154] In step S27 , when the cache level information indicates that the classification results of the first data table are to be cached at the column level, the classification results of each column in the first data table are stored.

[0155] Among them, the relevant descriptions have been mentioned above and will not be repeated here.

[0156] According to some embodiments, Figure 9 This is an example diagram of a multi-level cache method provided by an embodiment of the present disclosure. Figure 9 As shown, multi-level caching can include table-level caching and column-level caching. Table-level caching can store the classification results of the entire data table and is suitable for data tables with high structural similarity. Column-level caching can, for example, store the classification results of a single column and is suitable for columns with the same definition across tables.

[0157] According to some embodiments, Figure 10 is a flowchart of the fourth data classification method provided by the embodiment of the present disclosure, such as Figure 10 As shown, the method includes:

[0158] 1. Task initialization: receive classification tasks, obtain table information and database connection;

[0159] 2. Metadata acquisition: read the table structure information and necessary context data;

[0160] 3. Feature map generation: calculate the structural feature map value of the table;

[0161] 4. Cache query: Find matching classification results in multiple layers of cache;

[0162] 5. Association analysis: If it is a view, perform association analysis and derive the results;

[0163] 6. Traditional analysis: If the first two steps fail to obtain complete results, perform traditional rule evaluation;

[0164] 7. Result caching: storing classification results in a multi-layer caching system;

[0165] 8. Result return: Return the classification results to the system.

[0166] In one or related embodiments, when it is determined that the classification results corresponding to the first data table are to be cached at multiple levels, cache level information corresponding to the classification results of the first data table can be obtained. When the cache level information indicates that the classification results of the first data table are to be cached at the table level, the classification results of the first data table are stored; when the cache level information indicates that the classification results of the first data table are to be cached at the column level, the classification results of each column in the first data table are stored. Therefore, the classification results of the first data table can be cached accordingly, the accuracy of subsequent data table classification can be improved, and incremental classification capabilities can be achieved through table structure change detection and cache management mechanisms. Only newly added and changed tables can be processed to improve classification efficiency, and the system can be applied to the periodic classification needs of large databases.

[0167] A block diagram of a data classification device according to an exemplary embodiment is shown. Figure 11 , the apparatus 1100 comprises:

[0168] A table type acquisition unit 1101 is configured to acquire a table type corresponding to a first data table in a database, wherein the first data table is any data table in at least one unclassified data table in the database;

[0169] A method acquisition unit 1102 is used to acquire a data analysis method corresponding to the table type;

[0170] A result acquisition unit 1103 is configured to acquire a multi-level cache result corresponding to the first data table, wherein the multi-level cache result includes a classification result of at least one second data table, and the at least one second data table is a data table for which a classification result has been determined before the first data table;

[0171] The data classification unit 1104 is configured to analyze the first data table using the data analysis method according to the priority corresponding to the data analysis method and the multi-level cache result, and obtain a classification result corresponding to the first data table.

[0172] According to some embodiments, the data classification unit 1104 is configured to analyze the first data table using a data analysis method to obtain a classification result corresponding to the first data table, specifically for:

[0173] When the table type of the first data table is the first data table type, determining to analyze the first data table in a fingerprint analysis manner to obtain fingerprint information of the first data table;

[0174] According to the similarity threshold requirement, the fingerprint information is queried in the multi-level cache results to obtain the first query result;

[0175] In a case where the first query result indicates that a classification result set corresponding to the fingerprint information is queried, determining a classification result of the first data table according to the classification result set, wherein the classification result set includes at least one classification result;

[0176] or,

[0177] When the first query result indicates that no classification result corresponding to the fingerprint information is found, the first data table is scanned to obtain the classification result corresponding to the first data table, and the classification result corresponding to the first data table is cached in multiple levels.

[0178] According to some embodiments, obtaining fingerprint information of the first data table is specifically performed as follows:

[0179] extracting key features from metadata of the first data table;

[0180] Processing the key features to obtain the processed key features;

[0181] A hash value operation is performed on the processed key features to obtain a feature mapping value of a preset length, and the feature mapping value of the preset length is used as the fingerprint information of the first data table.

[0182] According to some embodiments, the data classification unit 1104 is configured to analyze the first data table using a data analysis method to obtain a classification result corresponding to the first data table, specifically for:

[0183] In a case where the table type of the first data table is the second data table type, determining to analyze the first data table in a lineage analysis manner to obtain dependent tables and column mapping information of the first data table;

[0184] According to the similarity threshold requirement, query the dependent table and column mapping information in the multi-level cache results to obtain the second query result;

[0185] When the second query result indicates that a dependent table result corresponding to the dependent table is found, obtaining a classification result of the first data table according to the column mapping information;

[0186] or,

[0187] When the second query result indicates that no dependent table result corresponding to the dependent table is found, the first data table is scanned to obtain the classification result corresponding to the first data table, and the classification result corresponding to the first data table is cached in multiple levels.

[0188] According to some embodiments, the data classification unit 1104 is configured to analyze the first data table using a lineage analysis method to obtain dependent tables and column mapping information of the first data table, specifically to:

[0189] Obtaining a definition statement corresponding to the first data table;

[0190] Parsing the definition statement corresponding to the first data table to obtain structured query language SQL information corresponding to the first data table;

[0191] Obtaining column mapping information according to structured query language SQL information and reference information of a SELECT statement, wherein the column mapping information is used to indicate mapping information between view columns and source table columns;

[0192] Analyze and process SQL statements to obtain dependent tables.

[0193] According to some embodiments, the data classification unit 1104 is configured to obtain the classification result of the first data table according to the column mapping information, specifically to:

[0194] When the second query result indicates that the dependent table result corresponding to the dependent table is queried, and multiple historical classification results are obtained according to the column mapping information, the multiple historical classification results are analyzed according to the column information corresponding to each historical classification result to obtain the classification result of the first data table.

[0195] According to some embodiments, the data classification unit 1104 is configured to perform multi-level caching on the classification results corresponding to the first data table, specifically to:

[0196] Obtaining cache level information corresponding to the classification result of the first data table;

[0197] When the cache level information indicates that the classification result of the first data table is to be cached at the table level, storing the classification result of the first data table;

[0198] When the cache level information indicates that column-level caching is performed on the classification results of the first data table, the classification results of each column in the first data table are stored.

[0199] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0200] In some or related embodiments, a table type acquisition unit is used to acquire a table type corresponding to a first data table in a database, wherein the first data table is any data table in at least one data table that has not been classified in the database; a method acquisition unit is used to acquire a data analysis method corresponding to the table type; a result acquisition unit is used to acquire a multi-level cache result corresponding to the first data table, wherein the multi-level cache result includes a classification result of at least one second data table, and the at least one second data table is a data table for which a classification result has been previously determined for the first data table; and a data classification unit is used to analyze the first data table using the data analysis method according to a priority corresponding to the data analysis method and the multi-level cache result to acquire a classification result corresponding to the first data table. Therefore, the data processing method can be determined by the table type of the data table, thereby improving the accuracy of data processing. The data table is classified by the multi-level cache result without traversing the data tables in the database, which can reduce the time required for data table classification and the resources required for traversing all data tables. Moreover, the multi-level cache result and the data analysis method corresponding to the data table can reduce the situation where the same classification rule is used for all types of tables, resulting in inaccurate data table classification, thereby improving the accuracy of data classification and improving the efficiency of data classification.

[0201] Figure 12 A schematic block diagram of an example electronic device 1200 that can be used to implement embodiments of the present disclosure is shown. The electronic device 1200 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0202] like Figure 12 As shown, the electronic device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. Various programs and data required for the operation of the electronic device 1200 can also be stored in the RAM 1203. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0203] Multiple components in the electronic device 1200 are connected to the I / O interface 1205, including an input unit 1206, such as a keyboard, a mouse, etc.; an output unit 1207, such as various types of displays, speakers, etc.; a storage unit 1208, such as a magnetic disk, an optical disk, etc.; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1209 allows the electronic device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0204] The computing unit 1201 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1201 performs the various methods and processes described above, such as the data classification method. For example, in some embodiments, the data classification method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the method described above can be performed. Alternatively, in other embodiments, the computing unit 1201 can be configured to perform the data classification method by any other appropriate means (e.g., by means of firmware).

[0205] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0206] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0207] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0208] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0209] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0210] A computer system may include a client and a server. The client and server are typically remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.

[0211] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0212] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A data classification method, characterized in that: include: Obtaining a table type corresponding to a first data table in a database, wherein the first data table is any data table in at least one unclassified data table in the database; Obtaining a data analysis method corresponding to the table type; Acquire a multi-level cache result corresponding to the first data table, wherein the multi-level cache result includes a classification result of at least one second data table, and the at least one second data table is a data table for which a classification result has been determined before the first data table; According to the priority corresponding to the data analysis method and the multi-level cache result, the first data table is analyzed using the data analysis method to obtain a classification result corresponding to the first data table.

2. The method according to claim 1, characterized in that The adopting the data analysis method to analyze the first data table to obtain a classification result corresponding to the first data table includes: In a case where the table type of the first data table is the first data table type, determining to adopt a fingerprint analysis method to analyze the first data table to obtain fingerprint information of the first data table; According to the similarity threshold requirement, query the fingerprint information in the multi-level cache results to obtain a first query result; If the first query result indicates that a classification result set corresponding to the fingerprint information is found, determining a classification result of the first data table according to the classification result set, wherein the classification result set includes at least one classification result; or, If the first query result indicates that no classification result corresponding to the fingerprint information is found, the first data table is scanned to obtain the classification result corresponding to the first data table, and the classification result corresponding to the first data table is cached in multiple levels.

3. The method according to claim 2, characterized in that The obtaining of fingerprint information of the first data table includes: Extracting key features from metadata of the first data table; Processing the key features to obtain processed key features; A hash value operation is performed on the processed key features to obtain a feature mapping value of a preset length, and the feature mapping value of the preset length is used as the fingerprint information of the first data table.

4. The method according to claim 1, wherein The adopting the data analysis method to analyze the first data table to obtain a classification result corresponding to the first data table includes: In a case where the table type of the first data table is the second data table type, determining to analyze the first data table in a lineage analysis manner to obtain dependent tables and column mapping information of the first data table; According to the similarity threshold requirement, query the dependent table and column mapping information in the multi-level cache result to obtain a second query result; When the second query result indicates that a dependent table result corresponding to the dependent table is found, obtaining a classification result of the first data table according to the column mapping information; or, When the second query result indicates that no dependency table result corresponding to the dependency table is found, the first data table is scanned to obtain the classification result corresponding to the first data table, and the classification result corresponding to the first data table is cached in multiple levels.

5. The method according to claim 4, characterized in that The adopting the lineage analysis method to analyze the first data table to obtain dependent tables and column mapping information of the first data table includes: Obtaining a definition statement corresponding to the first data table; Parsing the definition statement corresponding to the first data table to obtain structured query language SQL information corresponding to the first data table; Acquire column mapping information according to the structured query language SQL information and reference information of the SELECT statement, wherein the column mapping information is used to indicate mapping information between view columns and source table columns; The structured query language SQL information is analyzed and processed to obtain a dependency table.

6. The method according to claim 4, characterized in that The obtaining, according to the column mapping information, a classification result of the first data table includes: When the second query result indicates that the dependent table result corresponding to the dependent table is queried, and multiple historical classification results are obtained according to the column mapping information, the multiple historical classification results are analyzed according to the column information corresponding to each historical classification result to obtain the classification result of the first data table.

7. The method according to claim 2 or 4, characterized in that The multi-level caching of the classification results corresponding to the first data table includes: Obtaining cache level information corresponding to the classification result of the first data table; When the cache level information indicates that the classification result of the first data table is to be cached at the table level, storing the classification result of the first data table; When the cache level information indicates that column-level caching is performed on the classification results of the first data table, the classification results of each column in the first data table are stored.

8. A data classification device, characterized in that: include: A table type acquiring unit, configured to acquire a table type corresponding to a first data table in a database, wherein the first data table is any data table in at least one unclassified data table in the database; A method acquisition unit, configured to acquire a data analysis method corresponding to the table type; a result acquisition unit, configured to acquire a multi-level cache result corresponding to the first data table, wherein the multi-level cache result includes a classification result of at least one second data table, and the at least one second data table is a data table for which a classification result has been determined before for the first data table; A data classification unit is configured to analyze the first data table using the data analysis method according to the priority corresponding to the data analysis method and the multi-level cache result, and obtain a classification result corresponding to the first data table.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the data classification method according to any one of claims 1 to 7.

10. A storage medium storing instructions, characterized in that: When the instruction is executed on an electronic device, the electronic device is caused to execute the data classification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data classification method and device and computer equipment

    CN112948370A

  • Intelligent data asset storage data evaluation method

    CN118503236A

  • Method and apparatus for automatically classifying data

    US20100030781A1

  • System and method for data classification centric sensitive data discovery

    US20200057864A1

  • Automated sensitive data classification in computerized databases

    US20210056219A1