Data Classification and Grading Method and Device

By obtaining the application scenarios of the target file in the data classification and grading method, and using preset matching rules and classification rules for label matching and grading, the problem of great limitations in the application scenarios in the existing technology is solved, and the flexible application of the data classification and grading method in multiple scenarios is realized.

CN119293673BActive Publication Date: 2025-07-11BEIJING HORIN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411826917.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-07-11
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

The existing data classification and grading methods are insufficient to expand the application scenarios, and are difficult to meet the needs of multiple scenarios, and have great limitations.

Method used

By obtaining the application scenarios of the target file, the target data is labeled and classified and graded using preset matching rules and classification rules, and combined with sensitive level sets, the classification and grading of the target file is achieved.

Benefits of technology

It improves the flexibility of application scenarios of data classification and grading methods, can accurately classify and gradle data in different scenarios, and broadens the scope of application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119293673B_ABST
    Figure CN119293673B_ABST
Patent Text Reader

Abstract

This application relates to the field of computer technology, and in particular, to a method and device for data classification and grading. The method includes: obtaining target data corresponding to a target file and the application scenario of the target file; performing label matching on the target data according to a preset matching rule corresponding to the application scenario to obtain a label set; classifying and grading the label set according to a preset classification and grading rule and a sensitive level set corresponding to the application scenario to obtain a classification and grading result corresponding to the target file. This solves the problem that the existing classification and grading methods have great limitations in application scenarios and improves the flexibility of data classification and grading application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method and device for data classification and grading. Background Art

[0002] With the development of social informatization, data has increasingly become an important production factor, among which there are different sensitive data in different scenarios. Along with the need to manage sensitive data and non-sensitive data differently, data classification and grading have emerged. Automated data classification and grading is an important part of modern data governance. It helps organizations better manage their data assets, ensure that sensitive data is properly protected, and meet regulatory requirements.

[0003] Currently, most data classification and grading methods are for data in specific scenarios. However, such data classification and grading methods can only be used in a single scenario. Since the sensitive data and the levels of sensitive data in different scenarios are different, that is, the scalability of traditional data classification and grading methods is significantly insufficient, making it difficult to meet multi-scenario situations and having great limitations in the application scenarios of the method. Summary of the Invention

[0004] Embodiments of this application provide a method and device for data classification and grading to at least solve the problem of great limitations of existing classification and grading methods in application scenarios.

[0005] According to one aspect of the embodiments of this application, a method for data classification and grading is provided, including: obtaining target data corresponding to a target file and the application scenario of the target file; performing label matching on the target data according to a preset matching rule corresponding to the application scenario to obtain a label set; classifying and grading the label set according to a preset classification and grading rule and a sensitive level set corresponding to the application scenario to obtain a classification and grading result corresponding to the target file.

[0006] According to another aspect of the embodiments of this application, a device for data classification and grading is further provided, including: an obtaining unit, configured to obtain target data corresponding to a target file and the application scenario of the target file; a label matching unit, configured to perform label matching on the target data according to a preset matching rule corresponding to the application scenario to obtain a label set; a classification and grading unit, configured to classify and grade the label set according to a preset classification and grading rule and a sensitive level set corresponding to the application scenario to obtain a classification and grading result corresponding to the target file.

[0007] Optionally, the above classification and grading unit includes: an acquisition subunit, configured to acquire a preset classification rule set, a preset grading rule set, and a set of sensitive levels corresponding to the preset grading rule set, where the preset classification and grading rules include the preset classification rule set and the preset grading rule set; a classification unit, configured to classify the label set according to the preset classification rule set to obtain a classification rule set; a grading unit, configured to grade the classification rule set according to the preset grading rule set to obtain a grading rule set; and a determination unit, configured to determine the classification and grading result according to the set of sensitive levels and the grading rule set.

[0008] Optionally, the above determination unit includes: a first determination bullet subunit, configured to determine a reference sensitive level set corresponding to the grading rule set from the set of sensitive levels; a first acquisition module, configured to acquire the file category of the target file, where the file category includes a structured file category and an unstructured file category; and a second determination subunit, configured to determine the classification and grading result according to the reference sensitive level set and the file category.

[0009] Optionally, the above second determination subunit includes: a second acquisition module, configured to acquire the priority corresponding to each grading rule in the grading rule set; a first determination module, configured to determine the target sensitive level according to the multiple priorities and the reference sensitive level set; and a second determination module, configured to, when the file category of the target file is the unstructured file category, determine the target sensitive level as the classification and grading result; and when the file category of the target file is the structured file category, determine the target sensitive level and the multiple reference sensitive levels in the reference sensitive level set as the classification and grading result.

[0010] Optionally, the above label matching unit includes: an acquisition submodule, configured to acquire a preset matching rule, where the preset matching rule includes multiple matching methods, and the multiple matching methods include: keyword matching, similarity matching, named entity recognition, regular matching, and metadata matching; a matching module, configured to perform label matching on the target data respectively by using the multiple matching methods to obtain a label list; and a determination submodule, configured to determine the label set according to the label list.

[0011] Optionally, the above acquisition unit includes: a first acquisition subunit, configured to acquire the file category of the target file, where the file category includes a structured file category and an unstructured file category; and a second acquisition subunit, configured to, when the file category of the target file is the structured file category, acquire the file name, file remarks, field name, and field remarks; and when the file category of the target file is the unstructured file, acquire the file size and file type.

[0012] According to another aspect of the embodiments of the present application, there is provided a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer instructions for causing a computer to execute the data classification and grading method as described above.

[0013] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory and a processor, characterized in that the memory stores a computer program, and the processor is configured to execute the data classification and grading method as described above through the computer program.

[0014] Through the above data classification and grading method of the present application, the problem that the existing classification and grading methods have great limitations on application scenarios is solved, and the flexibility of the application scenarios of data classification and grading is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a schematic diagram of the hardware environment of an optional data classification and grading method according to an embodiment of the present invention;

[0017] Figure 2 It is a flowchart of an optional data classification and grading method according to an embodiment of the present invention;

[0018] Figure 3 It is a schematic diagram of an optional data classification and grading method according to an embodiment of the present invention;

[0019] Figure 4 It is a schematic diagram of another optional data classification and grading method according to an embodiment of the present invention;

[0020] Figure 5 It is a schematic diagram of yet another optional data classification and grading method according to an embodiment of the present invention;

[0021] Figure 6 It is a schematic diagram of the structure of an optional data classification and grading device according to an embodiment of the present invention;

[0022] Figure 7 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion.

[0025] It should be noted that all the data required in the embodiments of this application and the embodiments are obtained and used by legitimate means with the consent of users and relevant personnel, and the information security of users will also be ensured during the process of obtaining and using.

[0026] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will detail this application with reference to the drawings and in conjunction with the embodiments.

[0027] To solve the problem that the existing classification and grading methods have great limitations in application scenarios, this application provides a data classification and grading method. As an alternative implementation, as an alternative implementation, the above data classification and grading method can be applied, but is not limited to, for example Figure 1 shown in the data classification and grading system composed of the terminal device 102 and the server 104. As Figure 1 shown, the server 104 is connected to the terminal device 102 through the network 116. The above network can include, but is not limited to: wired networks, wireless networks. Among them, the wired network includes: local area networks, metropolitan area networks, and wide area networks. The wireless network includes: Bluetooth, WIFI, and other networks that implement wireless communication. The above terminal device can include, but is not limited to, at least one of the following: mobile phones (such as Android mobile phones, iOS mobile phones, etc.), laptop computers, tablet computers, handheld computers, MIDs (Mobile Internet Devices), PADs, desktop computers, smart TVs, etc. An application client is running on the above terminal device, and an application program can be installed in the client.

[0028] The above terminal device 102 is also provided with a display 106, a processor 108, and a memory 110. The display 106 can be used to display the program interface of the above application program. The above processor 108 can be used to process the acquired data, for example, encryption processing, compression processing, filtering processing, etc.; the memory 110 is used to store the data to be processed, the processed data, application programs, etc. It can be understood that the above terminal device 102 sends a data classification and grading request to the server 104 through the network 116. When the server 104 receives the data classification and grading request, it classifies and grades the data according to the target file carried in the data classification and grading request; the terminal device 102 can receive the classification and grading result returned by the server 104 through the network 116.

[0029] The above server 104 can be a single server, or a server cluster composed of multiple servers (such as a server cluster for executing multiple different tasks), or a cloud server. The above server includes a database 112 and a processing engine 114. Among them, the above processing engine 114 is used to process the above data classification and grading request; the above database 112 is used to store the data classification and grading request, the target file, the target data, the application scenario of the target file, and the classification and grading result corresponding to the target file.

[0030] According to one aspect of the embodiments of the present invention, the above data classification and grading system can also perform the following steps: The terminal device 102 executes step S102 and sends a data classification and grading request to the server 104 through the network 116; the server 104 executes steps S104 to S108 to obtain the target data corresponding to the target file and the application scenario of the target file; perform label matching on the target data according to the preset matching rules corresponding to the application scenario to obtain a label set; classify and grade the label set according to the preset classification and grading rules and the sensitive level set corresponding to the application scenario to obtain the classification and grading result corresponding to the target file. Then the server 104 executes S110 and sends the classification and grading result to the terminal device 102 through the network 116.

[0031] In the above embodiments of the present invention, the preset matching rules, preset classification and grading rules, and sensitive level set corresponding to the application scenario of the target file in the present application are used to classify and grade the target file, overcoming the problem that the classification and grading methods in the related art have great limitations on the application scenario, broadening the application scenario of the data classification and grading method, and improving the flexibility of the application scenario of the data classification and grading.

[0032] The above is only an example, and no limitation is made thereto in this embodiment.

[0033] As an optional implementation manner, please refer to Figure 2, which shows a flowchart of the data classification and grading method according to an embodiment of the present application. The method includes the following steps:

[0034] S202, obtain the target data corresponding to the target file and the application scenario of the target file;

[0035] S204, perform label matching on the target data according to the preset matching rules corresponding to the application scenario to obtain a label set;

[0036] S206, classify and grade the label set according to the preset classification and grading rules and the sensitive level set corresponding to the application scenario to obtain the classification and grading result corresponding to the target file.

[0037] It should be noted that the target file in S202 above can be a structured file and an unstructured file. The structured file is, for example, a structured library table; the unstructured file is, for example: text file, image file, audio file, video file, PDF file, XML file, HTML file, report, and other free-form files. When the target file is a structured file, the above target data is structured data; when the target file is an unstructured file, the above target data is unstructured data. Structured data has a clear data structure and a standardized format, and can be organized and stored through tables and databases; the organization form of unstructured data is relatively free and flexible, without a fixed structure and format. The above application scenario (which can also be referred to as the data application scenario) can be understood, but not limited to, as the environment and context in which the data is actually applied. The application scenario includes, but is not limited to, data storage (such as data storage location, data storage form, etc.), data query, data analysis, big data application, cache database application (for example, cache databases are used to reduce server load and improve query efficiency, such as Redis and Memcached), search engine database application (used to index and retrieve website content or enterprise internal documents and data, such as Elasticsearch), big data query system application (used to store and process large-scale unstructured data, such as Hadoop HDFS and Apache Spark).

[0038] The preset matching rule in S204 above is a tag matching rule preset corresponding to the application scenario of the target file. The specific preset matching rule can be understood, but not limited to, as multiple reference tags preset according to the application scenario. The multiple reference tags are used to represent different data types (for example, personal name, gender, age, etc. are all reference tags). Through the preset matching rule to perform tag matching on the target data, multiple tags that match successfully (the multiple tags included in the above tag set) can be obtained; the above successful matching can be understood, but not limited to, as the data tags corresponding to each specific data in the target data determined according to the preset matching rule. For example, if the target data includes "Li XX, male, 2012 / 6 / 6, 12 years old, 412XXXXXX.M", the multiple tags stored in the tag set obtained according to the preset matching rule are "name, gender, birthday, age, ID card".

[0039] Essentially, the above tag matching is to convert the target data into a series of tags for subsequent processing. The preset tag matching rule (i.e., the above preset matching rule) and various tag matching methods constitute a tag library. The preset matching rule can be preset according to business requirements, or can be set according to national standards / industry standards / expert experience, etc. in various industries.

[0040] The preset classification and grading rule in S206 above is a classification and grading rule preset corresponding to the application scenario of the target file. The preset classification and grading rule specifically includes a preset classification rule set and a preset grading rule set. The preset classification rule set includes multiple classification rules, and each classification rule can be understood, but not limited to, as corresponding to a tag combination; the preset grading rule set includes multiple grading rules, and each grading rule can be understood, but not limited to, as corresponding to a category combination; the above sensitive level set includes multiple sensitive levels, and each grading rule corresponds to a sensitive level.

[0041] Through the above implementation manners of the present application, the target file can be classified and graded according to the preset matching rule, preset classification and grading rule, and sensitive level set corresponding to the application scenario of the target file, so that data classification and grading can be applied to different application scenarios, broadening the application scenarios of data classification and grading.

[0042] As an optional implementation manner, according to the preset classification and grading rule and sensitive level set corresponding to the application scenario, the tag set is classified and graded to obtain the classification and grading result corresponding to the target file, including:

[0043] S1, obtain the preset classification rule set, preset grading rule set, and the sensitive level set corresponding to the preset grading rule set. The preset classification and grading rule includes the preset classification rule set and the preset grading rule set;

[0044] S2. Classify the tag set according to the preset classification rule set to obtain a classification rule set;

[0045] S3. Grade the classification rule set according to the preset grading rule set to obtain a grading rule set;

[0046] S4. Determine the classification and grading result according to the sensitivity level set and the grading rule set.

[0047] It can be understood that the preset classification and grading rules include a preset classification rule set and a preset grading rule set corresponding to the application scenario. Each grading rule in the preset grading rules corresponds to a sensitivity level, that is, multiple sensitivity levels in the sensitivity level set are pre-set and corresponding to the grading rules, indicating the level of data sensitivity. Each preset classification rule in the preset classification rule set is designed using a composite rule of data feature + condition, where the data feature is a tag combination (both the data feature and the condition are optional).

[0048] Now, the preset matching rules, preset classification rule set, and preset grading rule set corresponding to a certain business application scenario of personal information shown in Tables 1 to 3 below (the preset matching rules, preset classification rule set, and preset grading rule set corresponding to other business application scenarios of personal information may be different from the specific rules shown in Tables 1 to 3 below) are used to specifically illustrate the above S1 to S4:

[0049] As shown in the first column of Table 1 below, the content is multiple classification rules included in the above preset classification rule set, and the content in the second column is multiple reference tags pre-set corresponding to the application scenario (i.e., the above preset matching rules). It can be seen from Table 1 below that each classification rule corresponds to a set of reference tags; by performing tag matching on the target data according to the preset matching rules, multiple tags that match successfully (tags included in the tag set) can be obtained. For example, the personal name, personal phone number, ID card, fingerprint, more than 1000 pieces of data, etc. shown in Table 1 below. Then, multiple preset classification rules in the preset classification rule set are used to classify the multiple tags in the tag set. After the classification of the multiple tags in the tag set is completed, a classification rule set including multiple classification rules (i.e., multiple preset classification rules that are classified and hit in the preset classification rule set) can be obtained. For example, classification rules such as personal basic information, personal identity information, three key elements of personal information, large data volume, etc. shown in Table 1 below.

[0050] Table 1

[0051] It should be noted that each classification rule in the above preset classification rule set not only corresponds to a set of tags, but may also correspond to a preset condition, that is, each preset classification rule in this application adopts a composite design of data features (tag combination) + preset condition. As shown in Table 2 below, each classification rule corresponds to a set of tags and may also correspond to a pre-set condition. The "at least hitting three or more" shown in Table 2 can be, but is not limited to, understood as that in a set of tags corresponding to the three elements of a message obtained through tag matching, the number of tags hits at least three or more.

[0052] Table 2

[0053] After obtaining the classification rule set, the classification rule set is classified using a preset grading rule set (each grading rule corresponds to a set of classification rules) to obtain multiple final grading rules. For example, the multiple grading rules shown in Table 3 below (the multiple grading rules included in the above preset grading rule set), both the preset grading rule set and the preset classification rule set are designed with composite rules of data features + conditions. The difference is that the data feature of the classification rule in the preset classification rule set is a tag combination, and the data feature of the grading rule in the preset grading rule set is a category combination. Each grading rule also corresponds to a priority level and a sensitivity level. The preset grading rule set includes multiple preset grading rules. After classifying the classification rule set according to the preset grading rule set, the multiple grading rules shown in Table 3 below can be obtained. For example, non-sensitive personal information (non-sensitive personal information), sensitive personal information (sensitive personal information), multi-factor, and non-sensitive personal information with a large amount of data. Each grading rule corresponds to its unique data features, conditions, priority levels, and sensitivity levels. Finally, the final classification and grading results can be obtained by comprehensively calculating according to the sensitivity level (i.e., the above sensitivity level set) and priority level corresponding to each grading rule.

[0054] Table 3

[0055] It should be noted that the specific steps of classification and grading are similar. The following takes Figure 3 as an example to specifically illustrate the classification in S2 or the grading in S3 above. As Figure 3 shown, it specifically includes the following steps:

[0056] S302, obtain the classification / grading rule (i.e., the above preset classification rule set / preset grading rule set);

[0057] S304, determine whether the classification / grading rule contains data features;

[0058] The operation in S304 above can be, but is not limited to, understood as determining whether the preset classification rule set / preset grading rule set includes data features that can classify or grade the data in the previous step (i.e., whether it includes a label combination or a category combination). If not, execute S306; if so, execute S308.

[0059] S306, if the preset classification rule set or the preset grading rule set does not contain data features, take the intermediate result (label / category combination), that is, the label set obtained through label matching or the classification rule set obtained through classification (which can be, but is not limited to, understood as taking the current label / category combination, that is, taking multiple labels in the label set before classification / multiple classification rules in the classification rule set before grading), and then perform a judgment on the preset conditions, that is, execute S312;

[0060] S308, if data features are included, obtain the data features in the preset classification rule set / preset grading rule set;

[0061] S310, take the intersection of the intermediate result and the data features, which can be, but is not limited to, understood as the process of performing the above classification / grading to obtain the classification rule set / grading rule set, that is, determining the classification rule set corresponding to the label set, or determining the grading rule set corresponding to the classification rule set;

[0062] S312, then determine whether the preset classification rule set / preset grading rule set contains preset conditions. If the preset conditions are not included, execute S316; if the preset conditions are included, execute S314;

[0063] S314: Determine whether the preset conditions are met. If the preset conditions are met, execute S318; if the preset conditions are not met, execute S316;

[0064] S316, determine the general classification / grading rule as the classification rule in the classification rule set / the grading rule in the grading rule set;

[0065] S318, determine the final classification / grading rule set.

[0066] Through the above implementation manners of the present application, the target file can be accurately classified and graded according to the classification and grading rules set for different scenarios, thereby greatly broadening the application scenarios of classification and grading on the basis of ensuring accurate classification and grading.

[0067] As an alternative implementation, determining the classification and grading result according to the sensitive level set and the grading rule set includes: determining the reference sensitive level set corresponding to the grading rule set from the sensitive level set; obtaining the file category of the target file, where the file category includes a structured file category and an unstructured file category; and determining the classification and grading result according to the reference sensitive level set and the file category.

[0068] It should be noted that after grading (or called the grading operation), multiple final grading rules can be obtained, and then the reference sensitive level corresponding to each grading rule is determined from the sensitive level set, that is, the reference sensitive level set is determined; then the file category of the target file is obtained, and the file category includes a structured file category and an unstructured file category. The structured file category indicates that the target file is a structured file, and the unstructured file category indicates that the target file is an unstructured file; finally, the classification and grading result is determined according to the multiple reference sensitive levels in the reference sensitive level set and the file category.

[0069] The specific steps for determining the classification and grading result according to the reference sensitive level set and the file category include: obtaining the priority corresponding to each grading rule in the grading rule set, and multiple priorities can be obtained. The target sensitive level is determined according to the multiple priorities and the reference sensitive level set; when the file category of the target file is the unstructured file category, the target sensitive level is determined as the classification and grading result; when the file category of the target file is the structured file category, the target sensitive level and the multiple reference sensitive levels in the reference sensitive level set are determined as the classification and grading result.

[0070] The specific steps for determining the target sensitive level are as follows: when multiple priorities are the same, the highest sensitive level in the reference sensitive level set is determined as the target sensitive level (that is, the principle of taking the higher one for calculation); when multiple priorities are different, the sensitive level corresponding to the highest priority among the multiple priorities is determined as the target sensitive level; then, when the file category of the target file is the unstructured file category (which can be but is not limited to understood as the target file is an unstructured file), the target sensitive level can be directly determined as the classification and grading result corresponding to the target file; when the file category of the target file is the structured file category (which can be but is not limited to understood as the target file is a structured file), at this time, each field in the target file corresponds to a reference sensitive level, and the metadata of the target file (also known as mediation data, relay data, which is data describing data, mainly information describing data attributes) also corresponds to a reference sensitive level, as shown in Table 4 below:

[0071]

[0072] Table 4

[0073] At this time, the reference sensitivity level corresponding to each field in the target file and the above-mentioned target sensitivity level are determined as the classification and grading result corresponding to the target file. Specifically, it can be understood, but not limited to, that each reference sensitivity level is determined as the classification and grading result of each field and metadata in the target file, and the target sensitivity level is determined as the classification and grading result of the entire target file. The determination process of the above-mentioned target sensitivity level can be understood, but not limited to, as the process of determining the level of the whole file (such as the process of determining the level of the table of a structured file). The determination method of the target sensitivity level has been specifically described above and will not be elaborated here. The target sensitivity level can be the sensitivity level finally determined according to the priority corresponding to each grading rule and the reference sensitivity level. For example, the overall sensitivity level of the target file (i.e., the above-mentioned target sensitivity level) determined according to the information listed in Table 4 above is L4 (when the priorities corresponding to each grading rule are the same, the highest priority is taken as the overall sensitivity level of the target file).

[0074] Through the above implementation manners of the present application, it is not necessary to adopt different classification and grading methods for different file types, and the classification and grading results of different file types can be accurately determined, the limitation of the classification and grading method on the file type is lifted, and the overall utilization rate of data classification and grading is improved.

[0075] As an optional implementation manner, the above-mentioned obtaining the label set by performing label matching on the target data according to the preset matching rule corresponding to the application scenario includes: obtaining the preset matching rule, where the preset matching rule includes multiple matching manners, and the multiple matching manners include: keyword matching, similarity matching, named entity recognition, regular expression matching, metadata matching; respectively performing label matching on the target data by using the multiple matching manners to obtain a label list; and determining the label set according to the label list.

[0076] The following takes Figure 4 as an example to specifically illustrate the process of determining the label set. The steps of determining the label set include:

[0077] S402, obtaining a structured database table / unstructured file (i.e., obtaining the target file);

[0078] S404, data preprocessing; the data preprocessing operation here can be understood, but not limited to, as the process of obtaining the target data according to the target file;

[0079] S406, performing label matching on the target data according to the label library, and the specific matching manners include keyword matching, similarity matching, NER matching (i.e., the above-mentioned named entity recognition), regular expression matching, metadata matching.

[0080] It can be understood that multiple tags can be obtained through each of the above matching methods, that is, a list of tags can be obtained after each matching method is used. There are duplicate tags among the multiple tag lists. Therefore, duplicate tags need to be removed to finally obtain a tag set (that is, the process of determining the tag set based on the tag list finally).

[0081] Since the data forms after data preprocessing may be diverse, various forms of matching methods are required to associate diverse data with tags. This application uses methods such as keyword matching, similarity matching, NER matching, regular matching, and metadata matching. For example, for the "personal name" tag: Keyword matching: Keywords such as "name, xingming, name" can be set for keyword matching; Similarity matching: Since "name" is generally similar to "personal name", "name" can be associated with the "personal name" tag through similarity matching; NER: According to the context semantics, specific names such as "Zhang XX", "Li XX", etc. can be recognized through the NER model and associated with the "personal name" tag; Regular matching: For example, for "ID card", a regular expression "\b[1-9]\d{5}(18|19|20)\d{2}(0[1-9]|1[0-2])(0[1-9]|

[12] \d|3

[01] )\d{3}[\dXx]\b" can be set for the ID card to capture each ID card and associate it with the "ID card" tag; Metadata matching: For example, if "more than 1000 data" is set, then by comparing the metadata information of the table, if it is judged that the number of table samples > 1000, the tag "more than 1000 data" is matched.

[0082] Determining the tag set through the above method can ensure that each part of the data in the target data can be tag-matched, avoid missing data during the tag-matching process, and improve the accuracy of tag matching.

[0083] As an optional implementation manner, obtaining the target data corresponding to the target file includes:

[0084] S1. Obtain the file category of the target file, where the file category includes structured file categories and unstructured file categories;

[0085] S2. When the file category of the target file is a structured file category, obtain the file name, file remarks, field name, and field remarks; when the file category of the target file is an unstructured file, obtain the file size and file type.

[0086] It can be understood that the operations in S1 to S2 above are Figure 4The data preprocessing process shown in S404, that is, for target files of different file categories, the target data obtained is different. The data preprocessing stage is essentially to obtain as detailed data information as possible for subsequent processing. For example, for structured library tables, possible data information includes: table name, table notes, table metadata information (table file size, number of table records), field name, field notes, field sampling data, metadata information of each field (field type, field length, field precision), etc.; for unstructured files, it is mainly to parse and obtain text, as well as metadata information such as file size, file type, etc.

[0087] By obtaining the target data corresponding to the target file in the above manner, detailed data corresponding to different file types can be obtained in a targeted manner, thereby more appropriately classifying and grading the data corresponding to different files.

[0088] The following Figure 5 The above data classification and grading method is generally described by taking the example of FIG. 1 as an example. In general, the present application provides a data classification and grading method applicable to a variety of application scenarios, such as Figure 5 As shown, the specific steps include:

[0089] S502, (obtaining) a structured database table / unstructured file; that is, obtaining the above-mentioned target file;

[0090] S504, data preprocessing: that is, performing data preprocessing on the target file to obtain target data corresponding to the target file;

[0091] S506, label matching; specifically, it can be understood as using a preset label library corresponding to the application scenario (i.e., the preset matching rule) to perform label matching on the target data to obtain a label set;

[0092] S508, classification; specifically, it can be understood as using a preset classification rule library corresponding to the application scenario (i.e., the preset classification rule set) to classify multiple tags in the tag set to obtain multiple classification rules (i.e., the classification rules included in the classification rule set);

[0093] S510, grading; specifically, it can be understood as using a preset grading rule library corresponding to the application scenario (i.e., the preset grading rule set) to grade multiple classification rules in the classification rule set to obtain multiple grading rules (i.e., the grading rules included in the grading rule set);

[0094] S512, comprehensive calculation: that is, the process of performing comprehensive calculation according to the sensitivity level and priority corresponding to each classification rule to obtain the final classification and grading result.

[0095] Through the above data classification and grading method, classification and grading results of target files in different application scenarios of the same data can be obtained, which broadens the application scenarios of the classification and grading method while ensuring the accuracy of the classification and grading results.

[0096] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0097] According to another aspect of the embodiments of the present invention, there is also provided a data classification and grading device for implementing the above data classification and grading method. As Figure 6 shown, the device includes:

[0098] An acquisition unit 602, configured to acquire target data corresponding to a target file and the application scenario of the target file;

[0099] A label matching unit 604, configured to perform label matching on the target data according to a preset matching rule corresponding to the application scenario to obtain a label set;

[0100] A classification and grading unit 606, configured to classify and grade the label set according to a preset classification and grading rule corresponding to the application scenario and a set of sensitive levels to obtain a classification and grading result corresponding to the target file.

[0101] Optionally, in this embodiment, for the embodiments to be implemented by the above respective unit modules, reference may be made to the above respective method embodiments, which will not be elaborated here.

[0102] According to yet another aspect of the embodiments of the present invention, there is also provided an electronic device for implementing the above data classification and grading method. The electronic device may be Figure 7 the terminal device or server shown. This embodiment takes the electronic device as the terminal device as an example for illustration. As Figure 7 shown, the electronic device includes a memory 702 and a processor 704. A computer program is stored in the memory 702, and the processor 704 is configured to execute the steps in any one of the above method embodiments through the computer program.

[0103] Optionally, in this embodiment, the above electronic device may be at least one network device among multiple network devices in a computer network.

[0104] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the above-mentioned data classification and grading method through a computer program.

[0105] Optionally, those of ordinary skill in the art can understand that Figure 7 the structure shown is only schematic, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and a mobile Internet device (Mobile Internet Devices, MID), a PAD, and other terminal devices. Figure 7 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown in Figure 7 Figure 7, or have a different configuration from that shown in Figure 7 Figure 7.

[0106] Among them, the memory 702 can be used to store software programs and modules, such as the program instructions / modules corresponding to the data classification and grading method and device in the embodiments of the present invention. The processor 704 executes various functional applications and data processing by running the software programs and modules stored in the memory 702, that is, implements the above-mentioned data classification and grading method. The memory 702 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 702 may further include a memory remotely provided with respect to the processor 704, and these remote memories may be connected to the terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise internal network, a local area network, a mobile communication network, and combinations thereof. Among them, the memory 702 may specifically but not limitedly be used to store file information such as a target data set. As an example, as Figure 7 shown in Figure 7, the above-mentioned memory 702 may include but are not limited to the acquisition unit 602, the tag matching unit 604, and the classification and grading unit 606 in the above-mentioned data classification and grading device. In addition, it may also include but are not limited to other module units in the above-mentioned data classification and grading device, which will not be elaborated in this example.

[0107] Optionally, the above-mentioned transmission device 706 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one instance, the transmission device 706 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, so as to communicate with the Internet or a local area network. In one instance, the transmission device 706 is a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet wirelessly.

[0108] In addition, the above electronic device further includes: a display 708, and a connection bus 710 for connecting each module component in the above electronic device.

[0109] In other embodiments, the above terminal device or server may be a node in a distributed system. Among them, the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes in a form of network communication. Among them, the nodes may form a peer-to-peer (P2P) network, and any form of computing device, such as an electronic device like a server or a terminal, can become a node in the blockchain system by joining the peer-to-peer network.

[0110] According to one aspect of the present application, there is provided a computer program product, which includes computer programs / instructions, and the computer programs / instructions contain program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it executes various functions provided by the embodiments of the present application.

[0111] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0112] According to one aspect of the present application, there is provided a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above data classification and grading method.

[0113] Optionally, in this embodiment, the above computer-readable storage medium may be set to store a computer program for executing the above data classification and grading method.

[0114] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0115] If the integrated units in the above embodiments are implemented in the form of software function units and sold or used as independent products, they can be stored in the above computer-readable storage media. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the above methods in various embodiments of the present invention.

[0116] In the above embodiments of the present invention, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0117] The client disclosed in several embodiments provided in this application can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0118] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0119] In addition, the functional units in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software function units.

[0120] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A data classification and grading method, characterized in that, Including: Obtain the target data corresponding to the target file and the application scenario of the target file, where the target data is the specific value stored in the target file; Perform label matching on the target data according to the preset matching rules corresponding to the application scenario to obtain a label set, including: obtaining the preset matching rules, where the preset matching rules include multiple matching methods, and the multiple matching methods include: keyword matching, similarity matching, named entity recognition, regular matching, metadata matching; respectively perform label matching on the target data using multiple matching methods to obtain a label list; determine the label set according to the label list; Obtain a preset classification rule set, a preset grading rule set, and a sensitivity level set corresponding to the preset grading rule set. The preset classification and grading rules include the preset classification rule set and the preset grading rule set. Each grading rule in the preset grading rules corresponds to a sensitivity level, and the multiple sensitivity levels in the sensitivity level set are pre-set and corresponding to the grading rules, indicating the sensitivity level of the data. Each preset classification rule in the preset classification rule set is designed using a composite rule of data features + conditions; Classify the label set according to the preset classification rule set to obtain a classification rule set; Grade the classification rule set according to the preset grading rule set to obtain a grading rule set; Determine the reference sensitivity level set corresponding to the grading rule set from the sensitivity level set; Obtain the file category of the target file, where the file category includes a structured file category and an unstructured file category; Determine the classification and grading result according to the reference sensitivity level set and the file category.

2. The method according to claim 1, characterized in that, Determining the classification and grading result according to the reference sensitivity level set and the file category includes: Obtain the priority corresponding to each grading rule in the grading rule set; Determine the target sensitivity level according to the multiple priorities and the reference sensitivity level set; When the file category of the target file is the unstructured file category, determine the target sensitivity level as the classification and grading result; When the file category of the target file is the structured file category, determine the target sensitivity level and the multiple reference sensitivity levels in the reference sensitivity level set as the classification and grading result.

3. The method according to claim 1, wherein Obtaining the target data corresponding to the target file includes: Obtain the file category of the target file, where the file category includes a structured file category and an unstructured file category; When the file category of the target file is the structured file category, obtain the file name, file remarks, field name, and field remarks; When the file category of the target file is the unstructured file category, obtain the file size and file type.

4. A data classification and grading device, characterized in that Including: An acquisition unit for obtaining the target data corresponding to the target file and the application scenario of the target file, where the target data is the specific value stored in the target file; A label matching unit, configured to perform label matching on the target data according to a preset matching rule corresponding to the application scenario to obtain a label set, including: obtaining the preset matching rule, where the preset matching rule includes multiple matching methods, and the multiple matching methods include: keyword matching, similarity matching, named entity recognition, regular expression matching, and metadata matching; respectively performing label matching on the target data by using the multiple matching methods to obtain a label list; determining the label set according to the label list; A classification and grading unit, configured to obtain a preset classification rule set, a preset grading rule set, and a sensitivity level set corresponding to the preset grading rule set. The preset classification and grading rules include the preset classification rule set and the preset grading rule set. Each grading rule in the preset grading rules corresponds to a sensitivity level, and the multiple sensitivity levels in the sensitivity level set are pre-set and corresponding to the grading rules, indicating the level of data sensitivity. Each preset classification rule in the preset classification rule set is designed by using a composite rule of data feature + condition; classifying the label set according to the preset classification rule set to obtain a classification rule set; grading the classification rule set according to the preset grading rule set to obtain a grading rule set; determining a reference sensitivity level set corresponding to the grading rule set from the sensitivity level set; obtaining the file category of the target file, where the file category includes a structured file category and an unstructured file category; determining the classification and grading result according to the reference sensitivity level set and the file category.

5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to execute the method described in any one of claims 1 to 3.

6. An electronic device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 3 through the computer program.

Citation Information

Patent Citations

  • Data classification and grading method and device and related equipment

    CN117493976A