Method and apparatus for identifying data category, electronic device, and storage medium
By identifying data categories in parallel using multiple identification units, and combining the identification ratio and weight, the result of the identification unit with the highest weight is determined as the data category. This solves the problems of low efficiency and low accuracy in existing technologies, and achieves accurate identification and refined management of data categories.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA TELECOM CORP LTD
- Filing Date
- 2022-11-23
- Publication Date
- 2026-07-31
AI Technical Summary
Existing data classification methods are inefficient and have low recognition accuracy, making it difficult to achieve refined management of data categories.
Multiple identification units are used to identify data categories in parallel. By obtaining the identification ratio and weight of each identification unit, the identification result of the identification unit with the highest weight is determined as the final data category.
It improves the accuracy and efficiency of data category identification, can objectively and truthfully determine data categories, and supports refined management of sensitive data.
Smart Images

Figure CN115758286B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network technology and security, and in particular to a method, apparatus, electronic device, storage medium, and program product for identifying data categories. Background Technology
[0002] In the practical operation of data security governance, data classification can avoid a one-size-fits-all control approach and also facilitates the adoption of more refined measures for data security management, so as to achieve a balance between data sharing and secure use.
[0003] Related data classification methods involve inputting data into a classification model, which then predicts the data's category. However, this method suffers from drawbacks such as low efficiency and low recognition accuracy. Summary of the Invention
[0004] In view of the above problems, embodiments of the present invention provide a method, apparatus, electronic device, storage medium, and program product for identifying data categories, so as to overcome the above problems or at least partially solve the above problems.
[0005] A first aspect of the present invention provides a method for identifying data categories, comprising:
[0006] The data to be identified is input into multiple identification units, and each identification unit is used to identify a data category.
[0007] Obtain the recognition ratio of each of the plurality of recognition units for the data to be recognized;
[0008] Based on the recognition ratio of each of the multiple recognition units for the data to be recognized, and the weight of each of the multiple recognition units, the final weight of each of the multiple recognition units is obtained;
[0009] The data category identified by the identification unit with the highest final weight is determined as the data category of the data to be identified.
[0010] Optionally, obtaining the recognition ratio of each of the plurality of recognition units for the data to be recognized includes:
[0011] The recognition ratio of each recognition mode contained in each of the plurality of recognition units to the data to be recognized is obtained, and the weight of each recognition mode is obtained. The recognition mode includes at least one or more of the following: keyword dictionary mapping, regular expression matching, and AI recognition.
[0012] For each identification unit, the identification ratio of the identification mode included in the identification unit to the data to be identified is weighted and summed according to the weight of each identification mode, so as to obtain the identification ratio of the identification unit to the data to be identified.
[0013] Optionally, inputting the data to be identified into multiple identification units includes:
[0014] Obtain multi-dimensional information of the data to be identified, wherein the multi-dimensional information includes at least one or more of the following: length, whether it is empty, and whether it is a primary key;
[0015] Obtain multiple candidate identification units from the data to be identified;
[0016] Based on the multi-dimensional information of the data to be identified, the multiple candidate identification units are filtered to obtain the identification unit;
[0017] The data to be identified is input into the identification unit.
[0018] Optionally, obtaining the final weight of each of the plurality of identification units based on their respective recognition ratios of the data to be identified and their respective weights includes:
[0019] Remove the recognition units whose recognition ratio is less than the ratio threshold to obtain multiple effective recognition units;
[0020] Based on the recognition ratio and weight of the multiple effective recognition units, the final weight of each of the multiple effective recognition units is obtained;
[0021] The step of determining the data category of the data to be identified as the data category of the identification unit with the highest final weight includes:
[0022] The data category identified by the effective identification unit with the highest final weight is determined as the data category of the data to be identified.
[0023] Optionally, it also includes:
[0024] Obtain the target data category to which the data to be identified belongs;
[0025] If the data category of the data to be identified and the target data category are inconsistent, the identification mode of the identification unit, the weight of the identification mode, and / or the weight of the identification unit shall be adjusted.
[0026] Optionally, it also includes:
[0027] Obtain the target data category to which the data to be identified belongs;
[0028] If the target data category of the data to be identified indicates that the data to be identified is sensitive data, a new identification pattern is constructed based on the data to be identified.
[0029] Based on the new recognition pattern, a new recognition unit is constructed.
[0030] Optionally, it also includes:
[0031] Based on the data category of the data to be identified, the sensitive category of the data to be identified is determined, and the sensitive category includes at least one or more of the following: personal sensitive category and enterprise sensitive category.
[0032] A second aspect of the present invention provides an apparatus for identifying data categories, comprising:
[0033] The input module is configured to input the data to be identified into multiple identification units, with each identification unit used to identify a data category.
[0034] The acquisition module is configured to acquire the recognition ratio of each of the plurality of recognition units for the data to be identified;
[0035] The weighting module is configured to obtain the final weight of each of the multiple identification units based on the identification ratio of each of the multiple identification units to the data to be identified, and the weight of each of the multiple identification units.
[0036] The determination module is configured to determine the data category of the data to be identified as the data category of the identification unit with the highest final weight.
[0037] Optionally, the acquisition module includes:
[0038] The ratio acquisition unit is configured to acquire the recognition ratio of each recognition mode contained in each of the plurality of recognition units for the data to be recognized, and to acquire the weight of each recognition mode, wherein the recognition mode includes at least any one or more of the following: keyword dictionary mapping, regular expression matching, and AI recognition.
[0039] The summation unit is configured to, for each of the identification units, perform a weighted summation of the identification ratios of the identification modes contained in the identification unit to the data to be identified, according to the weights of each of the identification modes, to obtain the identification ratio of the identification unit to the data to be identified.
[0040] Optionally, the input module includes:
[0041] The information acquisition unit is configured to acquire multi-dimensional information of the data to be identified, wherein the multi-dimensional information includes at least one or more of the following: length, whether it is empty, and whether it is a primary key;
[0042] The candidate acquisition unit is configured to acquire multiple candidate identification units of the data to be identified.
[0043] The filtering unit is configured to filter the multiple candidate identification units based on the multi-dimensional information of the data to be identified, and obtain the identification unit;
[0044] The input unit is configured to input the data to be identified into the identification unit.
[0045] Optionally, the weighting module includes:
[0046] The removal unit is configured to remove recognition units whose recognition ratio is less than a ratio threshold, thereby obtaining multiple valid recognition units.
[0047] The final weighting unit is configured to obtain the final weight of each of the multiple effective identifications based on the identification ratio and weight of the multiple effective identification units;
[0048] The determining module includes:
[0049] The determining unit is configured to determine the data category identified by the effective identification unit with the highest final weight as the data category of the data to be identified.
[0050] Optionally, it also includes:
[0051] The first acquisition module is configured to acquire the target data category to which the data to be identified belongs;
[0052] The adjustment module is configured to adjust the recognition mode, the weight of the recognition mode, and / or the weight of the recognition unit when the data category of the data to be identified and the target data category are inconsistent.
[0053] Optionally, it also includes:
[0054] The second acquisition module is configured to acquire the target data category to which the data to be identified belongs;
[0055] The pattern building module is configured to build a new recognition pattern based on the data to be identified when the target data category of the data to be identified indicates that the data to be identified is sensitive data.
[0056] The unit construction module is configured to construct new identification units based on the new identification pattern.
[0057] Optionally, it also includes:
[0058] The sensitive category determination module is configured to determine the sensitive category of the data to be identified based on the data category of the data to be identified, wherein the sensitive category includes at least one or more of the following: personal sensitive category and enterprise sensitive category.
[0059] A third aspect of the present invention provides an electronic device, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the method for identifying data categories as described in the first aspect.
[0060] A fourth aspect of the present invention provides a computer-readable storage medium that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform a method for identifying data categories as described in the first aspect.
[0061] A fifth aspect of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method for identifying data categories as described in the first aspect.
[0062] The embodiments of the present invention have the following advantages:
[0063] In this embodiment, one identification unit is used to identify one data category, thus the identification is targeted. Based on the identification ratio of each identification unit and the weight of each identification unit, the final weight of each identification unit is obtained. The data category identified by the identification unit with the highest final weight is determined as the data category of the data to be identified. In this way, the data category of the data to be identified is determined by the final weight, ensuring objectivity and accuracy. Multiple identification units identify the data to be identified in parallel to obtain the identification ratio, guaranteeing both accuracy and efficiency. Attached Figure Description
[0064] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0065] Figure 1 This is a flowchart illustrating the steps of a method for identifying data categories according to an embodiment of the present invention;
[0066] Figure 2 This is a schematic diagram of the business process of the recognition model in an embodiment of the present invention;
[0067] Figure 3 This is a schematic diagram of the data recognition process in an embodiment of the present invention;
[0068] Figure 4 This is a system architecture diagram for data recognition in an embodiment of the present invention;
[0069] Figure 5 This is a schematic diagram of the structure of a data category identification device according to an embodiment of the present invention. Detailed Implementation
[0070] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0071] Reference Figure 1 The diagram illustrates a flowchart of a method for identifying data categories according to an embodiment of the present invention. Figure 1 As shown, the method for identifying this data category may specifically include steps S11 to S14.
[0072] Step S11: Input the data to be identified into multiple identification units. Each identification unit is used to identify a data category.
[0073] Data categories can include mobile phone numbers, landline numbers, usernames, etc. The data to be identified is input into the identification unit, which then identifies the data and outputs the recognition rate. The recognition rate of a unit for a given piece of data represents the similarity between the data and the data category identified by that unit. For example, if an identification unit identifies a data category as mobile phone numbers, and its recognition rate for a given piece of data is 50%, then the data has a 50% similarity to mobile phone numbers, meaning there is a 50% probability that the data is a mobile phone number.
[0074] Optionally, the data to be identified can be directly input into each identification unit to obtain the identification ratio of the data to be identified in each identification unit, or the data to be identified can be input into the filtered identification units.
[0075] Each identification unit has its own applicable data types and usage conditions. For example, the applicable data type for mobile phone number identification units is numeric, and the usage condition is 11 digits.
[0076] It can obtain multi-dimensional information about the data to be identified, including at least one or more of the following: length, whether it is nullable, and whether it is a primary key. It obtains multiple candidate identification units for the data to be identified, each of which can be determined empirically. Based on the multi-dimensional information of the data to be identified, it filters the multiple candidate identification units to obtain the identification unit; then it inputs the data to be identified into the identification unit.
[0077] For example, if the data to be identified is numerical data, multiple candidate identification units for identifying numerical data can be determined. If the length of the data to be identified is 11 digits, and among the multiple candidate identification units, two candidate identification units are used to identify 11-digit numbers, then other candidate identification units can be excluded, and these two candidate identification units can be used as the identification units for the data to be identified, and the data to be identified can be input into these two identification units.
[0078] When the data comes from a communication service provider's database, it can first be broken down and categorized to obtain the data to be identified. This data is then roughly classified into three types: field names, field comments, and sample data. Based on which of these categories the data belongs to, candidate identification units are determined. For example, when the data to be identified is field comments, the identification mode used is typically Chinese dictionary and AI recognition, rather than regular expression matching.
[0079] In this way, based on the multi-dimensional information of the data to be identified and the usage rules of each identification unit, the identification units were screened, avoiding the input of the data to be identified into each identification unit, reducing the amount of calculation and effectively improving efficiency.
[0080] Step S12: Obtain the recognition ratio of each of the multiple recognition units for the data to be recognized.
[0081] The identification unit will identify the input data to be identified and obtain the identification ratio of the data to be identified. The higher the identification ratio of the data to be identified, the more likely the data category of the data to be identified is the data category that the identification unit can identify.
[0082] Each identification unit may contain one or more identification patterns, which shall include at least one or more of the following: keyword dictionary mapping, regular expression matching, and AI (Artificial Intelligence) recognition. When an identification unit contains multiple identification patterns, the unit combines these patterns according to a specific matching algorithm to perform data recognition on the data to be identified.
[0083] The keyword dictionary mapping incorporates a keyword dictionary to determine whether the data to be identified contains keywords from the keyword dictionary, thereby determining the similarity between the data category of the data to be identified and the data category identified by the identification unit. Regular expression matching incorporates the format of the data category identified by the identification unit, used to identify the similarity between the format of the data to be identified and the format of the data category identified by the identification unit. AI recognition is used for automatic machine identification of the similarity between the data to be identified and the data category identified by the identification unit. Specifically, the keyword dictionary in the identification mode can have multiple different keyword dictionaries, such as a Chinese dictionary and an English dictionary; the regular expression in the identification mode can have multiple different regular expressions, such as a regular expression for phone numbers and a regular expression for usernames; and the AI recognition in the identification mode can have AI that identifies different data categories.
[0084] Each recognition mode can have a weight, which can be initially set and continuously adjusted. When a recognition unit performs data recognition on the data to be recognized, each recognition mode within that unit recognizes the data separately, resulting in the recognition ratio of each recognition mode. The recognition ratio of a recognition unit on the data to be recognized is obtained by weighted summing the recognition ratios of each recognition mode within that unit. For example, if a recognition unit contains two recognition modes, recognition mode A and recognition mode B, with recognition mode A having a weight of 0.4 and recognition mode B having a weight of 0.6, and recognition mode A has a recognition ratio of 70% for the data to be recognized, while recognition mode B has a recognition ratio of 60%, then the recognition ratio of the unit for the data to be recognized is 0.4 × 70% + 0.6 × 60% = 64%.
[0085] Thus, an identification unit can contain one or more identification modes, making it more flexible in data identification and effectively improving the accuracy of data identification.
[0086] Step S13: Based on the recognition ratio of each of the multiple recognition units for the data to be recognized, and the weight of each of the multiple recognition units, obtain the final weight of each of the multiple recognition units.
[0087] Each identification unit can have a weight, which can be initially set and continuously adjusted subsequently. The weights of multiple identification units are obtained, and the final weight of an identification unit is the product of the identification ratio of the data to be identified by that unit and the weight of that unit.
[0088] Step S14: Determine the data category of the data to be identified as the data category identified by the identification unit with the highest final weight.
[0089] The final weight of each identification unit is compared, and the data category identified by the identification unit with the highest final weight is determined as the data category of the data to be identified. For example, if the data category identified by the identification unit with the highest final weight is a mobile phone number, then the data category of the data to be identified is a mobile phone number.
[0090] The technical solution of this invention involves one identification unit identifying a single data category, thus ensuring targeted identification. Based on the identification ratios of each identification unit and their respective weights, the final weights of each identification unit are obtained. The data category identified by the identification unit with the highest final weight is determined as the data category of the data to be identified. In this way, determining the data category of the data to be identified through the final weights is objective and accurate. The parallel identification of multiple identification units to obtain the identification ratio ensures both accuracy and efficiency.
[0091] Compared to related technologies where a model determines which data category the data leans towards based on its features, the data category identification method disclosed in this invention identifies only one data category per identification unit. The method determines whether the identified data belongs to the data category corresponding to that identification unit based on the identification ratio of that unit, making it more targeted and accurate.
[0092] Optionally, based on the above technical solution, when calculating the final weight of the identification unit, each identification unit can be screened first.
[0093] After obtaining the recognition ratio of each recognition unit in the data to be recognized, recognition units with recognition ratios lower than the recognition threshold can be removed, resulting in multiple effective recognition units. Effective recognition units are those with recognition ratios higher than the recognition threshold. The recognition threshold can be set according to requirements, for example, it can be 90%.
[0094] After determining the effective identification units, the final weights of each effective identification unit are obtained based on its identification ratio and weight. Specifically, the final weight of an effective identification unit is the product of its identification ratio of the data to be identified and its weight. The data category identified by the effective identification unit with the highest final weight is determined as the data category of the data to be identified.
[0095] Therefore, screening each identification unit beforehand avoids situations where some identification units have a low recognition rate for the data to be identified, but a high weight for the identification unit itself, resulting in a high final weight for that identification unit. This leads to a more accurate determination of the data category for the data to be identified.
[0096] Optionally, based on the above technical solution, after determining the data category of the data to be identified, the sensitive category of the data to be identified can be determined according to the data category. The sensitive category includes at least one or more of the following: personal sensitive category and enterprise sensitive category.
[0097] The correspondence between each data category and the sensitive category can be established in advance. After the data category of the data to be identified is determined, the correspondence between the data category and the sensitive category can be queried to determine the sensitive category of the data to be identified.
[0098] For example, if the data category to be identified is mobile phone number, then the data to be identified can be a sensitive category for individuals; if the data category to be identified is landline phone number, then the data to be identified can be a sensitive category for enterprises.
[0099] In this way, the sensitive categories of the data to be identified can be determined based on the data categories themselves. Determining the sensitive categories of the data to be identified allows for more refined measures to be taken for the security management of the data.
[0100] Optionally, based on the above technical solution, the target data category to which the data to be identified belongs can be obtained. The target data category is the correct data category obtained after analyzing the data to be identified, and can be categories such as mobile phone number, landline phone number, username, etc.
[0101] When the data category of the data to be identified is inconsistent with the target data category, it indicates that the identification of the data to be identified by each identification unit is not accurate enough. Therefore, the identification modes contained in the identification unit, the weights of each identification mode contained in the identification unit and / or the weights of the identification unit itself can be adjusted so that the adjusted identification unit can obtain a more accurate identification ratio when identifying the data to be identified, and thus obtain a more accurate data category of the data to be identified.
[0102] Optionally, based on the above technical solution, the target data category to which the data to be identified belongs can be obtained. If the target data category indicates that the data to be identified is sensitive data, a new identification pattern can be constructed based on the data to be identified. The new identification pattern can refer to a new keyword dictionary mapping, a new regular expression matching, and / or a new AI recognition method.
[0103] Once a new recognition pattern is obtained, a new recognition unit can be constructed based on it. This can involve adding the new recognition pattern to an existing recognition unit, replacing some of the recognition patterns in an existing recognition unit, or constructing a recognition unit that only includes the new recognition pattern.
[0104] For example, if the data to be identified includes keyword A, keyword A can be added to the keyword dictionary corresponding to the target data category of the data to be identified to obtain a new keyword dictionary.
[0105] In this way, we can enrich the recognition patterns, accumulate keyword dictionaries and regular expression libraries, and train AI recognition capabilities through machine learning, thereby ensuring that the keyword dictionaries and regular expression libraries are updated in a timely manner and the AI recognition capabilities are continuously enhanced.
[0106] Optionally, the data category identification method described in the embodiments of the present invention can be implemented using an identification model. Figure 2 This is a schematic diagram of the business process of the recognition model in an embodiment of the present invention. The recognition model may include a filtering module, a recognition module, a judgment module, a weight calculation module, and a model construction module.
[0107] The filtering module is used to break down and summarize the data, obtain multiple candidate recognition units of the data to be identified, and filter the multiple candidate recognition units based on the multi-dimensional information of the data to be identified to obtain the recognition unit.
[0108] The recognition module is used to input the data to be recognized into the recognition unit and obtain the recognition ratio of the data to be recognized by the recognition unit. Figure 3 This is a schematic diagram of the data recognition process in an embodiment of the present invention. The recognition unit may include one or more recognition modes. Each recognition mode performs data recognition on the data to be recognized, obtaining the recognition ratio of each mode. Based on the weight of each recognition mode, the recognition ratios of each recognition mode are weighted and summed to obtain the weight of the recognition unit.
[0109] The determination module is used to determine whether each identification unit is valid. Specifically, it compares the identification ratio of each identification unit with a ratio threshold, and identifies identification units whose identification ratio is greater than the ratio threshold as valid identification units. The identification ratio of the valid identification units is then transmitted to the weight calculation module.
[0110] The weight calculation module calculates the final weight of each effective identification unit based on its identification ratio and weight. The final weight of an identification unit equals the identification ratio multiplied by its weight. By comparing the final weights of each effective identification unit, the data category identified by the effective identification unit with the highest final weight is determined as the data category of the data to be identified. Based on this data category, the sensitive data category of the data to be identified is then determined.
[0111] The model building module can obtain the target data category to which the data to be identified belongs, and the data category of the data to be identified obtained by the weight calculation module. If the data category of the data to be identified and the target data category are inconsistent, the module adjusts the identification mode, the weight of the identification mode, and / or the weight of the identification unit. Furthermore, if the target data category indicates that the data to be identified is sensitive data, a new identification mode is constructed based on the data to be identified; and a new identification unit is constructed based on the new identification mode.
[0112] If no matching identification unit is found when acquiring the identification unit of the data to be identified, the data identification process can be terminated directly.
[0113] Figure 4 This is a system architecture diagram for data recognition in an embodiment of the present invention. The filtering module identifies candidate recognition units based on multi-dimensional information of the data, obtaining the recognition units. In the recognition module, the recognition units identify the data, obtaining the recognition ratio. The judgment module determines whether the recognition unit is valid based on its recognition ratio. The weight calculation module calculates the data category through weight calculation and defines the sensitive data category. The model building module accumulates a keyword dictionary and regular expression library for the recognition module based on a large amount of data and the output results of the recognition module, and trains the AI recognition capability through machine learning.
[0114] The filtering module filters out interference factors based on multiple dimensions of data, such as data type, length, and whether it is empty, to determine the appropriate recognition unit and improve recognition efficiency. The recognition unit in the recognition module combines keyword dictionary mapping, regular expression matching, and AI recognition capabilities. Each recognition unit has its applicable data type, usage conditions, and assigned weights. Data recognition is both targeted and flexible in using various rules. The recognition results are reflected through weighting, ensuring objectivity and accuracy while maintaining data recognition efficiency. The model building module accumulates a keyword dictionary and regular expression library for the matching module based on a large amount of data and the output results of the matching module. It also trains AI discovery capabilities through machine learning, ensuring the material library is updated in a timely manner and continuously improving recognition capabilities. The weight calculation module can handle multiple recognition results and derives the optimal recognition result through weight calculation.
[0115] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0116] Figure 5 This is a schematic diagram of the structure of a data category identification device according to an embodiment of the present invention, as shown below. Figure 5 As shown, the device includes an input module, an acquisition module, a weighting module, and a determination module, wherein:
[0117] The input module is configured to input the data to be identified into multiple identification units, with each identification unit used to identify a data category.
[0118] The acquisition module is configured to acquire the recognition ratio of each of the plurality of recognition units for the data to be identified;
[0119] The weighting module is configured to obtain the final weight of each of the multiple identification units based on the identification ratio of each of the multiple identification units to the data to be identified, and the weight of each of the multiple identification units.
[0120] The determination module is configured to determine the data category of the data to be identified as the data category of the identification unit with the highest final weight.
[0121] Optionally, the acquisition module includes:
[0122] The ratio acquisition unit is configured to acquire the recognition ratio of each recognition mode contained in each of the plurality of recognition units for the data to be recognized, and to acquire the weight of each recognition mode, wherein the recognition mode includes at least any one or more of the following: keyword dictionary mapping, regular expression matching, and AI recognition.
[0123] The summation unit is configured to, for each of the identification units, perform a weighted summation of the identification ratios of the identification modes contained in the identification unit to the data to be identified, according to the weights of each of the identification modes, to obtain the identification ratio of the identification unit to the data to be identified.
[0124] Optionally, the input module includes:
[0125] The information acquisition unit is configured to acquire multi-dimensional information of the data to be identified, wherein the multi-dimensional information includes at least one or more of the following: length, whether it is empty, and whether it is a primary key;
[0126] The candidate acquisition unit is configured to acquire multiple candidate identification units of the data to be identified.
[0127] The filtering unit is configured to filter the multiple candidate identification units based on the multi-dimensional information of the data to be identified, and obtain the identification unit;
[0128] The input unit is configured to input the data to be identified into the identification unit.
[0129] Optionally, the weighting module includes:
[0130] The removal unit is configured to remove recognition units whose recognition ratio is less than a ratio threshold, thereby obtaining multiple valid recognition units.
[0131] The final weighting unit is configured to obtain the final weight of each of the multiple effective identifications based on the identification ratio and weight of the multiple effective identification units;
[0132] The determining module includes:
[0133] The determining unit is configured to determine the data category identified by the effective identification unit with the highest final weight as the data category of the data to be identified.
[0134] Optionally, it also includes:
[0135] The first acquisition module is configured to acquire the target data category to which the data to be identified belongs;
[0136] The adjustment module is configured to adjust the recognition mode, the weight of the recognition mode, and / or the weight of the recognition unit when the data category of the data to be identified and the target data category are inconsistent.
[0137] Optionally, it also includes:
[0138] The second acquisition module is configured to acquire the target data category to which the data to be identified belongs;
[0139] The pattern building module is configured to build a new recognition pattern based on the data to be identified when the target data category of the data to be identified indicates that the data to be identified is sensitive data.
[0140] The unit construction module is configured to construct new identification units based on the new identification pattern.
[0141] Optionally, it also includes:
[0142] The sensitive category determination module is configured to determine the sensitive category of the data to be identified based on the data category of the data to be identified, wherein the sensitive category includes at least one or more of the following: personal sensitive category and enterprise sensitive category.
[0143] It should be noted that the device embodiments are similar to the method embodiments, so the description is relatively simple. For relevant details, please refer to the method embodiments.
[0144] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0145] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0146] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0147] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0148] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0149] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0150] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0151] The present invention has provided a detailed description of a method, apparatus, electronic device, storage medium, and program product for identifying data categories. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for identifying data categories, characterized in that, include: The data to be identified is input into multiple identification units, and each identification unit is used to identify a data category. The data categories include at least one of mobile phone numbers, landline phone numbers, and usernames; Obtain the recognition ratio of each of the plurality of recognition units for the data to be recognized; Based on the recognition ratio of each of the multiple recognition units for the data to be recognized, and the weight of each of the multiple recognition units, the final weight of each of the multiple recognition units is obtained; The data category identified by the identification unit with the highest final weight is determined as the data category of the data to be identified; The step of obtaining the recognition ratio of each of the plurality of recognition units for the data to be recognized includes: The recognition ratio of each recognition mode contained in each of the plurality of recognition units to the data to be recognized is obtained, and the weight of each recognition mode is obtained. The recognition mode includes at least one or more of the following: keyword dictionary mapping, regular expression matching, and AI recognition. For each identification unit, the identification ratio of the identification mode included in the identification unit to the data to be identified is weighted and summed according to the weight of each identification mode, so as to obtain the identification ratio of the identification unit to the data to be identified. The step of inputting the data to be identified into multiple identification units includes: Obtain multi-dimensional information of the data to be identified, wherein the multi-dimensional information includes at least one or more of the following: length, whether it is empty, and whether it is a primary key; Obtain multiple candidate identification units from the data to be identified; Based on the multi-dimensional information of the data to be identified, the multiple candidate identification units are filtered to obtain the identification unit; Input the data to be identified into the identification unit; It also includes: determining the sensitive category of the data to be identified based on the data category of the data to be identified, wherein the sensitive category includes at least one or more of the following: personal sensitive category and enterprise sensitive category.
2. The method according to claim 1, characterized in that, The step of obtaining the final weight of each of the multiple identification units based on their respective identification ratios of the data to be identified and their respective weights includes: Remove the recognition units whose recognition ratio is less than the ratio threshold to obtain multiple effective recognition units; Based on the recognition ratio and weight of the multiple effective recognition units, the final weight of each of the multiple effective recognition units is obtained; The step of determining the data category of the data to be identified as the data category of the identification unit with the highest final weight includes: The data category identified by the effective identification unit with the highest final weight is determined as the data category of the data to be identified.
3. The method according to claim 1, characterized in that, Also includes: Obtain the target data category to which the data to be identified belongs; If the data category of the data to be identified and the target data category are inconsistent, the identification mode of the identification unit, the weight of the identification mode, and / or the weight of the identification unit shall be adjusted.
4. The method according to claim 1, characterized in that, Also includes: Obtain the target data category to which the data to be identified belongs; If the target data category of the data to be identified indicates that the data to be identified is sensitive data, a new identification pattern is constructed based on the data to be identified. Based on the new recognition pattern, a new recognition unit is constructed.
5. A data category identification device, characterized in that, include: The input module is configured to input the data to be identified into multiple identification units, with each identification unit used to identify a data category. The data categories include at least one of mobile phone numbers, landline phone numbers, and usernames; The acquisition module is configured to acquire the recognition ratio of each of the plurality of recognition units for the data to be identified; The weighting module is configured to obtain the final weight of each of the multiple identification units based on the identification ratio of each of the multiple identification units to the data to be identified, and the weight of each of the multiple identification units. The determination module is configured to determine the data category of the data to be identified as the data category of the identification unit with the highest final weight. The acquisition module includes: The ratio acquisition unit is configured to acquire the recognition ratio of each recognition mode contained in each of the plurality of recognition units for the data to be recognized, and to acquire the weight of each recognition mode, wherein the recognition mode includes at least any one or more of the following: keyword dictionary mapping, regular expression matching, and AI recognition. The summation unit is configured to, for each of the identification units, perform a weighted summation of the identification ratios of the identification modes contained in the identification unit to the data to be identified, according to the weights of each of the identification modes, to obtain the identification ratio of the identification unit to the data to be identified. The input module includes: The information acquisition unit is configured to acquire multi-dimensional information of the data to be identified, wherein the multi-dimensional information includes at least one or more of the following: length, whether it is empty, and whether it is a primary key; The candidate acquisition unit is configured to acquire multiple candidate identification units of the data to be identified. The filtering unit is configured to filter the multiple candidate identification units based on the multi-dimensional information of the data to be identified, and obtain the identification unit; An input unit is configured to input the data to be identified into the identification unit; It also includes: a sensitive category determination module, configured to determine the sensitive category of the data to be identified based on the data category of the data to be identified, wherein the sensitive category includes at least one or more of the following: personal sensitive category and enterprise sensitive category.
6. The apparatus according to claim 5, characterized in that, The weighting module includes: The removal unit is configured to remove recognition units whose recognition ratio is less than a ratio threshold, thereby obtaining multiple valid recognition units. The final weighting unit is configured to obtain the final weight of each of the multiple effective identifications based on the identification ratio and weight of the multiple effective identification units; The determining module includes: The determining unit is configured to determine the data category identified by the effective identification unit with the highest final weight as the data category of the data to be identified.
7. The apparatus according to claim 5, characterized in that, Also includes: The first acquisition module is configured to acquire the target data category to which the data to be identified belongs; The adjustment module is configured to adjust the recognition mode, the weight of the recognition mode, and / or the weight of the recognition unit when the data category of the data to be identified and the target data category are inconsistent.
8. The apparatus according to claim 5, characterized in that, Also includes: The second acquisition module is configured to acquire the target data category to which the data to be identified belongs; The pattern building module is configured to build a new recognition pattern based on the data to be identified when the target data category of the data to be identified indicates that the data to be identified is sensitive data. The unit construction module is configured to construct new identification units based on the new identification pattern.
9. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the method for identifying data categories as described in any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the method for identifying data categories as described in any one of claims 1 to 4.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements a method for identifying data categories as described in any one of claims 1 to 4.