A method for intelligent identification of sensitive data based on label distribution learning
Through the method based on label distribution learning, a neural network model is established and iteratively trained, the problem of low recognition accuracy of sensitive data in the existing technology is solved, and higher recognition accuracy and recognition rate are achieved.
Patent Information
- Application Number
- CN202111223201.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-10-20
AI Technical Summary
In the prior art, the accuracy of sensitive data identification methods is low, making it difficult to quickly and accurately identify sensitive data in massive enterprise data.
Using a method based on label distribution learning, a collection of label distributions of training samples with multiple known results is generated, the parameters of the preset neural network are determined, the neural network model is established, and the sensitive data recognition model is obtained through iterative training.
The accuracy and recognition rate of sensitive data recognition are improved, especially in complex data scenarios, and the problem of low recognition accuracy in the prior art is solved.
Smart Images

Figure CN113962302B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data security, and specifically relates to a method for intelligently identifying sensitive data based on label distribution learning. Background Art
[0002] With the rapid development of information technology, all walks of life are highly dependent on information systems. How to ensure the security of information systems, especially how to ensure the security of data that reflects the core value of the enterprise, has become the most concerned issue for enterprises. Enterprise data contains a lot of user personal privacy information, commercial sensitive data, etc. Once leaked, it will bring huge economic losses to the enterprise, and it will have to bear relevant legal responsibilities and huge fines for violations. Therefore, how to ensure the security of enterprise users' personal privacy information, commercial sensitive data, etc. has become the top priority of enterprise information security work, and how to quickly identify sensitive data in massive amounts of enterprise information has become a key issue that needs to be solved.
[0003] At present, data leakage incidents occur frequently around the world, and a large number of internal information of enterprises and user information of Internet websites have been leaked by hackers. Faced with the frequent occurrence of data leakage incidents around the world, in order to better protect citizens' personal privacy information, countries have successively introduced relevant data protection laws and regulations. In May 2018, the European Union issued the General Data Protection Regulation (GDPR), which applies to all organizations that process and use personal data in the EU. Once violated, they will face extremely high fines. my country has also issued a series of relevant laws and regulations, including: "Cybersecurity Law of the People's Republic of China", "Critical Information Infrastructure Protection Regulations", "Network Data Security Management Measures", "Personal Information and Important Data Outbound Security Assessment Measures", "GB / T 35273 Personal Information Security Specification", "Telecommunications and Internet User Personal Information Protection Regulations", etc.
[0004] Faced with frequent data security incidents and increasingly stringent data security protection requirements, enterprises have realized the importance of data security protection, but the first thing enterprises face is how to confirm which data is sensitive data among the massive internal data of the enterprise. Common sensitive data matching methods usually use rule matching or machine learning methods to find sensitive fields in sensitive data, and judge the sensitivity of the data based on the type of sensitive fields, so as to establish corresponding protection measures.
[0005] However, the existing methods have a great deal of misjudgment, the main reason being that some data fields are ambiguous. For example, the telephone number field is a public field for customer service, but the personal telephone number is a sensitive field. Therefore, simply relying on rule matching or data type identification cannot effectively determine whether the field is a sensitive field, which in turn affects the judgment of the sensitivity of the entire data. Furthermore, the entire data does not only contain one field, but as a multi-field set, it contains multiple data fields. The identification of its sensitive attributes requires the identification of multiple sensitive fields, but the importance of different sensitive fields to the data is often different. For example, a user's electricity consumption data is marked with multiple tags such as "user name", "user location", "power consumption" and "power consumption time", but the degree to which these tags specifically describe the information is different; in enterprise data, the information of complex data is often the result of a mixture of multiple basic information (such as time, location, business and user), and these basic information often express different strengths in a specific data, thus presenting a complex information meaning. For such data containing complex information, existing methods often have a large misjudgment rate.
[0006] Machine learning is essentially the process of mapping instances to labels by computers. According to the different label marking methods, it can be divided into single-label learning and multi-label learning (MLL). As the name implies, single-label learning assigns a unique label to each instance. However, real objects do not have unique semantics. An object instance should be a collection of multiple features and multiple categories. Multi-label learning can assign multiple labels to each instance. Compared with a single label, it greatly enriches the label information of the instance, which helps the computer learn the instance more comprehensively and make more accurate judgments. Summary of the invention
[0007] Therefore, the technical problem to be solved by the present invention is that the accuracy of the sensitive data identification method in the prior art is low, and thus a sensitive data intelligent identification method based on label distribution learning is provided.
[0008] According to a first aspect, an embodiment of the present invention discloses a method for training a sensitive data identification model based on label distribution learning, comprising:
[0009] Obtain multiple training samples with known results;
[0010] Generate a label distribution set of training samples according to the label distribution learning algorithm and the training samples;
[0011] Determining parameters of a preset neural network according to the label distribution set to obtain a neural network model;
[0012] The neural network model is iteratively trained according to a plurality of training samples of known results to obtain a sensitive data recognition model.
[0013] Optionally, determining parameters of a preset neural network according to the label distribution set to obtain a neural network model includes:
[0014] Determining extraction feature parameters of a preset neural network according to the label distribution set;
[0015] Determine the loss function of the preset neural network based on the cross entropy loss;
[0016] The loss function can be expressed by the following formula:
[0017]
[0018] Among them, Loss represents the loss function, Represents the distribution value of the i-th sample data to the m-th label, represents the predicted probability that the i-th sample data belongs to the m-th label, N represents the number of samples, and q represents the number of labels;
[0019] A neural network model is determined according to the extracted feature parameters, the loss function, the approximation parameter and a preset approximation threshold.
[0020] Optionally, generating a label distribution set of training samples according to a label distribution learning algorithm and the training samples includes:
[0021] Get document vocabulary set;
[0022] Calculate the correlation between words and tags;
[0023] Calculate the correlation between vocabulary and samples;
[0024] Generate a set of label distributions for training samples.
[0025] Optionally, calculating the relevance between the vocabulary and the label includes:
[0026] Calculating a vocabulary label significance parameter; the vocabulary label significance parameter is the frequency of a vocabulary in a same-label vocabulary set, where the same-label vocabulary set is a vocabulary set marked with the same label in the document vocabulary set;
[0027] The vocabulary label significance parameter can be expressed by the following formula:
[0028]
[0029] in, Represents word w in document vocabulary set C jThe significance of the mth label, For word w j In X m The number of occurrences in X m is the set of words in the document vocabulary set C that are marked with label m;
[0030] Calculating a tag relevance parameter; the tag relevance parameter is the proportion of tags that can match the vocabulary in the tag set;
[0031] The tag relevance parameter can be expressed by the following formula:
[0032]
[0033] Where L is the set of labels and |L| is the number of elements in the set L. Represented as word w j Include w in the label set L j The number of labels;
[0034] Calculating a vocabulary tag relevance; the vocabulary tag relevance is the product of the vocabulary tag significance parameter and the vocabulary tag relevance parameter;
[0035] The vocabulary tag relevance can be expressed by the following formula:
[0036]
[0037] in, is the word w in the document vocabulary set C j Label relevance for the mth token.
[0038] Optionally, calculating the relevance between the vocabulary and the sample includes:
[0039] Calculating a vocabulary sample significance parameter; the vocabulary sample significance parameter is the frequency of a vocabulary in a same sample vocabulary set, the same sample vocabulary set being a vocabulary set of the same sample;
[0040] The vocabulary sample significance parameter can be expressed by the following formula:
[0041]
[0042] in, Represents word w in document vocabulary set C j The significance of the i-th training sample, For word w j In y i The number of occurrences in y i is the set of all words in the i-th training sample;
[0043] Calculating a sample relevance parameter; the sample relevance parameter is the proportion of samples that can match the vocabulary in the sample set;
[0044] The sample correlation parameter can be expressed by the following formula:
[0045]
[0046] Among them, S is the set of training samples, and |S| is the number of training samples in set S. Represented as word w j The training sample set S contains w j The number of training samples;
[0047] Calculating vocabulary sample relevance; the vocabulary sample relevance is the product of the vocabulary sample significance parameter and the sample relevance parameter;
[0048] The vocabulary sample relevance can be expressed by the following formula:
[0049]
[0050] in, is the word w in the document vocabulary set C j The relevance of the i-th training sample in the training sample set.
[0051] Optionally, generating a label distribution set of training samples includes:
[0052] Calculating a sample label relevance parameter; the sample label relevance parameter is the product of the vocabulary label relevance and the vocabulary sample relevance;
[0053] The sample label relevance parameter can be expressed by the following formula:
[0054]
[0055] Among them, ILR i,m is the correlation between the i-th sample and the m-th label;
[0056] Calculate a label distribution set; the label distribution set is the proportion of a single vocabulary sample label relevance parameter to all vocabulary sample label relevance parameters;
[0057] The label distribution set of the training samples can be expressed by the following formula:
[0058]
[0059]
[0060] Among them, Di is the label distribution set of the i-th sample, and q is the number of labels.
[0061] According to a second aspect, an embodiment of the present invention discloses a sensitive data identification method based on label distribution learning, comprising:
[0062] Obtaining samples to be tested;
[0063] The sample to be tested is input into the sensitive data identification model generated by the method for training a sensitive data identification model based on label distribution learning as described in the first aspect of the embodiment of the present invention, to obtain a sensitive data identification result of the sample to be tested.
[0064] Optionally, the sample to be tested is input into a sensitive data identification model generated by a method for training a sensitive data identification model based on label distribution learning, and a sensitive data identification result of the sample to be tested is obtained, including:
[0065] Inputting the sample to be tested into the sensitive data recognition model;
[0066] Extracting a label distribution set of the sample to be tested according to the sensitive data identification model;
[0067] Using the sensitive data recognition model to traverse the label distribution set of training samples, and determine the training sample that is closest to the label distribution set of the sample to be tested;
[0068] The sensitive data identification model is used to determine the sensitive data identification result that is closest to the training sample and output it as the sensitive data identification result of the sample to be tested.
[0069] Optionally, the step of traversing the label distribution set of training samples using the sensitive data identification model to determine a training sample that is closest to the label distribution set of the sample to be tested includes:
[0070] Traverse the label distribution set of training samples;
[0071] Calculate the approximation parameters between the sample to be tested and each training sample respectively;
[0072] When the similarity parameter is the smallest and smaller than a preset similarity threshold, the corresponding training sample is taken as the training sample closest to the sample to be tested.
[0073] Optionally, the approximation parameter is represented by a KL divergence value, which can be represented by the following formula:
[0074]
[0075] Among them, dis represents the KL divergence value, P j Represents the label distribution set of the samples to be tested, Q jRepresents the label distribution set of training samples.
[0076] The technical solution of the present invention has the following advantages:
[0077] 1. The present invention provides a method and device for training a sensitive data recognition model based on label distribution learning. Through the label distribution algorithm and preset parameters, a neural network model is established, which can use multiple labels to describe the detected data in a probabilistic manner. By training the neural network model with training samples, the detected data document can be matched with multiple labels related to sensitive data, and the document data can be converted into a mathematical model for machine recognition.
[0078] 2. The present invention provides a sensitive data identification method and device based on label distribution learning. By inputting the sample to be tested into the neural network model, the sensitive data identification process of the document data is converted into a comparison process of the mathematical model, and the characteristics of the neural network are used to achieve accurate identification of the sensitivity of the data being tested. It has a better recognition rate in difficult scenarios, and solves the problem of low accuracy of the sensitive data intelligent identification method in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0080] Figure 1 A flowchart of a specific example of a method for training a sensitive data identification model based on label distribution learning in an embodiment of the present invention;
[0081] Figure 2 A flowchart of another specific example of a method for training a sensitive data identification model based on label distribution learning in an embodiment of the present invention;
[0082] Figure 3 A flowchart of another specific example of a method for training a sensitive data identification model based on label distribution learning in an embodiment of the present invention;
[0083] Figure 4 A flowchart of another specific example of a method for training a sensitive data identification model based on label distribution learning in an embodiment of the present invention;
[0084] Figure 5 A flowchart of another specific example of a method for training a sensitive data identification model based on label distribution learning in an embodiment of the present invention;
[0085] Figure 6 A flowchart of another specific example of a method for training a sensitive data identification model based on label distribution learning in an embodiment of the present invention
[0086] Figure 7 This is a flowchart of a specific example of a sensitive data identification method based on label distribution learning in an embodiment of the present invention;
[0087] Figure 8 This is a flowchart of another specific example of a sensitive data identification method based on label distribution learning in an embodiment of the present invention;
[0088] Fig. 9 A principle block diagram of a specific example of a device for training a sensitive data recognition model in an embodiment of the present invention;
[0089] Fig.10 is a principle block diagram of a specific example of a sensitive data identification device in an embodiment of the present invention;
[0090] Fig.11 FIG. 4 is a specific example diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION
[0091] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0092] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
[0093] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, it can also be the internal connection of two components, it can be a wireless connection, or it can be a wired connection. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0094] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0095] The embodiment of the present invention discloses a method for training a sensitive data recognition model based on label distribution learning, such as Figure 1 As shown, the method comprises the following steps:
[0096] Step S11, obtaining a plurality of training samples of known results.
[0097] Specifically, the training samples with known results are document data that have been annotated with data sensitivity, that is, the sensitivity results of the training samples are known.
[0098] Step S12, generating a label distribution set of training samples according to the label distribution learning algorithm and the training samples.
[0099] Specifically, the label distribution learning algorithm is used to calculate the correlation between the preset label and the training sample, and the training sample is described by a probabilistic distribution set of the correlation between each preset label and the training sample. The probabilistic distribution set is the label distribution set.
[0100] Step S13, determining the parameters of the preset neural network according to the label distribution set to obtain a neural network model.
[0101] Specifically, the neural network parameters mainly include extracted feature parameters, loss function and approximation threshold. The extracted feature parameters represent the parameters compared when the neural network model works. The loss function represents the gap between the predicted calculation result of each iteration of the neural network and the true value, thereby guiding the next step of training to proceed in the right direction. The approximation threshold represents the gap between the extracted feature parameters of the sample to be tested and the extracted feature parameters of the training sample.
[0102] Among them, the extracted feature parameters are the label distribution set, the loss function is determined according to the cross entropy loss, and the approximation threshold is predetermined according to the empirical value. By combining the above-mentioned extracted feature parameters, loss function and approximation threshold with the above-mentioned label distribution learning algorithm, the neural network model can be determined.
[0103] Step S14, iteratively training the neural network model according to a plurality of training samples with known results to obtain a sensitive data recognition model.
[0104] The method for training a sensitive data identification model based on label distribution learning provided in an embodiment of the present invention can correspond data documents to multiple sensitive data labels through a label distribution algorithm, thereby realizing a probabilistic description of data by multiple labels, and determining a sensitive data identification model by iteratively training a determined neural network model. The sensitive data identification model can be used to realize accurate identification of data sensitivity.
[0105] As an optional implementation of the present invention, a label distribution set of training samples is generated according to the label distribution learning algorithm and the training samples, such as Figure 2 As shown, the following steps are included:
[0106] Step S121, obtaining a document vocabulary set.
[0107] Specifically, firstly, feature words are extracted from the sample data to generate a feature word matrix, and the feature word matrix is used as a document vocabulary set. Feature words can be extracted using a variety of known algorithms, which are not limited in the present invention.
[0108] Step S122, calculating the correlation between the vocabulary and the label.
[0109] Specifically, the calculated correlation between the vocabulary and the label is obtained by calculating the vocabulary label significance parameter and the label correlation parameter. The vocabulary label significance reflects the significance of the vocabulary to a specific label. If a vocabulary often appears in instances of a specific label, then this vocabulary will play an important role in the specific label. The label correlation parameter reflects the ability of a single vocabulary to distinguish different labels. If a single vocabulary often appears in multiple labels, then the label correlation of the vocabulary will be very low.
[0110] Step S123, calculating the correlation between the vocabulary and the sample.
[0111] Specifically, the correlation between the calculated vocabulary and the training sample is calculated by the vocabulary sample significance parameter and the sample correlation parameter. The vocabulary sample significance parameter reflects the significance of the vocabulary to a specific training sample. If a single vocabulary often appears in a sample, then the vocabulary will play an important role in the sample. The sample correlation parameter reflects the ability of a single vocabulary to distinguish different samples. If a single vocabulary often appears in multiple samples, then the sample correlation of the vocabulary will be very low.
[0112] Step S124, generating a label distribution set of training samples.
[0113] Specifically, the generated label distribution set is obtained by calculating the sample label relevance parameter. The sample label relevance parameter reflects the degree of relevance between the sample and the label.
[0114] In one embodiment, step S122 calculates the correlation between the vocabulary and the tag, such as Figure 3 As shown, the following steps are included:
[0115] Step S1221, calculating the vocabulary tag significance parameter WLS.
[0116] Exemplarily, the vocabulary tag significance parameter can be expressed by the following formula:
[0117]
[0118] in, Represents word w in document vocabulary set C j The significance of the mth label, For word w j In X m The number of occurrences in X m is the set of words in the document vocabulary set C that are labeled with label m.
[0119] Step S1222, calculating the tag relevance parameter LR.
[0120] Exemplarily, the tag relevance parameter can be expressed by the following formula:
[0121]
[0122] Where L is the set of labels and |L| is the number of elements in the set L. Represented as word w j Include w in the label set L j The number of labels.
[0123] Step S1223, calculating the vocabulary tag relevance WLR.
[0124] Exemplarily, the vocabulary tag relevance can be expressed by the following formula:
[0125]
[0126] in, is the word w in the document vocabulary set C j Label relevance for the mth token.
[0127] In one embodiment, step S123 calculates the correlation between the vocabulary and the sample, such as Figure 4As shown, the following steps are included:
[0128] Step S1231, calculating the vocabulary sample significance parameter WIS.
[0129] Exemplarily, the vocabulary sample significance parameter can be expressed by the following formula:
[0130]
[0131] in, Represents word w in document vocabulary set C j The significance of the i-th training sample, For word w j In y i The number of occurrences in y i is the set of all words in the i-th training sample.
[0132] Step S1232, calculating the sample correlation parameter IR.
[0133] Exemplarily, the sample correlation parameter can be expressed by the following formula:
[0134]
[0135] Among them, S is the set of training samples, and |S| is the number of training samples in set S. Represented as word w j The training sample set S contains w j The number of training samples.
[0136] Step S1233, calculating the vocabulary sample relevance WIR.
[0137] Exemplarily, the vocabulary sample relevance can be expressed by the following formula:
[0138]
[0139] in, is the word w in the document vocabulary set C j The relevance of the i-th training sample in the training sample set.
[0140] In one embodiment, step S124 generates a label distribution set of training samples, such as Figure 5 As shown, the following steps are included:
[0141] Step S1241, calculating the sample label relevance parameter ILR.
[0142] Exemplarily, the sample label relevance parameter can be expressed by the following formula:
[0143]
[0144] Among them, ILR i,m is the correlation between the i-th sample and the m-th label.
[0145] Step S1242, calculate the label distribution set D.
[0146] Exemplarily, the label distribution set of the training samples can be expressed by the following formula:
[0147]
[0148]
[0149] Among them, D i is the label distribution set of the i-th sample, and q is the number of labels.
[0150] As an optional implementation mode of the present invention, the parameters of the preset neural network are determined according to the label distribution set to obtain a neural network model, such as Figure 6 As shown, the following steps are included:
[0151] Step S131, determining the extraction feature parameters of the preset neural network according to the label distribution set.
[0152] Step S132, determining a loss function LOSS of a preset neural network according to the cross entropy loss.
[0153] Exemplarily, the loss function can be expressed by the following formula:
[0154]
[0155] in, Represents the distribution value of the i-th sample data to the m-th label, It represents the predicted probability that the i-th sample data belongs to the m-th label.
[0156] Step S133, determining a neural network model according to the extracted feature parameters, the loss function and a preset approximation threshold. Specifically, the approximation threshold of the neural network is predetermined based on experience, and the approximation threshold can be any number greater than 0 and less than 1.
[0157] The embodiment of the present invention also discloses a sensitive data identification method based on label distribution learning, such as Figure 7 As shown, the method comprises the following steps:
[0158] Step S21, obtaining a sample to be tested.
[0159] Step S22: input the sample to be tested into the sensitive data identification model generated by the method of training the sensitive data identification model based on label distribution learning in the above embodiment to obtain the sensitive data identification result of the sample to be tested.
[0160] The sensitive data identification method provided by the embodiment of the present invention inputs the sample to be tested into the sensitive data identification model, converts the sensitive data identification process of the document data into a comparison process of the mathematical model, and utilizes the characteristics of the neural network to achieve accurate identification of the sensitivity of the data being tested. It has a better recognition rate in difficult scenarios, and solves the problem of low accuracy of the sensitive data intelligent identification method in the prior art.
[0161] As an optional implementation of the present invention, step S22 inputs the sample to be tested into the sensitive data identification model generated by the method of training the sensitive data identification model based on label distribution learning in the above embodiment to obtain the sensitive data identification result of the sample to be tested, such as Figure 8 As shown, the following steps are included:
[0162] Step S221: input the sample to be tested into the sensitive data recognition model.
[0163] Step S222: extracting a label distribution set of the sample to be tested according to the sensitive data identification model.
[0164] Step S223: traverse the label distribution set of training samples using the sensitive data identification model to determine the training sample that is closest to the label distribution set of the sample to be tested.
[0165] Specifically, the label distribution set of training samples is traversed, and the approximation parameters between the sample to be tested and each training sample are calculated respectively. When the approximation parameter is the smallest and less than a preset approximation threshold, the corresponding training sample is taken as the training sample closest to the sample to be tested.
[0166] Exemplarily, the KL divergence value may be used to represent the approximation parameter, and the KL divergence value may be represented by the following formula:
[0167]
[0168] Among them, dis represents the KL divergence value, P j Represents the label distribution set of the samples to be tested, Q j Represents the label distribution set of training samples.
[0169] Step S224: Use the sensitive data identification model to determine the sensitive data identification result that is closest to the training sample, and output it as the sensitive data identification result of the sample to be tested.
[0170] The embodiment of the present invention discloses a device for training a sensitive data recognition model based on label distribution learning, such as Fig. 9 As shown, including:
[0171] Communication module 901, used to obtain a plurality of training samples of known results;
[0172] A first calculation module 902, configured to generate a label distribution set of training samples according to a label distribution learning algorithm and the training samples;
[0173] A second calculation module 903 is used to determine the parameters of a preset neural network according to the label distribution set to obtain a neural network model;
[0174] The training module 904 is used to iteratively train the neural network model according to a plurality of training samples of known results to obtain a sensitive data recognition model.
[0175] The embodiment of the present invention provides a device for training a sensitive data recognition model based on label distribution learning. Through the label distribution algorithm and preset parameters, a neural network model is established, which can use multiple labels to describe the detected data in a probabilistic manner. By training the neural network model with training samples, the detected data document can correspond to multiple labels related to sensitive data, and the document data can be converted into a mathematical model for machine recognition.
[0176] The embodiment of the present invention also discloses a sensitive data identification device based on label distribution learning, such as Fig.10 As shown, including:
[0177] Communication module 1001, used to obtain samples to be tested;
[0178] The identification module 1002 is used to input the sample to be tested into the sensitive data identification model generated by the method for training the sensitive data identification model based on label distribution learning described in the above method embodiment to obtain the sensitive data identification result of the sample to be tested.
[0179] The present invention provides a sensitive data identification device based on label distribution learning. By inputting the sample to be tested into the neural network model, the sensitive data identification process of the document data is converted into a comparison process of the mathematical model, and the characteristics of the neural network are utilized to realize accurate identification of the sensitivity of the data being tested. It has a better recognition rate in difficult scenarios, and solves the problem of low accuracy of the sensitive data intelligent identification method in the prior art.
[0180] The embodiment of the present invention further provides an electronic device, such as Fig.11As shown, the electronic device may include a processor 1101 and a memory 1102, wherein the processor 1101 and the memory 1102 may be connected via a bus or other means. Fig. 9 The example of connecting through bus is taken in the following.
[0181] The processor 1101 may be a central processing unit (CPU). The processor 1101 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.
[0182] The memory 1102, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as the method for training a sensitive data identification model based on label distribution learning and the program instructions / modules corresponding to the sensitive data identification method based on label distribution learning in the embodiments of the present invention. The processor 1101 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory 1102, that is, the method for training a sensitive data identification model based on label distribution learning and the sensitive data identification method based on label distribution learning in the above method embodiments are implemented.
[0183] The memory 1102 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required by at least one function; the data storage area may store data created by the processor 1101, etc. In addition, the memory 1102 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 1102 may optionally include a memory remotely arranged relative to the processor 1101, and these remote memories may be connected to the processor 1101 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0184] The one or more modules are stored in the memory 1102, and when executed by the processor 1101, the following is performed: Figure 1 The method for training a sensitive data recognition model based on label distribution learning in the illustrated embodiment and Figure 7The sensitive data identification method based on label distribution learning in the illustrated embodiment.
[0185] For details of the above electronic equipment, please refer to Figure 1 The corresponding related descriptions and effects in the illustrated embodiments can be understood and will not be repeated here.
[0186] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memory.
[0187] Although the exemplary embodiments and their advantages have been described in detail, those skilled in the art may make various changes, substitutions and modifications to these embodiments without departing from the spirit of the present invention and the scope of protection defined by the appended claims, and such modifications and variations all fall within the scope defined by the appended claims. For other examples, those of ordinary skill in the art should readily understand that the order of the process steps may be changed while maintaining the scope of protection of the present invention.
[0188] In addition, the scope of application of the present invention is not limited to the processes, mechanisms, manufactures, material compositions, means, methods and steps of the specific embodiments described in the specification. From the disclosure of the present invention, it will be easily understood by those skilled in the art that the processes, mechanisms, manufactures, material compositions, means, methods or steps that currently exist or will be developed in the future, which perform substantially the same functions as the corresponding embodiments described in the present invention or obtain substantially the same results, can be applied according to the present invention. Therefore, the claims attached to the present invention are intended to include these processes, mechanisms, manufactures, material compositions, means, methods or steps within their scope of protection.
Claims
1. A method for training a sensitive data recognition model based on label distribution learning, characterized in that: include: Obtain multiple training samples with known results; The training samples with known results are document data that have been annotated with data sensitivity, that is, the sensitivity results of the training samples are known; Generate a label distribution set of training samples according to the label distribution learning algorithm and the training samples; calculate the correlation between the preset labels and the training samples by using the label distribution learning algorithm, and describe the training samples by a probabilistic distribution set of the correlation between each preset label and the training samples, the probabilistic distribution set being the label distribution set; Determining parameters of a preset neural network according to the label distribution set to obtain a neural network model; Iteratively training the neural network model according to a plurality of training samples with known results to obtain a sensitive data recognition model; Determining parameters of a preset neural network according to the label distribution set to obtain a neural network model includes: Determining extraction feature parameters of a preset neural network according to the label distribution set; Determine the loss function of the preset neural network based on the cross entropy loss; The loss function is expressed by the following formula: Among them, Loss represents the loss function, Represents the distribution value of the i-th sample data to the m-th label, It represents the predicted probability that the i-th sample data belongs to the m-th label, N represents the number of samples, and q represents the number of labels; Determine a neural network model according to the extracted feature parameters, the loss function, the approximation parameter and a preset approximation threshold; Generating a label distribution set of training samples according to the label distribution learning algorithm and the training samples includes: Get document vocabulary set; Calculate the correlation between words and tags; Calculate the correlation between vocabulary and samples; Generate a label distribution set of training samples; The label distribution set for generating training samples includes: Calculating a sample label relevance parameter; the sample label relevance parameter is the product of the vocabulary label relevance and the vocabulary sample relevance; The sample label relevance parameter is expressed by the following formula: Among them, ILR i,m is the correlation between the i-th sample and the m-th label; Calculate a label distribution set; the label distribution set is the proportion of a single vocabulary sample label relevance parameter to all vocabulary sample label relevance parameters; The label distribution set of the training samples is expressed by the following formula: Among them, D i is the label distribution set of the i-th sample, and q is the number of labels.
2. The method for training a sensitive data recognition model based on label distribution learning according to claim 1, characterized in that: The calculation of the relevance between the vocabulary and the label includes: Calculating a vocabulary label significance parameter; the vocabulary label significance parameter is the frequency of a vocabulary in a same-label vocabulary set, where the same-label vocabulary set is a vocabulary set marked with the same label in the document vocabulary set; The vocabulary tag significance parameter is expressed by the following formula: in, Represents word w in document vocabulary set C j The significance of the mth label, For word w j In X m The number of occurrences in X m is the set of words in the document vocabulary set C that are marked with label m; Calculating a tag relevance parameter; the tag relevance parameter is the proportion of tags that can match the vocabulary in the tag set; The tag relevance parameter is expressed by the following formula: Where L is the set of labels, |L| is the number of elements in the set L, Represented as word w j Include w in the label set L j The number of labels; Calculating a vocabulary tag relevance; the vocabulary tag relevance is the product of the vocabulary tag significance parameter and the vocabulary tag relevance parameter; The vocabulary tag relevance is expressed by the following formula: in, is the word w in the document vocabulary set C j Label relevance for the mth token.
3. The method for training a sensitive data recognition model based on label distribution learning according to claim 1, characterized in that: The calculating the relevance between the vocabulary and the sample comprises: Calculating a vocabulary sample significance parameter; the vocabulary sample significance parameter is the frequency of a vocabulary in a same sample vocabulary set, the same sample vocabulary set being a vocabulary set of the same sample; The vocabulary sample significance parameter is expressed by the following formula: in, Represents word w in document vocabulary set C j The significance of the i-th training sample, For word w j In y i The number of occurrences in y i is the set of all words in the i-th training sample; Calculating a sample relevance parameter; the sample relevance parameter is the proportion of samples that can match the vocabulary in the sample set; The sample correlation parameter is expressed by the following formula: Among them, S is the set of training samples, |S| is the number of training samples in the set S, Represented as word w j The training sample set S contains w j The number of training samples; Calculating vocabulary sample relevance; the vocabulary sample relevance is the product of the vocabulary sample significance parameter and the sample relevance parameter; The vocabulary sample relevance is expressed by the following formula: in, is the word w in the document vocabulary set C j The relevance of the i-th training sample in the training sample set.
4. A sensitive data identification method based on label distribution learning, characterized in that: Obtaining samples to be tested; The sample to be tested is input into the sensitive data identification model generated by the method for training a sensitive data identification model based on label distribution learning as described in any one of claims 1 to 3 to obtain a sensitive data identification result of the sample to be tested.
5. The sensitive data identification method according to claim 4, characterized in that: The sample to be tested is input into the sensitive data identification model generated by the method of training the sensitive data identification model based on label distribution learning, and the sensitive data identification result of the sample to be tested is obtained, including: Inputting the sample to be tested into the sensitive data recognition model; Extracting a label distribution set of the sample to be tested according to the sensitive data identification model; Using the sensitive data recognition model to traverse the label distribution set of training samples, and determine the training sample that is closest to the label distribution set of the sample to be tested; The sensitive data identification model is used to determine the sensitive data identification result that is closest to the training sample and output it as the sensitive data identification result of the sample to be tested.
6. The sensitive data identification method according to claim 5, characterized in that: The step of traversing the label distribution set of training samples using the sensitive data identification model to determine the training sample closest to the label distribution set of the sample to be tested comprises: Traverse the label distribution set of training samples; Calculate the approximation parameters between the sample to be tested and each training sample respectively; When the similarity parameter is the smallest and smaller than a preset similarity threshold, the corresponding training sample is taken as the training sample closest to the sample to be tested.
7. The sensitive data identification method according to claim 6, characterized in that: The similarity parameter is represented by the KL divergence value, which is represented by the following formula: Among them, dis represents the KL divergence value, P j Represents the label distribution set of the samples to be tested, Q j Represents the label distribution set of training samples.
Citation Information
Patent Citations
Method and system for multi-label distribution learning in natural language processing classification model
CN111797234A
System and method for a convolutional neural network for multi-label classification with partial annotations
US20200160177A1