Sensitive data identification method and device and storage medium

By converting the text to be identified into a first sequence and using a sensitive data recognition model for identification, the problem of low accuracy of sensitive data recognition in the prior art is solved, and a more efficient sensitive data recognition and training process is achieved.

CN120217445APending Publication Date: 2025-06-27CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510387070.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing natural language processing methods are not accurate when identifying sensitive data, especially in dynamic recognition scenarios, especially when new sensitive data are involved, the amount of data is limited and it is difficult to effectively identify.

Method used

By converting the text to be identified into a first sequence and encoding it positionally, using the sensitive data identification model to identify features that match the sensitive data from the first sequence and the position sequence, the model only trains iteratively for the new type of sensitive data during training, and using distillation learning to retain the ability to identify sensitive data of known types until the accuracy of identifying sensitive data of the new type reaches the expected value.

Benefits of technology

The accuracy and efficiency of natural language recognition methods for identifying sensitive data is improved, and different types of sensitive and non-sensitive data can be effectively identified, shortening training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217445A_ABST
    Figure CN120217445A_ABST
Patent Text Reader

Abstract

The invention discloses a sensitive data recognition method and device and a storage medium, and the method comprises the steps: converting a to-be-recognized text into a first sequence segmented according to vocabularies or characters, and carrying out the position coding of each element in the first sequence, so as to obtain a position sequence; identifying a plurality of elements which conform to the features of the sensitive data and are continuous in position from the first sequence and the position sequence by using a sensitive data identification model so as to determine the sensitive data of the to-be-identified text; the sensitive data identification model only performs iterative training on the sensitive data of the new type during training, and the sensitive data identification model keeps the ability of identifying the sensitive data of the known type based on distillation learning in the iterative training process until the accuracy of identifying the sensitive data of the new type reaches an expected value; the sensitive data of the known type is the sensitive data which can be identified by the sensitive data identification model before the sensitive data of the new type appears.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing, and in particular, to a method, apparatus, and storage medium for identifying sensitive data. Background Art

[0002] With the rapid development of Internet technology and the sharp increase in the amount of data, we are gradually entering the big data era.

[0003] While big data technology creates value for all walks of life and brings convenience to people's lives, it also inevitably increases the risk of data leakage. Once sensitive data such as user identity information and medical information is leaked, it is very likely to cause mental and economic losses to individual users, and will also have a serious negative impact on the reputation of enterprises. The leakage of personal privacy data of public officials even threatens national security. In order to analyze and calculate data to fully explore the potential value of data without the data holder leaking sensitive data, various identification technologies for sensitive data have been proposed. Currently, the sensitive data identification technologies can be mainly divided into three categories: sensitive data identification based on pattern matching, sensitive data identification based on machine learning, and identification technology based on natural language processing.

[0004] The sensitive data identification method based on natural language processing depends on a large amount of labeled sensitive data during model training. In actual applications, sensitive data often shows an obvious long-tail phenomenon, and its quantity is much less than that of normal data. Especially in the dynamic identification scenario, this situation is particularly serious. Especially when it comes to newly added sensitive data, its data volume is even more limited, and the newly added sensitive data needs to be trained together with the previous sensitive data in the model. And when the quantity of sensitive data is small, it is also very difficult for the existing identification methods to effectively identify these sensitive data.

[0005] In view of this, how to improve the accuracy of identifying sensitive data by using the identification method of natural language has become an urgent technical problem to be solved. Summary of the Invention

[0006] The present invention provides a method, apparatus, and storage medium for identifying sensitive data to solve the technical problem that the accuracy of identifying sensitive data by using the identification method of natural language in the prior art is not high.

[0007] In a first aspect, to solve the above technical problem, the technical solution of a method for identifying sensitive data provided by an embodiment of the present invention is as follows:

[0008] Convert the text to be identified into a first sequence; wherein, the elements in the first sequence represent a word or character in the text to be identified, and the arrangement order of the elements in the first sequence is the same as the reading order of the text to be identified;

[0009] Encode the position of each element in the first sequence within the first sequence to obtain the position sequence corresponding to the first sequence;

[0010] Use a sensitive data recognition model to identify multiple consecutive elements in the first sequence and the position sequence that conform to the characteristics of sensitive data, and determine the sensitive data of the text to be recognized based on the text corresponding to the multiple elements; wherein, during training, the sensitive data recognition model only performs iterative training for new types of sensitive data, and during the iterative training process, the sensitive data recognition model retains the ability to recognize known types of sensitive data based on distillation learning until the accuracy of the sensitive data recognition model in recognizing the new type of sensitive data reaches the expected value; the known type of sensitive data is the sensitive data that the sensitive data recognition model can recognize before the new type of sensitive data appears.

[0011] A possible implementation, the sensitive data recognition model includes:

[0012] A student model and a classifier connected in sequence;

[0013] Before the first iterative training, the parameters of the student model are inherited from the teacher model to acquire the capabilities of the teacher model; after each iterative training, the parameter update of the student model is constrained by the distillation loss between the teacher model and the student model and the classification loss of the classifier; the teacher model is a model that can correctly recognize the known type of sensitive data;

[0014] The classifier is used to determine the sensitive data in the text to be recognized according to the output result of the student model.

[0015] A possible implementation, the classifier includes:

[0016] An original part, which is used to classify and recognize the known type of sensitive data;

[0017] A new part, which is used to classify and recognize the new type of sensitive data.

[0018] A possible implementation, before using the sensitive data recognition model to identify multiple consecutive elements in the first sequence and the position sequence that conform to the characteristics of sensitive data, further includes:

[0019] Iteratively train the initial sensitive data recognition model with a data set containing the new type of sensitive data until the accuracy of the trained initial sensitive data recognition model in recognizing the new type of sensitive data reaches the expected value, thereby obtaining the sensitive data recognition model; wherein, the initial sensitive data recognition model includes an initial student model and an initial classifier, the initial student model inherits the parameters of the teacher model, and the parameter values of the new part are randomly set before the first iterative training.

[0020] During the process of iteratively training the initial sensitive data recognition model at any time, the parameters in the initial student model corresponding to the iterative training at any time and the parameters in the classifier are updated by backpropagation in the initial sensitive data recognition model corresponding to the iterative training at any time; wherein, the joint loss value is the sum of the distillation loss and the classification loss.

[0021] A possible implementation manner, before iteratively training the initial sensitive data recognition model with a data set containing the new type of sensitive data, further includes:

[0022] Construct a pre-trained language model.

[0023] Perform masked language pre-training on the pre-trained language model with the second sequences corresponding to the unlabeled sample data in the unlabeled sample data set until the number of times the pre-trained language model is trained with each of the second sequences reaches a preset value, thereby obtaining the teacher model that can express the unlabeled sample data in natural language; wherein, the unlabeled sample data set includes non-sensitive data and sensitive data.

[0024] A possible implementation manner, before iteratively training the initial sensitive data recognition model with a data set containing the new type of sensitive data, further includes:

[0025] Group the collected sensitive data samples of multiple types into pairs of two sensitive data samples, and randomly form a sample set composed of pairs of sensitive data samples.

[0026] Use the sample set to train the initial classifier to identify whether the types of the two sensitive data samples are the same until the recognition accuracy reaches a preset value, thereby obtaining the classifier; wherein, the initial classifier determines whether the types of the two sensitive data samples are the same based on the distribution of words or characters in the two sensitive data samples by learning with a conditional random field.

[0027] A possible implementation manner, the student model includes:

[0028] An encoding layer and a feature extraction layer connected in sequence; the encoding layer is used to calculate the features of each first element to be encoded in the first sequence itself and the relationship features between the first element and the second element, where the distance between the positions of the respective corresponding characters of the second element and the first element in the text to be recognized is less than or equal to a preset distance, and the relationship features include word order, grammar, and lexical relationships;

[0029] The feature extraction layer includes a bidirectional long short-term memory network, a residual network, and a feed-forward neural network connected in sequence; wherein, the input end and the output end of the bidirectional long short-term memory network are respectively connected to different input ends of the residual network, the output end of the residual network is connected to the input end of the feed-forward neural network, and the output end of the feed-forward neural network is used as the output end of the pre-trained language model; the bidirectional long short-term memory network is used to obtain the context features of each element from multiple features of each element in the received first sequence, the residual network is used to fuse the input and output of the bidirectional long short-term memory network and then transmit them to the feed-forward neural network, and the feed-forward neural network determines, according to the output of the residual network, multiple consecutive elements in position from the first sequence, and the data formed by the multiple elements conforms to the natural language expression habit of the unlabeled data.

[0030] A possible implementation manner, converting the text to be recognized into a first sequence, includes:

[0031] Segmenting out the same vocabulary as in the corpus from the text to be recognized; wherein, the vocabulary recorded in the corpus is known common vocabulary;

[0032] If there is a part of the text to be recognized that has not been segmented into the vocabulary, then segment the part of the data to be recognized that has not been segmented into the vocabulary by characters;

[0033] Converting the segmented vocabulary and / or characters into numerical vectors, and forming the numerical vectors into the first sequence according to the positions of the vocabulary and / or characters in the text to be recognized.

[0034] A possible implementation manner, converting the text to be recognized into a first sequence, includes:

[0035] If the text to be recognized only contains numbers and letters, then segment the data to be recognized by characters; and convert the segmentation result into numerical vectors, and form the numerical vectors into the first sequence according to the positions of the characters in the text to be recognized.

[0036] In a second aspect, an embodiment of the present invention provides a sensitive data recognition device, including:

[0037] A conversion unit for converting the text to be recognized into a first sequence; wherein, the elements in the first sequence represent a word or a character in the text to be recognized, and the arrangement order of the elements in the first sequence is consistent with the reading order of the text to be recognized;

[0038] An encoding unit for encoding the position of each element in the first sequence in the first sequence to obtain a position sequence corresponding to the first sequence;

[0039] A recognition unit for using a sensitive data recognition model to recognize, from the first sequence and the position sequence, multiple consecutive elements that conform to the characteristics of sensitive data and determining the sensitive data of the text to be recognized based on the characters corresponding to the multiple elements; wherein, during training, the sensitive data recognition model only performs iterative training for new types of sensitive data, and during the iterative training, the sensitive data recognition model retains the ability to recognize known types of sensitive data based on distillation learning until the accuracy of the sensitive data recognition model in recognizing the new type of sensitive data reaches an expected value; the known type of sensitive data is the sensitive data that the sensitive data recognition model can recognize before the new type of sensitive data appears.

[0040] A possible implementation, the sensitive data recognition model includes:

[0041] A student model and a classifier connected in sequence;

[0042] Before the first iterative training, the parameters of the student model are inherited from the teacher model to acquire the capabilities of the teacher model; after each iterative training, the parameter update of the student model is constrained by the distillation loss between the teacher model and the student model and the classification loss of the classifier; the teacher model is a model that can correctly recognize the known type of sensitive data;

[0043] The classifier is used to determine the sensitive data in the text to be recognized according to the output result of the student model.

[0044] A possible implementation, the classifier includes:

[0045] An original part for classifying and recognizing the known type of sensitive data;

[0046] A new part for classifying and recognizing the new type of sensitive data.

[0047] A possible implementation, the device further includes a training unit for:

[0048] Iteratively train the initial sensitive data recognition model with a data set containing the new type of sensitive data until the accuracy rate of the trained initial sensitive data recognition model in recognizing the new type of sensitive data reaches the expected value, thereby obtaining the sensitive data recognition model; wherein, the initial sensitive data recognition model includes an initial student model and an initial classifier, the initial student model inherits the parameters of the teacher model, and the parameter values of the new part are randomly set before the first iterative training.

[0049] During the process of iteratively training the initial sensitive data recognition model at any time, the parameters in the initial student model corresponding to the iterative training at any time and the parameters in the classifier are updated by backpropagation in the initial sensitive data recognition model corresponding to the iterative training at any time; wherein, the combined loss value is the sum of the distillation loss and the classification loss.

[0050] A possible implementation manner, the training unit is further configured to:

[0051] Construct a pre-trained language model;

[0052] Perform masked language pre-training on the pre-trained language model with the second sequences corresponding to the unlabeled sample data in the unlabeled sample data set until the number of times the pre-trained language model is trained with each of the second sequences reaches a preset value, thereby obtaining the teacher model capable of expressing the unlabeled sample data in natural language; wherein, the unlabeled sample data set includes non-sensitive data and sensitive data.

[0053] A possible implementation manner, the training unit is further configured to:

[0054] Group the collected sensitive data samples of multiple types into pairs of two sensitive data samples, and randomly form a sample set composed of pairs of sensitive data samples;

[0055] Use the sample set to train the initial classifier to recognize whether the types of the two sensitive data samples are the same until the recognition accuracy rate reaches a preset value, thereby obtaining the classifier; wherein, the initial classifier determines whether the types of the two sensitive data samples are the same based on the distribution of words or characters in the two sensitive data samples by learning with a conditional random field.

[0056] A possible implementation manner, the student model includes:

[0057] An encoding layer and a feature extraction layer connected in sequence; the encoding layer is used to calculate the features of each first element to be encoded in the first sequence itself and the relationship features between the first element and the second element, where the distance between the positions of the respective corresponding characters of the second element and the first element in the text to be recognized is less than or equal to a preset distance, and the relationship features include word order, grammar, and lexical relationships;

[0058] The feature extraction layer includes a bidirectional long short-term memory network, a residual network, and a feedforward neural network connected in sequence; wherein, the input end and the output end of the bidirectional long short-term memory network are respectively connected to different input ends of the residual network, the output end of the residual network is connected to the input end of the feedforward neural network, and the output end of the feedforward neural network is used as the output end of the pre-trained language model; the bidirectional long short-term memory network is used to obtain the context features of each element from multiple features of each element in the received first sequence, the residual network is used to fuse the input and output of the bidirectional long short-term memory network and then transmit them to the feedforward neural network, and the feedforward neural network determines, according to the output of the residual network, multiple consecutive elements in position from the first sequence, and the data formed by the multiple elements conforms to the natural language expression habits of the unlabeled data.

[0059] A possible implementation, the conversion unit is further used for:

[0060] Segment out the same vocabulary as in the corpus from the text to be recognized; wherein, the vocabulary recorded in the corpus is known common vocabulary;

[0061] If there is a part of the text to be recognized that has not been segmented into the vocabulary, then segment the part of the data to be recognized that has not been segmented into the vocabulary by characters;

[0062] Convert the segmented vocabulary and / or characters into numerical vectors, and form the first sequence with the numerical vectors according to the positions of the vocabulary and / or characters in the text to be recognized.

[0063] A possible implementation, the conversion unit is further used for:

[0064] If the text to be recognized only contains numbers and letters, segment the data to be recognized by characters; and convert the segmentation result into a numerical vector, and form the first sequence with the numerical vectors according to the positions of the characters in the text to be recognized.

[0065] In a third aspect, an embodiment of the present invention further provides a sensitive data recognition device, including:

[0066] At least one processor, and

[0067] A memory connected to the at least one processor;

[0068] Wherein, the memory stores instructions executable by the at least one processor, and the at least one processor executes the method described in the first aspect above by executing the instructions stored in the memory.

[0069] In a fourth aspect, an embodiment of the present invention further provides a readable storage medium, including:

[0070] A memory,

[0071] The memory is used to store instructions, and when the instructions are executed by a processor, the device including the readable storage medium completes the method described in the first aspect above.

[0072] Through the technical solutions in the above one or more embodiments of the embodiments of the present invention, the embodiments of the present invention have at least the following technical effects:

[0073] In the embodiments provided by the present invention, by allowing the sensitive data recognition model to perform iterative training only on new types of sensitive data during training, and during the iterative training process, based on distillation learning, the sensitive data recognition model retains the ability to recognize known types of sensitive data until the accuracy rate of the sensitive data recognition model in recognizing new types of sensitive data reaches the expected value. The known types of sensitive data are the sensitive data that the sensitive data recognition model can recognize before the new types of sensitive data appear. In this way, the sensitive data recognition model can acquire the ability to recognize new types of sensitive data while training on new types of sensitive data, and at the same time, it can not forget the ability to recognize the known types of sensitive data that it could recognize before the new types of sensitive data appear. And because the training is only for new types of sensitive data, the data volume is less than that in the prior art when the model is trained by jointly training new types of sensitive data and known types of sensitive data on the model, thus shortening the training time. Since the sensitive data recognition model adopts the above training method, the sensitive data recognition model does not simply distinguish sensitive data from non-sensitive data, but can distinguish different types of sensitive data and non-sensitive data, enabling the sensitive data model to recognize different types of sensitive data, and further enabling the sensitive data model to complete training through a small number of samples of new types of sensitive data, thereby improving the accuracy and efficiency of recognizing sensitive data using natural language recognition methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 It is a flowchart of a sensitive data recognition method provided by an embodiment of the present invention;

[0075] Figure 2 It is a structural schematic diagram of a sensitive data recognition model provided by an embodiment of the present invention;

[0076] Figure 3 Schematic diagram for the training sensitive data recognition model provided by the embodiment of the present invention to recognize new types of sensitive data;

[0077] Figure 4 Schematic diagram of the structure of a student model provided by the embodiment of the present invention;

[0078] Figure 5 Schematic diagram of the structure of a sensitive data recognition device provided by the embodiment of the present invention. Detailed implementation manners

[0079] The embodiment of the present invention provides a method, a device and a storage medium for sensitive data recognition, so as to solve the technical problem that the accuracy of recognizing sensitive data by using the natural language recognition method in the prior art is not high.

[0080] In order to better understand the above technical solutions, the technical solutions of the present invention will be described in detail below through the accompanying drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present invention are detailed descriptions of the technical solutions of the present invention, rather than limitations on the technical solutions of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.

[0081] Please refer to Figure 1 , the embodiment of the present invention provides a method for sensitive data recognition, and the processing process of this method is as follows.

[0082] Step 101: Convert the text to be recognized into a first sequence; wherein, the elements in the first sequence represent a word or a character in the text to be recognized, and the arrangement order of the elements in the first sequence is the same as the reading order of the text to be recognized;

[0083] Step 102: Encode the position of each element in the first sequence in the first sequence to obtain a position sequence corresponding to the first sequence;

[0084] In some embodiments, converting the text to be recognized into a first sequence can be achieved by the following method:

[0085] Cut out the same words as those in the corpus from the text to be recognized; wherein, the corpus records known common words; if there is a part in the text to be recognized that has not been cut into words, then cut the part of the data to be recognized that has not been cut into words by characters; convert the cut words and / or characters into numerical vectors, and form the numerical vectors into a first sequence according to the positions of the words and / or characters in the text to be recognized.

[0086] For example, common place names, personal names, common terms, etc. can be recorded in the corpus. Taking the corpus including "Nanjing, Yangtze River, railway, power, bridge, world" and the text to be recognized as "The Nanjing Yangtze River Bridge is the first double-deck railway and highway bridge designed and built by our country on the Yangtze River, which is of great significance in the history of bridges in our country and the world history of bridges" as an example.

[0087] By comparing the above-mentioned words in the corpus with the above text to be recognized one by one, the same words in the corpus are segmented from the above text to be recognized. The segmentation result is: Nanjing, Yangtze River, The bridge is, Yangtze River, the first one on the upper layer designed and built by our country, double-deck, railway, highway, dual-purpose, bridge, in our country, bridge, history and, world, bridge, history has great significance. For the parts that are not segmented into words above (i.e., The bridge is, the first one on the upper layer designed and built by our country, dual-purpose, in our country, history and, history has great significance), they are segmented by characters. For example, "The bridge is" is segmented by characters into: big, bridge, is; similarly, the remaining parts that are not segmented into characters are segmented in this way. The parts segmented into words and characters above are combined into a character sequence according to their positions in the text to be recognized

Nanjing, Yangtze River, big, bridge, is, Yangtze River, upper, first, one, by, our, country, self, designed, and, built, of, double, layer, railway, highway, dual, purpose, bridge, in, our, country, bridge, history, and, world, bridge, history, upper, has, great, significance

[0088] When segmenting the text to be recognized, by first performing lexical segmentation according to the words in the corpus and then segmenting the parts that are not segmented into words by characters, not only can the accurate semantic information contained in the words be retained, but also segmentation errors can be avoided and the complexity of modeling can be reduced.

[0089] To encode the positions of the corresponding characters of the elements in the first sequence in the text to be recognized, the following formula can be used:

[0090]

[0091]

[0092] Among them, pos represents the position of the text (vocabulary or character) corresponding to the element in the text to be recognized, d is the dimension of the position encoding (default value is 512), 2i identifies the even dimension, 2i + 1 identifies the odd dimension, and PE pos,2i represents the position encoding of the text corresponding to element i in the even dimension, and PE pos,2i+1 represents the position encoding of the text corresponding to element i in the odd dimension.

[0093] By encoding the position of the text corresponding to each element in the first sequence in the text to be recognized, a position sequence corresponding to the first sequence can be obtained, and the word order of the text corresponding to each element in the first sequence in the text to be recognized can be retained.

[0094] In some other embodiments, converting the text to be recognized into the first sequence can also be achieved in the following manner:

[0095] If the text to be recognized only contains numbers and letters, then segment the data to be recognized by character; and convert the segmentation result into a numerical vector, and form the numerical vector into the first sequence according to the position of the character in the text to be recognized.

[0096] For example, if the text to be recognized is: 123456789ABcdG, then segment it into a character sequence [1, 2, 3, 4, 5, 6, 7, 8, 9, A, B, c, d, G] by character, and convert each element in this character sequence into a numerical vector.

[0097] By segmenting the text to be recognized that only contains numbers and letters by character, it is possible to avoid transmitting incorrect information to the model and improve the accuracy of subsequent model recognition.

[0098] By converting the text to be recognized into the first sequence composed of the smallest units of vocabulary or characters, the character length of the text corresponding to the elements in the first sequence can be reduced, which is convenient for improving the processing speed of the model.

[0099] After obtaining the first sequence and the corresponding position sequence, step 103 can be executed.

[0100] Step 103: Use the sensitive data recognition model to identify multiple consecutive elements that meet the characteristics of sensitive data from the first sequence and the position sequence, and determine the sensitive data of the text to be recognized based on the text corresponding to the multiple elements; among them, the sensitive data recognition model only performs iterative training for new types of sensitive data during training, and retains the ability to recognize known types of sensitive data based on distillation learning during the iterative training process until the accuracy of the sensitive data recognition model in recognizing new types of sensitive data reaches the expected value; the known types of sensitive data are the sensitive data that the sensitive data recognition model can recognize before the new types of sensitive data appear.

[0101] In some embodiments, the sensitive data recognition model includes:

[0102] A student model and a classifier connected in sequence;

[0103] Before the first iteration training, the parameters of the student model are inherited from the teacher model to acquire the capabilities of the teacher model; after each iteration training, the parameter update of the student model is constrained by the distillation loss between the teacher model and the student model and the classification loss of the classifier; the teacher model is a model that can correctly identify sensitive data of known types; the classifier is used to determine the sensitive data in the text to be recognized according to the output result of the student model.

[0104] The classifier includes an original part for classifying and recognizing sensitive data of known types; and a new part for classifying and recognizing sensitive data of new types.

[0105] Please refer to Figure 2 the structural schematic diagram of a sensitive data recognition model provided by an embodiment of the present invention.

[0106] As Figure 2 shown, the sensitive data recognition model is composed of a student model and a classifier connected in sequence. After the first sequence and the corresponding position sequence are input into the student model, the student model can analyze the relationship features between the characteristics of each element in the first sequence itself and other elements, and divide the words and characters segmented from the text to be recognized into multiple data that conform to the expression habits of sensitive data or non-sensitive data according to their word order in the text to be recognized. Then, the classifier analyzes the distribution of words and characters in these data to determine the data type corresponding to each data and the probability of being this data type. The data types include non-sensitive data and the types of sensitive data, so as to determine the sensitive data in the text to be recognized.

[0107] In some embodiments, before using the sensitive data recognition model to identify multiple consecutive elements that conform to the characteristics of sensitive data from the first sequence and the position sequence, it further includes:

[0108] Iteratively training the initial sensitive data recognition model with a data set containing sensitive data of new types until the accuracy rate of the trained initial sensitive data recognition model in recognizing sensitive data of new types reaches the expected value to obtain the sensitive data recognition model; wherein, the initial sensitive data recognition model includes an initial student model and an initial classifier, the parameters of the initial student model are inherited from the teacher model, and the parameter values of the new part are randomly set before the first iteration training;

[0109] During the process of initially training the sensitive data recognition model in any iteration, the parameters in the initial student model corresponding to any iteration training and the parameters in the classifier are updated by backpropagation in the sensitive data recognition model initially trained in any iteration based on the combined loss value; where the combined loss value is the sum of the distillation loss and the classification loss.

[0110] In the embodiment provided by the present invention, during the process of training the sensitive data model to recognize new types of sensitive data, the student model inherits the parameters of the teacher model before the first iteration training to acquire the capabilities of the teacher model. After each iteration training, the parameter update of the student model is constrained by the distillation loss between the teacher model and the student model and the classification loss of the classifier. In this way, when the student model uses new types of sensitive data to increase corresponding capabilities, it does not forget the capabilities acquired from the teacher model, and the classifier increases the ability to classify and recognize new types of sensitive data through the newly added part, so that the sensitive data recognition model has the ability to continue learning to recognize new types of sensitive data and does not forget the ability to recognize known types of sensitive data during the learning process.

[0111] Please refer to Figure 3 which is a schematic diagram of training the sensitive data recognition model provided by the embodiment of the present invention to recognize new types of sensitive data.

[0112] For example, assume that the existing sensitive data recognition model (i.e., the initial sensitive data recognition model) can recognize non-sensitive data and 3 known types of sensitive data. Now, a 4th type of sensitive data appears, and the sensitive data recognition model needs to be trained to be able to recognize both the previous 3 known types of sensitive data and the newly emerged 4th type of sensitive data. The traditional training scheme is to form a training set with the previous 3 known types of sensitive data, the newly emerged 4th type of sensitive data, and non-sensitive data to retrain the model, but this training takes a long time, and the sample size of sensitive data is usually small, and the recognition accuracy is not high.

[0113] In this embodiment, only the newly emerged fourth type of sensitive data is used to continue training the initial sensitive data recognition model. However, if only the newly emerged fourth type of sensitive data is relied on to train the sensitive data recognition model, the initial sensitive data recognition model will be optimized in the direction of only being able to recognize the newly emerged fourth type of sensitive data, resulting in the initial sensitive data recognition model forgetting the three known types of sensitive data that it could previously recognize. To prevent this situation from occurring, in this embodiment, when training the initial sensitive data recognition model to recognize the fourth type of sensitive data, a teacher model that can process the three known types of sensitive data and non-sensitive data is introduced. By having the student model inherit the parameters of the teacher model, it enables the student model to acquire the capabilities of the teacher model. During training, the teacher model and the initial student model simultaneously receive the sequences corresponding to the same samples of the fourth new type of sensitive data. After each sample is used to complete one recognition by the initial sensitive data recognition model, it is necessary to calculate the distillation loss between the teacher model and the initial student model, as well as the classification loss of the initial classifier. The sum of the distillation loss and the classification loss is used to perform backpropagation on the initial sensitive data recognition model, thereby updating the parameters of the initial student model and the initial classifier, while the parameters in the teacher model remain unchanged. In this way, not only does the initial student model acquire the response ability for the new type of sensitive data, but also the teacher model enables the initial student model not to forget the capabilities corresponding to the three known types of sensitive data. Moreover, the newly added part in the initial classifier can also gradually increase the classification ability for the fourth type of sensitive data through backpropagation. By repeating the above process with different samples of the fourth type of sensitive data, the accuracy of the initial sensitive data recognition model in recognizing the fourth type of sensitive data reaches the expected value, and a sensitive data recognition model that can recognize both the three known types of sensitive data and the newly emerged fourth type of sensitive data can be obtained.

[0114] If after a period of time, a fifth type of sensitive data emerges, then the student model in the sensitive data recognition model that can recognize four types of sensitive data is used to update the teacher model, and a training method similar to the previous training for recognizing the fourth type of sensitive data is adopted to obtain a sensitive data recognition model that can recognize five types of sensitive data.

[0115] Among them, the calculation formula of the distillation loss is as follows:

[0116]

[0117] Among them, t represents the teacher model, s represents the student model, and W×H represents the dimension of the feature space that the teacher model and the student model can represent. represents the output result of the teacher model in the w×h feature space. represents the output result of the student model in the w×h feature space, and l kd represents the distillation loss.

[0118] The classification loss of the classifier is denoted as l crf , and the sum of the distillation loss and the classification loss can be the combined loss (denoted as l total ):

[0119] l total = α × l kd + l crf (4);

[0120] where α is a hyperparameter.

[0121] In the embodiments provided by the present invention, by allowing the sensitive data recognition model to perform iterative training only on new types of sensitive data during training, and during the iterative training, enabling the sensitive data recognition model to retain the ability to recognize known types of sensitive data based on distillation learning until the accuracy rate of the sensitive data recognition model in recognizing new types of sensitive data reaches the expected value. The known types of sensitive data are the sensitive data that the sensitive data recognition model could recognize before the new types of sensitive data appeared. In this way, when the sensitive data recognition model is trained on new types of sensitive data, it can not only acquire the ability to recognize new types of sensitive data, but also not forget the ability to recognize the known types of sensitive data that it could recognize before the new types of sensitive data appeared. And because the training is only carried out on new types of sensitive data, the amount of data is less than that in the prior art where the new types of sensitive data and the known types of sensitive data need to be used together to retrain the model, thereby being able to shorten the training time. Since the sensitive data recognition model adopts the above training method, the sensitive data recognition model can not only simply distinguish sensitive data from non-sensitive data, but also distinguish different types of sensitive data and non-sensitive data, enabling the sensitive data model to recognize different types of sensitive data. Furthermore, the sensitive data model can complete training through a small number of samples of new types of sensitive data, thereby improving the accuracy rate and efficiency of recognizing sensitive data using the natural language recognition method.

[0122] In some embodiments, before iteratively training the initial sensitive data recognition model with a dataset containing new types of sensitive data, it further includes:

[0123] Construct a pre-trained language model; perform masked language pre-training on the pre-trained language model with the second sequences corresponding to the unlabeled sample data in the unlabeled sample dataset until the number of times the pre-trained language model is trained with each second sequence reaches a preset value, obtaining a teacher model that can express the unlabeled sample data in natural language; where the unlabeled sample dataset includes non-sensitive data and sensitive data.

[0124] For example, the unlabeled sample data is segmented by characters to obtain a second sequence; for the character sequences corresponding to a preset proportion (such as 15%) of the unlabeled sample data in the unlabeled sample dataset of the pre-trained language model, a strategy of randomly masking, replacing, or leaving unchanged is adopted for the selected character sequences, and the masked elements are predicted to achieve masked language training. When the number of times the pre-trained language model is trained with each second sequence reaches a preset value, a teacher model that can express the unlabeled sample data in natural language is obtained, that is, the teacher model learns to express the unlabeled sample data in natural language.

[0125] By performing masked language pre-training on the pre-trained language model with the second sequences corresponding to the unlabeled sample data in the unlabeled sample dataset until the number of times the pre-trained language model is trained with each second sequence reaches a preset value, the teacher model can learn to understand the grammar structure, lexical relationships, and context information of the unlabeled sample data, and then the teacher model can be capable of expressing the unlabeled sample data in natural language.

[0126] In some embodiments, before iteratively training the initial sensitive data recognition model with a dataset containing new types of sensitive data, it further includes:

[0127] The collected sensitive data samples of multiple types are grouped in pairs of two sensitive data samples and randomly formed into a sample set composed of sensitive data sample pairs;

[0128] The sample set is used to train the initial classifier to identify whether the types of two sensitive data samples are the same until the recognition accuracy reaches a preset value to obtain the classifier; wherein, the initial classifier determines whether the types of two sensitive data samples are the same based on the distribution of words or characters in the two sensitive data samples by learning with a conditional random field.

[0129] For example, three types of sensitive data samples have been collected so far, denoted as type 1 to type 3. Each of these three types of sensitive data samples has two sensitive data samples. The sensitive data sample 1 of type 1 and the sensitive data sample 1 of type 2 are combined into a pair of sensitive data samples. The sensitive data sample 1 of type 1 and the sensitive data sample 1 of type 3 are combined into a pair of sensitive data samples. The sensitive data sample 1 of type 2 and the sensitive data sample 1 of type 3 are combined into a pair of sensitive data samples. The sensitive data sample 2 of type 1 and the sensitive data sample 1 of type 3 are combined into a pair of sensitive data samples... and so on to form pairs of sensitive data samples composed of different types of sensitive data samples. At the same time, two sensitive data samples of the same type are also combined into a pair of sensitive data samples. These pairs of sensitive data samples are combined into a sample set, and the initial classifier is continuously trained with the sample set to identify whether the types of the two sensitive data samples in the pair of sensitive data samples are the same, so that the initial classifier can distinguish the differences between different types of sensitive data samples. Based on the conditional random field, the distribution of words or characters in different types of sensitive data samples is learned. When the recognition accuracy of the initial classifier reaches the preset value, the training is completed, and a classifier is obtained. When using the sensitive data recognition model containing this classifier to identify sensitive data in the text to be recognized, after the classifier obtains the data determined according to the natural language expression habits of characters and words that conform to the unlabeled sample data from the student model, according to the learned distribution of words and characters of various types of sensitive data, it can determine what type of data the received data is.

[0130] Please refer to Figure 4 which is a schematic structural diagram of a student model provided by an embodiment of the present invention. The student model includes:

[0131] An encoding layer and a feature extraction layer connected in sequence; the encoding layer is used to calculate the features of each first element to be encoded in the first sequence itself and the relationship features associated with the first element and the second element. The distance between the positions of the respective corresponding characters of the second element and the first element in the text to be recognized is less than or equal to a preset distance. The relationship features include word order, grammar, and lexical relationships.

[0132] The feature extraction layer includes a bidirectional long short-term memory network, a residual network, and a feed-forward neural network connected in sequence; among them, the input end and the output end of the bidirectional long short-term memory network are respectively connected to different input ends of the residual network, the output end of the residual network is connected to the input end of the feed-forward neural network, and the output end of the feed-forward neural network serves as the output end of the pre-trained language model; the bidirectional long short-term memory network is used to obtain the context features of each element from multiple features of each element in the received first sequence, the residual network is used to fuse the input and output of the bidirectional long short-term memory network and then transmit them to the feed-forward neural network, and the feed-forward neural network determines multiple elements with continuous positions from the first sequence according to the output of the residual network, and the data composed of the multiple elements conforms to the natural language expression habits of unlabeled data.

[0133] The encoding layer can be composed of multiple layers of Transformer encoders, such as 12 layers of Transformer encoders; since the first sequence contains semantic words, by using multiple layers of Transformer encoders to form the encoding layer, deeper features of the words corresponding to the elements in the first sequence can be extracted, such as the features of the words corresponding to the elements themselves and the relationship features associated between the elements, so as to realize the processing of long-distance dependence relationships in the first sequence, which is particularly effective for the recognition of sensitive data composed of letters and numbers.

[0134] The bidirectional long short-term memory network can obtain the context features of each element from the features output by the encoding layer, and use the residual network to add (i.e., fuse) the input and output of the bidirectional long short-term memory network and then provide them to the feed-forward neural network, so as to be able to alleviate the problem of gradient disappearance during the training process, help to spread the gradient more effectively, and improve the training efficiency of the model and the training stability of the deep network. The fusion of the input and output of the bidirectional long short-term memory network by the residual network can be realized by the following formula:

[0135] R = Bi-LSTM(H 12 ) + H 12 (5);

[0136] where, R is the output of the residual network, H12 is the output of the encoding layer (i.e., the input of the bidirectional long short-term memory network), and Bi-LSTM(H 12 ) is the output of the bidirectional long short-term memory network.

[0137] The feed-forward neural network consists of two fully connected layers, connected by an activation function in the middle, and the activation function used is the rectified linear unit. The calculation formula of the feed-forward neural network is as follows:

[0138] FFN(R) = ReLU(R·W1 + b1)W2 + b2 (6);

[0139] Among them, R is the input of the feedforward neural network, ReLU() is the activation function, W1 and W2 are weight parameters, and b1 and b2 are bias parameters. During the training phase, through the learning of the above weight parameters and bias parameters, the feedforward neural network can map in a non-linear space, mapping multiple consecutive elements in position into data that conforms to the natural language expression habits of unlabeled data. This facilitates subsequent classification and recognition of the data output by the feedforward neural network by a classifier to determine the corresponding type of the data.

[0140] According to the previous training process, it can be known that the pre-trained language model, the teacher model, and the student model have the same structure. Therefore, the structures of the pre-trained language model and the teacher model will not be elaborated here.

[0141] Based on the same inventive concept, an embodiment of the present invention provides a sensitive data recognition device. For the specific implementation manner of the sensitive data recognition method of this device, reference can be made to the description in the method embodiment part. Repeated parts will not be elaborated here. Please refer to Figure 5 , the device includes:

[0142] A conversion unit 501, configured to convert the text to be recognized into a first sequence; wherein, the elements in the first sequence represent a word or a character in the text to be recognized, and the arrangement order of the elements in the first sequence is the same as the reading order of the text to be recognized;

[0143] An encoding unit 502, configured to encode the position of each element in the first sequence in the first sequence to obtain a position sequence corresponding to the first sequence;

[0144] A recognition unit 503, configured to use a sensitive data recognition model to recognize multiple consecutive elements that conform to the characteristics of sensitive data and are consecutive in position from the first sequence and the position sequence, and determine the sensitive data of the text to be recognized based on the words corresponding to the multiple elements; wherein, during training, the sensitive data recognition model only performs iterative training on new types of sensitive data, and during the iterative training process, the sensitive data recognition model retains the ability to recognize known types of sensitive data based on distillation learning until the accuracy rate of the sensitive data recognition model in recognizing the new type of sensitive data reaches an expected value; the known type of sensitive data is the sensitive data that the sensitive data recognition model can recognize before the new type of sensitive data appears.

[0145] A possible implementation manner, the sensitive data recognition model includes:

[0146] A student model and a classifier connected in sequence;

[0147] Before the first iterative training, the parameters of the student model are inherited from the teacher model to acquire the capabilities of the teacher model; after each iterative training, the update of the parameters of the student model is constrained by the distillation loss between the teacher model and the student model and the classification loss of the classifier; the teacher model is a model that can correctly identify sensitive data of known types;

[0148] The classifier is used to determine the sensitive data in the text to be recognized according to the output result of the student model.

[0149] A possible implementation, the classifier includes:

[0150] The original part, which is used to classify and identify the sensitive data of the known type;

[0151] The new part, which is used to classify and identify the sensitive data of the new type.

[0152] A possible implementation, the device further includes a training unit 504 for:

[0153] Iteratively training the initial sensitive data recognition model with a data set containing the sensitive data of the new type until the accuracy of the trained initial sensitive data recognition model in recognizing the sensitive data of the new type reaches the expected value, obtaining the sensitive data recognition model; wherein, the initial sensitive data recognition model includes an initial student model and an initial classifier, the initial student model inherits the parameters of the teacher model, and the parameter values of the new part are randomly set before the first iterative training;

[0154] During the process of any iterative training of the initial sensitive data recognition model, the parameters in the initial student model corresponding to any iterative training and the parameters in the classifier are updated by backpropagation in the initial sensitive data recognition model based on the joint loss value; wherein, the joint loss value is the sum of the distillation loss and the classification loss.

[0155] A possible implementation, the training unit 504 is further used for:

[0156] Construct a pre-trained language model;

[0157] Perform masked language pre-training on the pre-trained language model with the second sequences corresponding to the unlabeled sample data in the unlabeled sample data set until the number of times the pre-trained language model is trained with each second sequence reaches a preset value, obtaining the teacher model that can express the unlabeled sample data in natural language; wherein, the unlabeled sample data set includes non-sensitive data and sensitive data.

[0158] A possible implementation, the training unit 504 is further configured to:

[0159] Group the collected multiple types of sensitive data samples in pairs of two sensitive data samples, and randomly form a sample set composed of pairs of sensitive data samples;

[0160] Use the sample set to train an initial classifier to identify whether the types of the two sensitive data samples are the same until the recognition accuracy reaches a preset value, and obtain the classifier; wherein, the initial classifier determines whether the types of the two sensitive data samples are the same based on learning the distribution of words or characters in the two sensitive data samples by means of a conditional random field.

[0161] A possible implementation, the student model includes:

[0162] An encoding layer and a feature extraction layer connected in sequence; the encoding layer is used to calculate the features of each first element to be encoded in the first sequence itself and the relationship features between the first element and the second element, where the distance between the positions of the respective corresponding characters of the second element and the first element in the text to be recognized is less than or equal to a preset distance, and the relationship features include word order, grammar, and vocabulary relationship;

[0163] The feature extraction layer includes a bidirectional long short-term memory network, a residual network, and a feed-forward neural network connected in sequence; wherein, the input end and the output end of the bidirectional long short-term memory network are respectively connected to different input ends of the residual network, the output end of the residual network is connected to the input end of the feed-forward neural network, and the output end of the feed-forward neural network is used as the output end of the pre-trained language model; the bidirectional long short-term memory network is used to obtain the context features of each element from the multiple features of each element in the received first sequence, the residual network is used to fuse the input and output of the bidirectional long short-term memory network and then transmit them to the feed-forward neural network, and the feed-forward neural network determines, according to the output of the residual network, multiple consecutive elements in position from the first sequence, and the data composed of the multiple elements conforms to the natural language expression habit of the unlabeled data.

[0164] A possible implementation, the conversion unit 501 is further configured to:

[0165] Segment out the same vocabulary as in the corpus from the text to be recognized; wherein, the corpus records known common vocabulary;

[0166] If there is a part of the text to be recognized that has not been segmented into the vocabulary, segment the part of the data to be recognized that has not been segmented into the vocabulary by characters;

[0167] Convert the segmented words and / or characters into numerical vectors, and form the numerical vectors into the first sequence according to the positions of the words and / or characters in the text to be recognized.

[0168] A possible implementation manner, the conversion unit 501 is further configured to:

[0169] If the text to be recognized only contains numbers and letters, segment the data to be recognized by characters; convert the segmentation result into a numerical vector, and form the numerical vector into the first sequence according to the positions of the characters in the text to be recognized.

[0170] It should be noted that the division of units in the embodiments of the present application is illustrative, only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0171] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0172] It should be noted here that the above device provided in the embodiments of the present invention can implement all the method steps implemented in the above method embodiments, and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments will not be specifically described in this embodiment.

[0173] Based on the same inventive concept, an apparatus for sensitive data recognition is provided in the embodiments of the present invention, including: at least one processor, and

[0174] a memory connected to the at least one processor;

[0175] Among them, the memory stores instructions executable by the at least one processor, and the at least one processor executes the method for identifying sensitive data as described above by executing the instructions stored in the memory.

[0176] Based on the same inventive concept, an embodiment of the present invention further provides a readable storage medium, including:

[0177] A memory,

[0178] The memory is used to store instructions, and when the instructions are executed by a processor, the device including the readable storage medium completes the method for identifying sensitive data as described above.

[0179] The readable storage medium can be any available medium or data storage device accessible by the processor, including volatile memory or non-volatile memory, or can include both volatile memory and non-volatile memory. By way of example and not limitation, non-volatile memory can include Read-Only Memory (ROM), Programmable ROM (PROM), Electrically Programmable ROM (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), or flash memory, Solid State Disk (SSD), magnetic memory (such as floppy disks, hard disks, magnetic tapes, Magneto-Optical discs (MO), etc.), optical memory (such as CDs, DVDs, BDs, HVDs, etc.). Volatile memory can include Random Access Memory (RAM), which can serve as an external cache memory. By way of example and not limitation, RAM can be obtained in various forms, such as Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random-Access Memory (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Sync Link DRAM (SLDRAM). The storage devices of the disclosed aspects are intended to include, but are not limited to, these and other suitable types of memory.

[0180] Those skilled in the art will understand that the embodiments of the present invention can be provided as a method, system, or program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a machine program product implemented on one or more readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer / processor-usable program code.

[0181] Embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0182] These program instructions can also be stored in a readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the readable memory generate a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0183] These program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer / processor-implemented process, so that the instructions executed on the computer / processor or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0184] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for identifying sensitive data, characterized in that: include: Converting the text to be recognized into a first sequence; wherein the elements in the first sequence represent a word or character in the text to be recognized, and the arrangement order of the elements in the first sequence is consistent with the reading order of the text to be recognized; Encoding the position of each element in the first sequence in the first sequence to obtain a position sequence corresponding to the first sequence; A sensitive data recognition model is used to identify multiple elements that meet the characteristics of sensitive data and are continuous in position from the first sequence and the position sequence, and the sensitive data of the text to be recognized is determined based on the text corresponding to the multiple elements; wherein, during training, the sensitive data recognition model is iteratively trained only for new types of sensitive data, and during the iterative training process, the sensitive data recognition model retains the ability to recognize known types of sensitive data based on distillation learning until the accuracy of the sensitive data recognition model in recognizing the new type of sensitive data reaches an expected value; the known type of sensitive data is the sensitive data that the sensitive data recognition model can recognize before the new type of sensitive data appears.

2. The method according to claim 1, characterized in that The sensitive data identification model includes: The student model and classifier are connected in sequence; Before the first iteration of training, the parameters of the student model are inherited from the teacher model to learn the capabilities of the teacher model; after each iteration of training, the parameter update of the student model is constrained by the distillation loss between the teacher model and the student model and the classification loss of the classifier; the teacher model is a model that can correctly identify the known type of sensitive data; The classifier is used to determine the sensitive data in the text to be identified based on the output result of the student model.

3. The method according to claim 2, characterized in that The classifier comprises: an original part, wherein the original part is used to classify and identify the sensitive data of the known type; A newly added part is used to classify and identify the new type of sensitive data.

4. The method according to claim 3, characterized in that Before using the sensitive data identification model to identify multiple elements that meet the characteristics of sensitive data and are continuous in position from the first sequence and the position sequence, the method further includes: Iteratively training the initial sensitive data recognition model using a data set containing the new type of sensitive data until the accuracy of the trained initial sensitive data recognition model in recognizing the new type of sensitive data reaches the expected value, thereby obtaining the sensitive data recognition model; wherein the initial sensitive data recognition model includes an initial student model and an initial classifier, the initial student model inherits the parameters of the teacher model, and the parameter values ​​of the newly added part are randomly set before the first iterative training; During any iterative training of the initial sensitive data recognition model, the parameters in the initial student model and the parameters in the classifier corresponding to the any iterative training are back-propagated and updated in the initial sensitive data recognition model of the any iterative training based on the joint loss value; wherein the joint loss value is the sum of the distillation loss and the classification loss.

5. The method according to claim 4, characterized in that Before iteratively training the initial sensitive data recognition model with a data set containing the new type of sensitive data, the method further includes: Build a pre-trained language model; The pre-trained language model is subjected to masked language pre-training using a second sequence corresponding to the unlabeled sample data in the unlabeled sample data set, until the number of times the pre-trained language model is trained using each of the second sequence reaches a preset value, thereby obtaining the teacher model capable of expressing the unlabeled sample data in natural language; wherein the unlabeled sample data set includes non-sensitive data and sensitive data.

6. The method according to claim 4, characterized in that Before iteratively training the initial sensitive data recognition model with a data set containing the new type of sensitive data, the method further includes: The collected sensitive data samples of various types are grouped into two sensitive data samples to randomly form a sample set consisting of sensitive data sample pairs; The sample set is used to train an initial classifier to identify whether the types of the two sensitive data samples are the same, until the recognition accuracy reaches a preset value, thereby obtaining the classifier; wherein the initial classifier is based on conditional random fields to learn the distribution of words or characters in the two sensitive data samples to determine whether the types of the two sensitive data samples are the same.

7. The method according to claim 5, characterized in that The student model includes: A coding layer and a feature extraction layer connected in sequence; the coding layer is used to calculate the features of each first element to be encoded in the first sequence and the relationship features associated with the first element and the second element, the distance between the positions of the corresponding characters of the second element and the first element in the text to be recognized is less than or equal to a preset distance, and the relationship features include word order, grammar, and vocabulary relations; The feature extraction layer includes a bidirectional long short-term memory network, a residual network and a feedforward neural network connected in sequence; wherein, the input end and the output end of the bidirectional long short-term memory network are respectively connected to different input ends of the residual network, the output end of the residual network is connected to the input end of the feedforward neural network, and the output end of the feedforward neural network serves as the output end of the pre-trained language model; the bidirectional long short-term memory network is used to obtain the context feature of each element from multiple features of each element in the received first sequence, the residual network is used to fuse the input and output of the bidirectional long short-term memory network and transmit them to the feedforward neural network, and the feedforward neural network determines multiple elements that are continuous in position from the first sequence according to the output of the residual network, and the data composed of the multiple elements conforms to the natural language expression habits of the unlabeled data.

8. The method according to any one of claims 1 to 6, characterized in that: Convert the text to be recognized into the first sequence, including: Segmenting the same words as those in the corpus from the text to be recognized; wherein the corpus records known common words; If the text to be recognized has a part that is not segmented into the vocabulary, segmenting the part of the text to be recognized that is not segmented into the vocabulary according to characters; The segmented words and / or characters are converted into numerical vectors, and the numerical vectors are organized into the first sequence according to the positions of the words and / or characters in the text to be recognized.

9. The method according to any one of claims 1 to 6, characterized in that: Convert the text to be recognized into the first sequence, including: If the text to be recognized only contains numbers and letters, the data to be recognized is segmented by characters; the segmentation result is converted into a numerical vector, and the numerical vector is organized into the first sequence according to the position of the character in the text to be recognized.

10. A device for identifying sensitive data, characterized in that: include: A conversion unit, configured to convert the text to be recognized into a first sequence; wherein the elements in the first sequence represent a word or character in the text to be recognized, and the arrangement order of the elements in the first sequence is consistent with the reading order of the text to be recognized; an encoding unit, configured to encode the position of each element in the first sequence in the first sequence to obtain a position sequence corresponding to the first sequence; An identification unit is used to use a sensitive data identification model to identify multiple elements that meet the characteristics of sensitive data and are continuous in position from the first sequence and the position sequence, and determine the sensitive data of the text to be identified based on the text corresponding to the multiple elements; wherein the sensitive data identification model is iteratively trained only for new types of sensitive data during training, and during the iterative training process, the sensitive data identification model retains the ability to identify known types of sensitive data based on distillation learning until the accuracy of the sensitive data identification model in identifying the new type of sensitive data reaches an expected value; the known type of sensitive data is sensitive data that the sensitive data identification model can identify before the new type of sensitive data appears.

11. A device for identifying sensitive data, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the at least one processor executes the method according to any one of claims 1 to 9 by executing the instructions stored in the memory.

12. A readable storage medium, characterized in that: Including memory, The memory is used to store instructions. When the instructions are executed by the processor, the device including the readable storage medium implements the method according to any one of claims 1 to 9.