Data desensitization processing method and system based on large model
Through a large-model-based data desensitization processing method, combined with multiple target sensitive data perception projects, the model is iteratively optimized, and the problems of inaccurate identification of sensitive data and low model calibration efficiency in the prior art are solved, and efficient and accurate sensitive data perception and data desensitization processing are achieved.
Patent Information
- Application Number
- CN202510194529.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-21
AI Technical Summary
Existing data desensitization methods are difficult to accurately identify complex and diverse sensitive data, and are prone to misjudgment or misjudgment, and the model adjustment efficiency is low, which increases training cost and time.
The data desensitization processing method based on large models is adopted, and the model is iteratively optimized by obtaining sample learning data sets and target-sensitive data perception projects by combining multiple target-sensitive data perception projects to improve the model's perception and tuning efficiency.
It realizes more efficient and accurate sensing of sensitive data, reduces model training costs and time, and improves the quality and efficiency of data desensitization processing.
Smart Images

Figure CN119670157B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and more specifically, to a data desensitization processing method and system based on a large model. Background Art
[0002] In today's digital age, data security and privacy protection are becoming increasingly important, especially the processing of sensitive data has become the focus of attention in various industries. In traditional data desensitization processing methods, the perception of sensitive data mainly relies on simple rule matching or single model judgment, lacking targeted guidance on target sensitive data types. These traditional methods are difficult to accurately identify complex and diverse sensitive data, and are prone to missed or misjudgment, resulting in incomplete or over-desensitization of data, affecting the availability and security of data. In addition, traditional methods are inefficient in model tuning, and require repeated training to learn the characteristics of different sensitive data, which increases the training cost and time of the model, and is difficult to meet the requirements of data processing speed and efficiency in practical applications. Therefore, there is an urgent need for a more efficient, accurate and sensitive data perception method that can adapt to a variety of sensitive data types to improve the quality and efficiency of data desensitization processing and better protect data security and privacy. Summary of the invention
[0003] The purpose of the present invention is to provide a data desensitization processing method and system based on a large model. The embodiment of the present application is implemented as follows:
[0004] In a first aspect, an embodiment of the present application provides a data desensitization processing method based on a large model, including: obtaining an example learning data set and a target sensitive data perception project of the example learning data set; the example learning data set is configured with at least one training supervision information, and any one of the training supervision information represents the sensitive data annotation result of the example learning data set under the corresponding sensitive data perception project; performing a sensitive data perception operation on the example learning data set in combination with the target sensitive data perception project to obtain a sensitive data perception result of the example learning data set under the target sensitive data perception project; determining the target training supervision information from at least one training supervision information configured with the example learning data set, and determining the sensitive data annotation result of the example learning data set under the target sensitive data perception project from the target training supervision information, wherein the target training supervision information is unified training supervision information for the corresponding sensitive data perception project and the target sensitive data perception project; iteratively optimizing the sensitive data perception model in combination with the sensitive data perception result of the example learning data set under the target sensitive data perception project and the sensitive data annotation result under the target sensitive data perception project; the optimized converged sensitive data perception model is used to perform sensitive data perception operations on data under the corresponding sensitive data perception project, so as to desensitize the perception results.
[0005] In a second aspect, the present application provides a computer system comprising: one or more processors; a memory; one or more computer programs; wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method as described above is implemented.
[0006] The present invention obtains an example learning data set containing at least one training supervision information and a target sensitive data perception project of the example learning data set, and any one training supervision information represents the sensitive data annotation result under the corresponding sensitive data perception project. In combination with the target sensitive data perception project, the example learning data set can be subjected to a sensitive data perception operation to obtain the sensitive data perception result of the example learning data set under the target sensitive data perception project. During the sensitive data perception operation, the target sensitive data perception project can guide the perception of sensitive information in the example learning data set, and the sensitive data perception operation of the example learning data set can be carried out in the right direction, promoting the perception of sensitive data in the example learning data set. Then, the target training supervision information is determined from at least one training supervision information contained in the example learning data set, and the sensitive data perception project corresponding to the target training supervision information is unified with the target sensitive data perception project. Then, the determination of the target training supervision information can also be based on the target sensitive data perception project, and the sensitive data perception model can be adjusted based on the sensitive data perception result of the example learning data set under the target sensitive data perception project and the sensitive data annotation result under the target sensitive data perception project, so as to enable the sensitive data perception model to have the performance of perceiving sensitive information under the target sensitive data perception project. The optimized and converged sensitive data perception model can perform sensitive data perception operations on data under the corresponding sensitive data perception project. Compared with the model without the target sensitive data perception project, the optimized and converged sensitive data perception model has a stronger perception of sensitive information, which improves the perception of the model. Furthermore, combined with the example learning data set containing no less than one training supervision information, it can provide support for the sensitive data perception model to be adjusted in combination with multiple target sensitive data perception projects, helping the sensitive data perception model to learn the knowledge of perceiving multiple sensitive information in a single adjustment process, and the adjustment cost and speed of the model are significantly optimized. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in describing the embodiments of the present application are briefly introduced below.
[0008] Figure 1 It is a flow chart of a data desensitization processing method based on a large model provided in an embodiment of the present application.
[0009] Figure 2 It is a schematic diagram of the composition of a computer system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0010] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation method part of the embodiments of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0011] The execution subject of the data desensitization processing method based on the big model in the embodiment of the present application is a computer system, including but not limited to a server, a personal computer, a laptop computer, a tablet computer, a smart phone, etc. The embodiment of the present application provides a data desensitization processing method based on the big model, which is applied to a server, such as Figure 1 As shown, the method includes:
[0012] Step 10: Obtain an example learning dataset and a target sensitive data perception project of the example learning dataset; the example learning dataset is configured with no less than one training supervision information, and any one of the training supervision information represents the sensitive data annotation result of the example learning dataset under the corresponding sensitive data perception project.
[0013] The example learning dataset is a dataset sample used for model training, which contains the data information required for training the sensitive data perception model. The target sensitive data perception project specifies which sensitive data types in the example learning dataset need to be perceived, for example, sensitive data types such as ID card number, mobile phone number, bank card number, etc. The training supervision information is the annotation of sensitive data in the example learning dataset, which provides supervision and guidance for model training, so that the model can learn how to accurately identify sensitive data.
[0014] In actual operation, a computer system can obtain an example learning data set in many ways. One feasible way is to extract relevant data from an existing database. For example, for a database of an e-commerce platform, a computer system can extract user order information, personal information and other data from it as an example learning data set. This data may contain various sensitive information, such as the user's name, address, contact information, etc.
[0015] The determination of target sensitive data perception projects is usually based on specific business needs and security requirements. For example, in the financial field, it may be necessary to focus on sensitive information such as bank card numbers and credit card numbers; in the medical field, it may be necessary to focus on sensitive information such as patients' medical records and ID numbers. The computer system can determine the target sensitive data perception project through configuration files or user input.
[0016] It should be noted that when obtaining relevant sensitive data, it is done within the scope permitted by laws and regulations and with the authorization of the stakeholders.
[0017] The generation of training supervision information can be achieved through manual labeling or automatic labeling. Manual labeling refers to the labeling of sensitive data in the example learning data set by professionals. This method has high accuracy but low efficiency. Automatic labeling is to label sensitive data through preset rules or machine learning algorithms. This method is more efficient, but the accuracy of labeling may be affected to a certain extent. For example, for text data containing ID numbers, the computer system can automatically identify the ID numbers through regular expressions and label them as sensitive data.
[0018] The following is a specific example to illustrate the execution process of step 10. Assume that a computer system trains a sensitive data perception model to identify sensitive information in medical data. First, the computer system extracts patient medical records, examination reports and other data from the hospital database as an example learning data set. This data may include the patient's name, age, ID number, medical record number and other information.
[0019] Then, based on business needs and security requirements, the computer system determines that the target sensitive data perception items are ID number, medical record number, and patient name. Next, the computer system manually labels the ID number, medical record number, and patient name in the example learning data set to generate training supervision information. For example, for a medical record information "Patient name: Zhang San, ID number: 123456789012345678, Medical record number: 0001", the computer system labels "Zhang San" as the patient's name, "123456789012345678" as the ID number, and "0001" as the medical record number.
[0020] In the process of acquiring sample learning data sets and target sensitive data perception projects, the computer system considers the quality and security of the data. In terms of data quality, the computer system cleans and preprocesses the acquired data to remove noise data and invalid data to improve the training effect of the model. For example, for data containing duplicate records or incorrect formats, the computer system performs operations such as deduplication and format conversion.
[0021] Computer systems can also use data augmentation methods to expand example learning data sets to improve the generalization ability of the model. Data augmentation refers to generating new data samples by transforming and augmenting the original data. For example, for text data, new text samples can be generated by replacing synonyms, changing sentence structures, etc.
[0022] In practical applications, step 10 may be affected by many factors. For example, the diversity and complexity of the data may make it more difficult to label the training supervision information. Data in different fields may have different characteristics and formats, requiring different labeling methods and rules. Data updates and changes may also affect the training effect of the model. Over time, the distribution and characteristics of the data may change, and the example learning data set and training supervision information need to be updated in a timely manner to ensure the accuracy and effectiveness of the model.
[0023] Step 20: Perform sensitive data perception operations on the example learning dataset in combination with the target sensitive data perception project to obtain the sensitive data perception results of the example learning dataset under the target sensitive data perception project.
[0024] When a computer system performs sensitive data perception operations, the target sensitive data perception project provides it with clear directions and standards, and the example learning dataset is the object of the operation. The target sensitive data perception project specifies the types of sensitive data that need to be identified. For example, in a financial data processing scenario, the target sensitive data perception project may include bank card numbers, credit card expiration dates, CVV codes, etc.; while the example learning dataset may contain a large amount of customer transaction records, account information, and other data. The task of the computer system is to find sensitive data that meets the requirements of these projects from the example learning dataset based on the target sensitive data perception project.
[0025] For example, for an example learning data set containing text information, the computer system can use lexical analysis, syntactic analysis and other technologies to parse the text and identify sensitive information that may be contained therein. Taking the ID card number as an example, the computer system can find a string that conforms to the ID card number format from the text by matching regular expressions. Alternatively, using a machine learning algorithm, the computer system can use a pre-trained classification model to classify the data in the example learning data set and determine whether it is sensitive data under the target sensitive data perception project.
[0026] In actual operation, the computer system selects corresponding technical means to perform sensitive data perception operations according to the different characteristics and requirements of the target sensitive data perception items. For example, if the target sensitive data perception items have clear formats and rules, such as ID card numbers, telephone numbers, etc., regular expression matching methods can be used to perform sensitive data perception operations; and for some sensitive data with more complex semantics, such as disease diagnosis information in medical records, deep learning models can be used for semantic understanding and classification to complete sensitive data perception operations.
[0027] Step 30: Determine target training supervision information from at least one piece of training supervision information configured in the self-example learning data set, and determine the sensitive data annotation result of the example learning data set under the target sensitive data perception project from the target training supervision information, wherein the target training supervision information is unified training supervision information for the corresponding sensitive data perception project and the target sensitive data perception project.
[0028] In step 30, the computer system selects the target training supervision information from the multiple training supervision information configured in the example learning data set. The training supervision information configured in the example learning data set is like a series of reference guides, each of which corresponds to a specific sensitive data perception project and records the labeling of sensitive data in the example learning data set under the project. The target training supervision information is the reference guide that fully matches the target sensitive data perception project currently being focused on. For example, in an example learning data set containing multiple sensitive data types (such as ID number, telephone number, email address, etc.), the target sensitive data perception project is the ID number, then the computer system finds the training supervision information specifically for the sensitive data perception project of the ID number from all the training supervision information as the target training supervision information.
[0029] In order to determine the target training supervision information, the computer system can use a variety of technical means. One method is through data matching. The computer system compares the identification information of the target sensitive data perception project with the identification information of the sensitive data perception project corresponding to each training supervision information to find out the completely consistent training supervision information. For example, if the target sensitive data perception project uses "ID number" as the identification, the computer system will traverse all the training supervision information to see which training supervision information also has the identification of "ID number" and determine it as the target training supervision information. It is also possible to use the index search method to establish an index for each training supervision information, and quickly locate the corresponding target training supervision information according to the index information of the target sensitive data perception project. This method can significantly improve the search efficiency when processing a large amount of training supervision information. After determining the target training supervision information, the computer system extracts the sensitive data annotation results of the example learning data set under the target sensitive data perception project. These annotation results are the annotations of sensitive data in the example learning data set, and are an important reference for model training. For example, in the target training supervision information, it may be marked which data items in the example learning data set are ID numbers, and the specific locations of these ID numbers in the data set. In actual operation, the computer system ensures the accuracy and completeness of the extracted sensitive data annotation results. For complex example learning data sets, there may be multiple data items that are related to the target sensitive data perception project. The computer system carefully checks the annotation of each data item to avoid omissions or incorrect annotations. At the same time, the computer system also needs to deal with possible data inconsistency problems. For example, different training supervision information may have different annotations for the same data item. At this time, the computer system processes it according to certain rules, such as selecting a more credible annotation result. The sensitive data annotation results under the target sensitive data perception project indicated by the target training supervision information include the annotation results of each data item in the example learning data set under the target sensitive data perception project. The annotation result of any data item under the target sensitive data perception project is a true annotation result, which can indicate whether the corresponding data item is a data item in the sensitive information under the target sensitive data perception project.
[0030] When executing step 30, the computer system also needs to consider the diversity of the target sensitive data perception items. If there are multiple target sensitive data perception items, the computer system determines the corresponding target training supervision information for each item, and extracts the corresponding sensitive data annotation results from it. For example, the target sensitive data perception items include ID card number and telephone number. The computer system finds the training supervision information for these two items, and then extracts the sensitive data annotation results of the example learning data set under the two items of ID card number and telephone number.
[0031] In order to improve the execution efficiency of step 30, the computer system can adopt a parallel processing method. For multiple target sensitive data perception projects, the target training supervision information can be determined and the sensitive data annotation result can be extracted at the same time, thereby saving time and computing resources.
[0032] In some embodiments, the number of target sensitive data perception items is not less than one, and when the number of target sensitive data perception items is multiple, the multiple target sensitive data perception items are different sensitive data perception items. The number of sensitive data perception items corresponding to the training supervision information contained in the example learning data set may be greater than or equal to the number corresponding to the target sensitive data perception items. If the number of sensitive data perception items corresponding to the training supervision information contained in the example learning data set is equal to the number corresponding to the target sensitive data perception items, specifically, the total number of sensitive data perception items corresponding to each training supervision information is the same as the total number of target sensitive data perception items, then all the training supervision information contained in the example learning data set can be used as the target training supervision information. If the number of sensitive data perception items corresponding to the training supervision information contained in the example learning data set is greater than the number corresponding to the target sensitive data perception items, specifically, the total number of sensitive data perception items corresponding to each training supervision information is greater than the total number of each target sensitive data perception item, then the target training supervision information includes at least one of each training supervision information contained in the example learning data set.
[0033] Step 40: Based on the sensitive data perception results of the example learning dataset under the target sensitive data perception project and the sensitive data annotation results under the target sensitive data perception project, the sensitive data perception model is iteratively optimized; the optimized and converged sensitive data perception model is used to perform sensitive data perception operations on the data under the corresponding sensitive data perception project, so as to desensitize the perception results.
[0034] In step 40, the sensitive data perception result is obtained by the computer system after performing sensitive data perception operation on the example learning data set in step 20, which reflects the current recognition of sensitive data in the example learning data set by the model; while the sensitive data annotation result is determined by the computer system from the target training supervision information in step 30, which is the real annotation of sensitive data in the example learning data set. By comparing these two results, the computer system can find the deviations and errors in the model when identifying sensitive data.
[0035] In order to iteratively optimize the sensitive data perception model, the computer system measures the degree of difference between the sensitive data perception results and the sensitive data annotation results. A variety of technical means can be used to achieve this goal, such as using loss functions. The loss function is a function that quantifies the difference between the model's prediction results and the true label. Possible loss functions include cross entropy loss function, mean square error loss function, etc. Taking the cross entropy loss function as an example, for classification problems, it can measure the difference between the category probability distribution predicted by the model and the true category label. Assume that the probability that the model predicts a data item as category i is p i , and the probability corresponding to the true category label is y i (If it is a true category, y i =1, otherwise y i =0), then the calculation formula of the cross entropy loss function L is The computer system calculates the loss value of each data item, and then sums or averages the loss values of all data items to obtain the loss value of the entire example learning data set. This loss value reflects the overall degree of difference between the model's prediction results and the true annotations.
[0036] After calculating the loss value, the computer system adjusts the model internal variables of the sensitive data perception model in the direction of reducing the loss value. Model internal variables are parameters that the model continuously learns and adjusts during the training process, such as weights and biases in a neural network model. The computer system can use optimization algorithms to adjust the model internal variables. Viable optimization algorithms include stochastic gradient descent (SGD), Adagrad, Adadelta, Adam, etc. Taking the stochastic gradient descent algorithm as an example, its basic idea is to update these variables based on the gradient of the loss function to the model internal variables. Assume that the loss function is L and the model internal variables are , then in each iteration, the update formula of the model internal variables is ,in is the learning rate, which controls the step size of each update, is the loss function L in The gradient at .
[0037] The computer system performs multiple iterations when performing iterative model optimization. In each iteration, the computer system recalculates the loss value and updates the model's internal variables based on the loss value until the loss value converges to a smaller value or reaches a preset number of iterations. For example, in an example learning data set containing 1,000 data items, the computer system may perform 100 iterations, and each iteration will process all data items, update the model's internal variables, and gradually reduce the loss value.
[0038] In the iterative optimization process, the problems of overfitting and underfitting also need to be considered. Overfitting refers to the phenomenon that the model performs well on the training data but performs poorly on the new data; underfitting refers to the phenomenon that the model does not perform well on the training data. To avoid overfitting, the computer system can use regularization techniques, such as L1 regularization and L2 regularization, to limit the range of values of the model's internal variables by adding regularization terms to the loss function to prevent the model from being too complex. To avoid underfitting, the computer system can increase the complexity of the model, such as increasing the number of layers or neurons in the neural network.
[0039] The computer system can also use the validation set to evaluate the performance of the model. After each iteration, the computer system can use the validation set to evaluate the model and determine whether the performance of the model has improved based on indicators such as the loss value and accuracy on the validation set. If the performance on the validation set begins to decline, it means that the model may be overfitting. The computer system can stop the iteration in advance and select the model with the best performance as the final sensitive data perception model.
[0040] The optimized and converged sensitive data perception model has the ability to more accurately identify sensitive data under the target sensitive data perception project.
[0041] In one implementation scheme, there is no less than one target sensitive data perception item, and the number of sensitive data perception results obtained by performing sensitive data perception operations on an example learning data set in combination with the target sensitive data perception item is consistent with the number of target sensitive data perception items; the number of sensitive data perception items corresponding to the training supervision information configured in the example learning data set is greater than or equal to the number corresponding to the target sensitive data perception items; wherein, if the number of target sensitive data perception items is greater than 1, the number of sensitive data perception results obtained is also greater than 1, and at the same time, the number of sensitive data perception items corresponding to the target training supervision information obtained in combination with the training supervision information configured in the example learning data set is determined to be equal to the number of target sensitive data perception items.
[0042] In this implementation, the number of target sensitive data perception items faced by the computer system is no less than one, which means that the system focuses on multiple different types of sensitive data identification tasks at the same time. When the computer system performs sensitive data perception operations on the example learning data set in combination with these target sensitive data perception items, the number of sensitive data perception results obtained is consistent with the number of target sensitive data perception items. This is because each target sensitive data perception item corresponds to an independent sensitive data perception result, and each result reflects the sensitive data situation under this specific item in the example learning data set. For example, if the target sensitive data perception items include the three items of ID number, telephone number, and bank card number, then after the computer system operates on the example learning data set, it will obtain three sensitive data perception results corresponding to the ID number, telephone number, and bank card number, respectively.
[0043] The number of sensitive data perception items corresponding to the training supervision information configured in the example learning data set is greater than or equal to the number corresponding to the target sensitive data perception items. This is because the training supervision information needs to provide sufficient learning references for the model, and the range of sensitive data perception items it covers may be wider to ensure that the model can learn the characteristics of more types of sensitive data. For example, the training supervision information of the example learning data set may correspond to five sensitive data perception items: ID card number, telephone number, bank card number, email address, and license plate number, while the current target sensitive data perception items are only ID card number, telephone number, and bank card number. This setting allows the model to have richer information to learn during the training process, so that it can more accurately identify the target sensitive data perception items.
[0044] If the number of target sensitive data perception items is greater than 1, then the number of sensitive data perception results obtained must also be greater than 1, in order to accurately reflect the sensitive data situation under each target sensitive data perception item. At the same time, the number of sensitive data perception items corresponding to the target training supervision information obtained in combination with the training supervision information configured in the example learning data set is determined to be equal to the number of target sensitive data perception items. This is to ensure that the model can be optimized for specific target sensitive data perception items during training. For example, when the target sensitive data perception items are the three items of ID number, telephone number and bank card number, the computer system will determine the target training supervision information corresponding only to these three items from the training supervision information of the example learning data set, so that the model can focus on learning the characteristics of sensitive data under these three items.
[0045] In order to achieve the above functions, the computer system can adopt a variety of technical means. When performing sensitive data perception operations, the computer system can use a multi-task learning method to process multiple target sensitive data perception projects as different tasks at the same time. For each target sensitive data perception project, an independent classifier or neural network layer can be used for identification, and finally the identification results of each project are summarized to obtain multiple sensitive data perception results. When determining the target training supervision information, the computer system can compare the identification of the target sensitive data perception project with the identification of the sensitive data perception project in the training supervision information through data matching to find the corresponding target training supervision information.
[0046] When processing multiple target sensitive data perception projects, the computer system also considers the relationship between the projects. Some sensitive data may be associated with each other. For example, the ID card number and the phone number may belong to the same user. The computer system can use this association information to improve the accuracy of sensitive data perception. At the same time, the computer system evaluates and verifies the sensitive data perception results of each target sensitive data perception project to ensure the accuracy and reliability of the results. Some evaluation indicators, such as accuracy, recall rate, and F1 value, can be used to evaluate the perception effect of each project.
[0047] In one implementation, the process of obtaining the example learning data set includes the following steps:
[0048] Step 11: Obtain at least one backup indication data and project index; the project index indicates the number of projects of the target sensitive data perception project;
[0049] Step 12: in combination with the project indicators, obtaining candidate indication data corresponding to the target sensitive data perception projects of the corresponding project number from at least one candidate indication data, and using the obtained candidate indication data as project indication data of the example learning data set;
[0050] Step 13: Obtain the data set to be sensed, and fuse the project indication data and the data set to be sensed based on a preset fusion method to obtain an example learning data set.
[0051] In step 11, the computer system obtains no less than one candidate indication data and project index. The candidate indication data is data describing the sensitive data perception task, which can use the corresponding type label to reflect which types of sensitive data need to be perceived, and the project index represents the number of projects of the target sensitive data perception project. The computer system can obtain the candidate indication data from multiple data sources, such as extracting from an existing data dictionary, industry standard specification or a predefined rule base. These candidate indication data contain description information of various sensitive data types, such as type labels such as ID card number, mobile phone number, bank card number, etc. The determination of project indicators is usually based on specific business needs and data security requirements. For example, in financial business scenarios, it may be necessary to focus on sensitive data such as bank card number and credit card expiration date. At this time, the project index will be set to the number of projects corresponding to these sensitive data types. Assuming that the data processing task of a financial institution needs to perceive the three sensitive data of ID card number, bank card number and mobile phone number in user information, then the project index is 3, and the candidate indication data will contain the corresponding type labels such as ID card number, bank card number, mobile phone number, etc. To achieve this step, the computer system may use a data reading tool to read the standby indication data from a storage data source, and obtain project indicators through user input or configuration files.
[0052] In step 12, the computer system obtains the candidate indication data corresponding to the target sensitive data perception project of the corresponding number of projects from at least one candidate indication data in combination with the project index, and uses the obtained candidate indication data as the project indication data of the example learning data set. The computer system screens the candidate indication data according to the project index and selects the candidate indication data corresponding to the target sensitive data perception project. If the project index is 3, it means that there are three target sensitive data perception projects. The computer system will select the type labels corresponding to these three projects from the many candidate indication data as the project indication data. In actual operation, the computer system can use a matching algorithm to match the project index with the identifier of the candidate indication data to find the candidate indication data that meets the requirements. For example, by string matching, it is determined whether the type label in the candidate indication data is consistent with the identifier of the target sensitive data perception project. If it is consistent, it is selected as the project indication data.
[0053] In step 13, the computer system obtains the data set to be sensed, and fuses the project indication data and the data set to be sensed based on a preset fusion method (such as splicing) to obtain an example learning data set. The data set to be sensed is a collection of data that needs to be sensed for sensitive data. It can come from different business systems, such as a user information database, a transaction record database, etc. The preset fusion method is a method of combining the project indication data and the data set to be sensed. Splicing is a feasible fusion method, which is to connect the project indication data and the data set to be sensed in a certain order. For example, the project indication data is placed in front of the data set to be sensed and separated by a specific delimiter. Assume that the project indicator data is "ID number, bank card number, mobile phone number", and the data set to be sensed is "Zhang San 123456789012345678600000000000000000 100000000000", the example learning data set obtained after splicing and fusion is "ID number, bank card number, mobile phone number | Zhang San 123456789012345678 60000000000000000010000000000". In order to achieve this step, the computer system can use data processing tools to read and process the project indicator data and the data set to be sensed, and fuse them according to the preset fusion method.
[0054] In the process of obtaining the example learning data set, the computer system may encounter the problem of too much data. In order to improve the processing efficiency, the computer system can use the data sampling method to extract a part of the data from the standby indication data and the data set to be sensed for processing. Random sampling, stratified sampling and other methods can be used to ensure that the extracted data is representative. At the same time, the computer system can use distributed computing to distribute data processing tasks to multiple computing nodes for parallel processing to improve the processing speed.
[0055] In practical applications, the acquisition process of example learning data sets may be affected by many factors. Changes in business requirements may lead to changes in project indicators and target sensitive data perception projects, and the computer system adjusts the acquisition strategy in a timely manner. The reliability and stability of the data source will also affect the acquisition of data. The computer system establishes a data monitoring mechanism to detect and handle data anomalies in a timely manner.
[0056] In one implementation, the sensitive data perception model includes an implicit representation layer and a classification perception layer; the implicit representation layer is used to perform implicit data representation; the classification perception layer is used to perform sensitive data perception operations on example learning data sets to obtain corresponding sensitive data perception results.
[0057] In this implementation, the sensitive data perception model used by the computer system includes an implicit representation layer and a classification perception layer. These two layers each have different functions and work together to implement sensitive data perception operations on example learning data sets and obtain corresponding results.
[0058] The main function of the implicit representation layer is to perform implicit data representation. Implicit data representation converts the original data into a more abstract representation that better reflects the intrinsic characteristics and contextual relationships of the data. When processing text data, the original text is composed of characters or words. The semantic information of these characters or words is relatively scattered and difficult to be directly used for model learning and judgment. The implicit representation layer converts these text data into vector form. Each vector can contain the characteristics of the data item and the contextual relationship with other data items. For example, for the sentence "Zhang San's ID number is 123456789012345678", the implicit representation layer converts it into a vector. This vector not only contains the feature information of the words "Zhang San", "ID number" and "123456789012345678", but also reflects the semantic association between them. For example, "123456789012345678" is associated with "ID number" and belongs to "Zhang San". In order to achieve implicit data representation, the computer system can use a pre-trained language model, such as the BERT model. The BERT model can learn rich language knowledge and semantic information by performing unsupervised learning on a large-scale corpus. The computer system inputs the example learning data set into the BERT model, and after being processed by the model, the output data is the implicitly represented data.
[0059] The classification perception layer is responsible for performing sensitive data perception operations on the implicitly represented example learning data set, thereby obtaining the corresponding sensitive data perception results. This layer will determine whether the data items in the example learning data set are sensitive data and what type of sensitive data they belong to based on the target sensitive data perception items. For example, the target sensitive data perception items include ID card number, mobile phone number and bank card number. The classification perception layer will analyze the implicitly represented example learning data set to determine whether the data items therein are ID card number, mobile phone number or bank card number. The classification perception layer can use classification algorithms such as logistic regression, support vector machine or neural network. Taking neural network as an example, the computer system can build a multi-layer perceptron (MLP) as the classification perception layer. The implicitly represented data is input into the MLP. Each neuron of the MLP will perform weighted summation on the input data and perform nonlinear transformation through the activation function, and finally output the probability that each data item belongs to a different sensitive data type. The computer system determines the sensitive data type of the data item based on these probabilities, thereby obtaining sensitive data perception results.
[0060] In one implementation, step 20, performing a sensitive data perception operation on the example learning data set in combination with the target sensitive data perception project, and obtaining a sensitive data perception result of the example learning data set under the target sensitive data perception project, includes:
[0061] Step 21: perform implicit representation operations on the example learning data set in combination with the target sensitive data perception project to obtain the data implicit representation corresponding to the example learning data set; the data implicit representation is used to characterize the characteristics of the data items in the example learning data set and the contextual relationship between different data items;
[0062] Step 22: Perform sensitive data perception operations on the example learning dataset through data implicit representation to obtain the sensitive data perception results of the example learning dataset under the target sensitive data perception project.
[0063] In step 21, the computer system performs an implicit representation operation on the example learning data set in combination with the target sensitive data perception project to obtain the data implicit representation corresponding to the example learning data set, which is used to characterize the features of the data items in the example learning data set and the contextual relationship between different data items. The target sensitive data perception project provides a clear direction for the implicit representation operation, which guides the computer system to focus on the data features and contextual information related to these items in the example learning data set. For example, if the target sensitive data perception project is an ID number, a mobile phone number, and a bank card number, the computer system will focus on capturing the text features related to these sensitive data types in the example learning data set and their association in the context when performing the implicit representation operation.
[0064] In order to achieve implicit data representation, the computer system can use a pre-trained language model, such as BERT (Bidirectional Encoder Representations from Transformers). The BERT model can learn rich language knowledge and semantic information by performing unsupervised learning on a large-scale corpus. The computer system inputs the example learning data set into the BERT model, and the model processes the input data and converts it into a series of vector representations, which are the implicit representations of the data. Each vector not only contains the feature information of the corresponding data item, but also contains the contextual relationship between it and other data items. For example, for the sentence "Zhang San's ID number is 123456789012345678", after being processed by the BERT model, each word (such as "Zhang San", "ID number", "123456789012345678") will be converted into a vector, and the relationship between these vectors can reflect the semantic structure in the sentence, that is, "123456789012345678" is associated with "ID number" and belongs to "Zhang San".
[0065] Implicit data representation converts raw, complex data into a more abstract and expressive form, which facilitates subsequent sensitive data perception operations. Raw text data is often diverse and complex, and different expressions may express the same semantics. Implicit data representation can map these different expressions into similar vector spaces, allowing the model to better understand and process data. In addition, implicit data representation can also capture the contextual relationship between data items, which is very important for accurately identifying sensitive data. In the above example, if we only consider the string "123456789012345678" itself, it is difficult to determine whether it is an ID number, but combined with the context "Zhang San's ID number is", its identity can be clarified.
[0066] In step 22, the computer system performs sensitive data perception operations on the example learning data set through the data implicit representation, and obtains the sensitive data perception results of the example learning data set under the target sensitive data perception project. The data implicit representation obtained in step 21 provides more valuable input for this step, enabling the computer system to make more accurate sensitive data judgments based on the characteristics and context of the data.
[0067] In order to realize sensitive data perception operation, the computer system can use classification algorithms such as logistic regression, support vector machine or neural network. Taking neural network as an example, the computer system can build a multi-layer perceptron (MLP) as a classifier. The implicitly represented data is input into the MLP. Each neuron of the MLP will perform weighted summation on the input data and perform nonlinear transformation through the activation function, and finally output the probability that each data item belongs to different sensitive data types. The computer system determines whether the data item is sensitive data and what type of sensitive data it belongs to based on these probabilities, thereby obtaining sensitive data perception results. For example, for an implicitly represented data item, the MLP outputs that the probability of it belonging to an ID card number is 0.9, the probability of it belonging to a mobile phone number is 0.1, and the probability of it belonging to other types is 0.0, then the computer system can determine that the data item is an ID card number.
[0068] By reasonably selecting and using data implicit representation methods and sensitive data perception algorithms, computer systems can accurately identify sensitive data under the target sensitive data perception project in the example learning data set. In practical applications, computer systems continuously optimize and improve these steps to adapt to different business needs and data characteristics, improve the accuracy and efficiency of sensitive data perception, and provide reliable support for subsequent data desensitization operations. At the same time, computer systems also need to pay attention to data security and privacy protection to ensure that data leakage and abuse will not occur when processing sensitive data. In the face of ever-changing business environments and data security challenges, computer systems continue to explore and apply new technologies and methods to continuously improve the ability and level of sensitive data perception.
[0069] In one implementation, in step 22, before performing a sensitive data perception operation on the example learning data set through data implicit representation to obtain a sensitive data perception result of the example learning data set under the target sensitive data perception project, the method further includes:
[0070] Step 211: obtaining data classification of data segments corresponding to each data item in the example learning data set;
[0071] Step 212: generating a category attention identification sequence according to the data classification; the category attention identification sequence includes a plurality of attention identifications, and the data items included in the example learning data set correspond to the attention identifications in the category attention identification sequence;
[0072] Step 213: Perform a matrix multiplication operation on the data implicit representation of the example learning data set according to the category attention identification sequence to obtain the data implicit representation after the operation; the data implicit representation after the operation includes the data item implicit representation corresponding to each data item in the data set to be perceived, and the data implicit representation after the operation is used to perform sensitive data perception operations on the example learning data set.
[0073] In step 211, the computer system obtains the data classification of the data fragments corresponding to each data item in the example learning data set. The data classification is used to distinguish whether the data item is indicator data or data to be sensed. The example learning data set is usually a fusion of project indicator data and data to be sensed. The project indicator data describes the sensitive data perception task, while the data to be sensed contains data that needs to be sensitive data identified. The computer system clearly defines the data classification to which each data item belongs so that it can be processed in a targeted manner later. For example, in an example learning data set, "ID number, mobile phone number" belongs to project indicator data, while "Zhang San 123456789012345678 10000000000" belongs to data to be sensed. The computer system can classify data according to the source, format or pre-set rules of the data. If the data is a type label obtained from a data dictionary, then it is likely to be project indicator data; if the data is user information extracted from the database of a business system, then it is data to be sensed. In order to implement this step, the computer system can use a data parsing tool to parse the example learning data set and determine the data classification of each data item based on the characteristics and metadata information of the data.
[0074] In step 212, the computer system generates a category attention identification sequence according to the data classification. The category attention identification sequence is a sequence composed of masks, where 0 means no attention and 1 means attention. The sequence contains multiple attention identifications, and the data items contained in the example learning data set correspond to the attention identifications in the category attention identification sequence. The purpose of the computer system is to use this sequence to clarify which data items should be paid attention to in subsequent processing. In the above example, for the example learning data set "ID number, mobile phone number | Zhang San 12345678901234567810000000000", the corresponding category attention identification sequence may be "0 0 | 1 1 1", which means that the computer system will focus on the part of the data to be perceived, and will not pay too much attention to the part of the project indication data. In order to generate the category attention identification sequence, the computer system can assign a corresponding attention identification to each data item according to the data classification result obtained in step 211. A conditional judgment statement can be used. When the data item belongs to the data to be perceived, the attention identification 1 is assigned; when the data item belongs to the project indication data, the attention identification 0 is assigned.
[0075] In step 213, the computer system performs a matrix multiplication operation on the data implicit representation of the example learning data set according to the category attention identification sequence to obtain the data implicit representation after the operation, which includes the data item implicit representation corresponding to each data item in the data set to be perceived, and is used to perform sensitive data perception operations on the example learning data set. Matrix multiplication operation is a mathematical operation. By multiplying the category attention identification sequence with the data implicit representation, the information of the data items that do not need to be paid attention to can be filtered out, and only the information of the data items to be perceived can be retained. Assuming that the data implicit representation of the example learning data set is a matrix X, and the category attention identification sequence can be represented as a vector A, then the formula of the matrix multiplication operation is Y=A·X, where· represents element-by-element multiplication. In the above example, through the matrix multiplication operation, the computer system will only retain the data implicit representation of the data items to be perceived such as "Zhang San 123456789012345678 10000000000", and ignore the data implicit representation of the data items indicating the items such as "ID number, mobile phone number".
[0076] The generation of category attention identification sequences needs to be adjusted according to specific business needs and data characteristics. In some cases, it may be necessary to pay some attention to the project indicator data. In this case, the attention identification of some project indicator data items can be set to 1. Alternatively, different levels of attention may be required for different types of data to be perceived, which can be achieved by adjusting the value of the attention identification. For example, for some important sensitive data types, such as ID card numbers, their corresponding attention identification can be set to a higher value to increase the attention to them. In order to improve efficiency, the computer system can use parallel computing methods to distribute matrix multiplication operations to multiple computing nodes for parallel processing. Sparse matrix storage and operation methods can also be used to reduce unnecessary calculations.
[0077] Steps 211-213 screen and process the implicit data representation of the example learning data set through data classification, generation of category attention identification sequences and matrix multiplication operations, so that the computer system can perform sensitive data perception operations more targeted. In practical applications, the computer system flexibly adjusts the specific implementation methods of these three steps according to different business scenarios and data characteristics to improve the performance and effect of sensitive data perception. At the same time, the computer system also needs to be continuously optimized and improved to adapt to the ever-changing data environment and security requirements, and provide more reliable support for data desensitization processing. When faced with complex and diverse example learning data sets, the computer system comprehensively uses various technical means to ensure the accurate execution of steps 211-213, thereby laying a solid foundation for the smooth progress of the entire data desensitization processing process.
[0078] In one implementation, step 22, performing a sensitive data perception operation on the example learning data set through data implicit representation to obtain a sensitive data perception result of the example learning data set under a target sensitive data perception project, includes:
[0079] Step 221: combining the implicit representation of the data, estimating the sensitivity support of each data item in the example learning data set; the sensitivity support of any data item represents the confidence level of the data item classification of the corresponding data item belonging to the data item classification of sensitive information under the target sensitive data perception project;
[0080] Step 222: Generate the sensitive data perception result of the example learning dataset under the target sensitive data perception project through the sensitivity support corresponding to each data item in the example learning dataset.
[0081] In step 221, the computer system estimates the sensitivity support of each data item in the example learning data set in combination with the implicit representation of the data. The sensitivity support of any data item indicates the confidence level that the data item classification of the corresponding data item belongs to the data item classification of sensitive information under the target sensitive data perception item. The implicit representation of the data provides the computer system with the characteristics of the data items and the contextual relationships between different data items. Based on this information, the computer system can more accurately determine whether each data item is sensitive data and which sensitive data type it belongs to. For example, in an example learning data set containing user information, the target sensitive data perception items include ID number, mobile phone number, and bank card number. For the data item "123456789012345678", the computer system estimates its sensitivity support for the ID number based on its implicit representation of the data.
[0082] To achieve this step, the computer system can use classification models such as logistic regression, support vector machines or neural networks. Taking neural networks as an example, the computer system passes the implicit representation of data as input to the neural network, and the output layer of the neural network outputs the probability that each data item belongs to different sensitive data types. These probabilities are sensitive support. Assuming that the output layer of the neural network has three neurons, corresponding to the ID card number, mobile phone number and bank card number, for the data item "123456789012345678", the output of the neural network may be [0.9, 0.1, 0.0], which means that the sensitive support of the data item belonging to the ID card number is 0.9, the sensitive support of the mobile phone number is 0.1, and the sensitive support of the bank card number is 0.0. From a mathematical point of view, if x represents the data implicit representation of the data item, W represents the weight matrix of the neural network, b represents the bias vector, and f represents the activation function, then the sensitive support vector y output by the neural network can be calculated by the formula y=f(Wx + b).
[0083] In step 222, the computer system generates the sensitive data perception result of the example learning data set under the target sensitive data perception project by using the sensitivity support corresponding to each data item in the example learning data set. The sensitivity support reflects the possibility that each data item belongs to different sensitive data types. The computer system determines the final classification of each data item based on these possibilities, thereby generating the sensitive data perception result. In the above example, based on the sensitivity support [0.9, 0.1, 0.0], the computer system can determine that the data item "123456789012345678" is an ID card number.
[0084] To achieve this step, the computer system can use a simple threshold judgment method or a maximum probability selection method. The threshold judgment method is to set a threshold. When the sensitivity support of a data item belonging to a certain sensitive data type exceeds the threshold, the data item is classified as the sensitive data type. For example, the threshold is set to 0.5. For the data item "123456789012345678", its sensitivity support for the ID card number is 0.9, which exceeds the threshold, so it is classified as the ID card number. The maximum probability selection method directly selects the sensitive data type with the largest sensitivity support as the classification of the data item. In the above example, since the sensitivity support of the ID card number is 0.9, the data item is classified as the ID card number.
[0085] Steps 221-222: By reasonably selecting and using classification models and classification methods, the computer system can accurately estimate the sensitivity support of each data item and generate sensitive data perception results for the example learning data set under the target sensitive data perception project. In practical applications, the computer system continuously optimizes and improves these steps to adapt to different business needs and data characteristics, improve the accuracy and efficiency of sensitive data perception, and provide reliable support for subsequent data desensitization operations. At the same time, the computer system also pays attention to data security and privacy protection to ensure that data leakage and abuse will not occur when processing sensitive data.
[0086] In one implementation scheme, there is no less than one data item classification of sensitive information under the target sensitive data perception project, and the sensitivity support of any data item includes an estimated weight that the data item classification of the corresponding data item belongs to different data item classifications, and any estimated weight represents the confidence level that the data item classification of the corresponding data item belongs to the corresponding data item classification.
[0087] At this time, step 222 generates a sensitive data perception result of the example learning data set under the target sensitive data perception project through the sensitivity support corresponding to each data item in the example learning data set, including:
[0088] Step 2221: Obtain the maximum estimated weight from the estimated weights included in the sensitivity support of any data item;
[0089] Step 2222: Classify the maximum estimated weight and the data item corresponding to the maximum estimated weight as the target recognition result corresponding to any data item;
[0090] Step 2223: Combine the target recognition results of each data item in the example learning dataset to generate the sensitive data perception results of the example learning dataset under the target sensitive data perception item.
[0091] In step 2221, the computer system obtains the maximum estimated weight from the estimated weights included in the sensitive support of any data item. Sensitive support indicates the confidence level that the data item classification of the corresponding data item belongs to different data item classifications. Each different data item classification corresponds to an estimated weight, and these estimated weights reflect the possibility that the data item belongs to each classification. For example, in a scenario where the target sensitive data perception items include ID card number, mobile phone number and bank card number, for the data item "123456789012345678", the estimated weights included in its sensitive support may be 0.9 for ID card number, 0.05 for mobile phone number, and 0.05 for bank card number. The computer system finds the largest one among these three estimated weights, that is, 0.9. In order to implement this step, the computer system can use a traversal algorithm to compare the size of each estimated weight in turn and record the maximum value. In programming languages, a loop structure can usually be used to implement this process, and the maximum estimated weight is finally obtained by continuously updating the maximum value variable.
[0092] In step 2222, the computer system uses the maximum estimated weight and the data item classification corresponding to the maximum estimated weight as the target recognition result corresponding to any data item. In the above example, the maximum estimated weight is 0.9, and the corresponding data item classification is the ID card number, so the computer system uses the ID card number as the target recognition result of the data item "123456789012345678". This step is based on the principle that the maximum estimated weight represents the classification to which the data item is most likely to belong. By determining the maximum estimated weight and its corresponding classification, a clear sensitive data type identifier can be assigned to each data item. From a mathematical point of view, if w_i represents the estimated weight of the i-th data item classification, C_i represents the i-th data item classification, M represents the maximum estimated weight, and C_M represents the classification corresponding to the maximum estimated weight, then the target recognition result is (M, C_M).
[0093] In step 2223, the computer system combines the target recognition results of each data item in the example learning data set to generate the sensitive data perception results of the example learning data set under the target sensitive data perception project. After obtaining the target recognition results of each data item, the computer system aggregates these results to form the sensitive data perception results of the entire example learning data set. For example, the example learning data set contains three data items: "123456789012345678", "10000000000", and "6222021234567890". After steps 2221 and 2222, their target recognition results are ID card number, mobile phone number, and bank card number, respectively. Then the computer system integrates these results to obtain the sensitive data perception results of the example learning data set: data item "123456789012345678" is the ID card number, data item "10000000000" is the mobile phone number, and data item "6222021234567890" is the bank card number. To achieve this step, the computer system can use data structures to store the target recognition results of each data item, such as a list or dictionary, and then integrate these storage structures to form the final sensitive data perception results.
[0094] The computer system also needs to handle the situation where multiple estimated weights are equal and are the maximum value. In some cases, a data item may have multiple classifications with equal estimated weights and are all maximum values, which makes it difficult to determine the unique classification of the data item. For this situation, the computer system can use random selection, selection based on prior knowledge, or further analysis of the data context to solve it. For example, if it is known based on prior knowledge that the frequency of ID card numbers in the current data set is much higher than that of mobile phone numbers and bank card numbers, then when the estimated weights are equal, the ID card number can be selected as the target recognition result.
[0095] Steps 2221-2223 provide a basis for subsequent data desensitization operations by determining the maximum estimated weight of each data item, the corresponding target recognition result, and finally generating the sensitive data perception result of the example learning data set. In practical applications, computer systems continuously optimize these steps to improve the accuracy and efficiency of sensitive data perception, while fully considering the diversity, complexity and possible special circumstances of the data to ensure that the generated sensitive data perception results can accurately reflect the sensitive data in the example learning data set. In the face of ever-changing business needs and data security challenges, computer systems also need to continue to improve and innovate, introduce more advanced technologies and methods, such as the attention mechanism in deep learning, to further improve the performance of sensitive data perception. Through continuous adjustment and optimization, computer systems can better complete sensitive data perception tasks and provide more reliable support for data security and privacy protection.
[0096] In another implementation scheme, there is no less than one data item classification of sensitive information under the target sensitive data perception project, and the sensitivity support also indicates that when the data item classification of the corresponding data item belongs to different data item classifications, the confidence level of the data item classification of the adjacent data item belongs to different data item classifications, and the confidence level is expressed by the classification transition index of the adjacent data item; the sensitivity support includes the classification transition index and the estimated weight that the data item classification of the corresponding data item belongs to different data item classifications; any estimated weight indicates the confidence level that the data item classification of the corresponding data item belongs to the corresponding data item classification.
[0097] At this time, step 222 generates a sensitive data perception result of the example learning data set under the target sensitive data perception project through the sensitivity support corresponding to each data item in the example learning data set, including:
[0098] Step 222A: Generate at least one data item classification link corresponding to the example learning data set in combination with the sensitivity support; any data item classification link includes the data item classification of each data item in the example learning data set; any data item classification link corresponds to the estimated weight of the data item classification of each data item in the example learning data set under the corresponding data item classification link and the classification transition index of the adjacent data item;
[0099] Step 222B: determining a link score of any data item classification link according to the estimated weight and classification transition index corresponding to any data item classification link;
[0100] Step 222C: Obtain the data item classification link corresponding to the maximum link score from at least one data item classification link, and combine the data item classifications included in the data item classification link corresponding to the maximum link score to generate the sensitive data perception result of the example learning data set under the target sensitive data perception project.
[0101] In step 222A, the computer system generates at least one data item classification link corresponding to the example learning data set in combination with the sensitive support, and any data item classification link includes the data item classification of each data item in the example learning data set, and any data item classification link corresponds to the estimated weight of the data item classification of each data item in the example learning data set under the corresponding data item classification link and the classification transition index of the adjacent data item. The sensitive support here not only indicates the estimated weight of the data item classification of the corresponding data item belonging to different data item classifications, but also indicates the confidence level that the data item classification of the adjacent data item belongs to different data item classifications when the data item classification of the corresponding data item belongs to different data item classifications, and the confidence level is represented by the classification transition index of the adjacent data item. For example, in an example learning data set containing user information, the target sensitive data perception items include ID number, mobile phone number, and bank card number. For the data item sequence "Zhang San 123456789012345678 10000000000", the possible data item classification links generated are: Link 1 is "non-sensitive data-ID number-mobile phone number", Link 2 is "non-sensitive data-bank card number-mobile phone number", etc. Each data item in each link has a corresponding estimated weight and classification transition index. Assume that in link 1, the estimated weight of "123456789012345678" being classified as an ID number is 0.9, and its classification transition index with the previous data item "Zhang San" (non-sensitive data) is 0.8, indicating that the possibility of transitioning from non-sensitive data to ID number is 0.8. In order to generate these data item classification links, the computer system can use an exhaustive method or a heuristic search algorithm. The exhaustive method is to list all possible data item classification combinations, but this method will be very computationally intensive when the number of data items is large. The heuristic search algorithm selectively searches for possible data item classification links based on certain rules and experience, thereby reducing the amount of computation.
[0102] In step 222B, the computer system determines the link score of any data item classification link according to the estimated weight and classification transition index corresponding to any data item classification link. The link score comprehensively considers the classification possibility of each data item and the consistency of classification between adjacent data items. A reasonable link score can reflect the rationality and possibility of the data item classification link. The link score can be calculated using the weighted summation method. Assuming that there are n data items in a data item classification link, the estimated weight of the i-th data item is w i , the classification transition index between it and the previous data item is t i (For the first data item, t1 can be set to 1), the calculation formula of the link score S can be expressed as For example, for the above link 1, assuming that the estimated weight of "Zhang San" (non-sensitive data) is 0.95, the estimated weight of "123456789012345678" (ID number) is 0.9, the classification transition index is 0.8, and the estimated weight of "10000000000" (mobile phone number) is 0.92, and the classification transition index is 0.85, then the link score of this link is The computer system calculates the link score of each data item classification link, providing a quantitative basis for the subsequent selection of the optimal link.
[0103] In step 222C, the computer system obtains the data item classification link corresponding to the maximum link score from at least one data item classification link, and combines the data item classification included in the data item classification link corresponding to the maximum link score to generate the sensitive data perception result of the example learning data set under the target sensitive data perception project. After calculating the link scores of all data item classification links, the computer system compares these scores to find the link with the highest score. Because the link with the maximum link score indicates that the data item classification combination in the link is the most reasonable and most consistent with the contextual relationship and classification possibility of the data. For example, after calculating the link scores of link 1 and link 2, it is found that link 1 has a higher score, so the computer system selects link 1 as the optimal link. Then, according to the classification of each data item in link 1, that is, "non-sensitive data-ID number-mobile phone number", the sensitive data perception result of the example learning data set is generated, that is, "Zhang San" is identified as non-sensitive data, "123456789012345678" is the ID number, and "10000000000" is the mobile phone number. To achieve this step, the computer system may use a sorting algorithm to sort all link scores and find the link corresponding to the maximum value.
[0104] The accuracy of the classification transition index is also an important factor affecting the link score and the final result. The classification transition index reflects the consistency of the classification between adjacent data items. If the classification transition index is inaccurate, the link score may not correctly reflect the rationality of the link. In order to improve the accuracy of the classification transition index, the computer system can use a large amount of historical data for training to learn the transition rules between different data item classifications. Machine learning algorithms such as decision trees and random forests can be used to construct a classification transition index prediction model based on the data item classification and adjacent relationships in historical data.
[0105] In step 222B, the calculation method of the link score needs to be adjusted according to specific business needs and data characteristics. Different business scenarios may attach different importance to the estimated weight and classification transition index. In some scenarios with higher requirements for data consistency, the weight of the classification transition index can be set higher; and in some scenarios that pay more attention to the classification accuracy of a single data item, the weight of the estimated weight can be set higher. The computer system flexibly adjusts the calculation formula of the link score according to the actual situation to ensure that the generated link score can accurately reflect the rationality of the link.
[0106] When processing large-scale example learning data sets, the amount of calculation in steps 222A-222C will increase significantly. The increase in the number of data items will cause the number of data item classification links to grow exponentially, and the time and resource consumption for calculating link scores and selecting the optimal link will also increase significantly. In order to address this problem, the computer system can adopt distributed computing and parallel processing methods. The generation of data item classification links and the calculation tasks of link scores are assigned to multiple computing nodes and performed simultaneously, and finally the results are summarized for the selection of the optimal link. Pruning algorithms can also be used to exclude some obviously unreasonable links in advance in the process of generating data item classification links to reduce the amount of calculation.
[0107] Steps 222A-222C can more accurately identify sensitive data by considering the contextual relationship and classification coherence between data items. In practical applications, the computer system flexibly adjusts the specific implementation methods of these three steps according to different business needs and data characteristics, and continuously optimizes the methods of generating data item classification links, calculating link scores, and selecting the optimal link to improve the accuracy and efficiency of sensitive data perception and provide reliable support for data desensitization processing.
[0108] In one implementation, step 40, combining the sensitive data perception results of the example learning data set under the target sensitive data perception project and the sensitive data annotation results under the target sensitive data perception project, iteratively optimizes the sensitive data perception model, including:
[0109] Step 41: Determine a target debugging indicator under the target sensitive data perception project by combining the sensitive data perception result of the example learning data set under the target sensitive data perception project and the difference between the sensitive data annotation result under the target sensitive data perception project;
[0110] Step 42: Adjust the model internal variables of the sensitive data perception model in the direction of reducing the target debugging index to iteratively optimize the sensitive data perception model.
[0111] In step 41, the computer system determines the target debugging index under the target sensitive data perception project, that is, a loss value, by combining the sensitive data perception result of the example learning data set under the target sensitive data perception project and the difference between the sensitive data annotation result under the target sensitive data perception project. The sensitive data perception result is the identification result of the sensitive data obtained by the computer system after processing the example learning data set through the sensitive data perception model, while the sensitive data annotation result is a pre-labeled sensitive data identifier representing the actual situation. The difference between the two reflects the accuracy of the model prediction, and the target debugging index is a numerical value used to quantify this difference. For example, in a scenario where the target sensitive data perception project is to identify ID card numbers, mobile phone numbers, and bank card numbers, there is a record in the example learning data set "Zhang San 123456789012345678 10000000000", and the sensitive data annotation results show that "123456789012345678" is the ID card number and "10000000000" is the mobile phone number, while the sensitive data perception results may mistakenly identify "123456789012345678" as a bank card number. By comparing these two results, the computer system can find that the model's prediction has deviated.
[0112] In order to determine the target debugging index, the computer system can use a variety of loss functions. The feasible loss functions include cross entropy loss function, mean square error loss function, etc. Taking the cross entropy loss function as an example, assuming that the probability that the sensitive data perception result of the model example learning data set under the target sensitive data perception project is category i is p i , and the probability of sensitive data annotation results under the target sensitive data perception project is (If it is a real category, = 1, otherwise = 0), then the calculation formula of the cross entropy loss function L is , L is the target debugging index.
[0113] In step 42, the computer system adjusts the model internal variables of the sensitive data perception model in the direction of reducing the target debugging index to iteratively optimize the sensitive data perception model. Model internal variables are parameters that the model continuously learns and adjusts during the training process, such as weights and biases in a neural network model. The target debugging index reflects the difference between the model prediction result and the true annotation. The task of the computer system is to gradually reduce this difference by adjusting the model internal variables, that is, to reduce the target debugging index.
[0114] To achieve this goal, the computer system can use optimization algorithms such as stochastic gradient descent (SGD), Adagrad, Adadelta, Adam, etc. Taking the stochastic gradient descent algorithm as an example, its basic idea is to update these variables based on the gradient of the loss function to the internal variables of the model. The computer system will continue to repeat this update process until the target debugging indicator converges to a smaller value or reaches a preset number of iterations.
[0115] During the execution of steps 41-42, the computer system considers multiple factors. The calculation accuracy of the target debugging indicator directly affects the direction of model optimization. If the loss function is not selected properly or errors occur in the calculation process, the target debugging indicator may not accurately reflect the prediction error of the model, thereby causing the model optimization to proceed in the wrong direction. Therefore, the computer system selects a suitable loss function based on the specific problem and data characteristics. For classification problems, the cross entropy loss function is usually a better choice; for regression problems, the mean square error loss function may be more appropriate.
[0116] The choice of learning rate is also a key factor. If the learning rate is too large, the model may skip the optimal solution during the optimization process, resulting in the failure of the target debugging indicator to converge; if the learning rate is too small, the model will converge very slowly, increasing the training time. The computer system can adopt a learning rate decay strategy, using a larger learning rate at the beginning of training to quickly approach the optimal solution, and then gradually reducing the learning rate as training progresses to improve the stability of convergence. You can use an exponential decay strategy, with a learning rate of As the number of iterations t changes, the formula is ,in is the initial learning rate, is the attenuation coefficient.
[0117] When processing large-scale example learning data sets, the computer system can use batch training methods. The example learning data set is divided into multiple small batches, and only one small batch of data is used each time to calculate the target debugging indicators and update the internal variables of the model. This can reduce memory usage and improve training efficiency. The computer system can also use distributed training methods to distribute training tasks to multiple computing nodes at the same time, further improving the training speed.
[0118] Steps 41-42 can continuously optimize the sensitive data perception model by accurately calculating the target debugging index and adjusting the model's internal variables in the direction of reducing the index, so that the computer system can more accurately identify sensitive data under the target sensitive data perception project. In practical applications, the computer system continuously monitors the changes in the target debugging index and adjusts the optimization strategy according to the trend of changes. If the target debugging index no longer decreases or starts to increase during the training process, it means that the model may have problems such as overfitting or unreasonable learning rate settings, and the computer system takes timely measures to adjust it. At the same time, the computer system can also use the validation set to evaluate the performance of the model. After each iteration, the target debugging index and accuracy on the validation set are used to judge the generalization ability of the model, and the model with the best performance is selected as the final sensitive data perception model.
[0119] In one implementation, the method further includes an application phase after the model tuning converges, which may specifically include the following steps:
[0120] Step 100: Obtain the data to be desensitized and the project index; the project index represents the sensitive data perception project of the data to be desensitized and the number of projects corresponding to the sensitive data perception project;
[0121] Step 200: Acquire project indication data corresponding to sensitive data perception projects of a corresponding number of projects from at least one standby indication data in combination with project indicators, and generate target data to be processed by combining the project indication data and the data to be desensitized;
[0122] Step 300: Taking the optimized and converged sensitive data perception model and the sensitive data perception project indicated by the project indicator, a sensitive data perception operation is performed on the target data to be processed to obtain the sensitive data items of the target data to be processed under the sensitive data perception project;
[0123] Step 400: Based on a preset desensitization strategy, perform desensitization operations on sensitive data items.
[0124] In step 100, the computer system obtains the data to be desensitized and the project indicators, wherein the project indicators represent the sensitive data perception items of the data to be desensitized, and the number of items corresponding to these sensitive data perception items. The data to be desensitized refers to the data that needs to be sensitive data identified and desensitized, which can come from various business systems, such as user information databases, transaction record databases, etc. The project indicators clarify the types of sensitive data that need to be identified and their quantities, and provide targets and scopes for subsequent sensitive data perception operations. For example, in a financial business scenario, the data to be desensitized may be a data set containing user personal information and transaction records, and the project indicators indicate that the sensitive data perception items that need to be identified are ID card numbers, bank card numbers, and mobile phone numbers, and the number of items is 3. The computer system can obtain the data to be desensitized from the relevant business system through the data interface, and obtain the project indicators through user input or configuration files.
[0125] In step 200, the computer system obtains project indication data corresponding to the sensitive data perception projects of the corresponding project number from at least one candidate indication data in combination with the project indicator, and generates the target data to be processed by combining the project indication data and the data to be desensitized. The candidate indication data is pre-prepared data describing the sensitive data perception task, and contains description information of various sensitive data types. The computer system selects the project indication data corresponding to the target sensitive data perception project from the candidate indication data according to the project indicator. For example, the candidate indication data contains type labels such as ID card number, bank card number, mobile phone number, and email address, and the sensitive data perception projects corresponding to the project indicator are ID card number, bank card number, and mobile phone number. The computer system extracts "ID card number, bank card number, and mobile phone number" as project indication data. Then, the computer system generates the target data to be processed by combining the project indication data and the data to be desensitized in a preset fusion method (such as splicing). In order to achieve this step, the computer system can use a data matching algorithm to filter the project indication data from the candidate indication data, and then use a data processing tool to fuse the project indication data and the data to be desensitized.
[0126] In step 300, the computer system uses the optimized and converged sensitive data perception model combined with the sensitive data perception project indicated by the project indicator to perform sensitive data perception operations on the target data to be processed, and obtains the sensitive data items of the target data to be processed under the sensitive data perception project. The optimized and converged sensitive data perception model has the ability to accurately identify sensitive data under the target sensitive data perception project after previous training and optimization. The computer system inputs the target data to be processed into the model, and the model analyzes and judges the data according to the sensitive data perception project indicated by the project indicator to find out the sensitive data items. In the above example, the model will identify "123456789012345678" as the ID number, "6222021234567890" as the bank card number, and "10000000000" as the mobile phone number. These are the sensitive data items of the target data to be processed under the sensitive data perception project. To achieve this step, the computer system can directly call the prediction interface of the sensitive data perception model after optimization and convergence, take the target data to be processed as input, and obtain the sensitive data perception results output by the model.
[0127] In step 400, the computer system performs a desensitization operation on the sensitive data items based on a preset desensitization strategy. The preset desensitization strategy is a series of rules formulated according to business needs and data security requirements, which are used to process sensitive data so as to protect the privacy and security of the data without affecting the availability of the data. Feasible desensitization strategies include replacement, masking, encryption, etc. The replacement strategy is to replace sensitive data with other values, such as replacing the last four digits of the ID number with "XXXX"; the masking strategy is to cover part of the sensitive data with specific characters, such as replacing the middle digits of the bank card number with "*"; the encryption strategy is to encrypt sensitive data using an encryption algorithm, and only authorized personnel can decrypt and view it. In the above example, for the ID card number "123456789012345678", the mask strategy can be used to process it into "123456********78"; for the bank card number "6222021234567890", it can be replaced with a randomly generated card number with the same format; for the mobile phone number "10000000000", an encryption algorithm can be used to encrypt it. The computer system can select the appropriate desensitization strategy according to different sensitive data types and business requirements, and implement the desensitization operation of sensitive data items by writing corresponding processing programs.
[0128] During the entire application phase, the computer system considers multiple factors. The quality and security of data are crucial. For the data to be desensitized, the computer system must ensure that its source is reliable and that sensitive information will not be leaked during the acquisition and processing process. Encrypted transmission and storage methods can be used to protect the data. At the same time, the computer system must clean and preprocess the data to remove noise data and invalid data to improve the accuracy of sensitive data perception. For standby indication data and project indicators, the computer system must ensure its accuracy and completeness to avoid using erroneous or incomplete information.
[0129] The determination of project indicators needs to be reasonably set according to specific business needs and data security requirements. Different business scenarios may have different requirements for the definition and identification of sensitive data. The computer system flexibly adjusts project indicators according to actual conditions. In medical business scenarios, it may be necessary to focus on sensitive data such as patients' medical records, ID numbers, and medical records; while in e-commerce business scenarios, more attention may be paid to sensitive data such as users' mobile phone numbers, delivery addresses, and bank card numbers. The computer system can determine appropriate project indicators by communicating with business personnel and analyzing business processes.
[0130] The performance of the sensitive data perception model after optimization convergence directly affects the accuracy of sensitive data perception. The computer system regularly evaluates and updates the model to ensure that it maintains good performance in different data environments and business scenarios. The model can be evaluated using a test data set, and the model can be fine-tuned or retrained based on the evaluation results. Over time, the distribution and characteristics of the data may change. The computer system collects new data in a timely manner and updates the model to adapt to these changes.
[0131] The preset desensitization strategy needs to be reasonably selected according to the type of sensitive data and the usage scenario. Different desensitization strategies have different effects on the availability and security of data. The replacement strategy may cause the loss of data uniqueness, affecting the analysis and use of data; although the encryption strategy can well protect the security of data, it requires decryption operations when using data, which increases the complexity and processing time of the system. The computer system comprehensively considers these factors and selects the most appropriate desensitization strategy. For some data that needs to be analyzed but privacy must be protected, a masking strategy can be used to protect sensitive information and ensure a certain degree of data availability.
[0132] When processing large-scale data to be desensitized, computer systems also need to consider performance and efficiency issues. Distributed computing, parallel processing and other technical means can be used to improve the speed and efficiency of data processing. Using a distributed deep learning framework, sensitive data perception and desensitization operations can be distributed to multiple computing nodes for parallel processing, thereby shortening processing time. Computer systems can also use data caching technology to reduce repeated calculations and improve system response speed.
[0133] In one implementation, step 300, the sensitive data perception model after optimization convergence is combined with the sensitive data perception project indicated by the project indicator, and a sensitive data perception operation is performed on the target data to be processed to obtain the sensitive data items of the target data to be processed under the sensitive data perception project, including:
[0134] Step 310: using the optimized converged sensitive data perception model in combination with the sensitive data perception project indicated by the project indicator to perform a sensitive data perception operation on the target data to be processed, and obtaining a first sensitive data perception result; wherein the first sensitive data perception result includes a data item classification corresponding to the maximum estimated weight of each data item in the target data to be processed, and any estimated weight represents a confidence level that the data item classification of any data item belongs to the corresponding data item classification;
[0135] Step 320: According to the data item classification corresponding to the maximum estimated weight of each data item in the target data to be processed, determine the sensitive data items of the to-be-mass-desensitized data in the target data to be processed under the sensitive data perception item.
[0136] In step 310, the computer system uses the optimized and converged sensitive data perception model in combination with the sensitive data perception project indicated by the project indicator to perform sensitive data perception operations on the target data to be processed, and obtains a first sensitive data perception result, which includes the data item classification corresponding to the maximum estimated weight of each data item in the target data to be processed, and any estimated weight represents the confidence level that the data item classification of any data item belongs to the corresponding data item classification. The optimized and converged sensitive data perception model has been pre-trained and optimized, and has the ability to accurately identify sensitive data under the target sensitive data perception project. The project indicator indicates the type of sensitive data that needs to be identified, such as ID number, mobile phone number, bank card number, etc. The computer system inputs the target data to be processed into the model, and the model will analyze and judge the data according to the project indicators.
[0137] When processing the target data to be processed, the model will calculate the estimated weight of each data item belonging to different data item categories. For example, for the data item "123456789012345678" in the target data to be processed, if the sensitive data perception items indicated by the project indicator are ID number, mobile phone number and bank card number, the model may output that the estimated weight of the data item belonging to the ID number is 0.9, the estimated weight belonging to the mobile phone number is 0.05, and the estimated weight belonging to the bank card number is 0.05. The computer system will find the maximum value from these estimated weights. In this example, the maximum value is 0.9, and the corresponding category is ID number. Then the data item category corresponding to the maximum estimated weight of the data item is ID number. The computer system will perform such processing on each data item in the target data to be processed, and finally obtain the first sensitive data perception result containing the data item categories corresponding to the maximum estimated weights of each data item. To achieve this step, the computer system can directly call the prediction interface of the sensitive data perception model after optimization convergence, take the target data to be processed and project indicators as input, obtain the estimated weight of each data item output by the model belonging to different categories, and then find the maximum estimated weight and its corresponding category through comparison.
[0138] Step 320 is that the computer system determines the sensitive data items of the to-be-mass-desensitized data in the target data to be processed under the sensitive data perception project according to the data item classification corresponding to the maximum estimated weight of each data item in the target data to be processed. The first sensitive data perception result has clarified the data item classification corresponding to the maximum estimated weight of each data item. The computer system screens out the sensitive data items belonging to the sensitive data perception project based on these classification information. In the above example, it is known that the data item classification corresponding to the maximum estimated weight of the data item "123456789012345678" is the ID card number, and the ID card number belongs to the sensitive data perception project indicated by the project indicator, so "123456789012345678" is the sensitive data item of the to-be-mass-desensitized data in the target data to be processed under the sensitive data perception project. The computer system will judge all the data items in the target data to be processed one by one, and determine the data items that meet the classification of the sensitive data perception project as sensitive data items. To achieve this step, the computer system may use conditional judgment statements to compare the data item classification corresponding to the maximum estimated weight of each data item with the sensitive data perception item indicated by the project indicator, and if there is a match, the data item is determined to be a sensitive data item.
[0139] Steps 310-320, by using the optimized converged sensitive data perception model and project indicators, the computer system can accurately identify the sensitive data items in the target data to be processed, providing a basis for subsequent data desensitization operations. In practical applications, the computer system continuously optimizes these two steps to improve the accuracy and efficiency of sensitive data item determination, while fully considering the diversity, complexity and possible special circumstances of the data to ensure that the determined sensitive data items can accurately reflect the sensitive information in the target data to be processed.
[0140] In another implementation, step 300, the sensitive data perception model after optimization convergence is combined with the sensitive data perception project indicated by the project indicator to perform a sensitive data perception operation on the target data to be processed, and obtain the sensitive data items of the target data to be processed under the sensitive data perception project, including:
[0141] Step 300A: Take the optimized converged sensitive data perception model and combine it with the target sensitive data perception project to perform sensitive data perception operation on the target to-be-processed data to obtain a second sensitive data perception result; the second sensitive data perception result includes the data item classification included in the data item classification link corresponding to the maximum link score; wherein any data item classification link includes the estimated weight of the data item classification of each data item of the example learning data set under the corresponding data item classification link and the classification transition index of the data item classification corresponding to the adjacent data item; the link score of any data item classification link is determined by combining the estimated weight and classification transition index corresponding to any data item classification link;
[0142] Step 300B: Determine the sensitive data items of the to-be-demensitized data in the target data to be processed under the sensitive data perception item according to the data item classification included in the data item classification link corresponding to the maximum link score.
[0143] In step 300A, the computer system uses the optimized converged sensitive data perception model in combination with the target sensitive data perception project to perform sensitive data perception operations on the target data to be processed, and obtains a second sensitive data perception result, which includes the data item classification included in the data item classification link corresponding to the maximum link score. Any data item classification link here contains the estimated weight of the data item classification of each data item in the example learning data set under the corresponding data item classification link and the classification transition index of the data item classification corresponding to the adjacent data item. The link score of any data item classification link is determined in combination with the estimated weight and classification transition index corresponding to the data item classification link. Different from step 310, step 310 focuses on the data item classification corresponding to the maximum estimated weight of each data item, while step 300A pays more attention to the contextual relationship between data items and the coherence of classification.
[0144] When processing the target data to be processed, the computer system will generate multiple data item classification links. For example, the target data to be processed is "Zhang San 123456789012345678 10000000000", and the target sensitive data perception items are ID number, mobile phone number and non-sensitive data. The data item classification links that may be generated are: Link 1 is "non-sensitive data-ID number-mobile phone number", Link 2 is "non-sensitive data-mobile phone number-ID number", etc. Each data item in each link has a corresponding estimated weight and classification transition index of adjacent data items. Assume that in link 1, the estimated weight of "123456789012345678" being classified as an ID number is 0.9, and its classification transition index with the previous data item "Zhang San" (non-sensitive data) is 0.8, indicating that the possibility of transitioning from non-sensitive data to ID number is 0.8. The computer system will calculate the link score based on the estimated weight and classification transition index corresponding to each data item classification link, for example, using the weighted summation method. Assuming that there are n data items in a data item classification link, the estimated weight of the ith data item is w_i, and its classification transition index with the previous data item is t_i (for the first data item, t_1 can be set to 1), the calculation formula of the link score S is S=\sum_{i=1}^{n} (w_i\times t_i). After calculating the scores of all links, find the link corresponding to the maximum link score. The data item classification in this link constitutes the second sensitive data perception result. In order to achieve this step, the computer system can use an exhaustive method or a heuristic search algorithm to generate a data item classification link, and then calculate the link score according to the above formula, and finally find the data item classification link corresponding to the maximum link score by comparison.
[0145] Step 300B is that the computer system classifies the data items included in the data item classification link corresponding to the maximum link score, and determines the sensitive data items of the to-be-desensitized data in the target data to be processed under the sensitive data perception project. After obtaining the second sensitive data perception result, the computer system will screen out the sensitive data items belonging to the sensitive data perception project according to the data item classification information therein. In the above example, if the data item classification link corresponding to the maximum link score is "non-sensitive data-ID number-mobile phone number", then "123456789012345678" and "10000000000" are the sensitive data items of the to-be-desensitized data in the target data to be processed under the sensitive data perception project. This is different from step 320, which determines the sensitive data items according to the data item classification corresponding to the maximum estimated weight of each data item, while step 300B determines the sensitive data items based on the rationality of the entire data item classification link, and can more comprehensively consider the relationship between data items. To achieve this step, the computer system may use a conditional judgment statement to compare the data item classification in the data item classification link corresponding to the maximum link score with the target sensitive data perception item, and if there is a match, the data item is determined to be a sensitive data item.
[0146] Steps 300A-300B and steps 310-320 have their own advantages in different scenarios. Steps 300A-300B are suitable for scenarios with high requirements for contextual relationships between data items, and can more accurately identify sensitive data items with semantic associations. In a data set containing user information and transaction records, there are complex associations between data items. Using steps 300A-300B can better consider these associations and improve the accuracy of sensitive data identification. Steps 310-320 are suitable for scenarios with high requirements for classification accuracy of single data items and relatively weak associations between data items. They have fast processing speed and can quickly identify sensitive data items.
[0147] Steps 300A-300B provide a method that pays more attention to contextual relationships and classification consistency for the computer system to determine sensitive data items in the target data to be processed, which complements steps 310-320. In practical applications, the computer system selects a suitable implementation scheme according to specific business needs and data characteristics, or combines the two schemes to improve the accuracy and efficiency of sensitive data identification and provide reliable support for subsequent data desensitization operations. At the same time, the computer system also needs to continuously optimize the two implementation schemes and adjust them according to data changes and business needs to adapt to different application scenarios and data security challenges.
[0148] The present application embodiment provides a computer system, such as Figure 2As shown, the computer system 100 includes: a processor 101 and a memory 103. The processor 101 and the memory 103 are connected, such as through a bus 102. Optionally, the computer system 100 may also include a transceiver 104. It should be noted that in actual applications, the transceiver 104 is not limited to one, and the structure of the computer system 100 does not constitute a limitation on the embodiments of the present application.
[0149] An embodiment of the present application provides a computer system. The computer system in the embodiment of the present application includes: one or more processors; a memory; one or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more programs are executed by the processor, the above-mentioned data desensitization processing method based on the large model is implemented.
[0150] The above description is only a partial implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A data desensitization processing method based on a large model, characterized in that: include: Obtaining an example learning data set and a target sensitive data perception item of the example learning data set; The example learning data set is configured with at least one piece of training supervision information, and any piece of training supervision information represents a sensitive data annotation result of the example learning data set under a corresponding sensitive data perception project; Performing a sensitive data perception operation on the example learning data set in combination with the target sensitive data perception project to obtain a sensitive data perception result of the example learning data set under the target sensitive data perception project; Determine target training supervision information from at least one piece of training supervision information configured for the example learning data set, and determine sensitive data annotation results of the example learning data set under the target sensitive data perception project from the target training supervision information, wherein the target training supervision information is unified training supervision information for the corresponding sensitive data perception project and the target sensitive data perception project; In combination with the sensitive data perception result of the example learning data set under the target sensitive data perception project and the sensitive data annotation result under the target sensitive data perception project, the sensitive data perception model is iteratively optimized; the sensitive data perception model after optimization convergence is used to perform sensitive data perception operations on data under the corresponding sensitive data perception project, so as to perform desensitization operations on the perception results, wherein the sensitive data perception model includes an implicit representation layer and a classification perception layer; the implicit representation layer is used to perform implicit representation of data; the classification perception layer is used to perform sensitive data perception operations on the example learning data set to obtain corresponding sensitive data perception results; Obtaining the data to be desensitized and project indicators; the project indicators represent sensitive data perception projects of the data to be desensitized and the number of projects corresponding to the sensitive data perception projects; Acquire project indication data corresponding to sensitive data perception projects of a corresponding number of projects from at least one standby indication data in combination with the project indicators, and generate target data to be processed in combination with the project indication data and the data to be desensitized; Taking the optimized converged sensitive data perception model and the sensitive data perception project indicated by the project indicator, a sensitive data perception operation is performed on the target data to be processed to obtain the sensitive data items of the target data to be processed under the sensitive data perception project; Based on a preset desensitization strategy, a desensitization operation is performed on the sensitive data item.
2. The method according to claim 1, characterized in that The number of target sensitive data perception items is no less than one, and the number of sensitive data perception results obtained by performing sensitive data perception operations on the example learning data set in combination with the target sensitive data perception items is consistent with the number of target sensitive data perception items; the number of sensitive data perception items corresponding to the training supervision information configured in the example learning data set is greater than or equal to the number corresponding to the target sensitive data perception items; wherein, if the number of target sensitive data perception items is greater than 1, the number of sensitive data perception results obtained is also greater than 1, and at the same time, the number of sensitive data perception items corresponding to the target training supervision information obtained in combination with the training supervision information configured in the example learning data set is determined to be equal to the number of target sensitive data perception items.
3. The method according to claim 1, characterized in that The process of obtaining the example learning data set includes the following steps: Acquire at least one standby indication data and a project index; the project index represents the number of projects of the target sensitive data perception project; In combination with the project indicators, obtaining candidate indication data corresponding to the target sensitive data perception projects of the corresponding project number from the at least one candidate indication data, and using the obtained candidate indication data as the project indication data of the example learning data set; The data set to be sensed is obtained, and the project indication data and the data set to be sensed are fused based on a preset fusion method to obtain an example learning data set.
4. The method according to claim 3, characterized in that The step of performing a sensitive data perception operation on the example learning data set in combination with the target sensitive data perception project to obtain a sensitive data perception result of the example learning data set under the target sensitive data perception project includes: An implicit representation operation is performed on the example learning data set in combination with the target sensitive data perception project to obtain a data implicit representation corresponding to the example learning data set; the data implicit representation is used to characterize the characteristics of the data items in the example learning data set and the contextual relationship between different data items; A sensitive data perception operation is performed on the example learning data set through the data implicit representation to obtain a sensitive data perception result of the example learning data set under the target sensitive data perception project.
5. The method according to claim 4, characterized in that Before performing a sensitive data perception operation on the example learning data set through the data implicit representation to obtain a sensitive data perception result of the example learning data set under the target sensitive data perception project, the method further includes: Obtaining data classification of data segments corresponding to each data item in the example learning data set; Generating a category attention identification sequence according to the data classification; the category attention identification sequence includes a plurality of attention identifications, and the data items included in the example learning data set correspond to the attention identifications in the category attention identification sequence; Performing a matrix multiplication operation on the data implicit representation of the example learning data set according to the category attention identification sequence to obtain an implicit representation of the data after the operation; the implicit representation of the data after the operation includes an implicit representation of data items corresponding to each data item in the data set to be perceived, and the implicit representation of the data after the operation is used to perform a sensitive data perception operation on the example learning data set; The step of performing a sensitive data perception operation on the example learning data set through the data implicit representation to obtain a sensitive data perception result of the example learning data set under the target sensitive data perception project includes: In combination with the implicit representation of the data, the sensitivity support of each data item in the example learning data set is estimated; the sensitivity support of any data item represents the confidence level of the data item classification of the corresponding data item belonging to the data item classification of sensitive information under the target sensitive data perception project; The sensitive data perception result of the example learning data set under the target sensitive data perception project is generated by using the sensitivity support corresponding to each data item in the example learning data set.
6. The method according to claim 5, characterized in that The data item classification of sensitive information under the target sensitive data perception project is at least one, the sensitivity support of any data item includes an estimated weight that the data item classification of the corresponding data item belongs to different data item classifications, and any estimated weight represents a confidence level that the data item classification of the corresponding data item belongs to the corresponding data item classification; the sensitive data perception result of the example learning data set under the target sensitive data perception project is generated by the sensitivity support corresponding to each data item in the example learning data set, including: Obtaining the maximum estimated weight from the estimated weights included in the sensitivity support of any one of the data items; Classifying the maximum estimated weight and the data item corresponding to the maximum estimated weight as the target recognition result corresponding to any one of the data items; Combining the target recognition result of each data item in the example learning data set, generating a sensitive data perception result of the example learning data set under the target sensitive data perception item; There is no less than one data item classification of sensitive information under the target sensitive data perception project, and the sensitivity support also indicates that when the data item classification of the corresponding data item belongs to a different data item classification, the confidence level of the data item classification of the adjacent data item belongs to a different data item classification, and the confidence level is represented by the classification transition index of the adjacent data item; the sensitivity support includes the classification transition index and the estimated weight that the data item classification of the corresponding data item belongs to a different data item classification; any one of the estimated weights indicates the confidence level that the data item classification of the corresponding data item belongs to the corresponding data item classification; The step of generating a sensitive data perception result of the example learning data set under the target sensitive data perception project by using the sensitivity support corresponding to each data item in the example learning data set includes: In combination with the sensitivity support, at least one data item classification link corresponding to the example learning data set is generated; any data item classification link includes the data item classification of each data item in the example learning data set; any data item classification link corresponds to an estimated weight of the data item classification of each data item in the example learning data set under the corresponding data item classification link and a classification transition index of adjacent data items; Determine a link score of any data item classification link according to an estimated weight and a classification transition index corresponding to any data item classification link; Obtain a data item classification link corresponding to the maximum link score from the at least one data item classification link, and combine the data item classifications included in the data item classification link corresponding to the maximum link score to generate a sensitive data perception result of the example learning data set under the target sensitive data perception project.
7. The method according to claim 1, characterized in that The iterative optimization of the sensitive data perception model in combination with the sensitive data perception result of the example learning data set under the target sensitive data perception project and the sensitive data annotation result under the target sensitive data perception project includes: Determine a target debugging indicator under the target sensitive data perception project by combining the sensitive data perception result of the example learning data set under the target sensitive data perception project and the difference between the sensitive data annotation result under the target sensitive data perception project; In the direction of reducing the target debugging index, the model internal variables of the sensitive data perception model are adjusted to iteratively optimize the sensitive data perception model.
8. The method according to claim 1, characterized in that: The sensitive data perception model after optimization convergence is combined with the sensitive data perception project indicated by the project indicator to perform a sensitive data perception operation on the target data to be processed, and obtain the sensitive data items of the target data to be processed under the sensitive data perception project, including: The sensitive data perception model after optimization convergence is combined with the sensitive data perception project indicated by the project indicator to perform a sensitive data perception operation on the target data to be processed, and obtain a first sensitive data perception result; wherein the first sensitive data perception result includes a data item classification corresponding to the maximum estimated weight of each data item in the target data to be processed, and any estimated weight represents a confidence level that the data item classification of any data item belongs to the corresponding data item classification; According to the data item classification corresponding to the maximum estimated weight of each data item in the target data to be processed, determine the sensitive data item of the to-be-desensitized data in the target data to be processed under the sensitive data perception item; Alternatively, the sensitive data perception model after optimization convergence is combined with the sensitive data perception project indicated by the project indicator to perform a sensitive data perception operation on the target data to be processed, and obtain the sensitive data items of the target data to be processed under the sensitive data perception project, including: The sensitive data perception model after optimization convergence is combined with the target sensitive data perception project to perform sensitive data perception operation on the target data to be processed, and obtain a second sensitive data perception result; the second sensitive data perception result includes the data item classification included in the data item classification link corresponding to the maximum link score; wherein, any data item classification link contains the estimated weight of the data item classification of each data item of the example learning data set under the corresponding data item classification link and the classification transition index of the data item classification corresponding to the adjacent data item; the link score of any data item classification link is determined in combination with the estimated weight and classification transition index corresponding to any data item classification link; According to the data item classification included in the data item classification link corresponding to the maximum link score, the sensitive data item of the to-be-desensitized data in the target data to be processed under the sensitive data perception item is determined.
9. A computer system, characterized in that: include: one or more processors; Memory; one or more computer programs; The one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
APP sensitive feature detection method and system based on large-scale language model
CN118606937A
Data desensitization method, data desensitization device, electronic equipment and readable storage medium
CN119089504A