Multi-strategy data desensitization method

Through a multi-strategy data desensitization method combining information entropy with k-means clustering dynamic rating and static dynamic desensitization framework, the problems of rigid policy and lack of evaluation in the existing technology are solved, the flexibility and security of data desensitization are improved, and the adaptability and privacy preferences are adapted to changing business scenarios and privacy preferences are improved, and the security and efficiency of data sharing are improved.

CN120354458AActive Publication Date: 2025-07-22CENT SOUTH UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510843886.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

Existing data desensitization solutions have rigid strategies and cannot adapt to changing business scenarios and privacy preferences. They lack desensitization quality assessment for specific scenarios such as machine learning and collaborative analysis, making it difficult to quantify the balance between privacy protection and data value.

Method used

A multi-strategy data desensitization method is adopted to dynamically rated attribute security level through information entropy and k-means clustering, combined with a hybrid execution framework of static and dynamic desensitization, and a multi-factor privacy leakage risk assessment and desensitization are performed based on the privacy and utility dual evaluation model to generate desensitization files.

Benefits of technology

It realizes the flexibility and security of data desensitization, and can dynamically select desensitization parameters based on actual scenarios, improves the desensitization efficiency and privacy protection capabilities, and ensures the security and availability of data during the sharing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354458A_ABST
    Figure CN120354458A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-strategy data desensitization method, which belongs to the technical field of data processing and specifically comprises the following steps of: 1, acquiring a file to be desensitized and selecting parameters; step 2, performing dynamic grading on each attribute based on information entropy and k-means clustering to obtain a security level of each attribute; 3, performing multi-factor privacy disclosure risk assessment according to the selected parameters and the security level of each attribute to obtain a privacy disclosure risk coefficient; 4, desensitizing the to-be-desensitized file according to the privacy disclosure risk coefficient based on a static and dynamic desensitization hybrid execution framework to generate a desensitized file; step 5, evaluating the desensitized file based on the privacy and utility double evaluation model, and if an evaluation result does not meet a preset condition, returning to the step 4 until the evaluation result meets the preset condition; and step 6, outputting the desensitization file and the final evaluation result to the client. Through the scheme of the invention, the desensitization efficiency and safety are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of data processing, and in particular, to a multi-strategy data desensitization method. Background Art

[0002] Currently, with the frequent occurrence of data security issues and the gradual enhancement of personal privacy protection awareness, numerous data desensitization service providers have emerged in the market, offering customized solutions to meet the needs of different customers. However, the existing data desensitization solutions expose two major defects: 1. Rigid strategies: 82% of commercial desensitization systems adopt static rule libraries and cannot adapt to changing business scenarios and privacy preferences; 2. Lack of evaluation: There is a lack of a desensitization quality evaluation system for specific scenarios such as machine learning and collaborative analysis, resulting in difficulties in quantitatively balancing privacy protection and data value.

[0003] It can be seen that there is an urgent need for a multi-strategy data desensitization method with better desensitization efficiency and security. Summary of the Invention

[0004] In view of this, the embodiments of the present invention provide a multi-strategy data desensitization method, which at least partially solves the problem of poor desensitization efficiency and security in the prior art.

[0005] The embodiments of the present invention provide a multi-strategy data desensitization method, including: Step 1, obtain the file to be desensitized and perform parameter selection, where the parameters include the total amount of data in the file to be desensitized, the selected number of records, the total number of attribute types, and the selected number of attributes; Step 2, perform dynamic grading on each attribute based on information entropy and k-means clustering to obtain the security level of each attribute; Step 3, perform a multi-factor privacy leakage risk assessment according to the selected parameters and the security level of each attribute to obtain a privacy leakage risk coefficient; Step 4, based on a hybrid execution framework of static and dynamic desensitization, desensitize the file to be desensitized according to the privacy leakage risk coefficient to generate a desensitized file; Step 5, evaluate the desensitized file based on a privacy and utility dual-evaluation model. If the evaluation result does not meet the preset conditions, return to Step 4 until the preset conditions are met; Step 6, output the desensitized file and the final evaluation result to the client.

[0006] According to a specific implementation manner of the embodiments of the present invention, the specific content of Step 1 includes: Step 1.1, obtain the file to be desensitized uploaded by the user side, and select the number of records, the application scenario, and the desensitization method; Step 1.2, set the total amount of data in the file to be desensitized as , the number of selected records is , the total number of attribute types , select attributes.

[0007] According to a specific implementation manner of an embodiment of the present invention, step 2 specifically includes: Step 2.1, after performing one-hot encoding on each attribute, calculate the mutual information between pairwise attributes and normalize it, and accordingly form an attribute correlation matrix, where the expression of the mutual information is where and are marginal entropies, and are conditional entropies, is the attribute and joint entropy; The expression of the normalization is where is the mutual information; The expression of the attribute correlation matrix is where represents the normalized value of the joint entropy of the th attribute field and the th attribute field, that is, the correlation coefficient; Step 2.2, perform K-means clustering on the attributes according to the attribute correlation matrix, use the silhouette coefficient to find the optimal number of clusters, and obtain the security level of each attribute according to the cluster to which it belongs, the correlation coefficient, and the security level of the attribute data label.

[0008] According to a specific implementation manner of an embodiment of the present invention, step 3 specifically includes: Step 3.1, use the ratio between the total amount of data N in the file to be desensitized and the number of selected records n as the record selection quantity weight ; Step 3.2, use the ratio between the total amount of attribute types M and the number of selected attributes m as the field selection quantity weight ; Step 3.3, obtain the kinds of attributes of the record and the kinds of supported application scenarios, obtain the security level of each attribute denoted as , obtain the privacy level mark established for the corresponding application scenario denoted as , the th field of the The operation risk weight corresponding to a certain operation is denoted as , with attributes as row vectors and application scenarios as column vectors, and establish the first operation risk matrix ; Step 3.4, according to the data usage requirements, select the th attribute field's operation risk weight corresponding to a certain operation , construct the second operation risk matrix in this scenario, and fill the un-involved requirements with . Sum the rows of the second operation risk matrix and perform normalization to obtain the field operation risk weight ; Step 3.5, assume that among the selected m attributes, there are continuous data and categorical data. Calculate the first distance between every two pieces of data in the continuous data where is the distance length of the numerical range, and represents continuous data; Step 3.6, calculate the second distance between every two pieces of data in the categorical data where is the number of attribute value types included in the attribute set, and represents discrete records; Step 3.7, for the selected n records, continuous data, categorical data, and Step 3.8, calculate the relevance weight of the records according to the total distance; ; Step 3.9, calculate the privacy leakage risk coefficient according to the record selection quantity weight, field selection quantity weight, field operation risk weight, and record relevance weight.

[0009] According to a specific implementation manner of an embodiment of the present invention, the specific steps of step 4 include: Step 4.1, perform masking, hash transformation, and asymmetric encryption on the continuous data in the file to be desensitized in the static desensitization layer of the hybrid execution framework for static and dynamic desensitization; Step 4.2, desensitize the categorical data in the file to be desensitized in the dynamic desensitization layer of the hybrid execution framework for static and dynamic desensitization based on k-anonymity, differential privacy, and generative adversarial networks; Step 4.3, form a desensitized file from the outputs of the static desensitization layer and the dynamic desensitization layer.

[0010] According to a specific implementation manner of an embodiment of the present invention, the specific steps of step 5 include: Step 5.1, use a decision tree classifier, polynomial regression, and a multi-layer perceptron to evaluate the utility between the desensitized file and the file to be desensitized; Step 5.2, evaluate the privacy between the desensitized file and the file to be desensitized based on the minimum distance quantile value and the nearest neighbor ratio to form an evaluation result; Step 5.3, if the evaluation result does not meet the preset conditions, return to step 4 until the preset conditions are met.

[0011] The multi-strategy data desensitization solution in the embodiments of the present invention includes: Step 1, obtain the file to be desensitized and perform parameter selection, where the parameters include the total amount of data in the file to be desensitized, the selected number of records, the total amount of attribute types, and the selected number of attributes; Step 2, perform dynamic grading on each attribute based on information entropy and k-means clustering to obtain the security level of each attribute; Step 3, perform a multi-factor privacy leakage risk assessment based on the selected parameters and the security level of each attribute to obtain a privacy leakage risk coefficient; Step 4, based on the hybrid execution framework for static and dynamic desensitization, desensitize the file to be desensitized according to the privacy leakage risk coefficient to generate a desensitized file; Step 5, evaluate the desensitized file based on the privacy and utility dual-evaluation model. If the evaluation result does not meet the preset conditions, return to step 4 until the preset conditions are met; Step 6, output the desensitized file and the final evaluation result to the client.

[0012] The beneficial effects of the embodiments of the present invention are as follows: Through the solution of the present invention, with the help of information entropy and the k-means clustering algorithm, dynamic grading is performed on the combined fields, the privacy leakage risk of the data to be published is evaluated considering the number of records, attribute fields, and usage scenarios, desensitization parameters are dynamically selected according to the desensitization strategy, and at the same time, static desensitization technology is combined to automatically identify and desensitize the quasi-identifier, and the utility and privacy of the desensitized data are evaluated, thereby helping the data publisher to achieve more accurate and intelligent data desensitization, improving the desensitization efficiency and security. Description of the Drawings

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0014] Figure 1 It is a schematic flow chart of a multi-strategy data desensitization method provided by an embodiment of the present invention; Figure 2 It is a schematic specific implementation flow chart of a multi-strategy data desensitization method provided by an embodiment of the present invention; Figure 3 It is a schematic diagram of the category of attribute fields provided by an embodiment of the present invention; Figure 4 It is a schematic flow chart of the privacy leakage risk coefficient assessment provided by an embodiment of the present invention; Figure 5 It is a schematic diagram of the selection of data desensitization strategies provided by an embodiment of the present invention. Detailed implementation manners

[0015] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0016] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0017] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present invention, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. Additionally, this device and / or this method can be implemented using other structures and / or functionality in addition to one or more of the aspects described herein.

[0018] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. The diagrams only show the components related to the present invention, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0019] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0020] An embodiment of the present invention provides a multi-strategy data desensitization method, which can be applied to the information protection process in the Internet scenario.

[0021] See Figure 1 , which is a schematic flowchart of a multi-strategy data desensitization method provided by an embodiment of the present invention. As Figure 1 and Figure 2 shown, the method mainly includes the following steps: Step 1: Obtain the file to be desensitized and perform parameter selection. Among them, the parameters include the total amount of data in the file to be desensitized, the selected number of records, the total amount of attribute types, and the selected number of attributes; In specific implementation, the specific processes of uploading the file and selecting the parameters can be as follows: Step 1.1: The user uploads the file to be desensitized, and the supported formats are as follows:.txt,.csv,.excel, etc. At the same time, the user selects the number of records, the application scenario, and the desensitization method.

[0022] In specific implementation, on the main interface of multi-strategy data intelligent desensitization, the user clicks to select a file for uploading. In this embodiment, the Adult dataset is used; the system automatically obtains the file header information and returns it to the desensitization interface. The user selects the fields to be desensitized, the number of records, the application scenario, and the desensitization method according to the requirements, and clicks the "Start Desensitization Processing" button to enter the dynamic grading stage. The parameter selection is shown in the following table: Table 1

[0023] Step 1.2: Data preprocessing. Set the total number of data to be desensitized by the user as , and the selected number of records as ; there are a total of types of attribute fields, and attribute fields are selected from them. The attribute fields are as Figure 3 shown.

[0024] The total number of data in the Adult dataset is 45,223, and 10,000 data are selected from it. There are 15 attribute fields in total, 12 of which are selected, and 3 additional static desensitization fields are added (dynamic grading is not required).

[0025] Step 2: Based on information entropy and k-means clustering, perform dynamic grading between each attribute to obtain the security level of each attribute; When specifically implemented, the specific process of dynamic grading of combined attributes based on information entropy and k-means clustering is as follows: Step 2.1: Calculate the similarity matrix between attributes. For different types of attributes, one-hot encoding is performed on discrete attributes, and the mutual information between pairwise variables is calculated. The calculation formula is as shown in Equation (1): (1) Among them, and are marginal entropies, and are conditional entropies, and is and 's joint entropy. Normalize the information entropy between attributes. By normalizing the joint entropy, the value range is limited to The calculation formula is as shown in Equation (2): (2) Among them, is mutual information, and are and 's entropies. Obtain the attribute correlation matrix .

[0026] (3) Among them, represents the joint entropy normalization value of the th attribute field and the th attribute field, that is, the correlation coefficient.

[0027] Step 1.2: K-means clustering. Perform K-means clustering on the attributes, and use the silhouette coefficient to find the optimal number of clusters. Finally, based on the cluster to which it belongs, the correlation coefficient, and the security level of the attribute data label , obtain the security level of the attribute . The specific calculation process is as follows: In each cluster, obtain the security level of the data label of the attribute , and the correlation coefficient , the calculation formula for the relevant security level of every two attributes is as shown in (4): (4) For each attribute, perform summation to obtain . For example, use the similarity matrix to cluster the attribute fields, and the clustering result is [0 0 0 0 0 0 4 2 4 3 5 4 0 1 0] original security level of the attribute fields , after combined classification, .

[0028] Step 3, perform a multi-factor privacy leakage risk assessment based on the selected parameters and the security level of each attribute to obtain the privacy leakage risk coefficient; Specifically, in this step, analyze the data usage requirements and correlations to determine the privacy protection score of the data selected by the user. The requirements analysis includes two parts. The first part is to analyze the proportion of the number of records and the number of attributes of the data selected by the user, collectively referred to as quantity analysis; the second part is to analyze the application scenario and construct an operation risk matrix. The correlation analysis also includes two parts. The first part is to analyze the attribute correlation and re-define the security level of the attribute using mutual information for the construction of the operation risk matrix. The record correlation takes into account that the data requested by the user includes continuous data and discrete data, and different distance functions are used to calculate the correlation. As Figure 4 shown: Step 3.1, Requirements analysis ① Quantity analysis: Record selection quantity weight : The total number of records is , the number of requested record lines is , define the ratio of the two as , is the record selection quantity weight. Field selection quantity weight : There are a total of types of attribute fields, and attribute fields are selected and returned for query, define the ratio of the two as , is the field selection quantity weight.

[0029] ② Scenario analysis: Data operation: A total of types of application scenarios are supported, and the user selects types of application scenarios.

[0030] Operation risk matrix : Obtain the types of attribute fields of the record and the types of supported usage scenarios. Obtain the security level related to the attribute, denoted as ; Obtain the privacy level mark established for the usage scenario, denoted as Taking the fields as row vectors and the application scenarios as column vectors, establish the operation risk matrix , the th field's th operation's operation risk weight is denoted as , and the operation risk matrix is denoted as: (5) Operation risk weight vector : According to the data usage requirements, select the th attribute field's th operation's operation risk weight , construct the operation risk matrix in this scenario , fill the un-involved requirements with 0, and obtain: (6) Sum the rows of and perform normalization processing to obtain the vector , which is the operation risk weight vector of the fields.

[0031] In specific implementation, .

[0032] Step 3.2, Correlation analysis Suppose among the selected m attributes, there are continuous data and categorical data. Calculate the first distance between every two continuous data (7) where is the distance length of the numerical domain, and represents continuous data; Step 3.6, Calculate the second distance between every two categorical data (8) (9) where is the number of attribute value types included in the attribute set, and represents discrete records For the selected n records, continuous data, categorical data, and types of global combined records, the distance between records is calculated as: (10) Among them, the tags and represent continuous data and categorical data respectively.

[0033] Then the similarity between records is denoted as , which is the correlation weight coefficient of the record.

[0034] In specific implementation , .

[0035] Step 3.3, Privacy leakage risk assessment Evaluate the data privacy risk. By conducting requirement analysis and correlation analysis on the data to be released, evaluate the privacy leakage risk S of the data. If the score exceeds the set threshold, it means that the data sensitivity is low and the data can be directly returned; if the score is lower than the threshold, it indicates that the data privacy risk is high and further desensitization is required.

[0036] The overall privacy leakage risk coefficient is: (11) In specific implementation . According to the privacy leakage coefficient, configure relevant parameters for subsequent dynamic desensitization, as shown in the following table Table 2

[0037] Step 4, Based on the hybrid execution framework of static and dynamic desensitization, desensitize the file to be desensitized according to the privacy leakage risk coefficient to generate a desensitized file In specific implementation, as Figure 5 shown, based on the hybrid execution framework of static plus and dynamic desensitization. The present invention proposes a hybrid execution framework based on deep coupling of static-dynamic desensitization. Static desensitization layer: Use asymmetric encryption RSA and hash transformation to ensure the privacy security of display identifiers; Dynamic desensitization layer: Based on data real-time query protection of k-anonymity, differential privacy, and generative adversarial network, the hybrid strategy forms multiple defenses to resist background attacks and knowledge attacks.

[0038] Step 4.1, Static desensitization (1) Masking: Uniformly replace some content of the original data with general characters, so that only part of the sensitive data is made public.

[0039] (2) Hash transformation: It converts the original data (such as passwords, ID numbers, etc.) into a fixed-length hash value through a hash algorithm. Even if the hash value is leaked, the original data cannot be reversely restored.

[0040] (3) Asymmetric encryption RSA: Combine a set of large prime numbers with other parameters to generate a public key and a private key, encrypt data with the public key, and decrypt data with the private key to achieve secure transmission of information.

[0041] In specific implementation, the present invention automatically identifies the mobile phone number field of the patient, and masks the middle 4 digits with " ", and keeps the rest public, in the form of " "; adopt the SHA256 algorithm, automatically identify the ID number field, and perform hash transformation to protect personal identity; adopt asymmetric encryption, automatically identify the hospitalization ID number for asymmetric encryption to achieve privacy protection. The specific desensitization results are shown in Table 3.

[0042] Table 3

[0043] Step 4.2, Dynamic desensitization (1) : Under the framework of , each record in the dataset is the same as at least k - 1 other records in key attributes (such as age, gender, postal code, etc.).

[0044] In this specific implementation, the 2-anonymous data of some data in the Adult dataset is shown in Table 4: Table 4

[0045] (2) Differential privacy: Make it impossible for an attacker to infer any sensitive information of a record even if they know all records except one.

[0046] In specific implementation, after differential privacy is performed on the Adult dataset, it is shown in Table 5: Table 5

[0047] (3) Data synthesis: Simulate the statistical patterns and relationships in real data, but do not directly point to any "real" person, and resist re-identification attacks.

[0048] In specific implementation, after the Adult dataset is desensitized by a generative adversarial network, it is shown in Table 6: Table 6

[0049] Step 5, Evaluate the desensitized file based on the privacy and utility dual-evaluation model. If the evaluation result does not meet the preset conditions, return to Step 4 until the preset conditions are met; Specific implementation, the specific process of evaluating the desensitization result based on the privacy-utility dual evaluation model is as follows: Step 5.1, Multi-dimensional utility evaluation To evaluate the machine learning performance of the synthetic data, the present invention uses a decision tree classifier, polynomial regression, and a multi-layer perceptron to evaluate the original data and the synthetic data. The specific process is as follows: The model is trained using the original data and the desensitized data respectively, and the difference in accuracy, AUC, and F1 score between the original data and the desensitized data is used as a measure.

[0050] In specific implementation, the difference in indicators between the original data and the desensitized data in the desensitized data utility evaluation is shown in Table 7: Table 7

[0051] Step 5.2, Privacy leakage risk quantification model The present invention proposes a privacy leakage quantification model based on the Distance to Closest Rank (DCR) and the Nearest Neighbor Distance Ratio (NNDR) to achieve visual warning of the attack risk.

[0052] DCR (original data and desensitized data): The 5% quantile value of the minimum distance between the original data and the desensitized data. The smaller the value, the closer some synthetic data is to the real data, indicating a high leakage risk.

[0053] DCR (between original data): The 5% quantile value of the minimum distance within the original data, reflecting the density of the original data itself. A small value indicates that there are dense clusters in the original data and it is easy to be attacked.

[0054] DCR (between desensitized data): The 5% quantile value of the minimum distance within the desensitized data, reflecting the density of the desensitized data itself. A small value indicates that there are dense clusters in the desensitized data and it is easy to be attacked.

[0055] NNDR (real data and desensitized data): The 5% quantile value of the nearest neighbor ratio between the original data and the desensitized data, measuring the similarity of the local structure across datasets. A value approaching 1 indicates a high risk of structure replication.

[0056] NNDR (between real data): The 5% quantile value of the nearest neighbor ratio within the original data, reflecting the uniqueness of the local structure. A low value indicates the existence of repeated patterns and it is easy to be inferred.

[0057] NNDR (between desensitized data): The 5% quantile value of the nearest neighbor ratio within the desensitized data, reflecting the uniqueness of the local structure. A low value indicates the existence of repeated patterns and it is easy to be inferred. In specific implementation, the privacy evaluation of the desensitized data is shown in Table 8: Table 8

[0058] Step 6: Output the desensitized file and the final evaluation result to the client.

[0059] In specific implementation, after data desensitization is completed, the system returns the processed data and the data effect evaluation result to the user for data publishing and sharing. At this stage, the user can obtain a protected dataset, ensuring that no sensitive information is leaked, which helps to protect personal privacy and data security while promoting data sharing and utilization.

[0060] The multi-strategy data desensitization method provided in this embodiment dynamically grades combined fields by means of information entropy and k-means clustering algorithm, evaluates the privacy leakage risk of the data to be published considering the number of records, attribute fields, and usage scenarios, dynamically selects desensitization parameters according to the desensitization strategy, and at the same time combines static desensitization technology to automatically identify and desensitize quasi-identifiers, and conducts utility evaluation and privacy evaluation on the desensitized data, thereby helping data publishers to achieve more accurate and intelligent data desensitization, and improving the desensitization efficiency and security.

[0061] Aiming at the limitations of traditional data desensitization technology, the present invention has made significant improvements in terms of privacy protection, flexibility, intelligence, and strategy diversity, and specifically solves the following key problems: 1. Solve the problem of privacy leakage of combined attributes Traditional technologies only perform static classification and grading based on single attributes, ignoring the sensitivity of multi-attribute combinations, resulting in a privacy leakage risk after data publishing. The present invention realizes the dynamic grading of combined attributes based on information entropy and k-means clustering, avoiding privacy leakage caused by attribute correlation.

[0062] 2. Solve the problem of balancing user needs and data utility The traditional desensitization process is fixedly configured by the system, and data publishers cannot flexibly select the data features or desensitization strategies to be retained according to the actual scenario, resulting in reduced data utility or insufficient flexibility. The present invention allows data publishers to independently select the data fields and desensitization strategies to be published through the interaction interface, taking into account both data availability and privacy requirements.

[0063] 3. Solve the problem of insufficient automation of risk assessment and desensitization decision-making Traditional methods rely on manual experience to judge whether a dataset needs to be desensitized and select parameters, which is inefficient and error-prone. The present invention realizes intelligent dynamic data desensitization: based on dynamic factors such as data labels, usage scenarios, and record scales, automatically calculate the privacy risk score, and determine the necessity of desensitization and the noise scale through thresholds to achieve accurate decision-making.

[0064] 4. Solve the problems of single desensitization strategy and scenario adaptability Existing technologies usually only support a small number of desensitization methods (such as only static encryption or dynamic masking), making it difficult to meet the requirements of complex scenarios (such as dynamic desensitization for real-time queries and static encryption for long-term storage). The present invention integrates 6 desensitization methods (3 static + 3 dynamic), covering diverse scenarios: Static desensitization: applicable to stored data (such as asymmetric encryption to ensure long-term security); Dynamic desensitization: applicable to real-time interactions (such as K-anonymity to protect query privacy and generative adversarial networks to generate synthetic data).

[0065] Through the above improvements, the present invention provides a more refined and intelligent solution to the contradiction between data sharing and privacy protection, meeting the dual requirements of data security and value release in the digital age.

[0066] Specifically, compared with the existing technology, the improvement effects of the present invention are reflected in: 1. Multi-attribute combination dynamic grading: Before constructing a multi-strategy intelligent dynamic desensitization system, based on the traditional single-attribute static classification and grading, the security level of combined attributes is dynamically updated. Because even if a single attribute itself is not sensitive, its sensitivity significantly increases when combined. Multi-attribute combination dynamic grading improves the data privacy protection ability.

[0067] 2. Personalized selection of desensitized data: The data publisher uploads the desensitized data table to be published. The system automatically reads the attribute fields and returns them to the user interface through a multi-check dropdown form. The data publisher selects the data features to be published and the desensitization strategy according to personal preferences, performs desensitization on the data, and finally returns the desensitized data table to the user.

[0068] 3. Intelligent dynamic data desensitization: The desensitization system can integrate factors such as data labels, usage scenarios, the number of records, and the number of features to conduct a privacy risk assessment on the data to be published, and determine whether the data set needs to be desensitized through a threshold, and select an appropriate noise scale to enhance data availability.

[0069] 4. Support for multi-strategy data desensitization: The desensitization system adopts 6 desensitization methods, among which there are 3 static desensitization methods, namely: masking, hash transformation, and asymmetric encryption; there are 3 dynamic desensitization methods, namely: k-anonymity, differential privacy, and generative adversarial networks. Multi-strategy support meets the requirements of complex scenarios such as finance, healthcare, and cloud computing.

[0070] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof.

[0071] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A multi-strategy data desensitization method, characterized in that, include: Step 1, obtain the file to be desensitized and select parameters, wherein the parameters include the total amount of data in the file to be desensitized, the number of selected records, the total amount of attribute types and the number of selected attributes; Step 2: Dynamically grading each attribute based on information entropy and k-means clustering to obtain the security level of each attribute; Step 3: Perform a multi-factor privacy leakage risk assessment based on the selected parameters and the security level of each attribute to obtain a privacy leakage risk coefficient; Step 4: Based on the hybrid execution framework of static and dynamic desensitization, desensitize the desensitized file according to the privacy leakage risk coefficient to generate a desensitized file; Step 5: Evaluate the desensitized file based on the privacy and utility dual evaluation model. If the evaluation result does not meet the preset conditions, return to step 4 until it meets the preset conditions. Step 6: Output the desensitized file and the final evaluation results to the client.

2. The method according to claim 1, wherein The step 1 specifically includes: Step 1.1, obtain the files to be desensitized uploaded by the user, select the number of records, application scenarios and desensitization methods; Step 1.2, set the total amount of data in the file to be desensitized as , the number of selected records is , the total amount of attribute types , select attributes.

3. The method according to claim 2, wherein The step 2 specifically includes: Step 2.1, after one-hot encoding each attribute, calculate the mutual information between the two attributes and normalize them to form an attribute correlation matrix, where the expression of the mutual information is: Among them, and are the marginal entropies, and are the conditional entropies, is the attribute and is the joint entropy of; The normalized expression is: Among them, is the mutual information; The expression of the attribute correlation matrix is: Among them, represents the joint entropy normalization value of the th and th attribute fields, that is, the correlation coefficient; Step 2.2, perform K-means clustering on the attributes according to the attribute correlation matrix, use the silhouette coefficient to find the optimal number of clusters, and obtain the security level of each attribute based on the cluster and correlation coefficient, as well as the security level of the attribute data label.

4. The method according to claim 3, wherein The step 3 specifically includes: Step 3.1, use the ratio between the total amount of data N in the file to be desensitized and the selected number of records n as the record selection quantity weight ; Step 3.2: Use the ratio between the total number M of attribute types and the selected number m of attributes as the weight of the field selection quantity ; Step 3.3, obtain the recorded types of attributes and supported types of application scenarios, obtain the security level of each attribute, denoted as ; obtain the privacy level mark established for the corresponding application scenario, denoted as , the th field's th type of operation's corresponding operation risk weight, denoted as , taking the attributes as row vectors and the application scenarios as column vectors, establish the first operation risk matrix ; Step 3.4, according to the data usage requirements, select the operation risk weight corresponding to the th operation of the th attribute field, construct the second operation risk matrix in this scenario , and fill the un-involved requirements with . Sum the rows of the second operation risk matrix and perform normalization to obtain the field operation risk weight ; Step 3.

5. Among the selected m attributes, there are continuous data, categorical data. Calculate the first distance between every two continuous data: Among them, is the distance length size of the numerical range, indicating continuous data; Step 3.6, calculate the second distance between the two data in the classification data: Among them, is the number of types of attribute values included in the attribute set, represents a discrete record; Step 3.7, for the selected n records, consecutive data, categorical data, and classes of global combined records, calculate the total distance between the records: ; Step 3.8, calculate the relevance weight of the records according to the total distance : ; Step 3.9, calculate the privacy leakage risk coefficient based on the record selection quantity weight, field selection quantity weight, field operation risk weight and record relevance weight: 。 5. The method according to claim 4, characterized in that, The step 4 specifically includes: Step 4.1, masking, hashing and asymmetric encryption of continuous data in the file to be desensitized are performed in the static desensitization layer of the hybrid execution framework of static and dynamic desensitization; Step 4.2: Based on k-anonymity, differential privacy and adversarial generative network, the classified data in the desensitized file is desensitized in the dynamic desensitization layer of the hybrid execution framework of static and dynamic desensitization; Step 4.3, the outputs of the static desensitization layer and the dynamic desensitization layer are formed into a desensitization file.

6. The method according to claim 5, wherein The step 5 specifically includes: Step 5.1, using decision tree classifier, polynomial regression and multi-layer perceptron to evaluate the utility between the desensitized file and the file to be desensitized; Step 5.2, evaluating the privacy between the desensitized file and the file to be desensitized based on the minimum distance quantile value and the nearest neighbor ratio, and forming an evaluation result; Step 5.3: If the evaluation result does not meet the preset conditions, return to step 4 until it meets the preset conditions.

Citation Information

Patent Citations

  • Effect evaluation method during data desensitization

    CN114372271A

  • Data desensitization method, device and equipment and readable storage medium

    CN116933286A

  • Archive management method and system based on big data

    CN118551414A

  • Multi-dimensional data authority management and privacy protection method for electric power information network

    CN119538276A

  • Classification estimating system and classification estimating program

    US20120209134A1