Data processing method, system and equipment based on privacy and utility
By dynamically allocating privacy protection and data utility, using a random forest model and multi-objective optimization algorithm, combined with additive homomorphic encryption and federated learning, the problem of balancing protection and utility in privacy computing is solved, and a dynamic balance is achieved in different scenarios.
Patent Information
- Application Number
- CN202510656119.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-05
AI Technical Summary
Existing privacy computing technologies find it difficult to strike a balance between protecting privacy and maintaining data utility. Traditional methods often ignore the maintenance of data utility, and static allocation methods cannot adapt to the differentiated needs of different scenarios.
By dynamically allocating the degree of privacy protection and the degree of data utility retention, using random forest models and multi-objective optimization algorithms to screen desensitizing features, and combining additive homomorphic encryption and federated learning, a dynamic balance of data processing is achieved.
Dynamically balance privacy protection and data utility in different scenarios, effectively protect personal sensitive information, while retaining the valid information of the data to the greatest extent possible to meet strict privacy protection requirements.
Smart Images

Figure CN120597313A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data security processing technology, and in particular to a data processing method, system and device based on privacy and utility. Background Art
[0002] When data involves sensitive personal information, such as medical records, financial data, personal identity information, etc., and at the same time needs to be analyzed and mined to obtain valuable information to achieve certain specific uses, such as business intelligence analysis, scientific research, etc., it is necessary to perform privacy calculations on this data to retain the effective information of the data to the greatest extent while protecting the privacy of the data subject.
[0003] However, existing privacy-preserving computing technologies, such as differential privacy and homomorphic encryption, often face the following challenges in practical applications: Traditional approaches employ a single-objective optimization approach, focusing solely on improving privacy protection while neglecting the preservation of data utility. This results in data that has undergone privacy processing failing to meet the demands of real-world applications. Furthermore, existing technologies impose a fixed allocation between the degree of privacy protection and the degree of data utility preservation. This static allocation fails to dynamically adapt to the varying demands for privacy protection and data utility in different scenarios.
[0004] Therefore, existing privacy computing technologies are difficult to achieve a balance between privacy protection and data utility in practical applications. Summary of the Invention
[0005] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide a data processing method, system and device based on privacy and utility, which dynamically allocates the degree of privacy protection and the degree of data utility retention of data, so as to dynamically balance the degree of privacy protection and the degree of data utility retention in different scenarios.
[0006] In order to solve the above problems, the present invention is implemented according to the following scheme:
[0007] Provides a privacy- and utility-based approach to data processing, including:
[0008] Get the data to be processed and its corresponding data type;
[0009] Performing desensitization processing on the data to be processed to obtain multiple desensitization features of the data to be processed;
[0010] Determining, based on the desensitized feature, a dynamic parameter for indicating a privacy protection degree and a utility retention degree of the desensitized feature;
[0011] The desensitizing features are processed according to the data type and dynamic parameters to obtain target data.
[0012] Compared with the existing technology, the beneficial effects of the data processing method based on privacy and utility of the present invention are as follows: by determining dynamic parameters, the degree of privacy protection and the degree of data utility retention can be dynamically allocated according to the needs of different scenarios, which changes the limitations of the traditional static allocation method. It can not only effectively protect personal sensitive information in the data and meet strict privacy protection requirements, but also retain the effective information of the data to the greatest extent in different application scenarios, thereby achieving a dynamic balance between privacy protection and data utility.
[0013] Optionally, desensitizing the data to be processed to obtain multiple desensitizing features of the data to be processed, including:
[0014] Identifying multiple sensitive fields in the data to be processed, where the sensitive fields have corresponding sensitivity labels;
[0015] According to the sensitivity labels, multiple sensitive fields are processed respectively to obtain initial features for representing the sensitive fields;
[0016] A random forest model is used to screen out desensitization features that meet the first preset conditions from the initial features.
[0017] Optionally, based on the sensitivity labels, multiple sensitive fields are processed separately to obtain initial features for representing the sensitive fields, including:
[0018] Determining the sensitivity type of the sensitive label;
[0019] When the sensitive type is high sensitivity, converting the character code of the sensitive field into a byte sequence, adding a random number to the byte sequence to obtain a random byte sequence, and using a hash algorithm to generate an initial feature of a hash value for the random byte sequence;
[0020] When the sensitive type is low sensitivity, the sensitive field is interval-generalized to obtain an initial feature in the form of an interval.
[0021] Optionally, a random forest model is used to screen out desensitization features that meet the first preset condition from the initial features, including:
[0022] Acquiring historical data, and training a random forest model based on the historical data to obtain a privacy-random forest model for predicting the degree of data privacy protection and a utility-random forest model for predicting the degree of data utility retention;
[0023] Determining, according to the privacy-random forest model, a privacy score for quantifying the initial features on a privacy dimension;
[0024] Determining, according to the utility-random forest model, a utility score for quantifying the initial feature in the utility dimension;
[0025] When the ratio of the utility score to the privacy score is greater than the first preset condition, the corresponding initial feature is eliminated;
[0026] When the ratio of the utility score to the privacy score is less than or equal to the first preset condition, the corresponding initial feature is retained and the retained initial feature is used as the desensitized feature.
[0027] Optionally, determining, based on the desensitization feature, a dynamic parameter for indicating the privacy protection degree and utility retention degree of the desensitization feature includes:
[0028] Determining, based on the desensitization feature, an amount of information leakage of the sensitive field caused by the desensitization feature and a utility loss caused by the desensitization feature offsetting the sensitive field;
[0029] determining an objective function according to the information leakage amount and the utility loss;
[0030] Using a multi-objective optimization algorithm, determining a parameter set including a plurality of initial parameters, wherein the initial parameters satisfy the objective function;
[0031] An initial parameter in the parameter set that meets a second preset condition is determined as the dynamic parameter.
[0032] Optionally, the data type includes a precise statistical numerical type;
[0033] According to the data type and dynamic parameters, the desensitization features are processed to obtain target data, including:
[0034] Processing the desensitized features according to the dynamic parameters to obtain target features;
[0035] Performing additive homomorphic encryption on the target feature to convert it into ciphertext;
[0036] The target data is obtained by summing the ciphertexts.
[0037] Optionally, the data type includes a joint training numerical type;
[0038] Processing the desensitization feature according to the data type and the dynamic parameter to obtain target data includes:
[0039] Obtaining multiple model parameters for processing the desensitization features;
[0040] generating target noise according to the dynamic parameters;
[0041] Injecting the target noise into a plurality of model parameters respectively to obtain a plurality of noise model parameters;
[0042] Processing the desensitized features according to the dynamic parameters to obtain target features;
[0043] Optimize multiple noise model parameters through communication to obtain gradient model parameters;
[0044] The target features are processed according to the gradient model parameters to obtain the target data.
[0045] Optionally, multiple noise model parameters are optimized for communication to obtain gradient model parameters, including:
[0046] Determine the accuracy of each noise model parameter when training the model;
[0047] Sorting multiple noise model parameters according to their corresponding accuracies to obtain a sorting result;
[0048] According to the sorting result, the gradient model parameters are screened out from the multiple noise model parameters.
[0049] A data processing system based on privacy and utility is also provided, which is applied to the data processing method, including:
[0050] Application interface module, used to obtain the data to be processed and its corresponding data type;
[0051] A data preprocessing module is used to perform desensitization processing on the data to be processed to obtain multiple desensitization features of the data to be processed;
[0052] a joint optimization module, configured to determine, based on the desensitized feature, a dynamic parameter indicating a degree of privacy protection and a degree of utility retention of the desensitized feature;
[0053] The processing module is used to process the desensitizing features according to the data type and dynamic parameters to obtain target data.
[0054] A computer device is also provided, characterized in that the computer device includes a processor and a memory, the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the data processing method. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 Flowchart of the data processing method of the present invention. DETAILED DESCRIPTION
[0056] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0057] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances.
[0058] See also Figure 1 As shown, a data processing method based on privacy and utility of the present invention includes:
[0059] S1: Obtain the data to be processed and its corresponding data type; in one embodiment of the present invention, the data to be processed refers to data including personal sensitive information, and these data need to be privacy-calculated to ensure that multiple participants will not leak the personal sensitive information included in the data during the transmission of the data to be processed, while ensuring the information integrity (data utility) included in the data to be processed received by multiple participants.
[0060] The data to be processed includes file extensions used to indicate formats, such as .csv for representing a plain text file format with comma-separated fields, .json for a lightweight data exchange format, and .jpg for an image format. The data to be processed can be divided into structured data, semi-structured data, and unstructured data based on the file extensions. For example, the data to be processed with file extensions such as .csv, .xlsx, and .sql are structured data, the data to be processed with file extensions such as .json, .xml, and .log are semi-structured data, and the data to be processed with file extensions such as .txt, .pdf, .jpg, and .mp4 are unstructured data.
[0061] When the data to be processed is structured data, the data type corresponding to the data to be processed is a precise statistical numerical type; when the data to be processed is semi-structured data or unstructured data, the data type of the data to be processed is a joint training data type.
[0062] During the process of multiple parties transmitting data to be processed, format errors, data corruption, etc. may occur. If the data type corresponding to the data to be processed is determined directly based on the file extension included in the data to be processed, these situations will lead to data type errors.
[0063] Therefore, in one embodiment of the present invention, before obtaining the data type corresponding to the data to be processed, the data processing method based on privacy and utility of the present invention further includes:
[0064] If the file extension belongs to structured data, the data to be processed is parsed. If a table is obtained through parsing (successful parsing), the file extension of the data to be processed is correct and the data type is a precise statistical numerical type. If a table cannot be obtained through parsing (failed parsing), the file extension of the data to be processed is incorrect and the file extension is modified to semi-structured data or unstructured data.
[0065] If the file extension belongs to semi-structured data or unstructured data, the data to be processed whose file extension belongs to semi-structured data will be parsed. If the hierarchical label is obtained through parsing (successful parsing), the file extension of the data to be processed is correct. If the hierarchical label cannot be obtained through parsing (failed parsing), the file extension of the data to be processed is incorrect, and the file extension will be modified to unstructured data.
[0066] In one embodiment of the present invention, when obtaining the data to be processed and its corresponding data type, it also includes obtaining the privacy level corresponding to the data to be processed, and the privacy level includes a high privacy level, a medium privacy level, and a low privacy level. The data to be processed of different privacy levels have different degrees of privacy protection.
[0067] S2: Desensitize the data to be processed to obtain multiple desensitization features of the data to be processed, including:
[0068] First, multiple sensitive fields in the data to be processed are identified, and the sensitive fields have corresponding sensitivity labels. Based on the sensitivity labels, the multiple sensitive fields are processed separately to obtain initial features used to represent the sensitive fields. Finally, a random forest model is used to screen out desensitizing features that meet the first preset conditions from the initial features to filter out features that are not highly relevant but are prone to leaking personal sensitive information.
[0069] In one embodiment of the present invention, multiple sensitive fields in the data to be processed are identified, including: using a pre-stored "sensitive label: sensitive field" to identify sensitive fields with corresponding sensitive labels through regular expressions, for example, using the regular expression \d{17}[\dX] to identify ID card numbers, and the regular expression 1[3-9]\d{9} to identify mobile phone numbers, and combining semantic analysis in different application scenarios to detect sensitive words in unstructured data (such as "diagnosis result: positive", "bank card number: 6225880123456789", "transfer amount: 5,000 yuan").
[0070] In one embodiment of the present invention, multiple sensitive fields are processed separately according to the sensitivity labels to obtain initial features for representing the sensitive fields, including:
[0071] First, determine the sensitivity type of the sensitive label. Sensitive types include high sensitivity and low sensitivity. For example, sensitive labels such as diagnosis results and bank card numbers have a high sensitivity type, while sensitive labels such as age and transfer amount have a low sensitivity type.
[0072] When the sensitive type is high sensitivity, the character encoding of the sensitive field is converted into a byte sequence, a random number is added to the byte sequence to obtain a random byte sequence, and a hash algorithm is used to generate the initial feature of the hash value for the random byte sequence. Specifically, the UTF-8 character encoding is converted into a byte sequence, and the irreversible SHA-256 hash algorithm is used to encrypt it to obtain the initial feature. For example, "diagnosis result: positive" is converted into "diagnosis result: 6b86b273ff34"; when the sensitive type is low sensitivity, the sensitive field is interval-generalized to obtain the initial feature in the form of an interval. For example, the sensitive field "28 years old" is converted into the initial feature of "[25-30] years old".
[0073] In one embodiment of the present invention, a random forest model is used to screen out desensitization features that meet the first preset condition from the initial features, including:
[0074] First, historical data is obtained and a random forest model is trained based on the historical data to obtain a privacy-random forest model for predicting the degree of data privacy protection and a utility-random forest model for predicting the degree of data utility retention. The historical data includes the features obtained after privacy calculation of the data and their corresponding privacy scores and utility scores.
[0075] Next, based on the privacy-random forest model, a privacy score is determined for quantifying the initial features in the privacy dimension; based on the utility-random forest model, a utility score is determined for quantifying the initial features in the utility dimension.
[0076] When the ratio of the utility score to the privacy score is greater than the first preset condition, the feature is not highly relevant to the business objective and is prone to leaking personal sensitive information, so the corresponding initial feature will be eliminated; when the ratio of the utility score to the privacy score is less than or equal to the first preset condition, the corresponding initial feature will be retained and used as a desensitized feature. For example, when the first preset condition is 2.5 points, when the utility score / privacy score of the initial feature is greater than 2.5, the initial feature will be eliminated.
[0077] S3: Based on the desensitized features, determine dynamic parameters for indicating the privacy protection degree and utility retention degree of the desensitized features, including:
[0078] First, based on the desensitizing features, the amount of information leakage of the desensitizing features to the sensitive fields and the utility loss of the desensitizing features offsetting the sensitive fields are determined; based on the amount of information leakage and the utility loss, the objective function is determined; then, a multi-objective optimization algorithm is used to determine a parameter set including multiple initial parameters, and the initial parameters satisfy the objective function; finally, the initial parameters in the parameter set that meet the second preset condition are determined as dynamic parameters.
[0079] In one embodiment of the present invention, the amount of information leakage is determined by mutual information I(S;Y)=H(S)-H(S|Y), wherein I(S;Y) is the amount of information leakage, H() is the information entropy, Y is the desensitizing feature, and S is the sensitive field. The information leakage amount is used to measure the amount of information leaked from the desensitizing feature Y to the sensitive field S.
[0080] H(S) corresponds to the information entropy of the sensitive field S, reflecting the uncertainty of the sensitive field S. For example, if the sensitive label of sensitive field S is "diagnosis result," the possible values of sensitive field S are {negative, positive}. H(S) measures the uniformity of the distribution of these two values. H(S|Y) (the conditional entropy after knowing the desensitized features) corresponds to the residual uncertainty of sensitive field S inferred from the desensitized features Y. For example, if the desensitized feature is "[25-30] years old," it is impossible to accurately infer the uncertainty of the original age "28 years old." I(S;Y) measures the information about sensitive field S that is still contained in the desensitized feature Y. A larger value indicates a higher risk of privacy leakage.
[0081] In one embodiment of the present invention, KL divergence is used to determine the utility loss D KL (sensitive field||desensitized feature), KL divergence is used to measure the asymmetric difference between two probability distributions P and Q. Therefore, the utility loss D caused by the desensitized feature offsetting the sensitive field can be obtained by KL divergence. KL (sensitive field||desensitizing feature).
[0082] In one embodiment of the present invention, the objective function is min[0.75*λ*I(S;Y)+1.2*(1-λ)DKL (sensitive field||desensitized feature), where I(S;Y) is the amount of information leakage, D KL (sensitive field||desensitized feature) is the utility loss, λ is the privacy weight used to measure the degree of privacy protection, and its value ranges from 0.1 to 1.
[0083] In one embodiment of the present invention, the multi-objective optimization algorithm is the NSGA-II algorithm. When the multi-objective optimization algorithm is used to determine the parameter set, the privacy weight needs to be dynamically adjusted. The steps of executing the algorithm include: initializing the population (randomly generating an initial population P, the size of the population P is N, where each individual represents a configuration of privacy protection and data utility, including sensitive fields and desensitization features), and calculating the objective function value for each individual using the objective function, specifically by using the information leakage amount I(S; Y) and the utility loss D KL (Sensitive field||Desensitized feature) Perform non-dominated sorting and cross-mutation on each individual to generate a new population, and iteratively optimize to finally obtain a parameter set (Pareto optimal solution set).
[0084] In one embodiment of the present invention, the second preset condition includes a minimum information leakage of 0.1 bits, a solution set that satisfies the constraint of I(S; Y) < 0.1 bits is selected from the parameter set, and the utility loss D is selected from the solution set. KL The minimum initial parameter of (sensitive field || desensitized feature) is used as a dynamic parameter. The dynamic parameter includes a privacy weight used to measure the degree of data privacy protection of each desensitized feature and a utility weight used to measure the degree of data utility retention. The initial parameter is the privacy weight, where the sum of the privacy weight and the utility weight is 1. Therefore, after determining the privacy hit, the utility weight can be obtained through the privacy weight.
[0085] In one embodiment of the present invention, during the dynamic adjustment of the privacy weight, the adjustment range of the privacy weight is determined according to the privacy level of the data to be processed. For example, when the privacy level is a high privacy level, the adjustment range of the privacy weight is [0.8, 1]; when the privacy level is a medium privacy level, the adjustment range of the privacy weight is [0.4, 0.7]; when the privacy level is a low privacy level, the adjustment range of the privacy weight is [0.1, 0.3].
[0086] S4: Process the desensitized features according to the data type and dynamic parameters to obtain the target data.
[0087] In one embodiment of the present invention, the data type includes a precise statistical numerical type; and the desensitization feature is processed according to the data type and the dynamic parameter to obtain the target data, including:
[0088] First, the desensitized features are processed according to the dynamic parameters to obtain the target features; then, the target features are converted into ciphertexts by additive homomorphic encryption; finally, the ciphertexts are summed to obtain the target data.
[0089] In one embodiment of the present invention, the degree of data privacy protection of the desensitized feature is determined according to the privacy weight included in the dynamic parameters, and the degree of data utility retention of the desensitized feature is determined according to the utility weight included in the dynamic parameters. Based on the dynamic parameters, the privacy calculation of the desensitized feature is performed to obtain the target feature that balances the degree of data privacy protection and the degree of data utility retention.
[0090] When the data type is a precise statistical numerical type, the "encryption followed by calculation" calculation strategy is adopted. Specifically, multi-party secure computation (MPC) is used to encrypt the target features first, convert them into ciphertext, and then obtain the target data based on the sum of the ciphertexts to ensure that multiple participants can only receive the data itself and cannot obtain relevant information on how to determine the data. For example, the number of patients in each hospital is converted into ciphertext through additive homomorphic encryption, and the ciphertext is directly summed without decryption. Finally, only the sum result is decrypted to obtain the target data used to represent the number of patients, ensuring that a single hospital cannot obtain the specific data of other hospitals, such as the number of patients in which department.
[0091] In one embodiment of the present invention, the data type includes a joint training numerical type; and the desensitization feature is processed according to the data type and the dynamic parameter to obtain the target data, including:
[0092] First, multiple model parameters for processing desensitized features are obtained; target noise is generated according to the dynamic parameters; the target noise is injected into multiple model parameters respectively to obtain multiple noise model parameters; then, the desensitized features are processed according to the dynamic parameters to obtain target features; finally, the multiple noise model parameters are communication-optimized to obtain gradient model parameters; the target features are processed according to the gradient model parameters to obtain target data.
[0093] In one embodiment of the present invention, after obtaining the model parameters, the model parameters need to be encrypted to obtain an encrypted segment of the model parameters, and then the encrypted segment of the model parameters is processed; when the data type is a jointly trained numerical type, a federated learning (FL) framework is adopted. For example, when the image data has a file extension of .jpg or .mp4, a joint training model is required, so federated learning is used to perform privacy calculation on the desensitized features.
[0094] In one embodiment of the present invention, the model parameters are key parameters of the model that constitutes the privacy calculation of the desensitized features. The noise intensity of the target noise is determined according to the dynamic parameters. The calculation formula is as follows:
[0095]
[0096] Where ∈ is the noise intensity of the target noise, and λ is the privacy weight. The generated target noise is injected into the model parameters to obtain the noise model parameters. A larger privacy weight results in smaller target noise, which increases the accuracy of the model used for privacy-preserving computations, but weakens data privacy protection. A smaller privacy weight results in larger target noise, which increases the privacy protection of the model used for privacy-preserving computations, but reduces model accuracy. Privacy protection is achieved by adding noise to the model parameters, adjusting the noise intensity based on the privacy weight, thereby balancing the degree of data privacy protection and the degree of data utility preservation.
[0097] In one embodiment of the present invention, the degree of data privacy protection of the desensitized feature is determined according to the privacy weight included in the dynamic parameters, and the degree of data utility retention of the desensitized feature is determined according to the utility weight included in the dynamic parameters. Based on the dynamic parameters, the privacy calculation of the desensitized feature is performed to obtain the target feature that balances the degree of data privacy protection and the degree of data utility retention.
[0098] For example, multiple hospitals (participants) use CT image data to train models locally, and only upload encrypted fragments of the model parameters of the trained model to the central coordinator. The central coordinator injects Laplace noise (target noise) into the model parameters during aggregation, and then performs communication optimization to obtain gradient model parameters. The model used for privacy calculation of desensitized features is updated based on the gradient model parameters.
[0099] In one embodiment of the present invention, communication optimization is performed on multiple noise model parameters to obtain gradient model parameters, including:
[0100] First, the accuracy of each noise model parameter when training the model is determined; then, multiple noise model parameters are sorted according to their corresponding accuracies to obtain a sorting result; finally, based on the sorting result, the gradient model parameter is screened out from the multiple noise model parameters.
[0101] In a single model, multiple noise model parameters are included. After training the model with multiple noise model parameters, the accuracy corresponding to each noise model parameter is obtained. The noise model parameters are sorted according to the accuracy to form a gradient, and the absolute value of the gradient itself is used for judgment. A larger absolute value of the gradient means that the noise model parameter has a greater impact on the model update. The sorting result is to sort the noise model parameters from large to small according to the accuracy, indicating that the higher the noise model parameter is ranked, the greater the impact on the model update. Finally, among the sorting results, the noise model parameters of the top 10% important gradients are selected as gradient model parameters to achieve a single transmission data volume, and then only the gradients (gradient model parameters) that need to be retained have sufficient directional information, so that the model can still converge effectively, and can maintain model performance while reducing the consumption of a large amount of data, achieve communication optimization and reduce collaboration costs.
[0102] When the privacy level is high, the "calculation after encryption" strategy is adopted to ensure that multiple participants can only receive the data itself and cannot obtain relevant information on how to determine the data; when the privacy level is medium, the federated learning framework is adopted and the target noise is superimposed on the model parameters, and finally the model parameters are optimized for communication.
[0103] After obtaining the target data, generate a report that complies with regulations such as GDPR and HIPAA. You can use client.generate_audit_report(regulation="GDPR") to export a PDF document. The report includes the data obtained in each step, such as the data to be processed, data type, desensitization features, dynamic parameters, target features, and target data.
[0104] A data processing system based on privacy and utility of the present invention is applied to the above-mentioned data processing method, comprising:
[0105] Application interface module, used to obtain the data to be processed and its corresponding data type;
[0106] A data preprocessing module is used to perform desensitization processing on the data to be processed and obtain multiple desensitization features of the data to be processed;
[0107] A joint optimization module for determining, based on the desensitized features, dynamic parameters indicating the degree of privacy protection and utility retention of the desensitized features;
[0108] The processing module is used to process the desensitized features according to the data type and dynamic parameters to obtain the target data.
[0109] The present invention determines dynamic parameters including privacy weight and utility weight through a joint optimization module, so as to obtain target data that balances the degree of data privacy protection and the degree of data utility retention according to the dynamic parameters, and dynamically allocates the degree of privacy protection and the degree of data utility retention according to the needs of different scenarios. It changes the limitations of the traditional static allocation method, and can not only effectively protect personal sensitive information in the data and meet strict privacy protection requirements, but also retain the effective information of the data to the greatest extent in different application scenarios, thereby achieving a dynamic balance between privacy protection and data utility.
[0110] The present invention also provides a computer device, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the above-mentioned data processing method.
[0111] The processor can be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0112] The memory can be used to store the computer program or module, and the processor implements the various functions of the data processing method by running or executing the computer program or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0113] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A data processing method based on privacy and utility, characterized in that include: Get the data to be processed and its corresponding data type; Performing desensitization processing on the data to be processed to obtain multiple desensitization features of the data to be processed; Determining, based on the desensitized feature, a dynamic parameter for indicating a privacy protection degree and a utility retention degree of the desensitized feature; The desensitizing features are processed according to the data type and dynamic parameters to obtain target data.
2. A data processing method based on privacy and utility according to claim 1, characterized in that: Desensitizing the data to be processed to obtain multiple desensitizing features of the data to be processed, including: Identifying multiple sensitive fields in the data to be processed, where the sensitive fields have corresponding sensitivity labels; According to the sensitivity labels, multiple sensitive fields are processed respectively to obtain initial features for representing the sensitive fields; A random forest model is used to screen out desensitization features that meet the first preset conditions from the initial features.
3. A data processing method based on privacy and utility according to claim 2, characterized in that: According to the sensitivity labels, multiple sensitive fields are processed respectively to obtain initial features for representing the sensitive fields, including: Determining the sensitivity type of the sensitive label; When the sensitive type is high sensitivity, converting the character code of the sensitive field into a byte sequence, adding a random number to the byte sequence to obtain a random byte sequence, and using a hash algorithm to generate an initial feature of a hash value for the random byte sequence; When the sensitive type is low sensitivity, the sensitive field is interval-generalized to obtain an initial feature in the form of an interval.
4. A data processing method based on privacy and utility according to claim 3, characterized in that: A random forest model is used to screen out desensitization features that meet the first preset condition from the initial features, including: Acquiring historical data, and training a random forest model based on the historical data to obtain a privacy-random forest model for predicting the degree of data privacy protection and a utility-random forest model for predicting the degree of data utility retention; Determining, according to the privacy-random forest model, a privacy score for quantifying the initial features on a privacy dimension; Determining, according to the utility-random forest model, a utility score for quantifying the initial feature in the utility dimension; When the ratio of the utility score to the privacy score is greater than the first preset condition, the corresponding initial feature is eliminated; When the ratio of the utility score to the privacy score is less than or equal to the first preset condition, the corresponding initial feature is retained and the retained initial feature is used as the desensitized feature.
5. The data processing method based on privacy and utility according to claim 1, characterized in that: Determining, based on the desensitized feature, a dynamic parameter for indicating a privacy protection degree and a utility retention degree of the desensitized feature, including: Determining, based on the desensitization feature, an amount of information leakage of the sensitive field caused by the desensitization feature and a utility loss caused by the desensitization feature offsetting the sensitive field; determining an objective function according to the information leakage amount and the utility loss; Using a multi-objective optimization algorithm, determining a parameter set including a plurality of initial parameters, wherein the initial parameters satisfy the objective function; An initial parameter in the parameter set that meets a second preset condition is determined as the dynamic parameter.
6. The data processing method based on privacy and utility according to claim 1, characterized in that: The data type includes a precise statistical numerical type; According to the data type and dynamic parameters, the desensitization features are processed to obtain target data, including: Processing the desensitized features according to the dynamic parameters to obtain target features; Performing additive homomorphic encryption on the target feature to convert it into ciphertext; The target data is obtained by summing the ciphertexts.
7. The data processing method based on privacy and utility according to claim 1, characterized in that: The data type includes a joint training numerical type; Processing the desensitization feature according to the data type and the dynamic parameter to obtain target data includes: Obtaining multiple model parameters for processing the desensitization features; generating target noise according to the dynamic parameters; Injecting the target noise into a plurality of model parameters respectively to obtain a plurality of noise model parameters; Processing the desensitized features according to the dynamic parameters to obtain target features; Optimize multiple noise model parameters through communication to obtain gradient model parameters; The target features are processed according to the gradient model parameters to obtain the target data.
8. The data processing method based on privacy and utility according to claim 7, characterized in that: Optimize multiple noise model parameters to obtain gradient model parameters, including: Determine the accuracy of each noise model parameter when training the model; Sorting multiple noise model parameters according to their corresponding accuracies to obtain a sorting result; According to the sorting result, the gradient model parameters are screened out from the multiple noise model parameters.
9. A data processing system based on privacy and utility, applied to the data processing method according to claims 1-8, characterized in that: include: Application interface module, used to obtain the data to be processed and its corresponding data type; A data preprocessing module is used to perform desensitization processing on the data to be processed to obtain multiple desensitization features of the data to be processed; a joint optimization module, configured to determine, based on the desensitized feature, a dynamic parameter indicating a degree of privacy protection and a degree of utility retention of the desensitized feature; The processing module is used to process the desensitizing features according to the data type and dynamic parameters to obtain target data.
10. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement a data processing method as described in any one of claims 1 to 8.
Citation Information
Cited By
Database privacy network security protection method fusing AI dynamic desensitization
CN121302435A