A method and system for intelligent protection of user data based on differential privacy
Through adaptive privacy budget and multi-strategy fusion data perturbation strategy, combined with an improved differential privacy algorithm, the difficulties in parameter setting and large data loss in user health data protection in existing technologies are solved, and a balance between flexible privacy protection and data availability is achieved.
Patent Information
- Application Number
- CN202411473832.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-22
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-10-22
AI Technical Summary
Existing technologies have difficulty in parameter setting in user health data protection, resulting in large data loss, poor flexibility, difficulty in dealing with attackers with different background knowledge, and high computing resource requirements.
An intelligent user data protection method based on differential privacy is adopted. Through an adaptive privacy budget allocation mechanism and a multi-strategy fusion data perturbation strategy, perturbation data is generated. An improved differential privacy algorithm is used to evaluate the privacy leakage risk and data availability, and the privacy budget and perturbation strategy are iteratively adjusted.
It achieves flexible privacy protection, effectively reduces data loss, improves data availability and the flexibility of privacy protection, and can effectively protect user data under different background knowledge.
Smart Images

Figure CN119720263B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of data protection, and specifically relates to a user data intelligent protection method and system based on differential privacy. BACKGROUND
[0002] The method for intelligently protecting user data mainly involves multiple fields of information technology, privacy protection, data encryption and security protocols. With the development of the Internet, mobile computing and the Internet of Things, personal health data has become increasingly easy to obtain and share, which has also brought challenges to data privacy and security. Among them, data privacy protection is a key technology to ensure that user health data is not accessed, used or disclosed without authorization. It involves various aspects such as data collection, storage, use and sharing, aiming to protect the personal privacy rights and interests of users. In the early stage, the protection of user health data mainly relied on traditional data encryption and data desensitization technologies. Although these technologies can protect the privacy of data to a certain extent, their limitations gradually appear in the face of increasingly complex data usage scenarios and privacy leakage risks.
[0003] A distributed learning privacy protection method based on differential privacy is disclosed in a patent with the authorization announcement number CN111814189B, which includes: being applied to n user nodes in a network, each user having his own independently distributed group of data samples, and including the following steps: initialization stage; user node local learning stage; user node obtaining neighbor node information and updating stage; noise disturbance stage; broadcast stage. This technical solution can solve the privacy protection problem in current distributed learning, so that the user node updates its own local parameters through the neighbor node, and sends the noise-processed parameters to the neighbor node, thereby protecting the personal sensitive data of the user in a decentralized network environment.
[0004] A patent with publication number CN110598447B discloses a t-closeness privacy protection method satisfying epsilon-differential privacy, which comprises: preprocessing original data (hospital patient data including name, age, contact information, place of origin, health status), establishing a quasi-identifier and sensitive attribute association QIS data table, dividing the data table into a set of buckets according to the hierarchical tree of sensitive attributes in the data table, adding a dynamic programming algorithm to partition the data table, and finally generating an anonymous data table according to the partitioning result. In view of the problems of the traditional t-closeness privacy protection model, such as being unable to resist attackers with certain background knowledge, excessively relying on sensitive attribute distribution, and differential privacy possibly leading to loss of value of published data, a new distance function is introduced to optimize t-closeness, and the threshold t is optimized and adjusted, realizing the combination of t-closeness and differential privacy standard, and being able to reduce data loss as much as possible while ensuring data publishing utility and protecting the relationship between data.
[0005] The above prior art has the following problems: parameter setting is difficult and data loss is large; flexibility is poor, and there may be certain limitations when dealing with attackers with different background knowledge; and high technical level and computing resources are required. SUMMARY
[0006] In view of the deficiencies of the prior art, the present application provides a user data intelligent protection method and system based on differential privacy, which collects and preprocesses user health data, sets a dynamic privacy budget according to data sensitivity, analysis purpose and privacy requirements, adds random noise to generate perturbed data using a multi-strategy fusion data perturbation strategy, performs statistical analysis, and evaluates privacy disclosure risk, protection level and data availability using an improved differential privacy algorithm, iteratively adjusts the privacy budget and perturbation strategy based on the evaluation results through a feedback loop mechanism, and provides a health data analysis method that can protect user privacy and ensure data availability.
[0007] To achieve the above object, the present application provides the following technical solutions:
[0008] A user data intelligent protection method based on differential privacy, comprising:
[0009] Step S1: Collecting user health data and preprocessing;
[0010] Step S2: Setting an adaptive privacy budget allocation mechanism according to data sensitivity, analysis purpose and privacy protection requirements, and adding random noise to the preprocessed user health data using a multi-strategy fusion data perturbation strategy based on the adaptive privacy budget allocation mechanism to generate perturbed health data;
[0011] Step S3: Perform statistical analysis on the disturbed health data and use the improved differential privacy protection algorithm to evaluate the privacy leakage risk, privacy protection degree, and data availability of the differential privacy protection method by calculating the privacy loss;
[0012] Step S4: Based on the evaluation results of step S3, a feedback loop mechanism is used to iteratively adjust the privacy budget and data perturbation strategy.
[0013] Specifically, the specific steps of step S2 include:
[0014] S2.1: Assess the sensitivity of user health data, clarify the purpose of data analysis, and determine the specific privacy protection needs;
[0015] S2.2: Obtain pre-processed user health data and set up an adaptive privacy budget allocation mechanism based on data sensitivity, analysis purpose, and privacy protection requirements;
[0016] S2.3: Based on the privacy budget and data characteristics, a multi-strategy fusion data perturbation strategy is used to add random noise to the preprocessed user health data to generate perturbed health data.
[0017] Specifically, the specific steps of S2.2 include:
[0018] S2.21: Receive pre-processed user health data , obtain information on data sensitivity, analysis purpose, and privacy protection requirements, and set a global privacy budget ,in, represents the i-th preprocessed user health data, and N represents the number of preprocessed user health data;
[0019] S2.22: Based on the pre-processed user health data, the privacy budget allocation of each data item is dynamically adjusted using a weight allocation method based on data sensitivity. The formula is:
[0020] ;
[0021] in, represents the privacy budget allocated to the i-th preprocessed user health data, represents the importance weight of the i-th preprocessed user health data for the analysis purpose, represents the sensitivity level of the i-th preprocessed user health data, represents the correlation coefficient between the i-th preprocessed user health data and other preprocessed user health data, represents the dynamic adjustment factor of the i-th preprocessed user health data;
[0022] S2.23: Set the privacy budget consumption threshold In the data analysis process, monitor the consumption of the privacy budget in real time Loss;
[0023] If , adjust the privacy protection strategy or stop the analysis.
[0024] Specifically, the specific steps of S2.3 include:
[0025] S2.31: According to and , feature extraction is performed on to obtain user health feature data , wherein represents the nth user health feature data, and n represents the number of user health feature data.
[0026] S2.32: According to and the privacy protection requirement, determine the jth strategy the weight and the priority in the fusion, and fuse according to the weight and priority of to generate a data perturbation strategy , wherein represents the weight of the jth strategy in the fusion process, represents the priority of the jth strategy in the fusion process, , , represents the weight coefficient, represents the privacy protection contribution degree of the jth strategy, represents the availability loss of the jth strategy, represents the implementation cost of the jth strategy.
[0027] Specifically, the specific steps of S2.3 further include:
[0028] S2.33: Based on the data perturbation strategy , calculate the sensitivity of each data , and according to the noise distribution type and , calculate the amount of random noise to be added The calculation formula of sensitivity and random noise amount is:
[0029] ;
[0030] wherein and represent two adjacent data sets, i.e., only one data point is different, and express and The query function, Represents the sensitivity function f in two different data sets and The norm of the difference between the output values on , represents the Laplace noise function;
[0031] S2.34: Generate random noise using a random number generator based on the obtained random noise amount, and add the generated random noise to Generate perturbed health data ,in, represents the i-th disturbed health data.
[0032] Specifically, the specific steps of S3 include:
[0033] S3.1: Obtaining perturbation health data , combined with privacy protection requirement information, calculate The mean and variance of
[0034] S3.2: Use the improved differential privacy protection algorithm to calculate the privacy loss. The formula is:
[0035] ;
[0036] Where L represents the privacy loss, D represents the number of output sets considered, Represents the kth output set The weight of and Indicates in the dataset and Above, the parameters used are After the query function is queried, the output results fall into the set The probability of represents the time parameter associated with the k-th output set, represents the kth output set, b represents the adjustment parameter, represents the sensitivity function, represents the logarithmic function;
[0037] S3.3: Evaluate the privacy protection level of the differential privacy protection method based on the calculated results of privacy loss, and evaluate the data availability by comparing the statistical analysis results of the original data and the perturbed data.
[0038] Specifically, the strategy in S2.32 selects Laplace noise addition, k-anonymity, l-diversity and differential privacy.
[0039] The application discloses a user data intelligent protection system based on differential privacy, and relates to the technical field of user data intelligent protection.
[0040] The data processing module is used for collecting user health data and performing preliminary processing and cleaning.
[0041] The data perturbation module is used for protecting user privacy by adding random noise to the preprocessed user health data.
[0042] The differential privacy protection module is used for performing statistical analysis on the perturbed health data and quantitatively evaluating the differential privacy protection method by using an improved differential privacy protection algorithm.
[0043] The strategy adjustment module is used for iteratively adjusting the privacy budget and the data perturbation strategy according to the evaluation result.
[0044] Specifically, the data perturbation module comprises a privacy budget setting unit, a strategy selection unit and a noise adding unit.
[0045] The privacy budget setting unit is used for dynamically calculating and setting the privacy budget.
[0046] The strategy selection unit is used for selecting data perturbation strategies and noise based on the privacy budget.
[0047] The noise adding unit is used for adding noise to the preprocessed user health data according to the selected data perturbation strategy.
[0048] Specifically, the differential privacy protection module comprises an algorithm application unit, a risk evaluation unit, a privacy protection degree evaluation unit and a data availability evaluation unit.
[0049] The algorithm application unit is used for implementing the improved differential privacy protection algorithm.
[0050] The risk evaluation unit is used for evaluating the privacy leakage risk in the user health data processing process in real time and identifying potential security threats.
[0051] The privacy protection degree evaluation unit is used for quantitatively evaluating the effectiveness of the currently adopted privacy protection measures.
[0052] The data availability evaluation unit is used for evaluating whether the data processed by the privacy protection still has availability and accuracy.
[0053] Compared with the prior art, the application has the beneficial effects that:
[0054] 1.The application provides a user data intelligent protection method based on differential privacy, which collects and pre-processes user health data to ensure the accuracy and integrity of the data; according to the sensitivity of the data, the analysis purpose and the privacy protection requirement, a privacy budget is dynamically set, and a multi-strategy fusion data perturbation strategy is used to add random noise to the data, thereby effectively protecting the user privacy; the combination of dynamic adjustment and multiple protection strategies makes the privacy protection more flexible and effective.
[0055] 2.The application provides a user data intelligent protection method based on differential privacy, which also uses an improved differential privacy protection algorithm to statistically analyze the perturbed data and evaluate the privacy leakage risk, the privacy protection degree and the data availability; through a feedback loop mechanism, the privacy budget and the data perturbation strategy are iteratively adjusted according to the evaluation results, further optimizing the privacy protection effect and the data availability; this continuous optimization and improvement makes the process better balance the relationship between privacy protection and data utilization in practical applications. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 Fig. 1 is a schematic diagram of the user data intelligent protection method based on differential privacy of the application;
[0057] Figure 2 Fig. 3 is a principle flowchart of the user data intelligent protection method based on differential privacy of the application;
[0058] Figure 3 Fig. 4 is a flowchart of perturbed health data generation of the user data intelligent protection method based on differential privacy of the application;
[0059] Figure 4 Fig. 5 is a system architecture diagram of the user data intelligent protection system based on differential privacy of the application. DETAILED DESCRIPTION
[0060] Embodiment 1
[0061] Please refer to Figures 1-3 The application provides a user data intelligent protection method based on differential privacy, which includes the following steps:
[0062] Step S1: Collect user health data and pre-process it, wherein the user health data includes but is not limited to physiological indicators, disease history and living habits, specifically including age, gender, height, weight, sleep quality, eating habits, exercise amount, blood pressure and blood sugar.
[0063] Step S2: Set an adaptive privacy budget allocation mechanism according to data sensitivity, analysis purpose and privacy protection requirements, and use a multi-strategy integrated data perturbation strategy to add random noise to the preprocessed user health data based on the adaptive privacy budget allocation mechanism, to generate perturbed health data;
[0064] It should be noted that the adaptive privacy budget allocation mechanism needs to be set according to the sensitivity of the data, the specific purpose of the analysis and the actual requirements of privacy protection. This privacy budget is a key parameter that determines the magnitude of random noise that can be added during data perturbation. The higher the data sensitivity or the stricter the privacy protection requirements, the higher the privacy budget needs to be set to ensure that excessive personal privacy information is not leaked during data analysis. Then, the set privacy budget is used to guide the application of multiple data perturbation strategies, including but not limited to Laplace noise addition based on differential privacy, hybrid strategy combining k-anonymity and l-diversity, and synthetic data generated based on deep learning. The selection and parameter setting of these noise addition methods will be affected by the privacy budget. Specifically, the higher the privacy budget, the less or milder noise can be added to maintain data usability. The lower the privacy budget, the more or stronger noise needs to be added to enhance privacy protection. Therefore, after setting the dynamic privacy budget, the appropriate multi-data perturbation strategy is selected according to the budget and the corresponding parameters are set. Finally, these strategies are applied to the preprocessed user health data to add appropriate random noise, thereby generating perturbed health data that protects user privacy and maintains data characteristics.
[0065] Step S3: Perform statistical analysis on the perturbed health data and use an improved differential privacy protection algorithm to calculate privacy loss, evaluate the privacy leakage risk, privacy protection degree and data usability of the differential privacy protection method;
[0066] Step S4: Based on the evaluation results of step S3, use a feedback loop mechanism to iteratively adjust the privacy budget and data perturbation strategy.
[0067] The specific steps of step S2 include:
[0068] S2.1: Evaluate the sensitivity of user health data, clarify the purpose of data analysis, and determine the specific requirements of privacy protection;
[0069] Among them, (1) data sensitivity evaluation: evaluate the sensitivity of user health data, including the privacy level of data, the range of personal information involved; (2) clear analysis purpose: clear the purpose of data analysis, such as for medical research, health monitoring or other purposes; (3) determine the specific needs of privacy protection according to laws and regulations, industry standards or user agreement.
[0070] S2.2: Obtain preprocessed user health data, and set adaptive privacy budget allocation mechanism according to data sensitivity, analysis purpose and privacy protection requirement, wherein the adaptive privacy budget allocation mechanism not only considers the limitation of global privacy budget, but also dynamically adjusts the privacy budget allocation of each data item according to the sensitivity difference of data items;
[0071] S2.3: According to the privacy budget and data characteristics, add random noise to the preprocessed user health data using multi-strategy fusion data perturbation strategy to generate perturbed health data.
[0072] The specific steps of S2.2 include:
[0073] S2.21: Receive preprocessed user health data , obtain data sensitivity, analysis purpose and privacy protection requirement information, and set a global privacy budget , wherein, represents the ith preprocessed user health data, and N represents the number of preprocessed user health data;
[0074] S2.22: According to the preprocessed user health data, use the weight allocation method based on data sensitivity to dynamically adjust the privacy budget allocation of each data item, the formula is:
[0075] ;
[0076] Among them, represents the privacy budget allocated to the ith preprocessed user health data, represents the importance weight of the ith preprocessed user health data for the analysis purpose, represents the sensitivity level of the ith preprocessed user health data, which can be divided according to the type of data, the correlation degree of personal identity information and the time sensitivity of data, represents the correlation coefficient between the ith preprocessed user health data and other preprocessed user health data, which reflects the correlation degree between data items, which can be determined by calculating the correlation coefficient between data items or using other correlation measurement methods, in the formula, 1- allocating privacy budget for data items highly correlated with other data items to avoid over-protecting redundant information, represents the dynamic adjustment factor of the ith pre-processed user health data, which can adjust the allocation of privacy budget according to real-time feedback in the data analysis process, changes in user privacy preferences or data updates, and allows the system to dynamically optimize privacy protection strategies during the analysis process;
[0077] It should be noted that the present application is improved on the basis of the prior art. First, the formula in the present application introduces weights and sensitivity , and considers the privacy leakage risk and the dynamic adjustment factor , which realizes the fine allocation of privacy budget . This allocation method not only improves the efficiency of privacy protection, but also ensures the maximization of data availability, achieving a better balance between privacy protection and data utilization. Second, the formula in the present application has high flexibility and scalability. By adjusting the weights, sensitivity and adjustment factor, it can easily adapt to different application scenarios and privacy protection needs. At the same time, the structure of the formula provides convenience for its future expansion and optimization, enabling it to evolve with the development of technology and changes in application scenarios.
[0078] S2.23: Set privacy budget consumption threshold During the data analysis process, monitor the consumption of privacy budget Loss in real time;
[0079] If , adjust the privacy protection strategy or stop the analysis.
[0080] The specific steps of S2.3 include:
[0081] S2.31: According to and , feature extraction is performed on to obtain user health feature data , wherein represents the nth user health feature data, and n represents the number of user health feature data;
[0082] S2.32: According to and privacy protection requirements, determine the weight and priority of the jth strategy in fusion, and fuse according to the weight and priority of to generate data perturbation strategy , wherein represents the weight of the jth strategy in the fusion process, represents the priority of the jth strategy in the fusion process, , , represents the weight coefficient, represents the privacy protection contribution of the jth strategy, represents the availability loss of the jth strategy, represents the implementation cost of the jth strategy;
[0083] S2.32 Strategy selection Laplace noise addition, k-anonymity, l-diversity and differential privacy.
[0084] It should be noted that in determining the jth strategy weight in fusion , the size of the weight can be considered comprehensively according to the contribution of the strategy to privacy protection, data availability loss, and implementation cost, etc.
[0085] Further, the specific steps of S2.32 include:
[0086] (1) receiving the user's health feature data , including but not limited to physiological indicators, disease history, and life habits, and performing preprocessing operations such as cleaning, deduplication, and normalization on the data to improve the efficiency and accuracy of subsequent processing;
[0087] (2) analyze the user's privacy protection needs, including data sensitivity, use purpose, potential risks, and other factors, and determine the level and specific measures of privacy protection according to the needs, such as anonymization, encryption, and access control;
[0088] (3) select strategies, and determine the weight and priority of each strategy in the fusion according to the sensitivity of the user's health feature data and the privacy protection needs;
[0089] (4) according to the determined weight and priority, fuse the selected strategies to generate a data perturbation strategy, and apply the data perturbation strategy to the original health feature data to generate perturbed data.
[0090] S2.33: Based on the data perturbation strategy , calculate the sensitivity of each data , and according to the noise distribution type and , calculate the amount of random noise to be added, and the calculation formula of sensitivity and random noise amount is:
[0091] ;
[0092] wherein, and denote two adjacent data sets, i.e., only one data point is different, and denote and query functions, denote the norm of the difference between the output values of the sensitivity function f on two different data sets and denote the Laplace noise function, wherein the probability density function of the Laplace distribution is a prior art content in the art, and is not the creative scheme of the present application, and will not be described here.
[0093] S2.34: According to the obtained random noise quantity, a random number generator is used to generate random noise, and the generated random noise is added to to generate perturbed health data , and wherein, denote the i-th perturbed health data, and noise denotes random noise.
[0094] The specific steps of S3 include:
[0095] S3.1: Obtain perturbed health data , and calculate the mean and variance of in combination with the privacy protection requirement information, wherein the calculation formula of the mean and variance is a prior art content in the art, and is not the creative scheme of the present application, and will not be described here.
[0096] S3.2: An improved differential privacy protection algorithm is used to calculate the privacy loss, and the formula is:
[0097] ;
[0098] wherein, L denotes the privacy loss, D denotes the number of output sets considered, denotes the weight of the k-th output set , and and denote the probability that the output result falls in the set after querying using the query function with the parameter on the data sets and , and this probability value is used to measure the output distribution of the query function f under different data sets and parameter settings, denotes the time parameter related to the k-th output set, reflecting the change of the privacy loss with time, denotes the k-th output set, and b represents an adjustment parameter for controlling the size of the data set The impact on the privacy loss calculation is that when b = 0, the size of the data set does not affect the privacy loss, and when b > 0, the privacy loss decreases as the size of the data set increases. represents the sensitivity function, represents the symbol of the subset, and represents that the left set is a subset of the right set, represents the value range of the query function f, represents the logarithmic function,
[0099] It should be noted that the formula in the present application introduces a coefficient , which allows different weighting of the privacy loss for each query or event k; the introduction of b and also increases the flexibility of the formula, so that it can adjust the calculation of privacy loss according to the size of the data set and provides a parameterized method to adjust the calculation of privacy loss, which makes it possible to customize it according to different privacy protection needs and data characteristics; the output set is expanded to more comprehensively evaluate the privacy loss, and the introduction of the time factor potentially enhances privacy protection, which helps to further reduce the risk of privacy leakage.
[0100] S3.3: According to the calculation result of the privacy loss, the privacy protection degree of the differential privacy protection method is evaluated, and the data usability is evaluated by comparing the statistical analysis results of the original data and the perturbed data.
[0101] Further, when evaluating the privacy protection degree of the differential privacy protection method, the smaller the privacy budget, the higher the privacy protection degree, but the lower the data usability; the data usability is measured by the difference between the statistical quantities or the accuracy index of the query result.
[0102] Embodiment 2
[0103] Please refer to Figure 4 , the present application provides another embodiment: a user data intelligent protection system based on differential privacy, comprising:
[0104] a data processing module, a data perturbation module, a differential privacy protection module and a strategy adjustment module;
[0105] The data processing module is used to collect user health data and perform preliminary processing and cleaning to ensure data quality and consistency.
[0106] The data perturbation module is configured to protect user privacy by adding random noise to the preprocessed user health data, which involves introducing randomness into the data set, making the differences between query results for individuals ambiguous and unreliable, thereby preventing attackers from accurately restoring sensitive information of individuals.
[0107] The differential privacy protection module is configured to perform statistical analysis on the perturbed health data and quantitatively evaluate the differential privacy protection method using an improved differential privacy protection algorithm.
[0108] The strategy adjustment module is configured to iteratively adjust the privacy budget and data perturbation strategy based on the evaluation results to optimize the privacy protection effect and data usability.
[0109] The data perturbation module includes a privacy budget setting unit, a strategy selection unit, and a noise adding unit.
[0110] The privacy budget setting unit is configured to dynamically calculate and set the privacy budget to control the amount of noise added to the original data, and the size of the privacy budget directly affects the balance between the degree of privacy protection and data usability.
[0111] The strategy selection unit is configured to select a data perturbation strategy and noise based on the privacy budget, wherein the data perturbation strategy includes different noise adding methods and privacy budget allocation.
[0112] The noise adding unit is configured to add noise to the preprocessed user health data according to the selected data perturbation strategy to meet the requirements of differential privacy protection.
[0113] The differential privacy protection module includes an algorithm application unit, a risk assessment unit, a privacy protection degree assessment unit, and a data usability assessment unit.
[0114] The algorithm application unit is configured to implement an improved differential privacy protection algorithm to ensure that the data meets the preset privacy protection standard during processing.
[0115] The risk assessment unit is configured to assess the privacy leakage risk in the user health data processing process in real time and identify potential security threats.
[0116] The privacy protection degree assessment unit is configured to quantitatively evaluate the effectiveness of the current privacy protection measures to ensure that privacy protection meets the expected target.
[0117] The data usability assessment unit is configured to assess whether the data processed by the privacy protection still has usability and accuracy, ensuring that the data quality is not severely affected to support business needs.
[0118] The strategy adjustment module includes an evaluation result receiving unit, an adjustment unit, and a feedback loop control unit.
[0119] An evaluation result receiving unit is configured to receive evaluation results from the differential privacy protection module as the basis for subsequent policy adjustment.
[0120] An adjustment unit is configured to dynamically adjust the privacy budget and the data perturbation policy, including the noise addition method and parameters, according to the evaluation results.
[0121] A feedback loop control unit is configured to control the entire feedback loop process to ensure smooth iteration adjustment.
[0122] Through the data processing module, the data perturbation module, the differential privacy protection module, and the policy adjustment module and the units contained therein, the differential privacy-based user data intelligent protection system can comprehensively protect user health data while ensuring data usability and the accuracy of statistical analysis.
[0123] Those skilled in the art will appreciate that embodiments of the application can be provided as methods, systems, or computer program products. Accordingly, the application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage media, etc.) having computer-usable program code embodied in the medium.
[0124] The present application is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 An apparatus for performing a function specified by one or more flows and / or blocks in the flowcharts and / or block diagrams. Figure 1 An apparatus for performing a function specified by one or more flows and / or blocks in the flowcharts and / or block diagrams.
[0125] The embodiments of the application are described above with reference to the accompanying drawings, but the application is not limited to the specific embodiments described above, which are merely illustrative and not limiting. Those skilled in the art can make changes, modifications, replacements and variations to the above-described embodiments without departing from the spirit and scope of the application, and these are all within the scope of protection of the application.
Claims
1. A user data intelligent protection method based on differential privacy, characterized in that: include: Step S1: Collect user health data and perform preprocessing; Step S2: Based on the data sensitivity, analysis purpose, and privacy protection requirements, an adaptive privacy budget allocation mechanism is set up. Based on the adaptive privacy budget allocation mechanism, a multi-strategy fusion data perturbation strategy is used to add random noise to the pre-processed user health data to generate perturbed health data. Step S3: Perform statistical analysis on the disturbed health data and use the improved differential privacy protection algorithm to evaluate the privacy leakage risk, privacy protection degree, and data availability of the differential privacy protection method by calculating the privacy loss; Step S4: Based on the evaluation results of step S3, a feedback loop mechanism is used to iteratively adjust the privacy budget and data perturbation strategy; The specific steps of step S2 include: S2.1: Assess the sensitivity of user health data, clarify the purpose of data analysis, and determine the specific privacy protection needs; S2.2: Obtain pre-processed user health data and set up an adaptive privacy budget allocation mechanism based on data sensitivity, analysis purpose, and privacy protection requirements; S2.3: Based on the privacy budget and data characteristics, a multi-strategy fusion data perturbation strategy is used to add random noise to the pre-processed user health data to generate perturbed health data; The specific steps of S2.2 include: S2.21: Receive pre-processed user health data , obtain information on data sensitivity, analysis purpose, and privacy protection requirements, and set a global privacy budget ,in, represents the i-th preprocessed user health data, and N represents the number of preprocessed user health data; S2.22: Based on the pre-processed user health data, the privacy budget allocation of each data item is dynamically adjusted using a weight allocation method based on data sensitivity. The formula is: ; in, represents the privacy budget allocated to the i-th preprocessed user health data, represents the importance weight of the i-th preprocessed user health data for the analysis purpose, represents the sensitivity level of the i-th preprocessed user health data, represents the correlation coefficient between the i-th preprocessed user health data and other preprocessed user health data, represents the dynamic adjustment factor of the i-th preprocessed user health data; S2.23: Setting Privacy Budget Consumption Threshold ,During the data analysis process, the consumption of privacy budget Loss is monitored in real time; like , then adjust the privacy protection strategy or stop the analysis; The specific steps of S2.3 include: S2.31: According to and , and Perform feature extraction to obtain user health feature data ,in, represents the nth user health feature data, where n represents the number of user health feature data; S2.32: According to and privacy protection requirements, determine the jth strategy Weights in the ensemble and priority , and according to The weights and priorities are integrated to generate data perturbation strategies ,in, represents the weight of the j-th strategy in the fusion process, represents the priority of the jth strategy in the fusion process, 、 、 represents the weight coefficient, represents the privacy protection contribution of the j-th strategy, represents the availability loss of the j-th strategy, represents the implementation cost of the j-th strategy; The specific steps of S3 include: S3.1: Obtaining perturbation health data , combined with privacy protection requirement information, calculate The mean and variance of S3.2: Use the improved differential privacy protection algorithm to calculate the privacy loss. The formula is: ; Where L represents the privacy loss, D represents the number of output sets considered, Represents the kth output set The weight of and Indicates in the dataset and Above, the parameters used are After the query function is queried, the output results fall into the set The probability of represents the time parameter associated with the k-th output set, represents the kth output set, b represents the adjustment parameter, represents the sensitivity function, Representation data sensitivity, represents the logarithmic function; S3.3: Evaluate the privacy protection level of the differential privacy protection method based on the calculated results of privacy loss, and evaluate the data availability by comparing the statistical analysis results of the original data and the perturbed data.
2. The user data intelligent protection method based on differential privacy according to claim 1, characterized in that: The specific steps of S2.3 also include: S2.33: Data-based perturbation strategy , calculate each data Sensitivity , and according to the noise distribution type and , calculate the amount of random noise that needs to be added , the calculation formulas for sensitivity and random noise are: ; in, and Represents two adjacent data sets, that is, only one data point is different, and express and The query function, Represents the sensitivity function f in two different data sets and The norm of the difference between the output values on , represents the Laplace noise function; S2.34: Generate random noise using a random number generator based on the obtained random noise amount, and add the generated random noise to Generate perturbed health data ,in, represents the i-th disturbed health data.
3. The user data intelligent protection method based on differential privacy according to claim 2, characterized in that: The strategy in S2.32 selects Laplace noise addition, k-anonymity, l-diversity and differential privacy.
4. A user data intelligent protection system based on differential privacy, which is used to implement a user data intelligent protection method based on differential privacy according to any one of claims 1 to 3, characterized in that: include: Data processing module, data perturbation module, differential privacy protection module and policy adjustment module; The data processing module is used to collect user health data and perform preliminary processing and cleaning; The data perturbation module is used to protect user privacy by adding random noise to the pre-processed user health data; The differential privacy protection module is used to perform statistical analysis on the disturbed health data and quantitatively evaluate the differential privacy protection method using an improved differential privacy protection algorithm; The policy adjustment module is used to iteratively adjust the privacy budget and data perturbation strategy based on the evaluation results.
5. The user data intelligent protection system based on differential privacy according to claim 4, characterized in that: The data perturbation module includes: a privacy budget setting unit, a strategy selection unit and a noise addition unit; The privacy budget setting unit is used to dynamically calculate and set the privacy budget; The strategy selection unit is used to select a data perturbation strategy and noise based on a privacy budget; The noise adding unit is used to add noise to the preprocessed user health data according to the selected data perturbation strategy.
6. The user data intelligent protection system based on differential privacy according to claim 5, characterized in that: The differential privacy protection module includes: an algorithm application unit, a risk assessment unit, a privacy protection degree assessment unit and a data availability assessment unit; The algorithm application unit is used to implement an improved differential privacy protection algorithm; The risk assessment unit is used to assess the privacy leakage risk during the user health data processing process in real time and identify potential security threats; The privacy protection degree evaluation unit is used to quantitatively evaluate the effectiveness of the currently adopted privacy protection measures; The data availability evaluation unit is used to evaluate whether the data after privacy protection processing is still available and accurate.
Citation Information
Patent Citations
A t-closeness privacy protection method that satisfies ε-differential privacy
CN110598447B
A Distributed Learning Privacy Protection Method Based on Differential Privacy
CN111814189B
Privacy budget allocating and data publishing method and privacy budget allocating and data publishing system for protecting data query privacy
CN108537055A
Data protection method and device, electronic equipment and readable storage medium
CN117951741A