Data desensitization and integrity verification method and system based on differential privacy algorithm
The data desensitization method based on the differential privacy algorithm solves the problems of insufficient data desensitization accuracy and rough privacy budget allocation in existing technologies, realizes accurate quantification of desensitization needs and dynamic budget allocation, and improves data security and integrity.
Patent Information
- Application Number
- CN202511079593.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-08-04
AI Technical Summary
Existing technologies lack accuracy in data desensitization and privacy protection, have rough privacy budget allocation, cannot strike a balance between privacy protection and data utility, and the integrity verification mechanism lacks security and traceability.
A data desensitization method based on differential privacy algorithm is adopted to generate verification credentials through global sensitivity calculation, noise injection and blockchain evidence storage, to achieve accurate quantification of desensitization needs and dynamic and refined budget allocation, and to ensure data integrity through digital signatures and blockchain verification.
It achieves accurate quantification of data desensitization needs, balances privacy and data utility, improves the utility of data in multi-scenario applications, and ensures data security and integrity.
Smart Images

Figure CN120579227B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing and privacy protection, and more specifically, to a data desensitization and integrity verification method and system based on a differential privacy algorithm. Background Art
[0002] With the rapid development of the digital economy, the value of data as a core production factor in government administration, business operations, and other fields is becoming increasingly prominent. However, the risks of privacy leakage and integrity damage faced by data during sharing and circulation are becoming increasingly severe. This is especially true in scenarios such as government data openness and corporate data collaboration, which place extremely high demands on the accuracy of data desensitization and the reliability of data integrity verification.
[0003] Existing technologies have obvious shortcomings in data masking and privacy protection: on the one hand, the quantification of data masking needs is not accurate enough, relying more on empirical judgment, and lacks calculation of data leakage risks and the amount of privacy information contained, resulting in a lack of scientific basis for privacy budget allocation; on the other hand, the privacy budget allocation method is rough, and does not take into account the differences in sensitivity of different tasks to data utility. It is impossible to achieve dynamic optimization allocation of the budget for multi-task scenarios, and it is difficult to strike a balance between privacy protection and data utility. In addition, the integrity verification mechanism mostly relies on a single technology, and lacks security and traceability.
[0004] To address the above issues, the present invention proposes a data desensitization and integrity verification method and system based on a differential privacy algorithm to achieve maximum synergy between privacy protection and data utility. Summary of the Invention
[0005] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a data desensitization and integrity verification method and system based on a differential privacy algorithm.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A data desensitization and integrity verification method based on a differential privacy algorithm, the method comprising the following steps:
[0008] Step 1: Data application information processing.
[0009] S11. Accept the structured data application list submitted by the data applicant;
[0010] S12: Eliminate data that does not require desensitization based on the preset data sensitivity classification rules and output the data sequence to be processed ,in For a single sensitive data item, Indicates the total number of data categories;
[0011] Step 2: Calculate global sensitivity.
[0012] 定义并计算全局敏感度 为数据处理函数 The maximum output difference between adjacent datasets (datasets that differ by only one record);
[0013] 步骤三、脱敏核心处理。
[0014] S31、计算数据泄露风险系数:
[0015] S32、计算独立数据隐私信息量;
[0016] S33、计算联合数据隐私信息量;
[0017] S34、分配隐私预算:对于 , based on the calculated amount of joint data privacy information, calculate the total privacy budget ;
[0018] For a single task, the total privacy budget is the privacy budget of the task;
[0019] When multiple tasks use data, we prioritize maximizing the utility of each task and allocate more privacy budget to tasks that are sensitive to utility improvement (i.e., those with a fast utility improvement rate).
[0020] S35、噪声注入:
[0021] For different data types, the noise intensity is calculated based on the allocated privacy budget, and the data is differentially encrypted.
[0022] S36、输出脱敏数据序列: , 为数据类别数,其中, Indicates the differential privacy encryption result for this data type;
[0023] 步骤四、验证凭证生成。
[0024] Data fingerprint calculation: Use SHA-256 hash algorithm to calculate desensitized data sets and timestamps (accurate to milliseconds), the combined hash value of the operator ID ;
[0025] Digital signature generation: Use RSA algorithm to generate private key and public key pair, use private key 对数据指纹签名: ,公钥 公开用于后续验证;
[0026] 区块链存证:将{ , , time, ID} is written into the alliance chain (such as Hyperledger Fabric), and the smart contract automatically executes the evidence storage rules;
[0027] 步骤五、完整性验证。
[0028] 签名验证:接收方用公钥 解密签名: ,like ,则签名有效。
[0029] Data consistency verification: recalculate the fingerprint of the current desensitized dataset , compared with the stored fingerprint, if the difference rate is small enough, the data has not been tampered with;
[0030] Blockchain verification: Query the evidence record of the corresponding dataset ID in the blockchain to verify:
[0031] Block hash chain continuity (the current block hash contains the previous block hash);
[0032] 存证时间戳 在数据生成时间范围内;
[0033] 若均满足则通过完整性验证。
[0034] The present invention also provides a data desensitization and integrity verification system based on a differential privacy algorithm, the system comprising a data application information processing module, a global sensitivity calculation module, a desensitization core processing module, a verification credential generation module, and an integrity verification module;
[0035] Data application information processing module: This module receives and processes the structured data application list submitted by the data applicant, completes data screening and classification, and outputs the sensitive data sequence to be processed;
[0036] Global sensitivity calculation module: This module quantifies the sensitivity of different types of data and calculates the maximum output difference of data processing functions on adjacent data sets;
[0037] Desensitization core processing module: This module implements data desensitization based on a differential privacy algorithm. By quantifying the data desensitization requirements, it calculates the amount of privacy information in the joint data, determines the privacy allocation, and refines the privacy budget allocation plan based on the data usage tasks. Finally, it injects noise into the data based on the allocated privacy budget and encrypts it, outputting the encrypted desensitized data sequence.
[0038] Verification credential generation module: Generates unique identification and security credentials for desensitized data through data fingerprint calculation, digital signature generation, and blockchain evidence storage, ensuring data traceability and tamper-proofing;
[0039] Integrity verification module: Verifies the integrity and authenticity of desensitized data through digital signature verification, data consistency verification, and blockchain verification to prevent data tampering or forgery;
[0040] Furthermore, in step three, the data leakage risk coefficient is calculated as follows:
[0041] The data desensitization requirements are quantified using the data leakage risk coefficient based on multi-dimensional parameters. The specific calculation formula is as follows:
[0042] ;
[0043] in, ;
[0044] : Industry privacy level, which reflects the industry's demand for data desensitization. Based on the historical desensitized data of typical scenarios such as government affairs, enterprises, and scientific research, the effective information retention ratio reflects the data desensitization needs of different industries. The value range is 0~1;
[0045] : Historical leakage coefficient, reflecting the data desensitization needs caused by data leakage, through ,计算获得,最高取值1.0;
[0046] : Access frequency coefficient, which reflects the data desensitization requirements caused by data access frequency, and is represented by mapping the historical access frequency to a range of 0 to 1;
[0047] : Usage risk coefficient, reflecting the need for data desensitization induced by data usage, such as commercial transactions = 1.0, internal analysis = 0.6, and public statistics = 0.2;
[0048] : Correlation data coefficient, which reflects the data desensitization requirements caused by the correlation between data. Based on the feature similarity between the applicant's existing data and the application data, the correlation data coefficient is calculated using the cosine similarity algorithm and ranges from 0 to 1;
[0049] Weight 由专家经验确定。
[0050] Furthermore, in step 3, the amount of privacy information of independent data is calculated as follows:
[0051] Combine the data leakage risk coefficient and data entropy value to calculate the amount of independent data privacy information , the formula is as follows:
[0052] ;
[0053] in, , data entropy value is normalized and substituted into the formula , is the probability of the data item, is the number of data items; and is the weight of data leakage risk coefficient and data entropy value.
[0054] Further, the joint data privacy information quantity in step three is calculated, and the specific method is as follows:
[0055] By calculating the correlation of single-class data and other class data in the data application list, the cross-identification risk of multiple data items is quantified; for any one data item in the data application list , the number of data items associated with it is , , the joint data privacy information quantity of is calculated as follows:
[0056] ;
[0057] wherein, is the independent data privacy quantity of , is the independent data privacy information quantity of , is the conditional probability of when .
[0058] Further, the specific calculation method is as follows:
[0059] Search the data in the historical data set, respectively count the number of records and positions of and , compare the statistical results, filter and count the number of records that appear at the same time, and then calculate the ratio of to to calculate .
[0060] Further, more privacy budget is allocated to the task sensitive to utility improvement, and the specific method is as follows:
[0061] S341, total privacy budget average initial allocation: assuming the number of tasks is , then the initial average budget of each task is:
[0062] ;
[0063] S342, define the search range: the privacy budget search range of each task is based on Setting; referring to the actual tuning experience of differential privacy budget, the search range is set to: ;
[0064] S343、任务效用曲线拟合:
[0065] A certain number of sampling points are uniformly selected within the search range, denoted as ,in 为任务编号, 为采样点数量;
[0066] 针对每个采样点 ,计算任务 的效果指标 ;
[0067] 对每个任务 ,基于采样的 数据,采用二次多项式 回归拟合连续任务效用函数;
[0068] S344. Optimal allocation based on task utility curves: With the goal of maximizing total utility and combined with the total budget constraint, the final privacy budget allocation plan is determined through the marginal utility finite allocation method;
[0069] Furthermore, the final privacy budget allocation plan is determined by the marginal utility finite allocation method. The specific method is as follows:
[0070] Marginal utility calculation: Define the sensitivity of task utility improvement (i.e. marginal utility) as the derivative of the task utility function , The larger the value, the more significant the utility improvement of the task will be with each additional unit of budget under the current budget, and the task is a utility improvement sensitive task.
[0071] 设置约束条件:设任务 的最终分配预算为 ,需满足:
[0072] a.预算范围约束: ;
[0073] b.总预算约束: ;
[0074] Optimize allocation: Starting from the initial allocation baseline, the budget is iteratively adjusted until the total utility cannot be improved through adjustments. Finally, the solution is verified to meet the constraints.
[0075] Furthermore, the specific process of optimizing allocation is as follows:
[0076] (1) Initial allocation baseline: starting with the average budget, i.e. ,计算此时总预算 (刚好耗尽总预算);
[0077] (2) Iteratively adjust the budget: calculate the current budget for each task 下的边际效用 ,找出边际效用最大的任务 和最小的任务 ;从任务 的预算中减少 (微小步长,本实施例为0.01 ),分配给任务 ,得到新预算 Repeat the iteration until the total utility cannot be improved by adjustment (i.e., the marginal utility of all tasks is equal);
[0078] (3) Final solution verification: Check whether the adjusted total budget meets the constraints. If it exceeds, go back to the last step; output the total utility that meets the constraints. 最大的分配方案 .
[0079] Compared with the prior art, the present invention has the following beneficial effects:
[0080] 1. Accurately quantify data desensitization needs and balance privacy and data utility: By introducing the data leakage risk factor (integrating multi-dimensional parameters such as industry privacy level and historical leakage coefficient), independent data privacy information (combining data leakage risk factor and data entropy), and joint data privacy information (taking into account the correlation between data), we achieve accurate quantification of data desensitization needs and scientifically determine the total privacy budget. This innovative measure solves the problem of insufficient quantification of desensitization needs in existing technologies, making privacy protection more targeted. While ensuring privacy security, it effectively reduces the loss of data utility caused by excessive desensitization.
[0081] 2. Dynamically refine budget allocation to improve data utilization efficiency: For multi-task data usage scenarios, we first perform an average initial allocation of the privacy budget. Then, by fitting the task utility curves and applying the marginal utility finite allocation method, we allocate more budget to tasks that are sensitive to utility improvement. This approach addresses the coarse privacy budget allocation problem of existing technologies, maximizing the utility of each task within the total budget constraint, and improving the utility of data in multiple application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 This is a flowchart of the data desensitization and integrity verification method based on the differential privacy algorithm;
[0083] Figure 2 This is a flowchart of the desensitization core processing in step three of the present invention;
[0084] Figure 3 Flowchart of the method for allocating privacy budget in the multi-task data usage scenario of the present invention;
[0085] Figure 4 This is a module block diagram of the data desensitization and integrity verification system based on the differential privacy algorithm of the present invention. DETAILED DESCRIPTION
[0086] 实施例一,参照 Figure 1 The data desensitization and integrity verification method based on the differential privacy algorithm of this embodiment specifically includes the following processes:
[0087] Step 1: Data application information processing.
[0088] S11. Accept the structured data application list submitted by the data applicant, including the application data items (name, ID number, address, etc.), application purpose (such as statistical analysis, model training, etc.), applicant qualifications and data access frequency;
[0089] S12. Eliminate data that does not require desensitization based on preset data sensitivity classification rules, including:
[0090] Sensitivity classification: Based on industry background, data items are classified into high sensitivity, medium sensitivity, and low sensitivity;
[0091] Screening logic: Low-sensitivity data and applications for macro-statistics are marked as "no need for desensitization"; high-sensitivity data must be desensitized regardless of its use; medium-sensitivity data will have its screening results adjusted based on the applicant's qualifications;
[0092] Output: Get the original data set from the secure connection database, remove the "no need to desensitize" data, and form the data sequence to be processed ,in For a single sensitive data item, Indicates the total number of data categories;
[0093] Step 2: Calculate global sensitivity.
[0094] 定义全局敏感度 为数据处理函数 The maximum output difference between adjacent datasets (datasets that differ by only one record). The specific calculation rules are:
[0095] 数值型数据(如年龄,收入): ,in 为数据字段在数据集 上的取值范围;
[0096] 离散型数据(如职业、婚姻状态): ,如职业含医生、教师等5类,则 ;
[0097] Personal identification information (such as ID number): set directly (最高敏感度等级);
[0098] 步骤三、脱敏核心处理。
[0099] like Figure 2 The following is a flowchart of the core desensitization process. The specific process is as follows:
[0100] S31、计算数据泄露风险系数:
[0101] The data desensitization requirements are quantified based on multi-dimensional parameters and expressed as a data leakage risk coefficient. The specific calculation formula is as follows:
[0102] ;
[0103] in, ;
[0104] : Industry privacy level, which reflects the industry's demand for data desensitization. Based on the historical desensitized data of typical scenarios such as government affairs, enterprises, and scientific research, the effective information retention ratio reflects the data desensitization needs of different industries. The value range is 0~1;
[0105] : Historical leakage coefficient, reflecting the data desensitization needs caused by data leakage, through ,计算获得,最高取值1.0;
[0106] : Access frequency coefficient, which reflects the data desensitization requirements caused by data access frequency, and is represented by mapping the historical access frequency to a range of 0 to 1;
[0107] : Usage risk coefficient, reflecting the need for data desensitization induced by data usage, such as commercial transactions = 1.0, internal analysis = 0.6, and public statistics = 0.2;
[0108] : Correlation data coefficient, which reflects the data desensitization requirements caused by the correlation between data. Based on the feature similarity between the applicant's existing data and the application data, the correlation data coefficient is calculated using the cosine similarity algorithm and ranges from 0 to 1;
[0109] Weight 由专家经验确定,本实施例中 分别为0.3、0.2、0.2、0.15、0.15;
[0110] S32. Independent data privacy information volume: Calculate the independent data information volume by combining the data leakage risk coefficient and data entropy value. , the formula is as follows:
[0111] ;
[0112] in, ,数据熵值归一化后计算 , 为数据项取值的概率, 为不同数据项的数量; and is the data leakage risk coefficient and data entropy value. In this embodiment, , ;
[0113] S33. Joint Data Privacy Information Amount: Quantify the cross-identification risk of multiple data items by calculating the correlation between a single type of data and other types of data in the application data list, as follows:
[0114] For any data category in the data request list ,与其有关联性的 个数据类别为 , 的联合数据隐私信息量 The calculation formula is as follows:
[0115] ;
[0116] in, 为已知 hour The conditional probability is calculated based on historical data statistics. The specific calculation method is as follows:
[0117] ;
[0118] S34、分配隐私预算:对于 , based on the calculated amount of privacy information of the joint data, calculate the total privacy budget:
[0119] ;
[0120] in, 为常数,本实施例取1.2;
[0121] Since the privacy budget is usually positively correlated with the utility of the task using the data, in the case of a single task, the total privacy budget is the privacy budget of the task;
[0122] like Figure 3As shown in the figure, for the case of multi-task data usage, based on the sequence combination theorem of differential privacy, with the utility maximization of each task as the premise, tasks that are sensitive to utility improvement (with a fast utility improvement rate) are allocated more privacy budget, as follows:
[0123] S341. Average initial allocation of total privacy budget: Assume the number of tasks is , then the initial average budget for each task is:
[0124] ;
[0125] S342. Define the search scope: To cover the possible optimal allocation, the budget search scope of each task should be based on The setting should ensure flexibility while avoiding invalid computations caused by too wide a search range. Based on the actual tuning experience of differential privacy budget, the range is set as: In this embodiment ;
[0126] S343, Task Utility Curve Fitting:
[0127] Generate sampling points, each task in A certain number of sampling points (50 in this example) are uniformly selected within the interval and recorded as ,in Number the task;
[0128] Calculate the effect index for each sampling point , , computing tasks Effect indicators , needs to be defined according to the task type, for example:
[0129] Statistical analysis tasks: The indicator is the data aggregation error, such as the inverse of the mean square error (MSE). The higher the value, the better the effect.
[0130] Model training task: The indicator is prediction accuracy, the higher the value, the better;
[0131] Data query task: The indicator is the relevance of the query results, such as cosine similarity, the higher the value, the better;
[0132] Fitting utility curves for each task , based on 50 groups Data, using a quadratic polynomial Regression fits the continuous utility function, and the quadratic function can effectively capture the marginal diminishing characteristics of utility and budget;
[0133] S344. Optimal allocation based on utility curve: With the goal of maximizing total utility, combined with the total budget constraint, the final plan is determined through the marginal utility finite allocation method;
[0134] Marginal utility calculation: The sensitivity of utility improvement (i.e. marginal utility) is defined as the derivative of the utility function , The larger the value, the more significant the utility improvement of the task will be with each additional unit of budget under the current budget, and the task is a utility improvement sensitive task.
[0135] Constraints: Set the task The final allocated budget is , must meet the following requirements:
[0136] a. Budget range constraints: ;
[0137] b. Total budget constraint: ;
[0138] Optimize allocation:
[0139] (1) Initial allocation baseline: starting with the average budget, i.e. , calculate the total budget at this time (just enough to exhaust the total budget);
[0140] (2) Iteratively adjust the budget: calculate the current budget for each task Marginal utility under , find the task with the largest marginal utility and the smallest task ; From the task reduction in the budget (Small step size, in this example it is 0.01 ), assigned to the task , get the new budget ; Repeat the iteration until the total utility cannot be improved by adjustment (that is, the marginal utility of all tasks is equal); where, Indicates the number of iterations;
[0141] (3) Final solution verification: Check whether the adjusted total budget meets the constraints. If it exceeds, go back to the last step; output the total utility that meets the constraints. The largest allocation plan ;
[0142] S35, Noise Injection:
[0143] For continuous fields (numeric): Applicable to continuous numeric data such as age, income, temperature, etc. Adopt truncated Laplace noise injection, using the formula Calculate the scale parameter in the Laplace distribution probability density function to determine the noise intensity, where The privacy budget allocated for the current task;
[0144] For discrete fields (categorical data): Applicable to discrete categorical data such as occupation, marital status, and education level. An exponential mechanism is used to achieve privacy protection by randomly selecting a category that is "close to the original value". The core is to define a utility function to measure the value of category retention. For each possible value of the discrete field, , utility function Need to reflect The degree of match with the original value; the utility function needs to ensure The bigger, The higher the probability of being selected, the better the data availability. In the index mechanism, the category is selected. The probability of is proportional to, where is the global sensitivity of the utility function, A privacy budget for the current task;
[0145] For personal identification information (highly sensitive data): Applicable to highly sensitive data that uniquely identifies personal information, such as ID card numbers, mobile phone numbers, and precise addresses. Due to the highest privacy level, structural destruction and noise injection are required. For example, the address code of the ID card number is generalized to the provincial level, the birth date of the ID card number is randomly offset, the middle four digits of the mobile phone number are replaced with random numbers, and the precise address is generalized to the city address.
[0146] S36. Output desensitized data sequence: , is the number of data categories, where Indicates the differential privacy encryption result for this data type;
[0147] Step 4: Verify the generated credentials.
[0148] Data fingerprint calculation: Use SHA-256 hash algorithm to calculate desensitized data sets and timestamps (accurate to milliseconds), the joint hash value of the operator ID:
[0149] ;
[0150] Digital signature generation: Use RSA algorithm to generate private key and public key pair, use private key Sign the data fingerprint: , public key Public for subsequent verification;
[0151] Blockchain evidence: , , time, ID} is written into the alliance chain (such as Hyperledger Fabric), and the smart contract automatically executes the evidence storage rules;
[0152] Step 5: Integrity verification.
[0153] Signature verification: The recipient uses the public key Decrypted signature: ,like , the signature is valid.
[0154] Data consistency verification: recalculate the fingerprint of the current desensitized dataset , compared with the stored fingerprint, if the difference rate , the data has not been tampered with;
[0155] Blockchain verification: Query the evidence record of the corresponding dataset ID in the blockchain to verify:
[0156] Block hash chain continuity (the current block hash contains the previous block hash);
[0157] Evidence timestamp within the timeframe of data generation;
[0158] If all the conditions are met, the integrity verification is passed.
[0159] Example 2, refer to Figure 4 ,The data desensitization and integrity verification system based on the differential privacy algorithm of this embodiment includes a data application information processing module, a global sensitivity calculation module, a desensitization core processing module, a verification credential generation module, and an integrity verification module;
[0160] Data application information processing module: This module receives and processes the structured data application list submitted by the data applicant, completes data screening and classification, and outputs the sensitive data sequence to be processed;
[0161] Global sensitivity calculation module: This module quantifies the sensitivity of different types of data and calculates the maximum output difference of the data processing function on adjacent data sets (global sensitivity );
[0162] Desensitization core processing module: This module implements data desensitization based on a differential privacy algorithm, quantifies data desensitization requirements through multi-dimensional parameters, and further calculates the amount of joint data privacy information based on the amount of independent data privacy information. It also determines privacy allocation based on the amount of data privacy information and refines the privacy budget allocation scheme for different tasks in the data application, maximizing the data utility for each task using a limited privacy budget. Finally, it injects noise into the data based on the allocated privacy budget and encrypts it, outputting the encrypted desensitized data sequence.
[0163] Verification credential generation module: Generates unique identification and security credentials for desensitized data through data fingerprint calculation, digital signature generation, and blockchain evidence storage, ensuring data traceability and tamper-proofing;
[0164] Integrity verification module: Verifies the integrity and authenticity of desensitized data through digital signature verification, data consistency verification, and blockchain verification to prevent data tampering or forgery;
[0165] Through the detailed introduction of the above embodiments, the data desensitization and integrity verification method and system based on the differential privacy algorithm of the present invention determines the total privacy budget by calculating the global sensitivity, data leakage risk coefficient, and the amount of independent and joint data privacy information, and dynamically allocates the privacy budget for single-task or multi-task scenarios, and injects noise based on the allocation result to achieve differential privacy desensitization; at the same time, it uses hash functions, digital signatures and consortium chains to generate verification credentials, and ensures data integrity through signature verification, data consistency verification and blockchain verification. It is suitable for scenarios such as government affairs and enterprises that have high requirements for data privacy and integrity, and effectively improves the security and reliability of data processing.
[0166] The above formulas are all dimensionless and numerically calculated, and the preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0167] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0168] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0169] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0170] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0171] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0172] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0173] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A data desensitization and integrity verification method based on a differential privacy algorithm, characterized in that: The method flow is as follows: Step 1: Data application information processing: S11. Accept the structured data application list submitted by the data applicant; S12. Eliminate data that does not require desensitization based on preset data sensitivity classification rules, and output a sequence of data to be processed; Step 2: Global sensitivity calculation: Define and calculate global sensitivity as the maximum output difference of the data processing function on adjacent data sets; Step 3: Desensitization core processing: S31. Calculate the data leakage risk factor; S32. Calculate the amount of privacy information of independent data; the details are as follows: According to the data leakage risk coefficient and data entropy value, the amount of independent data privacy information is weightedly calculated; S33. Calculate the amount of joint data privacy information based on the amount of independent data privacy information. The specific method is as follows: For any data category in the data request list , which is related to Data categories , using the formula ,calculate The amount of joint data privacy information ;in, Known hour The conditional probability of for The amount of independent data privacy information, for The amount of independent data privacy information; The specific calculation method is as follows: Search the historical data of application data and count them separately and The number and position of the records that appear, compare the statistical results, filter and count the number of records that appear at the same time, and calculate and Ratio calculation ; S34. Calculate the total privacy budget based on the calculated amount of privacy information of the joint data The specific calculation formula is as follows: ; in, is a constant, for The amount of privacy information of the joint data; For a single task, the total privacy budget is the privacy budget of the task. For multiple tasks using data, tasks that are sensitive to utility improvement are allocated more privacy budget. S35. Calculate the noise intensity based on the allocated privacy budget and perform differential privacy encryption on the data. S36, outputting the desensitized data sequence; Step 4: Verification Credential Generation: Calculate Desensitized Dataset and Timestamp , the joint hash value of the operator ID ; Generate a private key and public key pair, and use the private key to sign the data fingerprint , the public key is made public for subsequent verification; , , time, ID} are written into the alliance chain, and the smart contract automatically executes the evidence storage rules; Step 5: Integrity Verification: Verify the digital signature, data consistency, and blockchain respectively.
2. The data desensitization and integrity verification method based on the differential privacy algorithm according to claim 1 is characterized in that: In step 3, during the desensitization core processing, the data leakage risk factor is calculated as follows: The data leakage risk coefficient is used to quantify the multi-dimensional data desensitization needs, specifically including five dimensions: industry privacy level, historical leakage coefficient, access frequency coefficient, usage risk coefficient, and associated data coefficient. The weighted summation method is used to quantify the data desensitization needs.
3. The data desensitization and integrity verification method based on the differential privacy algorithm according to claim 1 is characterized in that: Allocate more privacy budget to tasks that are sensitive to utility improvement. The specific method is as follows: S341, using formula , for the total privacy budget Perform an average initial allocation, where The number of tasks that use the data; S342. Define the privacy budget search scope for each task; S343, uniformly select a certain number of sampling points within the search range, denoted as ,in is the task number, is the number of sampling points; for each sampling point , computing tasks Effect indicators ; For each task , based on sampling Data, a quadratic polynomial was used to fit the continuous task utility function; S344. With the goal of maximizing total utility, under the constraint of the total budget, the final privacy budget allocation plan is determined through the marginal utility finite allocation method.
4. The data desensitization and integrity verification method based on the differential privacy algorithm according to claim 3 is characterized in that: The final privacy budget allocation plan is determined by the marginal utility finite allocation method. The specific method is as follows: Define marginal utility as the derivative of the task utility function and set constraints. Starting from the initial allocation baseline, iteratively adjust the budget and optimize the allocation until the total utility can no longer be improved through adjustments. Finally, verify whether the solution meets the set constraints.
5. A data desensitization and integrity verification system based on a differential privacy algorithm, used to implement the data desensitization and integrity verification method based on a differential privacy algorithm as described in any one of claims 1 to 4, characterized in that: The system includes a data application information processing module, a global sensitivity calculation module, a desensitization core processing module, a verification credential generation module, and an integrity verification module; Data application information processing module: This module receives and processes the structured data application list submitted by the data applicant, completes data screening and classification, and outputs the sensitive data sequence to be processed; Global sensitivity calculation module: This module quantifies the sensitivity of different types of data and calculates the maximum output difference of data processing functions on adjacent data sets; Desensitization core processing module: This module implements data desensitization based on a differential privacy algorithm. By quantifying the data desensitization requirements, it calculates the amount of privacy information in the joint data, determines the privacy allocation, and refines the privacy budget allocation plan based on the data usage tasks. Finally, it injects noise into the data based on the allocated privacy budget and encrypts it, outputting the encrypted desensitized data sequence. Verification credential generation module: Generates unique identification and security credentials for desensitized data through data fingerprint calculation, digital signature generation, and blockchain evidence storage, ensuring data traceability and tamper-proofing; Integrity verification module: Verifies the integrity and authenticity of desensitized data through digital signature verification, data consistency verification, and blockchain verification to prevent data tampering or forgery.
Citation Information
Patent Citations
Data blood relationship tracking method, system and device for privacy security protection
CN118536164A
Face sensitive data protection method based on elastic anti-interference privacy protection mechanism
CN119598517A