Multi-party data valuation analysis method and system based on differential privacy data generation
By combining differential privacy technology and data valuation methods in multi-party data collaboration scenarios, synthetic data consistent with the original data distribution is generated, and the problems of privacy risks and high calculation costs are solved through multi-dimensional difference analysis and dynamic penalty items, and efficient privacy protection and data valuation accuracy are achieved.
Patent Information
- Application Number
- CN202510203665.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
AI Technical Summary
Existing data valuation methods have privacy risks in multi-party data collaboration scenarios, especially due to the high accumulation of privacy budgets and high computational costs of differential privacy technology during multiple trainings.
By combining differential privacy technology and data valuation methods, synthetic data consistent with the distribution characteristics of the original data are generated, and dynamically adjusted valuation analysis results are constructed through multi-dimensional difference analysis and dynamic penalty items to balance privacy protection and data value evaluation.
It achieves efficient privacy protection and data valuation accuracy in multi-party data collaboration, reduces the accumulated and calculation costs of privacy budgets, and improves the accuracy and fairness of valuation results.
Smart Images

Figure CN120145440A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of privacy computing and data valuation, and particularly relates to a multi-party data valuation analysis method and system based on differential privacy data generation. Background Art
[0002] With the rapid development of artificial intelligence and machine learning, data has gradually become an important asset driving scientific and technological progress. High-quality data is not only the key to improving model performance but also provides support for data sharing and the prosperity of the data market. In recent years, with the rise of multi-party data collaboration and the data market, data valuation methods have gradually attracted attention. These methods help stakeholders achieve fair trading and reasonable utilization of data by quantifying the contribution of data provided by users to model performance.
[0003] Existing data valuation methods usually rely on model training and performance evaluation for different data subsets to quantify the contribution of each data point. Although these data valuation methods have promoted the reasonable valuation of data to a certain extent, they provide opportunities for malicious actors to attack using model outputs in the actual application process, thereby triggering privacy risks, especially in collaborative scenarios involving multiple data providers.
[0004] Privacy risks mainly include two common privacy attacks: membership inference attack and attribute inference attack. The membership inference attack allows an attacker to determine whether a data sample belongs to the training set by analyzing the model output, thereby exposing the membership of the training set and possible sensitive information. The attribute inference attack is even more serious. An attacker can reveal sensitive attributes outside the target model, such as personal characteristics and behavior patterns, through analyzing the model output, resulting in more serious privacy leakage.
[0005] In a multi-party data collaboration scenario, multiple data parties jointly provide data support for model training. Due to different privacy protection requirements of different data parties, the privacy protection requirements during data transmission, storage, and use are higher. The existing privacy risks mainly include: data may be maliciously intercepted during transmission, resulting in the leakage of privacy information; data stored on the server may be at risk of being stolen by hackers if there is a lack of effective encryption measures; when cooperating with a third party, the sharing of privacy data may violate relevant regulations and ethical standards. The existence of these problems may not only affect the enthusiasm of data providers for cooperation but also trigger legal and ethical disputes due to inadequate privacy protection.
[0006] Regarding the above problems, differential privacy technology provides an effective solution. Differential privacy introduces noise during the data processing and analysis process, making the impact of a single data point on the final output small enough to prevent attackers from inferring the specific information of individual data by observing the model output. Differential privacy stochastic gradient descent is a common application method of differential privacy technology, mainly achieving privacy protection through two steps: gradient clipping and noise addition. Gradient clipping restricts the excessive impact of each data sample on model updates, while noise addition hides the true gradient information, further protecting the privacy of user data.
[0007] However, although differential privacy stochastic gradient descent performs well in privacy protection, it still faces some challenges in practical applications. The cumulative problem of privacy budgets leads to a possible gradual decline in privacy protection effectiveness during multiple training processes, thus bringing unacceptable risks. In addition, due to the need to clip gradients and add noise, the computational cost increases significantly, which may become a performance bottleneck in large-scale dataset or multi-party collaboration scenarios. Therefore, it is particularly important to develop a data valuation method that can achieve privacy protection and high efficiency in multi-party data collaboration. Summary of the Invention
[0008] In view of the above, the purpose of the present invention is to provide a multi-party data valuation analysis method and system based on differential privacy data generation, which innovatively combines differential privacy technology and data valuation methods. It can use differential privacy as the core to generate synthetic data consistent with the distribution characteristics of the original data, while effectively balancing the relationship between privacy protection and data value evaluation. By constructing a valuation analysis result containing a dynamic penalty term, it dynamically adjusts and corrects the distribution difference between the differential privacy synthetic data and the original data, thereby improving the accuracy and fairness of the valuation result. This can not only promote the reasonable valuation and fair trading of data, but also effectively protect the privacy rights and interests of data providers, and promote the healthy development of artificial intelligence and machine learning technologies.
[0009] To achieve the above invention purpose, the technical solutions provided by the present invention are as follows:
[0010] In a first aspect, a multi-party data valuation analysis method based on differential privacy data generation provided by an embodiment of the present invention includes the following steps:
[0011] In the synthetic data generation stage, use the private evolutionary generation algorithm to generate differential privacy synthetic data based on the original data of each data provider;
[0012] In the data valuation stage, multi-dimensional difference analysis including downstream task accuracy analysis, distribution similarity analysis, and distance difference analysis is performed based on the original data and differentially private synthetic data, and a multi-dimensional difference score is constructed. Based on the multi-dimensional difference score, a dynamic penalty term is constructed and applied to the marginal contribution value of each data provider to construct and dynamically adjust the valuation analysis result of the differentially private synthetic data of each data provider.
[0013] Preferably, the multi-dimensional difference score is the weighted sum of the results of the multi-dimensional difference analysis of downstream task accuracy analysis, distribution similarity analysis, and distance difference analysis, and is expressed by the formula:
[0014] MDGS = α 1 ·PGR + α 2 ·FID + α 3 ·W
[0015] where MDGS represents the multi-dimensional difference score, PGR represents the performance difference ratio calculated in the downstream task accuracy analysis, FID represents the FID value calculated in the distribution similarity analysis, W represents the Wasserstein distance calculated in the distance difference analysis, and α 1 、α 2 and α 3 respectively represent the weight coefficients.
[0016] Preferably, a dynamic penalty term is constructed based on the multi-dimensional difference score, and the formula is expressed as:
[0017] P = λ·MDGS
[0018] where P represents the dynamic penalty term, MDGS represents the multi-dimensional difference score, and λ represents the penalty intensity coefficient, which is used to control the influence degree of the dynamic penalty term on the valuation analysis result.
[0019] Preferably, the dynamic penalty term is applied to the marginal contribution value of each data provider to construct the valuation analysis result of the differentially private synthetic data of each data provider, and the formula is expressed as:
[0020]
[0021] where, represents the valuation analysis result of the differentially private synthetic data of data provider i, and MC i represents the marginal contribution value of data provider i.
[0022] Preferably, the performance difference ratio is calculated through the downstream task accuracy analysis, including:
[0023] Select the downstream task according to the data type and application scenario and construct a benchmark model. Train the benchmark model using the original data and differentially private synthetic data respectively, and calculate the performance difference ratio PGR between the original data and the differentially private synthetic data. The formula is expressed as:
[0024]
[0025] where D origin represents the original data, D synth represents the differentially private synthetic data, and Perf(·) represents the performance metrics, including accuracy, F1 score, and mean squared error.
[0026] Preferably, calculate the FID value through distribution similarity analysis, including:
[0027] Extract the high-dimensional feature vectors of the original data and the differentially private synthetic data and obtain their feature distributions. Calculate the mean μ origin and covariance matrix Σ origin of the feature distribution of the original data, as well as the mean μ synth and covariance matrix Σ synth of the feature distribution of the differentially private synthetic data;
[0028] Evaluate the distribution consistency between the original data and the differentially private synthetic data through the FID metric. The formula is expressed as:
[0029]
[0030] where FID represents the FID value, ‖·‖ 2 represents the L2 norm, and Tr(·) represents the trace of the matrix.
[0031] Preferably, calculate the Wasserstein distance through distance difference analysis, including:
[0032] Extract the high-dimensional feature vectors of the original data and the differentially private synthetic data and obtain their feature distributions. Calculate the Wasserstein distance W between the feature distribution P origin of the original data and the feature distribution P synth of the differentially private synthetic data. The formula is expressed as:
[0033]
[0034] where inf represents the infimum, represents the expectation, Γ(P origin ,P synth ) represents the set of all joint distributions of P origin and P synth , and γ represents Γ(P origin ,P synth) Any joint distribution in (x, y) represents a sample pair sampled from the joint distribution γ, and ‖x - y‖ represents the Euclidean distance between the feature vectors x and y.
[0035] In a second aspect, an embodiment of the present invention further provides a multi-party data valuation analysis system based on differential privacy data generation, which is implemented by using the above-mentioned multi-party data valuation analysis method based on differential privacy data generation, and includes: a synthetic data generation module and a data valuation module;
[0036] The synthetic data generation module is used to generate differential privacy synthetic data respectively by using the private evolutionary generation algorithm based on the original data of each data provider;
[0037] The data valuation module is used to perform multi-dimensional difference analysis including downstream task accuracy analysis, distribution similarity analysis, and distance difference analysis based on the original data and the differential privacy synthetic data, construct a multi-dimensional difference score, construct a dynamic penalty term based on the multi-dimensional difference score, and apply it to the marginal contribution value of each data provider to construct and dynamically adjust the valuation analysis result of the differential privacy synthetic data of each data provider.
[0038] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory and a processor, the memory is used to store a computer program, and the processor is used to implement the above-mentioned multi-party data valuation analysis method based on differential privacy data generation when executing the computer program.
[0039] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a computer, the above-mentioned multi-party data valuation analysis method based on differential privacy data generation is implemented.
[0040] Compared with the prior art, the beneficial effects of the present invention at least include:
[0041] (1) By combining differential privacy technology and data valuation methods, the present invention more fairly reflects the actual value of differential privacy synthetic data by constructing a multi-dimensional difference score, and quantifies the contribution of each data provider to the overall model performance through the marginal contribution value, achieving a good balance between privacy protection and data valuation accuracy. It can not only generate high-quality differential privacy synthetic data, but also optimize the quantification method of data contribution by combining a dynamic valuation adjustment mechanism. In a multi-party collaboration scenario, a third party can complete data valuation without directly accessing sensitive information, providing a safe and effective technical solution for the healthy development of data sharing and the data market.
[0042] (2) By introducing a dynamic valuation adjustment mechanism based on a dynamic penalty term into the valuation analysis result, the present invention can further reduce the deviation and reduce the valuation error caused by the inconsistent distribution between the differentially private synthetic data and the original data; improve the valuation accuracy, and enable the valuation result of the differentially private synthetic data to more accurately reflect the actual value of the original data through the dynamic adjustment of the penalty term; enhance the fairness of cooperation among all parties, and ensure that the contributions of different data parties can be reasonably quantified in the multi-party data cooperation scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0044] Figure 1 is a schematic flowchart of a multi-party data valuation analysis method based on differentially private data generation provided by an embodiment of the present invention;
[0045] Figure 2 is a schematic flowchart of generating differentially private synthetic data using a private evolutionary generation algorithm provided by an embodiment of the present invention;
[0046] Figure 3 is a schematic diagram of the valuation comparison result between the differentially private synthetic data and the original data of the benchmark effectiveness evaluation provided by an embodiment of the present invention;
[0047] Figure 4 is a schematic structural diagram of a multi-party data valuation analysis system based on differentially private data generation provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the following further details the present invention with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the protection scope of the present invention.
[0049] The inventive concept of the present invention is as follows: Aiming at the problems of privacy leakage, privacy budget accumulation, high computational cost, and insufficient data valuation accuracy in the existing data valuation methods, the embodiments of the present invention provide a multi-party data valuation analysis method and system based on differential privacy data generation. Since in the practical applications of multi-party data collaboration and data markets, the value evaluation and privacy protection of data have always been a pair of contradictions, the embodiments of the present invention are based on differential privacy data generation technology and can avoid privacy budget accumulation and save computational cost through multi-dimensional difference analysis and a dynamic evaluation and adjustment mechanism introducing a dynamic penalty term, achieving an effective balance between privacy protection and data valuation accuracy, thereby providing a solution with both security and practicality for data sharing.
[0050] Figure 1 It is a schematic flowchart of the multi-party data valuation analysis method based on differential privacy data generation provided by the embodiments of the present invention. As Figure 1 shown, the embodiment provides a multi-party data valuation analysis method based on differential privacy data generation, including the following steps:
[0051] S1, in the synthetic data generation stage, use the private evolution generation algorithm to generate differential privacy synthetic data respectively based on the original data of each data provider.
[0052] In the embodiment, based on the differential privacy image generation technology open-sourced by Microsoft, the generation model DPSDA is adopted to generate through the private evolution (PE) generation algorithm, ensuring that the generated differential privacy synthetic image data is as consistent as possible with the original image data in distribution characteristics, generating high-quality differential privacy synthetic data for each data provider respectively, and avoiding the high cost and complexity of traditional model fine-tuning. Specifically, as Figure 2 shown, the private evolution generation algorithm initializes and generates synthetic samples for each party using the APIs interface of the generation model, constructs privacy samples for each party by dynamically adding differential privacy noise to protect sensitive information, and optimizes the quality of the synthetic samples for each party through the "parent selection", "parent selection", and "offspring generation (mutation)" mechanisms based on the parent samples and the privacy samples of each party, thereby generating high-quality differential privacy synthetic data. This design based on the "private evolution" algorithm avoids the complex training process of traditional models and can quickly generate high-quality differential privacy synthetic data, especially suitable for complex scenarios such as high-resolution image data sets.
[0053] S2. In the data valuation stage, based on the original data and the differentially private synthetic data, perform multi-dimensional difference analysis including downstream task accuracy analysis, distribution similarity analysis, and distance difference analysis, and construct a multi-dimensional difference score. Based on the multi-dimensional difference score, construct a dynamic penalty term and apply it to the marginal contribution value of each data provider to construct and dynamically adjust the valuation analysis result of the differentially private synthetic data of each data provider.
[0054] In the embodiment, in the data valuation stage, through the multi-dimensional difference analysis mechanism, systematically evaluate the difference between the original data and the differentially private synthetic data, and based on this, construct a dynamic penalty term to optimize the accuracy and fairness of the valuation result. The technical details of each analysis dimension and the dynamic adjustment mechanism are elaborated in detail below.
[0055] S2.1. Downstream task accuracy analysis.
[0056] Apply the original data and the differentially private synthetic data to downstream tasks such as classification and regression respectively, and evaluate the difference in their impact on the model performance. This dimension quantifies the data value by comparing the model performance differences between the original data and the differentially private synthetic data in specific downstream tasks. The specific implementation steps are as follows:
[0057] (1) Task selection and model construction: Select downstream tasks (such as image classification, object detection, regression prediction, etc.) according to the data type and application scenario, and construct a benchmark model (such as convolutional neural networks like ResNet and VGG).
[0058] (2) Performance evaluation metrics: In order to objectively quantify the model utility differences between the original data and the differentially private synthetic data, use the same test set as the unified evaluation benchmark, and use the original data D origin and the differentially private synthetic data D synth to train the benchmark model respectively, and record key performance metrics, including accuracy, F1-score, mean square error (MSE), etc.
[0059] (3) Difference quantification method: Calculate the performance gap ratio (PGR) between the original data and the differentially private synthetic data, and the formula is expressed as:
[0060]
[0061] where D origin represents the original data, D synth represents the differentially private synthetic data, and Perf(·) represents the key performance metric. When the PGR value approaches 0, it indicates that the contribution of the differentially private synthetic data to the downstream task is close to the original data, and the quality of the differentially private synthetic data is higher.
[0062] S2.2, FID distribution similarity analysis.
[0063] Based on the Fréchet Inception Distance (FID) metric, evaluate the distribution consistency between differentially private synthetic data and the original data. The specific implementation steps are as follows:
[0064] (1) Feature extraction: Use the pre-trained Inception-v3 model to extract the high-dimensional feature vectors of the original data and the differentially private synthetic data and obtain their feature distributions, and calculate the mean μ origin and covariance matrix Σ origin of the feature distribution of the original data, as well as the mean μ synth and covariance matrix Σ synth .
[0065] (2) FID calculation: Evaluate the distribution consistency between the original data and the differentially private synthetic data through the FID metric. The formula is expressed as:
[0066]
[0067] where FID represents the FID value, ‖·‖ 2 represents the L2 norm, and Tr(·) represents the trace of the matrix. The lower the value, the higher the distribution consistency and the higher the quality of the differentially private synthetic data.
[0068] S2.3, Wasserstein distance difference analysis.
[0069] To further quantify the distribution difference between the original data and the differentially private synthetic data, the Wasserstein distance is introduced as the third analysis metric. The Wasserstein distance can more intuitively reflect the difference between data distributions by calculating the minimum cost required to transform one distribution into another, especially when the distribution overlap is small or the structure is complex. The specific implementation steps are as follows:
[0070] (1) Feature extraction: Similar to the FID calculation, use the pre-trained CLIP or Inception-v3 model to extract the high-dimensional feature vectors of the original data D origin and the differentially private synthetic data D synth respectively, and obtain their feature distributions.
[0071] (2) Wasserstein distance calculation: Based on the extracted feature distributions, calculate the Wasserstein distance W between the feature distribution P origin of the original data and the feature distribution P synth of the differentially private synthetic data. The formula is expressed as:
[0072]
[0073] Among them, inf represents the infimum, represents the expectation, Γ(P origin ,P synth ) represents P origin and P synth the set of all joint distributions, γ represents any joint distribution in Γ(P origin ,P synth ), (x, y) represents a sample pair sampled from the joint distribution γ, and ‖x - y‖ represents the Euclidean distance between the feature vectors x and y. The smaller the Wasserstein distance value, the smaller the distribution difference between the original data and the differentially private synthetic data, and the higher the quality of the differentially private synthetic data.
[0074] S2.4, construct the multi-dimensional difference score and the dynamic penalty term.
[0075] Based on the evaluation results of the downstream task accuracy analysis, the FID distribution similarity analysis, and the Wasserstein distance difference analysis, construct a multi-dimensional difference score (Multi-Dimensional Gap Score, MDGS) by weighted summation, and the formula is expressed as:
[0076] MDGS = α 1 ·PGR + α 2 ·FID + α 3 ·W
[0077] Among them, MDGS represents the multi-dimensional difference score, α 1 , α 2 and α 3 respectively represent the weight coefficients of the downstream task accuracy, the FID distribution similarity, and the Wasserstein distance, satisfying α 1 + α 2 + α 3 = 1, and the weight coefficients can be dynamically adjusted according to the specific application scenario and data characteristics.
[0078] Based on the multi-dimensional difference score (MDGS), construct a dynamic penalty term P for optimizing the accuracy and fairness of data valuation, and the formula is expressed as:
[0079] P = λ·MDGS
[0080] Among them, λ represents the penalty intensity coefficient, which is used to control the influence degree of the dynamic penalty term on the valuation analysis result.
[0081] Since the MDGS value is negatively correlated with the quality of differentially private synthetic data (i.e., the smaller the MDGS, the higher the quality of differentially private synthetic data, and the penalty term P will also decrease accordingly), it can more fairly reflect the actual value of differentially private synthetic data and avoid valuation biases caused by distribution differences. The dynamic penalty term P, as an important input in the subsequent data valuation stage, can ensure to a certain extent that the valuation result of differentially private synthetic data is approximated to the valuation result of the original data.
[0082] S2.5, construct the valuation analysis result.
[0083] Adopt the Marginal Contribution method as the core module to quantify the contribution of each data provider to the overall model performance. The marginal contribution method calculates its marginal contribution value by comparing the changes in model performance with and without a certain data provider, so as to fairly evaluate the actual value of the data.
[0084] Based on the above multi-dimensional difference analysis, use the multi-dimensional difference analysis result as the basis for dynamically adjusting the valuation of differentially private synthetic data, and apply the dynamic penalty term to the marginal contribution value of each data provider to construct the valuation analysis result of the differentially private synthetic data of each data provider. The formula is expressed as:
[0085]
[0086] Among them, represents the valuation analysis result of the differentially private synthetic data of data provider i, and MC i represents the marginal contribution value of data provider i. In this way, when the quality of differentially private synthetic data is poor, the penalty term P increases, thereby increasing the valuation result of differentially private synthetic data to reflect its potential improvement space or additional value. By introducing a dynamic penalty term based on difference analysis, the valuation result of differentially private synthetic data can be dynamically optimized during the data valuation process, making the valuation result of differentially private synthetic data more accurately reflect the actual contribution of the data.
[0087] Based on the multi-party data valuation analysis method based on differentially private data generation provided in the above embodiments, the effectiveness and application value of the solution are further verified and revealed through the following three main evaluation scenarios, including benchmark effectiveness evaluation, different privacy budget evaluations, and different application scenario evaluations.
[0088] (1) Benchmark effectiveness evaluation.
[0089] The baseline validity assessment is used to verify the core functional performance of the method of the present invention under the condition of the standard privacy budget. During the experiment, based on the set typical experimental conditions, including the standard privacy budget (such as ε = 1.0), the public dataset (such as CIFAR-10), and the commonly used deep learning frameworks (such as PyTorch or TensorFlow), the private evolutionary generation algorithm is used to generate differentially private synthetic data. Subsequently, a detailed analysis is carried out on the differentially private synthetic data and the original data in multiple dimensions such as the accuracy of downstream tasks, the similarity of distributions, and the difference in distances. The core of the multi-dimensional difference analysis lies in quantifying the differences in the distributions of the two types of data themselves and their contributions to the model performance, providing a basis for subsequent valuation adjustment. Based on the above difference analysis, a multi-dimensional difference score is calculated, and according to the valuation analysis result formula constructed in the above embodiments, the possible biases in the valuation process of the differentially private synthetic data are corrected through a dynamic penalty term. For example, as Figure 3 shown, it is the comparison of the valuations of the differentially private synthetic data and the original data for the baseline validity assessment (the comparison before and after the penalty term is introduced by the difference analysis result). The existing experimental results show that under the baseline conditions, after the penalty term is introduced by the difference analysis result of the present invention, the valuation of the differentially private synthetic data can be adjusted, effectively solving the valuation bias problem between the original data and the differentially private synthetic data, thereby improving the accuracy of data valuation analysis.
[0090] (2) Different privacy budget assessments.
[0091] The different privacy budget assessments are used to study the impact of changes in the privacy budget on the performance of the method of the present invention and reveal the balance point between privacy protection and valuation accuracy. In this scenario, under the condition of keeping other parameters consistent, different privacy budget levels are set respectively (for example, ε = 1.0, ε = 4.0, ε = 8.0, and ε = 10.0). As the privacy budget increases, the noise intensity added to the differentially private synthetic data gradually decreases. In the present invention, by performing difference analysis and valuation adjustment on the differentially private synthetic data generated under different privacy budgets, a relationship curve between the privacy budget and the valuation accuracy is constructed, thereby providing a reference basis for selecting an appropriate privacy budget in practical applications and avoiding the accumulation of privacy budgets.
[0092] (3) Different application scenario assessments.
[0093] Evaluations in different application scenarios are used to verify the adaptability and generality of the method of the present invention under various task requirements and data distributions. The evaluation scenarios of the present invention cover multiple typical data sets, including general natural image data (such as CIFAR-10), simple image data (such as MNIST), and sensitive data in medical scenarios. In the simple image scenario, due to the relatively single data distribution, the test can more quickly verify the basic performance and valuation accuracy of the system; in the general natural image scenario and medical scenario, the data has high dimensions and diversity, and the focus of the test evaluation is to ensure that the generated data has high distribution similarity and valuation accuracy while protecting privacy by dynamically adjusting the penalty term mechanism, and an effective balance between privacy protection and data valuation accuracy can be achieved in different application scenarios.
[0094] In summary, the multi-party data valuation analysis method based on differential privacy data generation provided by the embodiments of the present invention generates differential privacy synthetic data and performs differential analysis on the original data and the differential privacy synthetic data in multiple dimensions. The penalty term is introduced from the differential analysis results to dynamically adjust the valuation analysis results of the synthetic data to make it closer to the valuation analysis results of the original data, thereby solving problems such as privacy leakage, privacy budget accumulation, and insufficient valuation accuracy in the existing methods, achieving an effective balance between privacy protection and data valuation accuracy, and being applicable to privacy protection and data valuation in multi-party collaboration scenarios.
[0095] Based on the same inventive concept, as Figure 4 shown, the embodiments of the present invention also provide a multi-party data valuation analysis system 400 based on differential privacy data generation, including: a synthetic data generation module 410 and a data valuation module 420.
[0096] The synthetic data generation module 410 is used to generate differential privacy synthetic data respectively by using the private evolutionary generation algorithm based on the original data of each data provider.
[0097] The data valuation module 420 is used to perform multi-dimensional differential analysis including downstream task accuracy analysis, distribution similarity analysis, and distance difference analysis on the original data and the differential privacy synthetic data, construct a multi-dimensional differential score, construct a dynamic penalty term based on the multi-dimensional differential score, and apply it to the marginal contribution value of each data provider to construct and dynamically adjust the valuation analysis results of the differential privacy synthetic data of each data provider.
[0098] Based on the same inventive concept, the embodiments of the present invention also provide an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor is used to implement the above-mentioned multi-party data valuation analysis method based on differential privacy data generation when executing the computer program.
[0099] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer, the above-mentioned multi-party data valuation analysis method based on differentially private data is implemented.
[0100] It should be noted that the above-mentioned multi-party data valuation analysis system, electronic device, and computer-readable storage medium provided by the above embodiments all belong to the same inventive concept as the multi-party data valuation analysis method based on differentially private data. For the specific implementation process, please refer to the embodiments of the multi-party data valuation analysis method based on differentially private data, which will not be elaborated here.
[0101] The above specific embodiments have elaborated in detail the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A multi-party data valuation analysis method based on differential privacy data generation, characterized in that: The following steps are involved: In the synthetic data generation stage, differentially private synthetic data are generated based on the original data of each data provider using a private evolutionary generation algorithm; In the data valuation stage, a multi-dimensional difference analysis including downstream task accuracy analysis, distribution similarity analysis and distance difference analysis is performed based on the original data and the differentially private synthetic data, and a multi-dimensional difference score is constructed. A dynamic penalty term is constructed based on the multi-dimensional difference score and applied to the marginal contribution value of each data provider to construct and dynamically adjust the valuation analysis results of the differentially private synthetic data of each data provider.
2. The multi-party data valuation analysis method based on differential privacy data generation according to claim 1 is characterized in that: The multi-dimensional difference score is the weighted sum of the results of multi-dimensional difference analysis of downstream task accuracy analysis, distribution similarity analysis, and distance difference analysis. The formula is expressed as: MDGS=α1·PGR+α2·FID+α3·W Among them, MDGS represents the multidimensional difference score, PGR represents the performance difference ratio calculated in the downstream task accuracy analysis, FID represents the FID value calculated in the distribution similarity analysis, W represents the Wasserstein distance calculated in the distance difference analysis, and α1, α2 and α3 represent weight coefficients respectively.
3. The multi-party data valuation analysis method based on differential privacy data generation according to claim 1 or 2, characterized in that: A dynamic penalty term is constructed based on multi-dimensional difference scores, and the formula is expressed as: P=λ·MDGS Among them, P represents the dynamic penalty term, MDGS represents the multi-dimensional difference score, and λ represents the penalty intensity coefficient, which is used to control the influence of the dynamic penalty term on the valuation analysis results.
4. The multi-party data valuation analysis method based on differential privacy data generation according to claim 3 is characterized in that: The dynamic penalty term is applied to the marginal contribution value of each data provider to construct the valuation analysis result of the differentially private synthetic data of each data provider. The formula is expressed as: in, represents the valuation analysis result of differentially private synthetic data of data provider i, MC i Represents the marginal contribution value of data provider i.
5. The multi-party data valuation analysis method based on differential privacy data generation according to claim 2 is characterized in that: The performance difference ratio is calculated through downstream task accuracy analysis, including: Select downstream tasks and build a benchmark model based on data types and application scenarios. Use original data and differentially private synthetic data to train the benchmark model, and calculate the performance difference ratio (PGR) between original data and differentially private synthetic data. The formula is: Among them, D origin represents the original data, D synth represents differentially private synthetic data, and Perf(·) represents performance indicators, including accuracy, F1 score, and mean square error.
6. The multi-party data valuation analysis method based on differential privacy data generation according to claim 2 is characterized in that: The FID value is calculated through distribution similarity analysis, including: Extract the high-dimensional feature vectors of the original data and the differentially private synthetic data and obtain their feature distributions, and calculate the mean μ of the feature distribution of the original data respectively. origin and the covariance matrix Σ origin And the mean μ of the feature distribution of differentially private synthetic data synth and the covariance matrix Σ synth ; The FID indicator is used to evaluate the distribution consistency of the original data and the differential privacy synthetic data. The formula is expressed as: Where FID represents the FID value, ‖·‖ 2 represents the L2 norm, and Tr(·) represents the trace of the matrix.
7. The multi-party data valuation analysis method based on differential privacy data generation according to claim 2 is characterized in that: Wasserstein distance is calculated by distance difference analysis, including: Extract the high-dimensional feature vectors of the original data and the differentially private synthetic data and obtain their feature distributions, and calculate the feature distribution P of the original data origin The feature distribution P of differentially private synthetic data synth The Wasserstein distance W between them is expressed as: Among them, inf represents the infimum, represents the expectation, Γ(P origin ,P synth ) indicates P origin and P synth The set of all joint distributions, γ represents Γ(P origin ,P synth ), (x,y) represents a sample pair sampled from the joint distribution γ, and ‖xy‖ represents the Euclidean distance between the feature vectors x and y.
8. A multi-party data valuation analysis system based on differential privacy data generation, implemented using the multi-party data valuation analysis method based on differential privacy data generation according to any one of claims 1 to 7, characterized in that: include: Synthetic data generation module and data valuation module; The synthetic data generation module is used to generate differential privacy synthetic data based on the original data of each data provider using a private evolutionary generation algorithm; The data valuation module is used to perform multi-dimensional difference analysis including downstream task accuracy analysis, distribution similarity analysis and distance difference analysis based on the original data and the differentially private synthetic data and construct a multi-dimensional difference score. A dynamic penalty term is constructed based on the multi-dimensional difference score and applied to the marginal contribution value of each data provider to construct and dynamically adjust the valuation analysis results of the differentially private synthetic data of each data provider.
9. An electronic device comprising a memory and a processor, wherein the memory is used to store a computer program, wherein: The processor is used to implement the multi-party data valuation analysis method based on differential privacy data generation as described in any one of claims 1-7 when executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a computer, the multi-party data valuation analysis method based on differential privacy data generation described in any one of claims 1 to 7 is implemented.