Data pricing method based on multi-dimensional value evaluation, terminal and storage medium
Through a data pricing method based on multi-dimensional value evaluation, combined with differential privacy mechanism, dual-objective optimization algorithm and meta-learning model, the problem of data value quantification and leakage risk balance in the existing technology is solved, and efficient and transparent data pricing and transactions are achieved.
Patent Information
- Application Number
- CN202510295427.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
Existing data product pricing methods lack a reliable value quantification mechanism, making it difficult to accurately estimate the actual application effect of data, and when sensitive information is involved, it is difficult to balance the real value of data and the risk of leakage.
The data pricing method based on multi-dimensional value evaluation is adopted, and the data is desensitized through a differential privacy mechanism, and privacy scores are generated. The dynamic weight coefficient is calculated using a dual-objective optimization algorithm, and the data pricing formula is generated in combination with the meta-learning model to output pricing results.
It achieves fair and reasonable data pricing, balances the data purchaser's expectations for data with the data provider's leak risk, improves the efficiency and transparency of data transactions, and reduces transaction risks and costs.
Smart Images

Figure CN120218976A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data product pricing, and in particular to a data pricing method based on multi-dimensional value evaluation, as well as a computer terminal and a computer-readable storage medium applying the method. Background Art
[0002] In the process of data trading, the data provider usually proposes a reference price based on the cost method, and the data service provider provides pricing guidance on this basis. However, data buyers often have an uncertain attitude towards whether the price of the data matches its value and whether it can bring value-added to the business. This uncertainty mainly stems from the differences in the expectations of both trading parties for the value of data products. For example, in the models used in bank credit business, real original data is usually required, but the data holder is afraid to sell the original data due to the risk of leakage, resulting in difficulties in reaching a deal between the buyer and the seller.
[0003] In the pricing of data products involving a large amount of personal privacy data, the core challenge lies in how to balance the contradiction between the data buyer's expectations for the data and the data holder's risk of leakage. The scarcity and privacy sensitivity of data directly affect its market value, and the actual application effect of the model highly depends on high-quality and complete data sets. However, due to legal and compliance issues, data holders often dare not easily sell complete original data, and this restriction forces data buyers to rely on manual review to make up for the deviation caused by incomplete or poor-quality data, thus affecting the efficiency and application experience of the model.
[0004] Existing pricing methods usually lack a reliable value quantification mechanism, and the traditional revenue-oriented pricing method is difficult to accurately estimate the actual application effect of data. Due to the lack of accurate evaluation of data value, existing pricing methods are often too subjective and arbitrary, especially when pricing data products involving sensitive information, it is more difficult to effectively evaluate their true value. Therefore, how to develop a pricing mechanism that can reflect the true value of data and balance the risk of leakage has become an urgent technical problem to be solved. Summary of the Invention
[0005] To solve the technical problems existing in the prior art, the present invention provides a data pricing method, a terminal and a storage medium based on multi-dimensional value evaluation, desensitizes the data of the data provider while reducing the impact on the model effect, adaptively adjusts the privacy content of the data according to the needs of the data provider, and the service provider prices according to the data privacy content. In this way, the present invention realizes fair and reasonable data pricing, balances the expectations of data buyers for the data and the leakage risk of data providers, and promotes the healthy development of the data trading market.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] The present invention discloses a data pricing method based on multi-dimensional value evaluation, including the following steps:
[0008] S1. Desensitize the data to be priced based on the differential privacy mechanism to generate the privacy budget value and privacy risk value of the data, and obtain the privacy score accordingly;
[0009] S2. Generate the quality score of the data according to the evaluation index of the desensitized data on the specified task;
[0010] S3. Use the bi-objective optimization algorithm to calculate the dynamic weight coefficient that satisfies the optimal privacy budget value, which is used to balance the privacy protection objective and the data quality objective;
[0011] S4. Based on the privacy score, the quality score, the dynamic weight coefficient, and the cost information of the data provider, use the meta-learning model to generate a data pricing formula and output the pricing result.
[0012] As a further improvement of the above solution, step S1 includes the following specific steps:
[0013] S11. Input data preprocessing: Standardize the original data using the Z-score standardization method to remove the dimension difference; among them, by calculating the mean and standard deviation of each feature, the data is converted into a standard normal distribution: X norm =(X - μ) / σ; in the formula, X norm represents the standardized data, X is the original data; μ and σ are the mean and standard deviation respectively;
[0014] Represent a differential privacy mechanism as M: X n ×Q → Y; in the formula, Q is the query space; Y is the output space of mechanism M; X n represents a data set composed of any n pieces of data;
[0015] Define the initial privacy budget value set and prepare to traverse;
[0016] S12. Define the individual privacy risk value:
[0017] PRV f =α·||q(x)-q(x -f )||1 + β·QueryFrequency(f)
[0018] In the formula, PRV f is the privacy risk value of individual f in the data set; α and β are weight coefficients used to balance the influence of data sensitivity and query frequency; q(x) represents the calculation query result of the complete data set; x -f is the data set after removing individual f; ||q(x)-q(x-f ) || 1 represents the data set x -f Impact on the query output; QueryFrequency(f) represents the frequency with which individual f is accessed or mentioned in a specific query;
[0019] S13. Determination condition for selecting the privacy budget: When Allocate the privacy budget; where PRV min and PRV max are the minimum and maximum privacy risk values under the influence of the current query period respectively; p is the privacy preference percentage, representing the acceptable privacy risk difference ratio;
[0020] S14. Define the Laplace mechanism: Given a query and a data set x; represents the k-dimensional real space; the Laplace mechanism M L is defined as:
[0021] M L (x, q, ∈) = q(x) + (Z1, …, Z k )
[0022] In the formula, Z1, …, Z k are k random variables drawn from the Laplace distribution, and each random variable has a probability density function:
[0023]
[0024] In the formula, p b (z) is the probability density function of the Laplace distribution, representing the probability that Z takes a value under this distribution; z is the noise value sampled from the Laplace distribution; b is the scale function in the Laplace mechanism, Δq is the global sensitivity of the query q; ∈ is the privacy budget value;
[0025] S15. Traverse the set of initial privacy budget values in descending order, calculate the query results with privacy protection added using the Laplace mechanism, and calculate the individual privacy risk values of each individual; according to the privacy preference percentage provided by the data provider, calculate the privacy budget value that meets the privacy considerations of the data provider, form a set of privacy budget candidates, and obtain the privacy score accordingly.
[0026] As a further improvement of the above solution, step S2 includes the following specific steps:
[0027] S21. Receive the structured de-sensitized data and preprocess different modalities of data for standardization;
[0028] S22. Initialize the set of privacy budget candidates {∈1, ∈2, …, ∈n}, desensitize the data using each candidate value of the privacy budget to obtain a desensitized data set;
[0029] S23. Perform a classification task on the classifier using the desensitized data set, traverse the candidate values of the privacy budget, and calculate the accuracy of the desensitized data set;
[0030] S24. Calculate the actual accuracy through cross-validation, and obtain the quality score accordingly.
[0031] As a further improvement of the above solution, step S3 includes the following specific steps:
[0032] S31. Construct an objective function
[0033]
[0034] In the formula, is the accuracy loss function, is the privacy loss function, and λ1 and λ2 are respectively and corresponding weight coefficients; Acc threshold is the preset accuracy threshold; Acc actual is the actual accuracy;
[0035] S32. Perform non-dominated sorting based on the NSGA-II algorithm, calculate the dominance relationship and crowding distance of the solutions, so as to select the optimal solution of the objective function, that is, the privacy budget value ∈ that minimizes the total loss of the objective function * , and output the dynamic weight coefficients and
[0036] As a further improvement of the above solution, in step S32, the non-dominated sorting includes the following specific steps:
[0037] S321. Define the dominance relationship between solutions, and the expression formula is:
[0038]
[0039] In the formula, i and j are any two candidate solutions, represents the dominance relationship; and are a set of privacy protection target values; and are a set of accuracy target values; {privacy, accuracy} is the target set of privacy protection loss and data accuracy loss, and k’ is the target variable to be optimized currently; is the loss value of solution i on target k’; is the loss value of solution j on target k'; indicates existence;
[0040] S322. Define the domination count and domination set; where, for each candidate solution i, its domination count is n i , and the domination set is S i ; the domination count n i represents the number of solutions that dominate solution i among all candidate solutions; the domination set S i represents all solutions dominated by solution i;
[0041] S323. Construct the non-dominated front in order according to the domination counts and domination sets of all solutions; where the k-th non-dominated front F k contains the solutions with a domination count of 0 after removing all solutions in the first k - 1 fronts;
[0042] S324. Define the crowding distance of each solution in the non-dominated front to reflect the distribution of solutions in the objective space; where the formula for calculating the crowding distance of a solution is:
[0043]
[0044] In the formula, d i represents the crowding distance of each solution i in the non-dominated front; and respectively represent the function values of the two adjacent solutions to solution i after sorting by function value on target m; and respectively represent the maximum and minimum values of all solutions on target m;
[0045] S325. After completing non-dominated sorting and crowding distance calculation, select the optimal solution of the objective function from the first non-dominated front F1, and the expression formula is:
[0046]
[0047] In the formula, represents the privacy loss under any privacy budget ∈ on the privacy protection objective; represents the accuracy loss under any privacy budget ∈ on the accuracy objective.
[0048] As a further improvement of the above solution, in step S15, the contribution of the privacy budget ∈ to the privacy score is: In step S24, the actual accuracy Acc actual 's contribution to the quality score is: Acc threshold is the minimum accuracy threshold specified by the data purchaser.
[0049] As a further improvement of the above solution, in step S4, the data pricing formula is as follows:
[0050]
[0051] In the formula, P final is the data pricing; C cost is the cost of the data provider; PrivacyScore is the privacy score; QualityScore is the quality score.
[0052] As a further improvement of the above solution, in step S4, the model-agnostic meta-learning algorithm is adopted to train the pricing model based on historical transaction data. After training, the meta-learning model can generate pricing models for different data sets. Among them, the loss function of the meta-learning model is defined as follows:
[0053]
[0054] In the formula, is the loss of the meta-learning model; N is the total number of samples in the historical transaction records; is the price of the i'-th data product predicted by the meta-learning model; is the market price of the i'-th data product in the historical transaction records.
[0055] The present invention also discloses a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the data pricing method based on multi-dimensional value evaluation as described above are implemented.
[0056] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the data pricing method based on multi-dimensional value evaluation as described above are implemented.
[0057] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0058] 1. Through privacy budget management and privacy quality trade-off, the present invention adopts a two-objective optimization method and a differential privacy mechanism to achieve a dynamic balance between privacy protection and data quality. While meeting the privacy protection requirements of data providers, it ensures the usability of desensitized data in tasks, thus providing an efficient and reliable solution for data transactions.
[0059] 2. The present invention realizes dynamic pricing through data pricing, comprehensively considering privacy protection, data quality, cost information, and market demand, and adopting a meta-learning model. The finally generated pricing result is not only scientific and reasonable, but also provides a transparent pricing basis through a detailed pricing report, effectively solving the problems of strong subjectivity and lack of fairness existing in traditional pricing methods.
[0060] 3. The present invention supports flexibly adjusting the priorities of privacy protection and data quality according to the actual needs of data purchasers, and generating multiple versions of data products through a privacy budget allocation and dynamic weight adjustment mechanism to meet the personalized needs of different scenarios, thereby improving the success rate and efficiency of data transactions.
[0061] 4. The present invention provides a modular and automated data pricing method, providing a standardized tool for the data trading market. Through an efficient privacy protection mechanism and a scientific pricing model, it reduces the risks and costs of data transactions and promotes the healthy development and large-scale application of the data factor market.
[0062] 5. The computer terminal and computer-readable storage medium disclosed by the present invention can produce the same beneficial effects as those obtained by applying the above method, and will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a flowchart of the data pricing method based on multi-dimensional value evaluation in Embodiment 1 of the present invention.
[0064] Figure 2 It is a schematic structural diagram of the computer terminal in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0066] Embodiment 1
[0067] Please refer to Figure 1 , this embodiment provides a data pricing method based on multi-dimensional value evaluation, including the following steps, namely S1 to S4.
[0068] S1. Desensitize the data to be priced based on the differential privacy mechanism, generate the privacy budget value and privacy risk value of the data, and obtain the privacy score accordingly.
[0069] Step S1 includes the following specific steps S14 to S15.
[0070] S11. Input data preprocessing: Use the Z-score normalization method to normalize the original data to be priced to remove the dimension difference; among them, by calculating the mean and standard deviation of each feature, the data is converted into a standard normal distribution: X norm =(X - μ) / σ; where X norm represents the normalized data, X is the original data; μ and σ are the mean and standard deviation respectively; X n represents a data set composed of any n pieces of data;
[0071] Represent a differential privacy mechanism as M: x n ×Q → Y; where Q is the query space; Y is the output space of mechanism M;
[0072] Define the initial privacy budget value set and prepare to traverse.
[0073] S12. Define the individual privacy risk value:
[0074] PRV f = α·||q(x) - q(x -f )||1 + β·QueryFrequency(f)
[0075] where PRV f is the privacy risk value of individual f in the data set; α and β are weight coefficients used to balance the influence of data sensitivity and query frequency; q(x) represents the calculation query result for the complete data set; x -f is the data set after removing individual f; ||q(x) - q(x -f )||1 represents the influence of data set x -f on the query output; QueryFrequency(f) represents the frequency with which individual f is accessed or mentioned in a specific query.
[0076] S13. Select the decision condition for the privacy budget: When , allocate the privacy budget; where PRV min and PRV max are the minimum and maximum privacy risk values under the influence of the current query period respectively; p is the privacy preference percentage, indicating the acceptable privacy risk difference ratio.
[0077] When the data purchaser and the data owner negotiate the privacy budget for a certain feature or individual, according to the quality evaluation of the processed data set under the selected privacy budget: (the calculated accuracy rate ≥ the specified minimum accuracy rate), then this ∈ can be selected. The goal is to find the ∈ that maximizes the data accuracy and still satisfies privacy protection under the conditions of both parties.
[0078] S14. Define the Laplace mechanism: Given a query and a data set x; denote the k-dimensional real space; the Laplace mechanism M L is defined as:
[0079] M L (x, q, ∈) = q(x) + (Z1, …, Z k )
[0080] where Z1, …, Z k are k random variables drawn from the Laplace distribution, and each random variable has a probability density function:
[0081]
[0082] where p b (Z) is the probability density function of the Laplace distribution, representing the probability that z takes a value under this distribution; z is the noise value sampled from the Laplace distribution; b is the scale function in the Laplace mechanism, Δq is the global sensitivity of the query q; ∈ is the privacy budget value.
[0083] S15. Traverse the set of initial privacy budget values in descending order, calculate the query results with privacy protection added using the Laplace mechanism, and calculate the individual privacy risk value for each individual; according to the privacy preference percentage provided by the data provider, calculate the privacy budget value that meets the privacy considerations of the data provider, form a set of privacy budget candidates, and obtain the privacy score accordingly.
[0084] In step S1, the privacy budget v is dynamically allocated through the differential privacy mechanism. A quantization method based on the individual privacy risk value (PRV) and privacy risk preference parameters are used to calculate the privacy protection effect of the desensitized data in combination with the Laplace mechanism, and the privacy budget allocation strategy is intelligently adjusted. While ensuring strict privacy protection, the accuracy of the data is maximized, providing a reasonable basis for the privacy budget for the subsequent generation of desensitized data.
[0085] S2. Generate a quality score for the data according to the evaluation metrics of the desensitized data on the specified task.
[0086] In this embodiment, the task is a classification task, and the evaluation metric is accuracy. Step S2 includes the following specific steps, namely S21 to S24.
[0087] S21. Receive the structured desensitized data and preprocess the data in different modalities for standardization.
[0088] S22. Initialize the set of privacy budget candidates {∈1, v2, …, v n}, desensitize the data using each privacy budget candidate value to obtain a desensitized dataset.
[0089] {∈1, ∈2, …, ∈ n} are usually selected within a reasonable range to fully explore the trade - off between high privacy protection (small ∈) and high data accuracy (large ∈). In the initial state, set the optimal loss function value to positive infinity and the optimal privacy budget value to the maximum. Next, the algorithm starts to iteratively evaluate each ∈ value in the privacy budget set. Test each ∈ value in ascending order. For each ∈, perform the following steps: desensitize the data using ∈ and calculate the accuracy Acc of the desensitized dataset actual ; calculate the privacy risk of the data provider using ∈ Calculate the overall loss
[0090] Select the optimal privacy budget: for all candidate ∈ that satisfy the constraints i , select the ∈ that minimizes the total loss i as the final optimal privacy budget value ∈ * . When outputting the results, the algorithm returns the optimal privacy budget value ∈ * , the corresponding actual accuracy Acc actual , the ratio of the privacy risk values and the total loss value.
[0091] S23. Perform a classification task on the classifier using the desensitized dataset, traverse the privacy budget candidate values, and calculate the accuracy of the desensitized dataset.
[0092] The specific process is as follows: use the desensitized dataset q ′ (x) to train a model task, such as a classifier, to calculate the accuracy Acc of the desensitized data on an actual task (such as a classification task). actual . This accuracy reflects whether the requirements of the data purchaser are met. If Acc actual is lower than the minimum accuracy threshold Acc threshold specified by the data purchaser, then the current candidate privacy budget value ∈ i will be directly discarded. Divide the original dataset into a training set and a test set according to a fixed ratio (usually 7:3 or 8:2). The training set is used to train the classifier model, and the test set is used to evaluate the model performance.
[0093] S24. Calculate the actual accuracy through cross - validation and obtain the quality score accordingly.
[0094] In this embodiment, the desensitized training set data is used to train a classifier. To ensure the reliability of the evaluation, the cross-validation method is adopted, and the training set is further divided into k subsets (usually k = 5 or 10). In each round of cross-validation, k - 1 subsets are used to train the model, and the remaining one subset is used for validation. This process is repeated k times, and each time a different validation set is selected. Finally, the average value of the k validation results is taken as the performance index of the model. The specific accuracy calculation formula is as follows:
[0095]
[0096] where k is the number of folds of cross-validation, is the i-th round of validation set, f θ is the trained classifier model, is the indicator function, which returns 1 when the predicted value is equal to the true value, and returns 0 otherwise. f θ (x) is the predicted label of the classifier model for the input sample x; y is the true label of the input sample x, that is, the correct classification result marked in the dataset.
[0097] In some embodiments, to comprehensively evaluate the model performance, in addition to accuracy, other evaluation metrics can also be calculated, such as Precision, Recall, and F1-score. Together, they constitute a complete performance evaluation system, which can more comprehensively reflect the practicality of the desensitized data. Finally, these evaluation metrics are compared with the thresholds specified by the data purchaser to determine whether the current privacy budget value meets the data quality requirements.
[0098] In this embodiment, a constraint test is performed on the accuracy to determine whether it meets the requirements of the purchaser. The expression formula is:
[0099]
[0100] In the formula, Quality pass = 1 indicates that the quality passes, otherwise it fails.
[0101] In step S2, through multi-modal data input processing and cross-validation technology, combined with the classifier model, the task accuracy of the desensitized data is evaluated. Through the accuracy calculation formula and the multi-index performance evaluation system, the actual value of the desensitized data is comprehensively quantified. This step ensures that the data purchaser's requirements for data quality are effectively met, providing a quality verification basis for the final selection of the privacy budget.
[0102] S3. Use the bi-objective optimization algorithm to calculate the dynamic weight coefficient that satisfies the optimal privacy budget value, which is used to balance the privacy protection goal and the data quality goal.
[0103] To balance the accuracy requirements of data buyers and the privacy protection requirements of data providers, in this embodiment, we propose a dual-objective privacy budget optimization algorithm, aiming to select an optimal privacy budget value ∈ by weighing the objectives of both parties. This algorithm is based on a dual-objective loss function, which is used to dynamically adjust whether the selected privacy budget can simultaneously meet the constraints of data accuracy and privacy protection.
[0104] Step S3 includes the following specific steps:
[0105] S31. Construct the objective function
[0106]
[0107] Among them, the accuracy loss function is:
[0108]
[0109] If the accuracy requirement has been met, the accuracy loss is 0; if not, the loss is the difference. Acc threshold is the preset accuracy threshold; Acc actual is the actual accuracy;
[0110] The privacy loss function is:
[0111]
[0112] If the privacy requirement has been satisfied, the privacy loss is 0; if not, the loss is the difference.
[0113] To dynamically adjust the priority weight coefficients λ1 and λ2, the Pareto optimization method and the NSGA-II algorithm are used to find the optimal balance point between privacy protection and accuracy.
[0114] S32. Based on the NSGA-II algorithm, perform non-dominated sorting, calculate the dominance relationship and crowding distance of the solutions, so as to select the optimal solution of the objective function, that is, the privacy budget value ∈ that minimizes the total loss of the objective function * , and output the dynamic weight coefficients and
[0115] Initialization: Select several points from the input privacy budget candidate set {∈1, ∈2, …, ∈ n}, and initialize the preference weights λ1, λ2.
[0116] Calculate the objective function value: For each privacy budget value ∈ i , calculate the privacy objective Accuracy target and the total objective function value
[0117] Non - dominated sorting: Use Pareto optimization to perform non - dominated sorting on all candidate points to find the Pareto - optimal front solutions.
[0118] Select the optimal balance point: Select a solution from the Pareto - optimal front to minimize the total objective function value while satisfying the basic constraints of privacy protection and accuracy.
[0119] In step S32, the non - dominated sorting includes the following specific steps:
[0120] S321. Define the dominance relationship between solutions. In the privacy - quality trade - off problem, it is necessary to balance the weights λ1 and λ2 between the two objectives of privacy protection and data accuracy. To this end, it is first necessary to clearly define the dominance relationship between solutions. Consider any two candidate solutions i and j, which each correspond to a set of objective function values: the privacy - protection objective value and and the accuracy target value and Solution i dominates solution j if and only if solution i is strictly better than solution j in at least one objective and not worse than solution j in other objectives. This dominance relationship can be strictly expressed in mathematical form as:
[0121]
[0122] where i and j are any two candidate solutions, represents the dominance relationship; and are a set of privacy - protection objective values; and are a set of accuracy objective values; {privacy, accuracy} is the objective set of privacy - protection loss and data - accuracy loss, and k’ is the currently optimized objective variable; is the loss value of solution i on objective k’; is the loss value of solution j on objective k’; means there exists.
[0123] S322. Define the dominance count and the dominance set; To systematically evaluate the status of each solution in the entire solution space, define the dominance count and the dominance set. For each candidate solution i, calculate two important metrics: the dominance count n i and the dominance set S i . The dominance count n i represents how many solutions dominate solution i among all candidate solutions, that is The dominance set S iThen it includes all the solutions that are i - dominated, that is, The calculation of these two indicators together constitutes the domination relationship network of solutions, providing a basis for the division of the non - dominated front.
[0124] S323. Construct the non - dominated front in order according to the domination count and domination set of all solutions. After obtaining the domination count and domination set of all solutions, construct the non - dominated front. The construction of the non - dominated front is a progressive process, starting from the first front and gradually building backward. The k - th non - dominated front F k includes the solutions with a domination count of 0 after removing all the solutions in the previous k - 1 fronts. This process can be expressed by the mathematical formula:
[0125]
[0126] Adopting a progressive construction method ensures that each solution is correctly assigned to the corresponding non - dominated level, thus forming a complete hierarchical structure of solutions. The solutions in each front represent the set of non - dominated solutions at the current level. These solutions are incomparable with each other, and they all represent the optimal choices under their respective trade - offs.
[0127] S324. Define the crowding distance of each solution in the non - dominated front to reflect the distribution of solutions in the objective space. To further distinguish the quality of solutions in the same non - dominated front, the concept of crowding distance is defined. The crowding distance reflects the distribution of solutions in the objective space. A larger crowding distance means that the solution is more different from other solutions, which helps to maintain the diversity of solutions. For each solution i in the front, its crowding distance d i The calculation needs to consider the differences in function values of adjacent solutions in all objective dimensions and perform normalization. The calculation formula is:
[0128]
[0129] In the formula, d i represents the crowding distance of each solution i in the non - dominated front; and respectively represent the function values of the two adjacent solutions of solution i after sorting by function value on objective m; and respectively represent the maximum and minimum values of all solutions on objective m, which are used for normalization. This calculation method ensures that the distances in different objective dimensions can be reasonably compared and accumulated.
[0130] S325. After completing non-dominated sorting and crowding distance calculation, select the optimal solution of the objective function from the first non-dominated front F1. Considering that the importance of privacy protection and data accuracy may vary, weight coefficients λ1 and λ2 are introduced to reflect this difference. The selection of the optimal solution is achieved by minimizing the weighted objective function:
[0131]
[0132] wherein, represents the privacy loss at any privacy budget ∈ for the privacy protection objective; represents the accuracy loss at any privacy budget ∈ for the accuracy objective.
[0133] After step S3, the output results include:
[0134] Optimal privacy budget value (∈ * ): The privacy budget value corresponding to the privacy-quality balance point found through Pareto optimization.
[0135] Dynamic weight coefficient Optimal weight allocation for privacy and quality objectives, reflecting the priorities of the data provider and the purchaser's requirements.
[0136] Values of each objective function: including privacy objective value, accuracy objective value, and total objective function value, for reference and explanation.
[0137] Step S3 adopts a two-objective optimization algorithm. Based on two objective functions of privacy protection and data quality, a dynamic weight adjustment mechanism is constructed. The balance point between privacy protection and data quality is found through non-dominated sorting and Pareto optimization, and the optimal privacy budget value is output. This step realizes the coordinated balance of the data provider's and purchaser's requirements by dynamically adjusting the weight parameters, providing a scientific input basis for the data pricing module.
[0138] S4. Based on the privacy score, the quality score, the dynamic weight coefficient, and the cost information of the data provider, use a meta-learning model to generate a data pricing formula and output the pricing result.
[0139] The objective of step S4 is to comprehensively consider multi-dimensional factors such as data privacy protection, data quality, and market demand, and dynamically generate reasonable and fair pricing results. Through the input information provided by the previous steps, including privacy scores (such as privacy budget, desensitization degree, etc.), data quality scores (mainly task accuracy, etc.), and the trade-off weights between privacy and quality, this step generates the final pricing result and report based on the meta-learning method, historical transaction data, and data cost information.
[0140] The above information will be normalized to eliminate the dimensional differences between different features:
[0141]
[0142] In the formula, Normalized(·) represents the normalization process; Value j is the original value of a certain score, and δ is a very small value used to prevent the denominator from being zero.
[0143] Data pricing needs to consider the cost factors of data providers, especially the costs of data collection, storage, and processing. Assume that the total cost of the data provider is C cost , then the pricing P final must satisfy the following constraint: P final ≥C cost , to ensure that the data pricing will not be lower than the cost. In addition, data privacy protection and quality scores will affect the final premium or discount, which will be specifically reflected in the subsequent formulas.
[0144] Construction of the meta-learning model:
[0145] To dynamically adapt to different datasets and trading scenarios, step S4 adopts the model-agnostic meta-learning algorithm (MAML, Model-Agnostic Meta-Learning) to train the pricing model based on historical trading data. The meta-learning model quickly generates a model adapted to the new pricing scenario by learning the relationship between price, quality, and privacy factors in historical data.
[0146] Specifically, the loss function of the meta-learning model is defined as:
[0147]
[0148] In the formula, is the loss of the meta-learning model; N is the total number of samples in the historical trading records; is the price of the i'-th data product predicted by the meta-learning model; is the market price of the i'-th data product in the historical trading records.
[0149] After training is completed, the meta-learning model can quickly generate pricing models for different datasets.
[0150] Generation of the data pricing formula:
[0151] Define the privacy score PrivacyScore, and the contribution of the privacy budget ∈ to the privacy score is: A lower ∈ value will make S budget close to 1 (high privacy protection), while a higher ∈ value will make it close to 0 (low privacy protection).
[0152] Define the quality score QualityScore and the actual accuracy Acc actual The contribution to the quality score is as follows: If the actual accuracy reaches or exceeds the threshold Acc set by the purchaser threshold , the score is 1; otherwise, the score is calculated according to the ratio of the actual accuracy to the threshold.
[0153] Based on the output of the meta-learning model and the input information in the foregoing steps, the final data pricing formula is:
[0154]
[0155] Through the above formula, the final pricing result P will be output final and a detailed pricing report. The pricing report includes: the privacy score and quality score of the data; the cost information of the data provider; the demand weights (λ1, λ2) of the buyer and the seller; the detailed calculation process of the pricing result (such as the contribution value of each score).
[0156] Step S4 is based on the meta-learning method, combines the privacy score (PrivacyScore) and the quality score (QualityScore), as well as the cost information of the data provider, to construct a dynamic pricing model. Dynamically adjust the impact of privacy protection and data quality on pricing through weight coefficients, generate a scientific and reasonable data pricing result and a detailed report, and ensure the fairness and transparency of data transactions.
[0157] Embodiment 2
[0158] This embodiment provides a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the pricing method described in Embodiment 1 are implemented.
[0159] As Figure 2 shown, the computer terminal provided in this embodiment includes: at least one processor 101, and a memory 102 connected to at least one processor 101. In this embodiment, the specific connection medium between the processor 101 and the memory 102 is not limited Figure 2 and it is taken as an example that the processor 101 and the memory 102 are connected through a bus 100. The bus 100 is Figure 2 shown in thick lines, and the connection manners between other components are only illustrative and not restrictive. The bus 100 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation Figure 2It is represented only by a thick line, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 101 may also be referred to as a controller, and there is no restriction on the name.
[0160] In this embodiment, the memory 102 stores instructions executable by at least one processor 101. By executing the instructions stored in the memory 102, the at least one processor 101 can execute the foregoing method.
[0161] Among them, the processor 101 is the control center of the device. It can connect various parts of the entire control device through various interfaces and lines. By running or executing the instructions stored in the memory 102 and calling the data stored in the memory 102, various functions of the device and process data, thereby performing overall monitoring of the device.
[0162] In a possible design, the processor 101 may include one or more processing units. The processor 101 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 101. In some embodiments, the processor 101 and the memory 102 may be implemented on the same chip. In some embodiments, they may also be separately implemented on independent chips.
[0163] The processor 101 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the pricing method disclosed in conjunction with Embodiment 1 can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor 101.
[0164] The memory 102, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 102 can include at least one type of storage medium. For example, it can include flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disc, and so on. The memory 102 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 102 in this embodiment can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0165] By programming the design of the processor 101, the code corresponding to the security verification method introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute Figure 1 the steps of the pricing method shown. How to program the design of the processor 101 is a well-known technology to those skilled in the art and will not be elaborated herein.
[0166] Embodiment 3
[0167] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the pricing method as described in Embodiment 1 are implemented.
[0168] The computer-readable storage medium may include flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the storage medium may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is generally used to store the operating system installed on the computer device and various application software. In addition, the memory may also be used to temporarily store various data that have been output or are to be output.
[0169] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A data pricing method based on multi-dimensional value assessment, characterized in that: The following steps are involved: S1. Desensitize the data to be priced based on the differential privacy mechanism, generate the privacy budget value and privacy risk value of the data, and obtain the privacy score based on this; S2. Generate a quality score for the data based on the evaluation indicators of the desensitized data on the specified task; S3. Use a dual-objective optimization algorithm to calculate a dynamic weight coefficient that satisfies the optimal privacy budget value, which is used to balance the privacy protection goal and the data quality goal. S4. Based on the privacy score, the quality score, the dynamic weight coefficient and the cost information of the data provider, a data pricing formula is generated using a meta-learning model, and a pricing result is output.
2. The data pricing method based on multi-dimensional value assessment according to claim 1 is characterized in that: Step S1 includes the following specific steps: S11. Input data preprocessing: The original data is standardized using the Z-score standardization method to remove dimensional differences; the data is converted to a standard normal distribution by calculating the mean and standard deviation of each feature: X norm =(X-μ) / σ; where x norm represents the standardized data, X is the original data; μ and σ are the mean and standard deviation respectively; Denote a differential privacy mechanism as M:X n ×Q→Y; where Q is the query space; Y is the output space of mechanism M; X n Represents a data set consisting of any n pieces of data; Define the initial privacy budget value set and prepare for traversal; S12. Define individual privacy risk value: PRV f =α·||q(x)-q(x -f )||1+β·QueryFrequency(f) Where PRV f is the privacy risk value of individual f in the data set; α and β are weight coefficients used to balance the impact of data sensitivity and query frequency; q(x) represents the query result of the complete data set; x -f is the data set after removing individual f; ||q(x)-q(x -f )||1 represents the data set x -f Impact on query output; QueryFrequency(f) indicates the frequency with which individual f is accessed or mentioned in a specific query; S13. Criteria for selecting privacy budget: When PRV min and PRV max are the minimum and maximum privacy risk values under the influence of the current query period, respectively; p is the privacy preference percentage, indicating the acceptable privacy risk difference ratio; S14. Define the Laplace mechanism: Given a query and a dataset x; represents k-dimensional real number space; Laplace mechanism M L Defined as: M L (x,q,∈)=q(x)+(Z1,…,Z k ) Where Z1,…,Z k are k random variables drawn from a Laplace distribution, each with a probability density function: In the formula, p b (z) is the probability density function of the Laplace distribution, which indicates the probability of z taking a value under this distribution; z is the noise value sampled from the Laplace distribution; b is the scale function in the Laplace mechanism, Δq is the global sensitivity of query q; ∈ is the privacy budget value; S15. Traverse the initial privacy budget value set in descending order, use the Laplace mechanism to calculate the query results with added privacy protection, and calculate the individual privacy risk value of each individual; according to the privacy preference percentage provided by the data provider, calculate the privacy budget value that meets the privacy considerations of the data provider, form a privacy budget candidate set, and obtain the privacy score based on this.
3. The data pricing method based on multi-dimensional value assessment according to claim 2 is characterized in that: Step S2 includes the following specific steps: S21. Receive structured desensitized data and preprocess data of different modalities for standardization; S22. Initialize the privacy budget candidate set {∈1,∈2,…,∈ n }, use each privacy budget candidate value to desensitize the data and obtain a desensitized dataset; S23. Perform the classification task on the classifier using the desensitized dataset, traverse the privacy budget candidate values and calculate the accuracy of the desensitized dataset; S24. Calculate the actual accuracy through cross-validation, and obtain the quality score accordingly.
4. The data pricing method based on multi-dimensional value assessment according to claim 3 is characterized in that: Step S3 includes the following specific steps: S31. Constructing the objective function In the formula, is the accuracy loss function, is the privacy loss function, λ1 and λ2 are and The corresponding weight coefficient; Acc threshold is the preset accuracy threshold; Acc actual is the actual accuracy; S32. Perform non-dominated sorting based on the NSGA-II algorithm, calculate the dominance relationship and crowding distance of the solution, and then select the optimal solution of the objective function, that is, the privacy budget value ∈ that minimizes the total loss of the objective function * , and output dynamic weight coefficient and 5. The data pricing method based on multi-dimensional value assessment according to claim 4 is characterized in that: In step S32, the non-dominated sorting includes the following specific steps: S321. Define the dominance relationship between solutions, expressed as: In the formula, i and j are any two candidate solutions, < indicates a dominance relationship; and is a set of privacy protection target values; and is a set of accuracy target values; {privacy,accuracy} is the target set of privacy protection loss and data accuracy loss, and k' is the target variable currently being optimized; is the loss value of solution i on target k'; is the loss value of solution j on target k'; Indicates existence; S322. Define dominance count and dominance set; where for each candidate solution i, its dominance count is n i , the dominating set is S i ; Dominance count n i Indicates the number of solutions that dominate solution i among all candidate solutions; the dominating set S i represents all solutions dominated by solution i; S323. Construct the non-dominated frontier in order according to the domination counts and domination sets of all solutions; wherein the kth non-dominated frontier F k Contains solutions with a dominance count of 0 after removing all solutions in the first k-1 fronts; S324. Define the crowding distance of each solution in the non-dominated frontier to reflect the distribution of the solution in the target space; wherein the crowding distance calculation formula of the solution is: Where, d i represents the crowding distance of each solution i in the non-dominated front; and They represent the function values of the two adjacent solutions of solution i after sorting by function value on target m; and Respectively represent the maximum and minimum values of all solutions on the target m; S325. After completing the non-dominated sorting and crowding distance calculation, the optimal solution of the objective function is selected from the first non-dominated front F1, and the expression formula is: In the formula, represents the privacy loss under any privacy budget ∈ on the privacy protection target; It represents the accuracy loss under any privacy budget ∈ on the accuracy target.
6. The data pricing method based on multi-dimensional value assessment according to claim 3 is characterized in that: In step S15, the contribution of the privacy budget ∈ to the privacy score is: In step S24, the actual accuracy Acc actual Contributions to the quality score are: Acc threshold The minimum accuracy threshold specified for data buyers.
7. The data pricing method based on multi-dimensional value assessment according to claim 6 is characterized in that: In step S4, the data pricing formula is: Where P final Pricing data; C cost is the cost of the data provider; PrivacyScore is the privacy score; QualityScore is the quality score.
8. The data pricing method based on multi-dimensional value assessment according to claim 7 is characterized in that: In step S4, a model-independent meta-learning algorithm is used to train the pricing model based on historical transaction data. The meta-learning model after training can generate pricing models for different data sets; wherein the loss function of the meta-learning model is defined as: In the formula, is the loss of the meta-learning model; N is the total number of samples in the historical transaction records; The price of the i'th data product predicted by the meta-learning model; is the market price of the i'th data product in the historical transaction record.
9. A computer terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of the data pricing method based on multi-dimensional value assessment as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the data pricing method based on multi-dimensional value assessment as described in any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Smart city multi-source data fusion quality evaluation and encryption method and system
CN120849404A
A method and system for quality assessment and encryption of multi-source data fusion in smart cities
CN120849404B