A fairness-based prediction method and system for multi-dimensional privacy data
Patent Information
- Application Number
- CN202310703296.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-14
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-06-14
AI Technical Summary
[0006]本发明针对现有多元数据预测方法中对隐私性保护及公平性保证考虑的不足问题,选择准确性更高的预测方法,对敏感数据同时保护隐私和保证公平的措施,提出了一种基于公平性的面向多元隐私数据的预测方法及系统,本发明所提出的方法在考虑多元数据预测的隐私这一基础上,研究数据预测隐私性和公平性保护结合、多元数据公平性保护、形成安全、公平且保证可用性的预测方法,最终建立面向应用需求的、基于多元数据的、同时考虑隐私性和公平性的预测模型
[0057]本发明在考虑多元数据预测的隐私这一基础上,研究数据预测隐私性和公平性保护结合、多元数据公平性保护、形成安全、公平且保证可用性的预测方法,最终建立面向应用需求的、基于多元数据的、同时考虑隐私性和公平性的预测模型。同时,本发明引入了切比雪夫多项式作为目标函数的展开方式,以便在降低误差方面具有更好的效果。与传统的泰勒展开方式相比,切比雪夫多项式可以更好地逼近目标函数,从而获得更好的展开精度和全局最优解;使用切比雪夫多项式作为目标函数的展开方式,可在降低误差方面具有显著成效。
Smart Images

Figure CN116894263B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of information security and privacy data mining technology, specifically, it relates to a prediction method and system for diverse privacy data based on fairness. Background Technology
[0002] Machine learning has achieved remarkable success in many fields. For example, applying machine learning algorithms in the data space can effectively promote data resource sharing and value release. However, because the success of machine learning largely depends on massive amounts of individual data, there are growing concerns about potential privacy breaches and unfairness in training and deploying machine learning algorithms. In the data space, multiple participants share large amounts of training data with the main server, which necessitates algorithm providers to offer privacy-preserving and fair algorithms.
[0003] Differential privacy (DP) has become the most commonly used method for protecting data privacy. There are many ways to achieve differential privacy, such as adding noise to the original data or results, adding noise to the gradient of each model iteration, or injecting noise into the objective function. The concept of objective perturbation was first proposed, which involves injecting noise into the polynomial coefficients of the objective function to achieve privacy protection. For prediction models, function mechanisms are very common; however, existing function mechanism methods truncate the last term after the binomial term in the Taylor series, which can lead to a large discrepancy between the results and the true values.
[0004] Meanwhile, unfair phenomena are emerging in various fields. For example, classification models in automated job recruitment systems tend to hire candidates from specific racial or gender groups, leading to increasing attention on fairness-aware learning in machine learning. Consequently, much work focuses on developing algorithms for designing fair classification models. Furthermore, considering the intrinsic relationship between privacy and fairness and striking a balance between the two is also a concern. For instance, analysis of US population privacy datasets from a fairness perspective reveals that adding noise to privacy-enhancing data can severely impact certain groups, and guidelines for mitigating these effects have been proposed.
[0005] However, existing work on combining fairness and privacy assumes that the attributes in the dataset are independent of each other, and only adds noise to a pre-selected attribute to be protected, without considering the influence of other attributes on the protected attribute. In real-world scenarios, the attributes involved in the dataset cannot be independent of each other. That is to say, even if sensitive attributes are protected, unfairness may occur because other attributes directly or indirectly affect the sensitive attributes. This is a problem that urgently needs to be solved. Summary of the Invention
[0006] This invention addresses the shortcomings of existing multivariate data prediction methods in considering privacy protection and fairness guarantees. It selects a more accurate prediction method and proposes a fairness-based prediction method and system for multivariate privacy-preserving data, taking into account the privacy aspects of multivariate data prediction. The proposed method, while considering privacy in multivariate data prediction, studies the combination of privacy and fairness protection in data prediction, the protection of fairness in multivariate data, and the formation of a secure, fair, and usable prediction method. Ultimately, it establishes a prediction model based on multivariate data that is application-oriented and simultaneously considers privacy and fairness.
[0007] This invention is achieved through the following technical solution:
[0008] A fairness-based prediction method for multi-source privacy data: The method specifically includes the following steps:
[0009] Step 1: Determine the prediction task and sensitive attributes for a given logistic regression multivariate dataset;
[0010] Step 2: Using the concept of a function mechanism, expand the logistic regression loss function into a polynomial form using Chebyshev polynomials;
[0011] Step 3: Use the decision tree algorithm to select the attribute that has the greatest impact on the sensitive attribute;
[0012] Step 4: Use a Bayesian network to select a group of attributes that have the greatest impact on the sensitive attributes;
[0013] Step 5: Combining Step 3 and Step 4, add Laplace noise with a fairness constraint penalty term to the coefficients of the polynomial function expanded in Step 2.
[0014] Step 6: Perform a prediction task on the dataset using a loss function that satisfies fairness and differential privacy guarantees.
[0015] Further, in step 1,
[0016] Given a dataset D = {X, X...} containing n tuples p Let X = {X1, X2, ..., XY}, where X = {X1, X2, ..., XY}. d} represents d unprotected attributes; X p Y represents protected attributes, i.e., sensitive attributes; Y is the tag.
[0017] Dataset D contains n tuples, each tuple is represented as t. i =(x i y i ), where the eigenvector x i It contains d unprotected attributes and one protected attribute, namely x.i =(x i1 x i2 , ···, x id ;x ip ), y i Indicates the corresponding label;
[0018] Protect X p Based on this, further consider other attributes X = {X1, X2, ..., X...} d} may be related to X p The close relationship between these attributes means they will also affect fairness. Therefore, an algorithm selects one or more other attributes that are most closely related to the protected attribute as the protected objects.
[0019] When choosing X p When choosing the attribute with the closest relationship, let's say the selected attribute is X. k (X k When multiple attributes are selected, let this set of attributes be G = {X}. k1 X k2 ,...},
[0020] Furthermore, step 2 includes the following steps:
[0021] Step 2.1: Use Chebyshev expansion to derive a polynomial approximation of the objective function for logistic regression;
[0022] Given an objective function E(D, W), where the training dataset is D and the model training parameters are W; since the objective function E(D, W) is a complex function of w, the function mechanism can be extended to a polynomial form and a Laplace mechanism can be deployed on the parameter terms; therefore, the function mechanism for the objective function E(D, W) can be formally represented as:
[0023]
[0024] in For each term of the polynomial, represents the parameters of each term.
[0025] Step 2.2: Truncate the polynomial to the square term as the expanded objective function, and extract the coefficients of each independent variable as the noise-adding objects for subsequent steps;
[0026] The number of terms in the Chebyshev polynomial can be specified as needed, usually n. This truncates the Chebyshev polynomial to the nth term. In this step, the formula obtained in step 2.1 is binarized; the function mechanism FM is applied to E... LSTM (D, W) is expanded using Chebyshev polynomials as follows: Where φ(w) represents a polynomial function. This represents the coefficient of the polynomial.
[0027] Furthermore, step 3 includes the following steps:
[0028] Step 3.1: Considering the impact of a single attribute on a sensitive attribute, a decision tree analysis algorithm is used to select the attribute that has the greatest impact on the sensitive attribute;
[0029] For choosing a pair X p The attribute with the greatest impact is selected using an attribute decision tree, denoted as X. k ;
[0030] The loss function for decision tree learning can be defined as follows: Where T represents the number of leaf nodes in the tree, H t (T) represents the entropy of the t-th leaf, i.e.: N t This indicates the number of training samples contained in the leaf node, and α represents the penalty coefficient;
[0031] In the use of attribute decision trees, the ultimate goal is to find the correct match for X. p The attribute with the greatest impact, rather than the attribute with the greatest impact on label Y, is chosen as the optimal score in the decision tree based on accurately predicting X. p The value;
[0032] Step 3.2: For the attribute selected in Step 3.1 that has the greatest impact on sensitive attributes, add stronger noise than other attributes to achieve fair constraints. The privacy budget allocated to this attribute is the total privacy budget multiplied by (number of attributes - 1) / number of attributes. The privacy budget allocated to other irrelevant attributes is the ratio of the total privacy budget to the number of attributes.
[0033] Furthermore, step 4 includes the following steps:
[0034] Step 4.1: Considering the combined influence of multiple attributes on the sensitive attribute, a weighted graph of the relationship between attributes is preprocessed using a Bayesian network, where each edge represents the relationship between two attributes and the edge weight reflects the degree of influence between attributes. A set (several) of attributes that have the greatest influence on the sensitive attribute is selected.
[0035] Step 4.2: For the weighted graph preprocessed in Step 4.1, select the first few attributes as protection attributes for fairness assurance. Based on the magnitude of the influence of each attribute on the sensitive attributes, and based on the allocation of different privacy budgets, add more noise to the attributes with greater influence to ensure fairness.
[0036] Choose a set of attributes G = {xk1 x k2 , ..., x kn As influencing factors, privacy budgets are allocated according to their weight.
[0037] Furthermore, step 5 includes the following steps:
[0038] Step 5.1: Based on the results obtained in Steps 3 and 4, use the idea of Lagrange multipliers to add a fairness constraint as a penalty term to the objective function;
[0039] The original objective function E(D, W), after adding fairness constraints, is defined as follows: Where λ is a Lagrange multiplier, setting λ = 1 and τ = 0, we can obtain:
[0040]
[0041] Step 5.2: For the objective functions obtained in Step 3 and Step 4, calculate the set of parameters that minimizes the loss to complete the training of the model.
[0042] The model parameters w = (w1, w2, ..., wd) are used to output the predicted value y by minimizing the empirical loss of the training dataset D in the parameter space P; that is, to solve the following optimization problem.
[0043]
[0044] Among them, w * These are the learned parameters, t i Let be the coefficients of each term in the polynomial, E be the loss function, E(D, W) be the objective function, and W be the model training parameters. Considering logistic regression as the loss function, we have:
[0045]
[0046] A fairness-based prediction system for diverse privacy-preserving data:
[0047] The system includes: a dataset module, a multinomial expansion module, a decision tree module, a Bayesian module, a noise module, and a prediction module;
[0048] The dataset module determines the prediction task and sensitive attributes of a given logistic regression multivariate dataset;
[0049] The polynomial expansion module uses the idea of a function mechanism to expand the logistic regression loss function into a polynomial form using Chebyshev polynomials.
[0050] The decision tree module uses a decision tree algorithm to select the attribute that has the greatest impact on the sensitive attribute;
[0051] The Bayesian module uses a Bayesian network to select a group of attributes that have the greatest impact on the sensitive attributes.
[0052] The noise module combines the decision tree module and the Bayesian module, adding Laplace noise with a fairness constraint penalty term to the coefficients of the polynomial function expanded by the polynomial expansion module.
[0053] The prediction module uses a loss function that satisfies fairness and differential privacy guarantees to perform prediction tasks on the dataset.
[0054] An electronic device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the above method.
[0055] A computer-readable storage medium for storing computer instructions that, when executed by a processor, implement the steps of the above-described method.
[0056] Beneficial effects of the invention
[0057] This invention, considering the privacy aspects of multivariate data prediction, researches a prediction method that combines privacy and fairness protection in data prediction, protects fairness in multivariate data, and forms a secure, fair, and usable prediction method. Ultimately, it establishes a prediction model based on multivariate data that is application-oriented and considers both privacy and fairness. Furthermore, this invention introduces Chebyshev polynomials as the expansion method for the objective function, resulting in better error reduction. Compared to traditional Taylor expansions, Chebyshev polynomials can better approximate the objective function, thus obtaining better expansion accuracy and a globally optimal solution; using Chebyshev polynomials as the expansion method for the objective function significantly reduces errors.
[0058] In summary, this invention introduces a joint privacy metric, fairness constraints, and Chebyshev polynomials to address data privacy and fairness issues more comprehensively, and provides new ideas for research in related fields. Attached Figure Description
[0059] Figure 1 This is a flowchart of the method of the present invention;
[0060] Figure 2 This is a schematic diagram illustrating the decision tree analysis algorithm used in step 3 of the present invention;
[0061] Figure 3 This is a schematic diagram illustrating the use of the Bayesian network algorithm in step 4 of this invention. Detailed Implementation
[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0063] Combination Figure 1 and Figure 3 ;
[0064] A fairness-based prediction method for multi-source privacy data: The method specifically includes the following steps:
[0065] Step 1: Determine the prediction task and sensitive attributes for a given logistic regression multivariate dataset;
[0066] The prediction task and sensitive attributes of the dataset were determined. For example, in a medical dataset, the prediction task is to determine whether a patient has a certain disease, and this disease is the sensitive attribute.
[0067] Step 2: Using the concept of a function mechanism, expand the logistic regression loss function into a polynomial form using Chebyshev polynomials;
[0068] The logistic regression loss function is expanded into a polynomial form for subsequent operations. Using Chebyshev polynomial expansion allows for a better approximation of the objective function, resulting in improved expansion accuracy and a closer global optimum.
[0069] Step 3: Use the decision tree algorithm to select the attribute that has the greatest impact on the sensitive attribute;
[0070] The decision tree algorithm is used to select attributes that are associated with the sensitive attribute, i.e., attributes that have the greatest impact on the sensitive attribute. These attributes will be used for subsequent fairness constraints to prevent unreasonable behavior.
[0071] Step 4: Use a Bayesian network to select a group of attributes that have the greatest impact on the sensitive attributes;
[0072] Using Bayesian networks to select attribute combinations that can simultaneously affect sensitive attributes and other attributes as input for subsequent steps helps narrow down the range of attributes, thereby reducing the complexity of the screening process.
[0073] Step 5: Combining Step 3 and Step 4, add Laplace noise with a fairness constraint penalty term to the coefficients of the polynomial function expanded in Step 2.
[0074] The selected attributes related to the sensitive attribute and combinations that can simultaneously affect the sensitive attribute and other attributes are used as inputs to the fairness constraint. Laplace noise with a fairness constraint penalty term is added to ensure that the prediction task and the fairness constraint are satisfied simultaneously.
[0075] Step 6: Perform a prediction task on the dataset using a loss function that satisfies fairness and differential privacy guarantees.
[0076] By using Laplace noise with a fairness constraint penalty term, a loss function that satisfies fairness and differential privacy guarantees is obtained, which is then used for prediction tasks on multivariate datasets.
[0077] Given a dataset D = {X, X...} containing n tuples p Let X = {X1, X2, ..., XY}, where X = {X1, X2, ..., XY}. d} represents d unprotected attributes; X p Y represents protected attributes, i.e., sensitive attributes; Y is the tag.
[0078] Dataset D contains n tuples, each tuple is represented as t. i =(x i y i ), where the eigenvector x i It contains d unprotected attributes and one protected attribute, namely x. i =(x i1 x i2 , ···, x id ;x ip ), y i Indicates the corresponding label;
[0079] Considering fairness, fairness issues mainly arise with a specific attribute, specifically X. p It will directly affect the tag y i The prediction, that is
[0080] Protect X p Based on this, further consider other attributes X = {X1, X2, ..., X...} d} may be related to X p The close relationship between these attributes means they will also affect fairness. Therefore, an algorithm selects one or more other attributes that are most closely related to the protected attribute as the protected objects.
[0081] When choosing X p When choosing the attribute with the closest relationship, let's say the selected attribute is X. k (X kWhen multiple attributes are selected, let this set of attributes be G = {X}. k1 X k2 , ...},
[0082] Step 2 includes the following steps:
[0083] Step 2.1: Use Chebyshev expansion to derive a polynomial approximation of the objective function for logistic regression;
[0084] Given an objective function E(D, W), where the training dataset is D and the model training parameters are W; since the objective function E(D, W) is a complex function of w, the function mechanism can be extended to a polynomial form and a Laplace mechanism can be deployed on the parameter terms; according to the Stone-Weierstrass theorem, any continuously differentiable function can be written as a polynomial of w, therefore, the function mechanism for the objective function E(D, W) can be formally represented as:
[0085]
[0086] in For each term of the polynomial, represents the parameters of each term.
[0087] Step 2.2: Truncate the polynomial to the square term as the expanded objective function, and extract the coefficients of each independent variable as the noise-adding objects for subsequent steps;
[0088] Each term of a Chebyshev polynomial is a function of the previous term, therefore it can be viewed as a recursive expression. The number of terms in a Chebyshev polynomial can be specified as needed, usually n. In this case, the Chebyshev polynomial will be truncated to the nth term. In this step, the formula obtained in step 2.1 is binarized; the function mechanism FM is applied to E... LSTM (D, W) is expanded using Chebyshev polynomials as follows: Where φ(w) represents a polynomial function. This represents the coefficient of the polynomial.
[0089] like Figure 2 As shown, step 3 includes the following steps:
[0090] Step 3.1: Considering the impact of a single attribute on a sensitive attribute, a decision tree analysis algorithm is used to select the attribute that has the greatest impact on the sensitive attribute;
[0091] For choosing a pair X p The attribute with the greatest impact is selected using an attribute decision tree, denoted as X. k ;
[0092] An attribute decision tree is a tree-like structure. Each non-leaf node represents a feature attribute, each branch represents the output of that feature attribute within a specific range, and each leaf node corresponds to the final decision result. The decision-making process using a decision tree starts from the root node, tests the corresponding feature attribute for the item to be classified, and selects the output branch based on its value, until a leaf node is reached. The class corresponding to the leaf node is then taken as the decision result. The goal of attribute decision tree learning is to generate decision trees with strong generalization performance, i.e., strong example processing capabilities.
[0093] The loss function for decision tree learning can be defined as follows: Where T represents the number of leaf nodes in the tree, H t (T) represents the entropy of the t-th leaf, i.e.: N t This indicates the number of training samples contained in the leaf node, and α represents the penalty coefficient;
[0094] In the use of attribute decision trees, the ultimate goal is to find the correct match for X. p The attribute with the greatest impact, rather than the attribute with the greatest impact on label Y, is chosen as the optimal score in the decision tree based on accurately predicting X. p The value of the value; as the splitting continues, the samples contained in each branch of the decision tree will increasingly belong to the same class, that is, the "purity" of the nodes will become higher and higher. However, in order to obtain a decision tree with strong generalization performance, the attribute that maximizes the "purity improvement" of the samples after the split should be selected as the optimal split.
[0095] Step 3.2: For the attribute selected in Step 3.1 that has the greatest impact on sensitive attributes, add stronger noise than other attributes to achieve fair constraints. The privacy budget allocated to this attribute is the total privacy budget multiplied by (number of attributes - 1) / number of attributes. The privacy budget allocated to other irrelevant attributes is the ratio of the total privacy budget to the number of attributes.
[0096] like Figure 3 As shown, step 4 includes the following steps:
[0097] Step 4.1: Considering the combined influence of multiple attributes on the sensitive attribute, a weighted graph of the relationship between attributes is preprocessed using a Bayesian network, where each edge represents the relationship between two attributes and the edge weight reflects the degree of influence between attributes. A set (several) of attributes that have the greatest influence on the sensitive attribute is selected.
[0098] In a directed acyclic graph G, each node corresponds to a variable in a K-dimensional random vector X, and there is a directed edge e. ij Represents random variable X i and X j There is a causal relationship between them, so these two points must be unconditionally independent. Let X...πk For variable X k The set of all parent node variables, P(X) k |X πk () represents the local conditional probability distribution of each random variable;
[0099] If the joint probability distribution of X can be decomposed into each random variable X k The product of local conditional probabilities, i.e. Then (G, X) constitutes a Bayesian network. Using the classic PC algorithm based on dependency statistics, the dependencies of each attribute are analyzed, connecting edges are added between two nodes with high dependencies, and then the direction of the edges is determined based on inclusion relationships and other methods to obtain the final directed acyclic graph.
[0100] By learning parameters to determine network parameters, in the prediction problem of multivariate data, all variables in the resulting directed acyclic graph are observable and do not contain latent variables. Therefore, a weighted directed acyclic graph can be directly calculated using the maximum likelihood estimation method. Each node in the graph represents an attribute, and the weights reflect the degree of relationship between two attributes.
[0101] Step 4.2: For the weighted graph preprocessed in Step 4.1, select the first few attributes as protection attributes for fairness assurance. Based on the magnitude of the influence of each attribute on the sensitive attributes, and based on the allocation of different privacy budgets, add more noise to the attributes with greater influence to ensure fairness.
[0102] Choose a set of attributes G = {x k1 x k2 , ..., x kn As influencing factors, privacy budgets are allocated according to their weight.
[0103] Step 5 includes the following steps:
[0104] Step 5.1: Based on the results obtained in Steps 3 and 4, use the idea of Lagrange multipliers to add a fairness constraint as a penalty term to the objective function;
[0105] A classification model If satisfied This classification model is considered to satisfy "Demographic Parity." Here, x is the attribute to be protected; that is, whether x is 0 or 1 has no impact on the model's prediction of y as 0 or 1. The model's discriminative power can be quantified by the risk difference (RD), formally expressed as:
[0106]
[0107] To meet the requirements of differential privacy and fairness, the idea of the mathematical Lagrange multiplier method is adopted, which treats the fairness constraint as a penalty term and adds it to the objective function of the function mechanism. This can effectively combine the function mechanism and decision boundary fairness, and transform a constrained optimization problem into an unconstrained problem.
[0108] The original objective function E(D, W), after adding fairness constraints, is defined as follows: Here, λ is the Lagrange multiplier, which plays a crucial role in maintaining the balance between model availability and fairness. For ease of discussion, let λ = 1 and τ = 0. Note that the following analysis still holds true when they are assigned other values. We can obtain:
[0109]
[0110] Step 5.2: For the objective functions obtained in Step 3 and Step 4, calculate the set of parameters that minimizes the loss to complete the training of the model.
[0111] For a binary classification task, assume Where x ij ≥0, y i Given a class ∈ {0, 1}, the goal is to construct a binary classification model P(x, y) with input x and parameters w = (w1, w2, ..., wd), outputting the predicted value y by minimizing the empirical loss of the training dataset D in the parameter space of P; that is, to solve the following optimization problem.
[0112]
[0113] Among them, w * Yes, t i Yes, E is the loss function, E(D, W) is the objective function, and W is the model training parameters. Considering logistic regression as the loss function, we have:
[0114]
[0115] A fairness-based prediction system for diverse privacy-preserving data:
[0116] The system includes: a dataset module, a multinomial expansion module, a decision tree module, a Bayesian module, a noise module, and a prediction module;
[0117] The dataset module determines the prediction task and sensitive attributes of a given logistic regression multivariate dataset;
[0118] The polynomial expansion module uses the idea of a function mechanism to expand the logistic regression loss function into a polynomial form using Chebyshev polynomials.
[0119] The decision tree module uses a decision tree algorithm to select the attribute that has the greatest impact on the sensitive attribute;
[0120] The Bayesian module uses a Bayesian network to select a group of attributes that have the greatest impact on the sensitive attributes.
[0121] The noise module combines the decision tree module and the Bayesian module, adding Laplace noise with a fairness constraint penalty term to the coefficients of the polynomial function expanded by the polynomial expansion module.
[0122] The prediction module uses a loss function that satisfies fairness and differential privacy guarantees to perform prediction tasks on the dataset.
[0123] This invention protects the privacy and fairness of multivariate datasets and ensures that prediction tasks can be successfully completed by using Chebyshev multinomials, decision tree algorithms, Bayesian networks, and Laplace noise with fairness constraint penalties.
[0124] An electronic device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the above method.
[0125] A computer-readable storage medium for storing computer instructions that, when executed by a processor, implement the steps of the above-described method.
[0126] The memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the methods described in this invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0127] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means such as coaxial cable, optical fiber, digital subscriber line, DSL, or wireless means such as infrared, wireless, microwave, etc. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium such as a floppy disk, hard disk, magnetic tape; an optical medium such as a high-density digital video disc, DVD; or a semiconductor medium such as a solid-state disk, SSD, etc.
[0128] In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software. The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or by a combination of hardware and software modules in the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are omitted here.
[0129] It should be noted that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiments can be completed by the integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied as execution by a hardware decoding processor, or as execution by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above methods.
[0130] The foregoing has provided a detailed description of the fairness-based prediction method and system for multi-source privacy data proposed in this invention, and has elucidated the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.
Claims
1. A fairness-based prediction method for multi-source privacy data, characterized in that: The method specifically includes the following steps: Step 1: Determine the prediction task and sensitive attributes for a given logistic regression multivariate dataset; Given a dataset D={X, ..., n tuples , Y}, where X={ This represents d unprotected attributes; Y represents protected attributes, i.e., sensitive attributes; Y is the tag. Dataset D contains n tuples, each tuple is represented as , where the feature vector It contains d unprotected attributes and one protected attribute, that is... = ( , , · · ·, ; ), Indicates the corresponding label; Protect Based on this, further consider other attributes X={ Possibly with If the relationship is close, select one or more other attributes that are most closely related to the protected attribute as the protected objects. When choosing a pair When choosing the attribute with the closest relationship, let the selected attribute be... ( When multiple attributes are selected, let this group of attributes be... , ( ); Step 2: Using the concept of a function mechanism, expand the logistic regression loss function into a polynomial form using Chebyshev polynomials; Step 2.1: Use Chebyshev expansion to derive a polynomial approximation of the objective function for logistic regression; Given the objective function The training dataset is The model training parameters are Due to the objective function yes The function is complex, so the function mechanism extends it to a polynomial form and deploys the Laplace mechanism on the parameter terms; therefore, for the objective function... The function mechanism can be formally represented as: in For each term of the polynomial, These are the parameters of each term; Step 2.2: Truncate the polynomial to the square term as the expanded objective function, and extract the coefficients of each independent variable as the noise-adding objects for subsequent steps; The number of terms in a Chebyshev polynomial can be specified as needed, and the number of terms can be chosen as follows: Then the Chebyshev polynomial will be truncated to the th... In this step, the formula obtained in step 2.1 is truncated by binomial calculation; the function mechanism FM is applied to... Using Chebyshev polynomial expansion ,in Represents a polynomial function. This represents the coefficient of the polynomial. Step 3: Use the decision tree algorithm to select the attribute that has the greatest impact on the sensitive attribute; Step 4: Use a Bayesian network to select a group of attributes that have the greatest impact on the sensitive attributes; Step 5: Combining Step 3 and Step 4, add Laplace noise with a fairness constraint penalty term to the coefficients of the polynomial function expanded in Step 2. Step 6: Perform a prediction task on the dataset using a loss function that satisfies fairness and differential privacy guarantees.
2. The method according to claim 1, characterized in that: Step 3 includes the following steps: Step 3.1: Considering the impact of a single attribute on a sensitive attribute, a decision tree analysis algorithm is used to select the attribute that has the greatest impact on the sensitive attribute; For choosing a pair The attribute with the greatest impact is selected using an attribute decision tree, denoted as . ; The loss function for decision tree learning can be defined as follows: Where T represents the number of leaf nodes in the tree. Let the entropy of the t-th leaf be: , This indicates the number of training samples contained in the leaf node. Indicates the penalty coefficient; In the use of attribute decision trees, the ultimate goal is to find the correct attribute. The biggest impact is not on the tags The attribute with the greatest impact is chosen, therefore the selection criterion for the optimal score in a decision tree is the accuracy of prediction. The value; Step 3.2: For the attribute selected in Step 3.1 that has the greatest impact on sensitive attributes, add stronger noise than other attributes to achieve fair constraints. The privacy budget allocated to this attribute is the total privacy budget multiplied by (number of attributes – 1) / number of attributes. The privacy budget allocated to other irrelevant attributes is the ratio of the total privacy budget to the number of attributes.
3. The method according to claim 2, characterized in that: Step 4 includes the following steps: Step 4.1: Considering the combined influence of multiple attributes on the sensitive attribute, a weighted graph of the relationship between attributes is preprocessed using a Bayesian network, where each edge represents the relationship between two attributes, and the edge weight reflects the degree of influence between attributes. A set of attributes that have the greatest influence on the sensitive attribute is selected. Step 4.2: For the weighted graph preprocessed in Step 4.1, select the first few attributes as protection attributes for fairness assurance. Based on the magnitude of the influence of each attribute on the sensitive attributes, and based on the allocation of different privacy budgets, add more noise to the attributes with greater influence to ensure fairness. Select a set of attributes As influencing factors, corresponding privacy budgets are allocated according to their weight.
4. The method according to claim 3, characterized in that: Step 5 includes the following steps: Step 5.1: Based on the results obtained in Steps 3 and 4, use the idea of Lagrange multipliers to add a fairness constraint as a penalty term to the objective function; Original objective function With the addition of fairness constraints, it is defined as , in As a Lagrange multiplier, let , We can obtain: Step 5.2: For the objective functions obtained in Step 3 and Step 4, calculate the set of parameters that minimizes the loss to complete the training of the model; The model parameters w = (w1, w2, ..., wd) are used to output the predicted value y by minimizing the empirical loss of the training dataset D in the parameter space P; that is, to solve the following optimization problem. in, These are the learned parameters. E is the loss function. It is the objective function. These are the model training parameters. Considering logistic regression as the loss function, we have: 。 5. A fairness-based prediction system for multi-source privacy data, characterized in that: The system is used to execute the fairness-based prediction method for multi-source privacy data as described in any one of claims 1 to 4; The system includes: a dataset module, a multinomial expansion module, a decision tree module, a Bayesian module, a noise module, and a prediction module; The dataset module determines the prediction task and sensitive attributes of a given logistic regression multivariate dataset; Given a dataset D={X, ..., n tuples , Y}, where X={ This represents d unprotected attributes; Y represents protected attributes, i.e., sensitive attributes; Y is the tag. Dataset D contains n tuples, each tuple is represented as , where the feature vector It contains d unprotected attributes and one protected attribute, that is... =( , , · · ·, ; ), Indicates the corresponding label; Protect Based on this, further consider other attributes X={ Possibly with If the relationship is close, select one or more other attributes that are most closely related to the protected attribute as the protected objects. When choosing a pair When choosing the attribute with the closest relationship, let the selected attribute be... ( When multiple attributes are selected, let this group of attributes be... , ( ); The polynomial expansion module uses the idea of a function mechanism to expand the logistic regression loss function into a polynomial form using Chebyshev polynomials. Chebyshev expansion is used to derive a polynomial approximation of the objective function of logistic regression; Given the objective function The training dataset is The model training parameters are Due to the objective function yes The function is complex, so the function mechanism extends it to a polynomial form and deploys the Laplace mechanism on the parameter terms; therefore, for the objective function... The function mechanism can be formally represented as: in For each term of the polynomial, These are the parameters of each term; The polynomial is truncated to the square term as the objective function after expansion, and the coefficients of each independent variable are extracted as the noise-adding objects in the subsequent steps. The number of terms in a Chebyshev polynomial can be specified as needed, and the number of terms can be chosen as follows: Then the Chebyshev polynomial will be truncated to the th... In this step, the formula obtained in step 2.1 is truncated by binomial calculation; the function mechanism FM is applied to... Using Chebyshev polynomial expansion ,in Represents a polynomial function. This represents the coefficient of the polynomial. The decision tree module uses a decision tree algorithm to select the attribute that has the greatest impact on the sensitive attribute; The Bayesian module uses a Bayesian network to select a group of attributes that have the greatest impact on the sensitive attributes. The noise module combines the decision tree module and the Bayesian module, adding Laplace noise with a fairness constraint penalty term to the coefficients of the polynomial function expanded by the polynomial expansion module. The prediction module uses a loss function that satisfies fairness and differential privacy guarantees to perform prediction tasks on the dataset.
6. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
7. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
System and method for privacy-preserving distributed training of neural network models on distributed datasets
CA3188608A1
Method for simultaneously realizing differential privacy and machine learning fairness in binary classification
CN115049072A