Anti-fact fair synthesis data generation method and device based on causal reasoning

Through the combination of causal reasoning and generative adversarial networks, the adversarial optimization mechanism between generator and discriminator solves the problem of lack of support for causal relationships in the existing technology, and generates fairness and interpretability synthetic data, improving the fairness and classification accuracy of the model.

CN120354938APending Publication Date: 2025-07-22JINAN UNIVERSITY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510398620.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing data generation methods lack causal reasoning support when generating fair synthetic data, and cannot ensure counterfactual fairness, resulting in bias in the model on sensitive attributes, affecting ethical and social issues.

Method used

Using a method based on causal reasoning, combined with a generative adversarial network, an adversarial optimization mechanism between the generator and the discriminator is constructed through a causal relationship diagram and a variational automatic codec, the causal path between sensitive attributes and target variables is simulated, and a fair and interpretable synthetic data is generated.

Benefits of technology

It realizes the generation of fair and interpretable synthetic data while maintaining the authenticity of the data, which significantly improves the fairness and classification accuracy of the model and reduces the impact of sensitive features on the model output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354938A_ABST
    Figure CN120354938A_ABST
Patent Text Reader

Abstract

The invention provides an anti-fact fair data synthesis method and device based on causal reasoning, and aims to generate high-quality synthesis data meeting the fairness requirement by mining the causal relationship between observable features. The synthesis method comprises the following steps: extracting observable features, sensitive features and labels from original data, extracting potential features through a variational automatic codec, and constructing a causal relationship graph; designing a generator according to a topological sequence of the causal relationship graph, connecting a causal path, inputting the potential features and the related features into the generator in sequence, and constructing a data generation process conforming to a causal structure; introducing a discriminator to carry out adversarial training on a generation result and original data, and optimizing generator parameter distribution; finally, synthetic data meeting fairness requirements are generated. According to the method, effective regulation and control on the influence of sensitive characteristics and strict constraint on a causal structure are realized, the generated data has higher fairness and interpretability, and the method can be applied to the fields with higher fairness requirements, such as finance, medical treatment and education.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data science and machine learning, and particularly relates to a counterfactual fairness synthetic data generation method, device, electronic device and storage medium based on causal inference, which is used to improve the fairness and accuracy of the model during the data generation process. Background Art

[0002] With the wide application of data-driven technologies, machine learning models have achieved remarkable results in fields such as finance, healthcare, and recruitment. However, the bias problems caused by sensitive attributes (such as race, gender, age, etc.) are always difficult to completely eliminate, and may even lead to discriminatory decisions, thus triggering a series of ethical and social issues. To address this challenge, the theory of counterfactual fairness emerged, aiming to generate fair synthetic data in different scenarios and reduce the model's dependence on sensitive attributes. Existing data generation methods mainly rely on modeling techniques such as adversarial networks. Although they alleviate the bias problem to a certain extent, they still lack the support of causal inference and cannot ensure the counterfactual fairness of the generated data from the perspective of causal relationships. Therefore, designing a method that can efficiently generate synthetic data that meets counterfactual fairness has become an important research direction for achieving model fairness. Summary of the Invention

[0003] The present invention aims to solve the above deficiencies in the prior art and proposes a counterfactual fairness synthetic data generation method, device, electronic device and storage medium based on causal inference. By combining the causal inference framework with the generative adversarial network, making full use of the causal dependencies in the data, and constructing an adversarial optimization mechanism for the generator and discriminator, it realizes the modeling and regulation of the causal path between sensitive attributes and target variables. This method can, on the premise of maintaining data authenticity, simulate the counterfactual data generation process based on sensitive attributes through a causal inference model, generate fair and interpretable synthetic data, thus meeting the requirements of counterfactual fairness and providing data support for various application scenarios (such as financial decision-making, medical diagnosis, recruitment screening, etc.).

[0004] To achieve the above object, the present invention adopts the following technical solutions:

[0005] The first object of the present invention is to provide a counterfactual fairness synthetic data generation method based on causal inference, including the following steps:

[0006] S1. Input recruitment screening training data Data train , identify the observable features X, sensitive features S and label values Y in the training data Data train , and combine them into a sample form and form a sample set C;

[0007] S2. Input the sample set C into the variational autoencoder for training and learning, extract the latent feature U from the sample set C. The latent feature U represents latent factors that cannot be directly observed but may affect the label value Y. Finally, the sample form is extended to

[0008] S3. Based on the dependencies among the observable feature X, the sensitive feature S, the latent feature U, and the label value Y, establish a causal relationship graph M. The causal relationship graph M is a directed acyclic graph, which is used to clarify the causal paths and dependencies among each feature. Each node in the graph represents a feature in the sample set C, and the directed edge in the graph represents the causal relationship between two features in the sample set C;

[0009] S4. Create an independent generator Gi for each node of the causal relationship graph M, where i = 1, 2,..., p, to obtain the generator set G = {G1, G2,..., G i ,..., G p}, where p is the total number of all features in the causal relationship graph M. Each node of the causal relationship graph M corresponds to a feature in the sample set C, and define a discriminator D;

[0010] S5. Input the sample set C into the generator set G, and generate variables in turn according to the topological order of the causal relationship graph M. First, generate the root node, and then generate its child nodes until all variables are generated, to obtain the generated sample set where, are the observable feature, the sensitive feature, the latent feature, and the label value respectively simulated and generated according to the causal relationship graph M and the generator set G;

[0011] S6. Input the sample set C and the generated sample set into the discriminator D. The generator set G and the discriminator D are jointly trained in an adversarial manner to obtain the optimized generator set G opt = {G 1,opt , G 2,opt ,..., G i,opt ,..., G p,opt}, where Gi,opt represents the final optimized generator for the i-th feature node, and the discriminator DOpt;

[0012] S7. Input the sample set C into the generator set G opt = {G 1,opt , G 2,opt ,..., G i,opt ,..., G p,opt}(DOpt) discriminator, for each different value of the sensitive feature S, generate corresponding counterfactual fair samples, and aggregate all counterfactual fair samples to obtain a counterfactual fair dataset Datafair that ensures the fairness requirements of each sensitive feature.

[0013] Further, the process of step S1 is as follows:

[0014] Represent the observable feature X in the training dataset Datatrain as a numerical matrix of dimension n×m where n is the number of samples and m is the feature dimension of each sample;

[0015] Represent the sensitive feature S in the training dataset Datatrain as a numerical matrix of dimension n×k where k is the dimension of the sensitive feature S; the purpose of this step is to extract and standardize the sensitive information separately for subsequent processing;

[0016] Represent the label value Y in the training dataset Datatrain as a vector of dimension n×1 where the values of the vector elements are 0 or 1. A value of 0 indicates that the data has an unfair risk, while a value of 1 indicates that there is no such potential unfair risk; this marking method can help the model identify potential unfair risks in the data in subsequent processing and lay a foundation for generating fair data;

[0017] Merge the observable feature X, the sensitive feature S, and the corresponding label value Y to form a sample set C. The sample C in the sample set C i =(X i , S i , Y i ) represents the complete information on the observable feature X i , the sensitive feature S i and the label value Y i of the i-th data record; this step ensures the standardization, structuring, and integrity of the data, and at the same time clarifies the analysis target of the unfair risk, laying a foundation for subsequent causal reasoning and generating a counterfactual fair dataset. Further, the process of step S2 is as follows:

[0018] Input the sample set C into a variational autoencoder for training and learning. The variational autoencoder consists of a decoder and an encoder; this step uses the variational autoencoder method for latent feature modeling, which can approximately infer the latent factors without explicitly determining the data distribution. This enables the model to still work stably in the face of diverse and complex data and try to extract key elements that are helpful for causal interpretation in the latent space;

[0019] Among them, the encoder takes the observable feature X and the label value Y in the sample set C as inputs, maps the inputs to the latent space containing the latent feature U through a neural network, and outputs the parameters describing the distribution q(U|X, Y) of the latent feature U, including the mean μ and the log variance logσ of the latent feature U 2 ,

[0020] Construct the latent feature U according to the following formula: U = μ + δ·ε

[0021] where δ = exp(0.5·logσ 2 ), and ε is the random noise sampled from the standard normal distribution N(0, 1);

[0022] The decoder takes the latent feature U and the sensitive feature S as inputs, reconstructs the observable feature X and the label value Y, and the output is the reconstructed observable feature X reconst and the label value Y reconst ;

[0023] Define the preliminary optimization loss function as follows:

[0024] L r = E q(U|X,Y) [-log(p(X, Y|U,S))] + KL[q(U|X, Y)||p(U)]

[0025] where p(X, Y|U, S) = P(Y|X, U, S)·p(X|U, S). The first term p(Y|X, U, S) describes the generation probability of the label value Y under the conditions of the observable feature X, the latent feature U, and the sensitive feature S, reflecting the influence of the observable feature X and the sensitive feature S on the label value Y. The second term p(x|U, S) describes the generation process of the observable feature X under the conditions of the latent feature U and the sensitive feature S, ensuring that the latent space has interpretability for the observable feature X. E q(U|X,Y) [-log(X, Y|U, S))] measures the similarity between the observable feature X and the label value Y and the reconstructed observable feature X generated by the decoder reconst and the label value Y reconst by the negative value of the log-likelihood. KL[q(U|X, Y)|p(U)] measures the difference between the distribution q(U|X, Y) of the latent feature U generated by the encoder and the prior distribution p(U) of the latent feature U. p(U) is the standard normal distribution N(0, 1), and KL[·||·] is the KL divergence. Introducing the KL divergence term in this step can ensure that the distribution of the latent feature U is smoother, has a certain regularization ability, reduces overfitting, and enhances the data extrapolation ability. This means that when facing new data or needing to expand the dataset, the model can still maintain good robustness.

[0026] The maximum mean discrepancy (MMD) is introduced as a regularization term. MMD is used to measure the difference between two distributions and and is defined as:

[0027]

[0028] where is a feature mapping function used to project the data into a reproducing kernel Hilbert space;

[0029] Finally, the expression of the loss function L for variational autoencoder training optimization is as follows:

[0030]

[0031] where α ≥ 0 is an important hyperparameter that controls the importance of the distribution averaging term, represents the number of different sensitive feature pairs (s, s′), |S| represents the number of different values of the sensitive feature S, p(U|s) represents the probability distribution of the latent feature U under the condition that the sensitive feature takes the value s, and p(U|s′) represents the probability distribution of the latent feature U under the condition that the sensitive feature takes the value s′;

[0032] This step successfully extracts the latent feature U required for causal fairness modeling and ensures that the latent feature can capture the data structure without carrying the information of the sensitive feature by optimizing the objective function. This step is directly related to the accuracy of subsequent causal relationship modeling and the fairness of the generated data.

[0033] Furthermore, in step S4, each generator G i simulates the distribution corresponding to the i-th feature by generating data. A discriminator D is defined to evaluate the generation quality of the generator set G and determine whether the data generated by the generator G i conforms to the feature distribution in the causal relationship graph M; through this step, the generation process is decomposed into multiple sub-generator modules, enabling each module to only focus on generating its corresponding feature and be conditioned on its parent node. This modular strategy is easier to control and debug. When it is necessary to adjust the distribution or generation method of a specific feature, only the corresponding module needs to be modified alone, without making large-scale changes to the entire generation process.

[0034] Furthermore, the process of step S5 is as follows:

[0035] According to the directed acyclic graph structure of the causal relationship graph M, features are generated layer by layer according to the following steps:

[0036] First, perform a topological sort on the causal relationship graph M to determine the node generation order, start generating from the root node, and generate dependent features layer by layer backward.

[0037] The root node has no parent node and is directly generated by a random noise variable:

[0038] Among them, represents the generation eigenvalue of the root node r in the causal relationship graph M; G r represents the generator for generating the root node features; Z r represents the random noise independently sampled from the standard normal distribution N(0, 1).

[0039] The generation of non-root nodes depends on the generation results of the parent nodes and random noise. The generation process is as follows:

[0040] Among them, represents the sample value generated by the i-th feature node in the causal relationship graph M, represents the set generated by the parent nodes of the i-th node; Z i is the random noise independently sampled from the standard normal distribution N(0, 1); the generator set G = {G1, G2,..., G i ,..., G p} is combined and modeled according to the node relationship;

[0041] According to the topological order, all nodes are generated in sequence until all the features in the causal relationship graph M are generated, and finally the generated sample set is output Through this step, the generation process is strictly carried out according to the structure of the causal relationship graph M, avoiding the problem that the causal relationship may be ignored in the traditional generation model. This makes the generated data not only have authenticity but also meet the requirements of causal fairness. When modifying the values of sensitive features later, the downstream features can be generated along the structure of the causal relationship graph M. This mechanism provides a reliable technical basis for realizing the balanced distribution of sensitive features and improving the fairness of the data set.

[0042] Furthermore, the process of step S6 is as follows:

[0043] The sample set C = (X, S, U, Y) and the generated sample set are used as the training inputs of the discriminator D. The discriminator D is used to distinguish real samples and generated samples and simultaneously evaluate the distribution consistency of the generated samples;

[0044] The loss function L of the discriminator D D is expressed as:

[0045] Among them, E real represents the expectation of the real sample set C, and E fake represents the generated sample set The expectation, D(C) represents the prediction probability of the discriminator for the authenticity of the real sample set C. represents the prediction probability of the discriminator for the generated sample set of the authenticity, and this loss function L D optimizes the discriminator D by maximizing the prediction probability of the real samples and minimizing the prediction probability of the generated samples;

[0046] The adversarial training process of the discriminator D is as follows:

[0047] Fix the generator set G and update the discriminator D: Use the gradient descent method to update the parameters to make the discriminator D improve the discrimination ability for the real sample set C and the generated sample set ;

[0048] Fix the discriminator D and update the generator set G: Minimize the following generator set loss function L G to maximize the probability of the generated sample set deceiving the discriminator D:

[0049] According to the adversarial training results, select the optimal solutions of the generator set G and the discriminator D: The generator set G opt ={G 1,opt , G 2,opt ,..., G i,opt ,..., G p,opt}, where G i,opt represents the final optimized generator and discriminator D for the i-th feature node Opt ;

[0050] Through this step, the ability of the generative adversarial network is extended to the field of causally fair data generation. And through adversarial training, the collaborative optimization of the generator and the discriminator is achieved, ensuring that the generated data is both authentic and meets the distribution requirements of causal fairness. The optimal selection of the final generator and discriminator lays the foundation for the generation of high-quality fair data.

[0051] Furthermore, it is characterized in that the process of step S7 is as follows:

[0052] Input the sample set C=(X, S, U, Y) into the optimized generator set G opt ={G 1,opt , G 2,opt ,..., G p,opt} and the discriminator D opt , and for each different value set {s1, s2,..., s k ,..., s K} of each sensitive feature S, generate the corresponding counterfactual sample sets one by one:

[0053] Among them, represents that the sensitive feature S takes the value s k when generating the sample set, s1, s2,..., s k ,..., s K are K different values of the sensitive feature S. During the generation process, other features are kept unchanged, and only the sensitive feature S is adjusted to evaluate the fairness impact;

[0054] Aggregate the counterfactual samples corresponding to different values of all sensitive features S, and calculate the final fairness data set by taking the mean:

[0055]

[0056] This step can effectively balance the data distributions corresponding to different values of sensitive features and eliminate the influence of a single value on the overall data set through the aggregation method of taking the average, thereby generating a more fair and representative data set. Average aggregation does not destroy the causal structure and data authenticity because the synthesis process of each sub-data set follows the constraints of causal relationships and adversarial training. Therefore, the finally synthesized Data fair ensures fairness without sacrificing data quality and causal rationality.

[0057] The second object of the present invention is to provide a counterfactual fairness synthetic data generation device based on causal inference for performing the above-mentioned counterfactual fairness synthetic data generation method based on causal inference. The counterfactual fairness synthetic data generation device based on causal inference includes:

[0058] A sample set generation module for inputting recruitment screening type training data Data train , identifying the observable feature X, sensitive feature S, and label value Y in the training data Data train , and combining them into a sample form and forming a sample set C;

[0059] A sample expansion module for inputting the sample set C into a variational autoencoder for training and learning, extracting the latent feature U from the sample set C. The latent feature U represents a latent factor that cannot be directly observed but may affect the label value Y. Finally, the sample form is expanded to

[0060] A causal relationship graph establishment module, based on the dependencies among the observable feature X, sensitive feature S, latent feature U, and label value Y, establishes a causal relationship graph M. The causal relationship graph M is a directed acyclic graph for clarifying the causal paths and dependencies between each feature. Each node in the graph represents a feature in the sample set C, and the directed edges in the graph represent the causal relationships between two features in the sample set C;

[0061] A generator discriminator creation module for creating an independent generator G for each node of the causal relationship graph M i , i = 1, 2, …, p, to obtain a generator set G = {G1, G2, ..., G i , ..., G p}, where p is the total number of all features in the causal relationship graph M, each node of the causal relationship graph M corresponds to a feature in the sample set C, and a discriminator D is defined;

[0062] A generated sample set acquisition module for inputting the sample set C into the generator set G, generating variables in sequence according to the topological order of the causal relationship graph M, first generating the root node, and then generating its child nodes until all variables are generated, to obtain a generated sample set wherein, are respectively the observable feature, sensitive feature, latent feature, and label value simulated and generated according to the causal relationship graph M and the generator set G;

[0063] A generator discriminator joint training module for inputting the sample set C and the generated sample set into the discriminator D, and jointly training the generator set G and the discriminator D in an adversarial manner to obtain an optimized generator set G opt ={G 1,opt , G 2,opt ,... G i,opt ,... G p,opt} and the discriminator D Opt , where G i,opt represents the final optimized generator for the i-th feature node;

[0064] A counterfactual fairness dataset generation module for inputting the sample set C into the generator set G opt ={G 1,opt , G 2,opt ,... G i,opt ,... G p,opt} and the discriminator D Opt , and generating corresponding counterfactual fairness samples for different values of each sensitive feature S, and aggregating all the counterfactual fairness samples to obtain a counterfactual fairness dataset Data that ensures the fairness requirements of each sensitive feature fair .

[0065] The third object of the present invention is to provide an electronic device, including a processor and a memory for storing a program executable by the processor. When the processor executes the program stored in the memory, the above-mentioned method for generating counterfactual fairness synthetic data based on causal inference is implemented.

[0066] The fourth object of the present invention is to provide a storage medium storing a program, which when executed by a processor, implements the above-mentioned method for generating counterfactual fair synthetic data based on causal reasoning.

[0067] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0068] 1. Fairness modeling driven by causal reasoning to ensure fairness and interpretability of data generation: The present invention constructs a causal relationship graph using a causal reasoning framework to clarify the causal dependence relationships between various features. This method not only captures the correlations between explicit features but also models the implicit influence of latent features on the label, thereby solving the problem of fairness deviation caused by ignoring potential causal mechanisms in the prior art. In addition, the causal relationship graph guides the data generation process through topological sorting, ensuring that the generated samples strictly follow causal logic and significantly enhancing the interpretability of the generated data.

[0069] 2. Deep combination of variational autoencoder (VAE) and generative adversarial network (GAN) to achieve full-process fair data modeling and optimization: The present invention adopts an organic combination of VAE and GAN in the overall framework, giving full play to the latent feature extraction ability of VAE and the adversarial training optimization ability of GAN. In the data preprocessing stage, VAE is used to capture latent factors that are unobservable but affect the label, and the independence between latent features and sensitive features is ensured through the maximum mean discrepancy (MMD) regularization term, fundamentally reducing the bias brought by sensitive features. In the data generation stage, the generator and discriminator of GAN further optimize the authenticity and fairness of the generated data through adversarial training. The generator strictly generates variables layer by layer according to the dependency structure of the causal relationship graph, while the discriminator dynamically evaluates the quality and fairness deviation of the generated samples. The combination of VAE and GAN not only realizes full-process control from latent feature modeling to adversarial optimization but also ensures high interpretability and fairness of the generated data, providing strong support for fair data generation in complex scenarios.

[0070] 3. Modular generation scheme design, strictly following the causal dependency path for data generation: The present invention generates data variables layer by layer in topological order based on the causal relationship graph, and each node corresponds to an independent generator. This modular design has the following technical advantages:

[0071] Flexibility: The modular structure facilitates the expansion or adjustment of the generation process to adapt to the requirements of different data scenarios.

[0072] Dependency control: Strictly following the causal dependency path for layer-by-layer generation ensures that the generated data conforms to the true causal structure, controlling data fairness from the source.

[0073] Interpretability: The generation process is clear and transparent, enabling tracing of the specific impact path of sensitive features on the results, thus avoiding the "black box" problem of traditional generation methods.

[0074] 4. Adversarial training mechanism enhances data generation quality and fairness control: The present invention adopts an adversarial training mechanism of a generator and a discriminator, and dynamically evaluates the distribution consistency between generated samples and real samples through the discriminator. While distinguishing real samples from generated samples, the discriminator is also responsible for the evaluation of fairness constraints to guide the generator to optimize the parameter distribution. This adversarial training method not only improves the authenticity of generated samples but also ensures that the generated samples minimize their dependence on sensitive features, further enhancing fairness.

[0075] 5. Counterfactual fairness verification mechanism for dynamic adjustment of sensitive features: The present invention generates a counterfactual sample set by dynamically adjusting the values of sensitive features, aggregates all counterfactual samples, and finally outputs a synthetic data set that meets the fairness requirements. This mechanism can adjust the values of sensitive features under the condition of keeping other features unchanged, evaluate their impact on the label results, and eliminate biases through mean calculation. This dynamic adjustment mechanism significantly improves the fairness control ability of the data generation process, fundamentally solving the defect that traditional methods only focus on distribution similarity while ignoring fairness.

[0076] 6. Double improvement in fairness and classification accuracy, with significant experimental verification results: The present invention shows significant improvements in fairness and classification accuracy in experimental verification: The classification accuracy rate is increased by an average of 18.47%, indicating that the generated data is more in line with the real label distribution and improves the model's prediction ability; the fairness index is increased by 78.89%, proving that the generated data significantly reduces the impact of sensitive features on the model output and reduces the potential bias and discrimination risks. This experimental result proves that the present invention not only maintains high data quality while ensuring fairness but also further improves the classification performance, providing reliable data support for high-demand scenarios. Description of the Drawings

[0077] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following-described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0078] Figure 1 It is a flowchart of a counterfactual fairness synthetic data generation method based on causal inference disclosed in the present invention;

[0079] Figure 2It is a flowchart for training a causal model to obtain hidden features in the present invention;

[0080] Figure 3 It is a flowchart for connecting multiple generator networks through the topological order of causal relationships and generating a counterfactual fair synthetic dataset;

[0081] Figure 4 It is a schematic diagram of an apparatus for generating counterfactual fair synthetic data based on causal inference disclosed in Embodiment 3 of the present invention;

[0082] Figure 5 It is a structural diagram of an electronic device disclosed in Embodiment 4 of the present invention. Detailed implementation manners

[0083] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0084] Referring to "embodiment" in the present application means that the specific features, structures or characteristics described in combination with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments.

[0085] Embodiment 1

[0086] This embodiment discloses a method and apparatus for generating counterfactual fair synthetic data based on causal inference, specifically including the following steps:

[0087] S1. Input recruitment screening type training data Data train , identify the observable features X, sensitive features S and label values Y in the training data Data train , and combine them into a sample form And form a sample set C. Among them, the recruitment screening data uses the law school admission dataset from the National Longitudinal Bar Passage Rate Study of the Law School Admission Council in the United States, consisting of 21,790 students from 163 law schools in the United States, used to study law school admission criteria, student performance, and the fairness of students from different backgrounds in the admission process. The number of features is 5. The observable feature X includes the grade point average, bar exam score, and legal knowledge mastery. The sensitive feature S is race, with a value range of {0, 1, 2}, representing different racial categories respectively. The dataset faces a regression prediction problem, and the label value Y is the first-year average grade. By merging, a sample set is formed and denoted as

[0088] S2. Input the sample set C into the variational autoencoder for training and learning, extract the latent feature U from the sample set C. The latent feature U represents the latent factors that cannot be directly observed but may affect the label value Y. Finally, the sample form is extended to

[0089] In this process, the encoder takes the observable feature X and the label value Y in the sample set C as inputs, maps the inputs to the latent space containing the latent feature U through a neural network, and outputs the parameters describing the distribution Q(U|X, Y) of the latent feature U, including the mean μ and the log variance logσ of the latent feature U 2 ,

[0090] Construct the latent feature U according to the following formula: U = μ + δ·ε

[0091] where δ = exp(0.5·logσ 2 ), and ε is the random noise sampled from the standard normal distribution N(0, 1);

[0092] From the above formula, the latent feature U = {1.430, -0.111, 1.038,..., 1.165} is calculated according to the calculation;

[0093] The decoder takes the latent feature U and the sensitive feature S as inputs, reconstructs the observable feature X and the label value Y, and the output is the reconstructed observable feature X reconst and the label value Y reconst ;

[0094] Define the preliminary optimization loss function as follows:

[0095] L r = E q(U|X,Y) [-log(p(X, Y|U, S))] + KL[q(U|X, Y)|p(U)]

[0096] Among them, p(X, Y|U, S) = p(Y|X, U, S)·p(X|U, S). The first term p(Y|X, U, S) describes the generation probability of the label value Y under the conditions of the observable feature X, the latent feature U, and the sensitive feature S, reflecting the influence of the observable feature X and the sensitive feature S on the label value Y. The second term p(X|U, S) describes the generation process of the observable feature X under the conditions of the latent feature U and the sensitive feature S, ensuring that the latent space has interpretability for the observable feature X, E q(U|X,Y) [-log(p(X, Y|U, S))] measures the similarity between the observable feature X and the label value Y and the reconstructed observable feature X generated by the decoder through the negative value of the log-likelihood reconst and the label value Y reconst The similarity of. KL[q(U|X, Y)|p(U)] measures the difference between the distribution q(U|X, Y) of the latent feature U generated by the encoder and the prior distribution p(U) of the latent feature U. p(U) is the standard normal distribution N(0, 1), and KL[·||·] is the KL divergence;

[0097] The maximum mean discrepancy MMD is introduced as a regularization term. MMD is used to measure the difference between two distributions and The difference is defined as:

[0098]

[0099] Among them, is a feature mapping function used to project data into a reproducing kernel Hilbert space;

[0100] Finally, the expression of the loss function L for variational autoencoder training optimization is as follows:

[0101]

[0102] Among them, α≥0 is an important hyperparameter that controls the importance of the distribution averaging term, represents the number of different sensitive feature pairs (s, s′), |S| represents the number of different values of the sensitive feature S, p(U|s) represents the probability distribution of the latent feature U under the condition that the sensitive feature takes the value s, and p(U|s′) represents the probability distribution of the latent feature U under the condition that the sensitive feature takes the value s′.

[0103] S3. Based on the dependencies among the observable feature X, the sensitive feature S, the latent feature U, and the label value Y, a causal relationship graph M is established. The causal relationship graph M is a directed acyclic graph used to clarify the causal paths and dependencies among each feature.

[0104] Construct the causal relationship graph M:

[0105] Node Definition: V1: Race is the root node with no parent node; V2: GPA is a non-root node and depends on race; V3: Bar exam score is a non-root node and depends on race and latent feature U; V4: First-year average grade is a non-root node and depends on race; V5: Mastery of legal knowledge is the root node with no parent node and affects the bar exam score.

[0106] Edges represent causal relationships: Race → GPA, Race → Bar exam score, Race → First-year average grade, Mastery of legal knowledge → Bar exam score

[0107] S4. Create an independent generator Gi for each node in the causal graph M, where i = 1, 2, 3, 4, 5, to obtain the generator set G = {G1, G2, G3, G4, G5}. Each node in the causal graph M corresponds to a feature in the sample set C, and a discriminator D is defined. i , i = 1, 2, 3, 4, 5, to obtain the generator set G = {G1, G2, G3, G4, G5}. Each node in the causal graph M corresponds to a feature in the sample set C, and a discriminator D is defined.

[0108] In this process, each generator Gi i simulates the distribution of the corresponding ith feature by generating data. A discriminator D is defined to evaluate the generation quality of the generator set G and determine whether the data generated by the generator Gi i conforms to the feature distribution in the causal graph M.

[0109] S5. Input into the generator set G and generate variables in the topological order of the causal graph M. First, generate the root nodes, and then generate their child nodes until all variables are generated, obtaining the generated sample set where is the observable feature simulated according to the causal graph M and the generator set G, where is the sensitive feature simulated according to the causal graph M and the generator set G, where is the latent feature simulated according to the causal graph M and the generator set G, where is the label value simulated according to the causal graph M and the generator set G;

[0110] According to the directed acyclic graph structure of the causal graph M, generate features layer by layer according to the following steps:

[0111] First, perform a topological sort on the causal graph M to determine the node generation order. Start generating from the root nodes and generate the dependent features layer by layer backward.

[0112] Root node generation:

[0113] Race and mastery of legal knowledge are root nodes and do not depend on other nodes, so they are directly generated:

[0114]

[0115] Non-root node generation:

[0116] Generation of grade point average:

[0117] Generation of judicial examination:

[0118] Generation of first-year average score:

[0119] According to the topological order, all nodes are generated in sequence until all features in the causal relationship graph M are generated, and finally the generated sample set is output

[0120] S6. Input the sample set C and the generated sample set into the discriminator D. The generator set G and the discriminator D are jointly trained in an adversarial manner to obtain the optimized generator set G opt ={G 1,opt , G 2,opt , G 3,opt , G 4,opt , G 5,opt}, and the discriminator D Opt ;

[0121] Input the sample set C=(X, S, U, Y) and the generated sample set as the training input of the discriminator D. The discriminator D is used to distinguish real samples and generated samples, and at the same time evaluate the distribution consistency of the generated samples;

[0122] The loss function L D of the discriminator D is expressed as:

[0123] where E real represents the expectation of the real sample set C, E fake represents the expectation of the generated sample set , D(C) represents the authenticity prediction probability of the discriminator for the real sample set C, represents the authenticity prediction probability of the discriminator for the generated sample set . This loss function L D optimizes the discriminator D by maximizing the prediction probability of real samples and minimizing the prediction probability of generated samples;

[0124] The adversarial training process of the discriminator D is as follows:

[0125] Fix the generator set G and update the discriminator D: Use the gradient descent method to update the parameters to make the discriminator D improve its ability to distinguish between the real sample set C and the generated sample set ;

[0126] Fix the discriminator D and update the set of generators G: Minimize the following loss function L of the set of generators G to maximize the probability of the generated sample set deceiving the discriminator D:

[0127] According to the adversarial training results, select the optimal solutions of the set of generators G and the discriminator D: the set of generators G opt ={G 1,opt , G 2,opt , G 3,opt , G 4,opt , G 5,opt} and the discriminator D Opt .

[0128] S7. Input the sample set C into the set of generators G opt ={G 1,opt , G 2,opt ,..., G i,opt ,..., G p,opt} and the discriminator D Opt , and for different values of each sensitive feature S, generate corresponding counterfactual fair samples, and aggregate all the counterfactual fair samples to obtain a counterfactual fair dataset Data that ensures the fairness requirements of each sensitive feature fair .

[0129] Input the sample set C=(X, S, U, Y) into the optimized set of generators G opt ={G 1,opt , G 2,opt , G 3,opt , G 4,opt , G 5,opt} and the discriminator D opt , and for different value sets {0, 1, 2} of each sensitive feature S, generate corresponding counterfactual sample sets one by one:

[0130]

[0131] Aggregate the counterfactual samples corresponding to different values of all sensitive features S, and calculate the final fairness dataset by taking the mean:

[0132] The present invention is compared with other algorithms, and the comparison metrics include Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Maximum Mean Discrepancy (MMD), and Wasserstein-1 distance. RMSE is a metric for measuring the deviation between predicted values and actual values. It is obtained by calculating the mean of the squared errors and then taking the square root. RMSE can magnify large errors, so it is more sensitive to large deviations in predicted values. The smaller the RMSE, the smaller the deviation between the predicted value and the actual value, and the higher the prediction accuracy of the model. MAE is the average absolute value of the difference between predicted values and actual values. It calculates the overall error of the model by directly calculating the absolute value of the difference between each predicted value and the actual value and then taking the average. The smaller the MAE, the smaller the deviation between the predicted value and the actual value, and the better the prediction effect of the model. The Maximum Mean Discrepancy (MMD) is a metric for measuring the distribution difference between generated data and real data. It evaluates whether the generated data is similar to the real data in terms of distribution by comparing different feature distributions. The smaller the MMD, the closer the generated data is to the real data in terms of overall feature distribution, thus ensuring the authenticity of the data. The Wasserstein-1 distance (also known as the Earth Mover's Distance) is a metric for measuring the distance between two probability distributions. It represents the "cost" required to "transport" one distribution to another. The smaller the WASS, the more similar the generated data is to the real data in terms of feature distribution. This similarity ensures that the generated data retains the structural characteristics of the real data and helps to generate high-quality data samples with fairness and authenticity. Table 1 is used to test the classification performance of the present invention. The second column is the original data, the third column TabFairGAN is a generative adversarial network proposed in the paper "Tabfairgan: Fair tabular data generation with generative adversarial networks." that generates a synthetic dataset that meets both authenticity and fairness criteria. The fourth column DECAF is a generative adversarial network proposed in the paper "Decaf: Generating fair synthetic data using causally-aware generative networks." that realizes unbiased data reconstruction by adopting conditional causal reweighting and focusing on causal parent variables. The last column is the method of the present invention. The comparison results are shown in Table 1. From the results in the table, it can be seen that the method proposed by the present invention maintains the highest accuracy, and the fairness of the model is greatly improved, which proves the effectiveness of the present invention.

[0133] Table 1. Comparison Results of Regression Performance and Fairness Metrics between the Public Method of the Present Invention and Existing Methods

[0134]

[0135] Example 2

[0136] This example further discloses a counterfactual fairness synthetic data generation method based on causal inference, which specifically includes the following steps:

[0137] S1. Input the recruitment screening type training data Datatrain, identify the observable features X, sensitive features S, and label values Y in the training data Datatrain, and merge them into a sample form and form a sample set C. Among them, the adult income dataset in the UCI public dataset of recruitment screening type data includes sample data from the 1994 US census, with 7 feature numbers. The observable features X include gender, marital status, education level, weekly working hours, occupation, and the sensitive feature S is race, with a value range of {0, 1, 2}, respectively representing different racial categories. The dataset faces a binary classification problem, and the label value Y is employment and non-employment. By merging, a sample set is formed and represented as

[0138] S2. Input the sample set C into the variational autoencoder for training and learning, and extract the latent features U from the sample set C. The latent features U represent latent factors that cannot be directly observed but may affect the label value Y. Finally, the sample form is extended to

[0139] In this process, the encoder takes the observable features X and label values Y in the sample set C as inputs, maps the inputs to the latent space containing the latent features U through a neural network, and outputs the parameters describing the distribution q(U|X, Y) of the latent features U, including the mean μ and log variance logσ of the latent features U 2 ,

[0140] Construct the latent features U according to the following formula: U = μ + δ·ε

[0141] where δ = exp(0.5·logσ 2 ), and ε is a random noise sampled from the standard normal distribution N(0, 1);

[0142] From the above formula, the latent features U = {0.332, 0.304, 0.283,..., 0.343} are calculated according to the calculation;

[0143] The decoder takes the latent features U and the sensitive features S as inputs, reconstructs the observable features X and the label values Y, and the output is the reconstructed observable features X reconst and the label values Y reconst ;

[0144] Define the preliminary optimization loss function as follows:

[0145] L r = E q(U|X,Y) [-log(p(X, Y|U, S))] + KL[q(U|X, Y)||p(U)]

[0146] where p(X, Y|U, S) = p(Y|X, U, S)·p(X|U, S). The first term p(Y|X, U, S) describes the generation probability of the label value Y under the conditions of the observable feature X, the latent feature U, and the sensitive feature S, reflecting the influence of the observable feature X and the sensitive feature S on the label value Y. The second term p(X|U, S) describes the generation process of the observable feature X under the conditions of the latent feature U and the sensitive feature S, ensuring that the latent space has explanatory power for the observable feature X, E q(U|X,Y) [-log(p(X, Y|U, S))] measures the similarity between the observable feature X and the label value Y and the reconstructed observable feature X generated by the decoder through the negative value of the log-likelihood reconst and the label value Y reconst . KL[q(U|X, Y)|p(U)] measures the difference between the distribution q(U|X, Y) of the latent feature U generated by the encoder and the prior distribution p(U) of the latent feature U. p(U) is the standard normal distribution N(0, 1), and KL[·||·] is the KL divergence;

[0147] The maximum mean discrepancy MMD is introduced as a regularization term. MMD is used to measure the difference between two distributions and , and is defined as:

[0148]

[0149] where is a feature mapping function used to project the data into a reproducing kernel Hilbert space;

[0150] Finally, the expression of the loss function L for variational autoencoder training optimization is as follows:

[0151]

[0152] where α ≥ 0 is an important hyperparameter controlling the importance of the distribution averaging term, represents the number of different sensitive feature pairs (s, s′), |S| represents the number of different values of the sensitive feature S, p(U|s) represents the probability distribution of the latent feature U under the condition that the sensitive feature takes the value s, and p(U|s′) represents the probability distribution of the latent feature U under the condition that the sensitive feature takes the value s′.

[0153] S3. Based on the dependencies among the observable feature X, the sensitive feature S, the latent feature U, and the label value Y, establish a causal relationship graph M, where the causal relationship graph M is a directed acyclic graph used to clarify the causal paths and dependencies among each feature.

[0154] Construct the causal relationship graph M:

[0155] Node definition: V1: Race is the root node with no parent node; V2: Gender is the root node with no parent node; V3: Marital status is a non-root node, dependent on race and gender; V4: Education level is a non-root node, dependent on race and gender; V5: Weekly working hours is a non-root node, dependent on race and education level; V6: Occupation is a non-root node, dependent on race, gender, and education level; V7: Employment situation is a non-root node, dependent on marital status, education level, occupation, and weekly working hours.

[0156] Edges represent causal relationships:

[0157] Race → Marital status → Employment situation,

[0158] Race → Marital status → Weekly working hours → Employment situation,

[0159] Race → Employment situation,

[0160] Race → Weekly working hours → Employment situation,

[0161] Race → Marital status → Education level → Employment situation,

[0162] Race → Education level → Employment situation,

[0163] Race → Education level → Occupation → Employment situation,

[0164] Gender → Marital status → Employment situation,

[0165] Gender → Marital status → Weekly working hours → Employment situation,

[0166] Gender → Employment situation,

[0167] Gender → Weekly working hours → Employment situation,

[0168] Gender → Marital status → Education level → Employment situation,

[0169] Gender → Education level → Employment situation,

[0170] Gender → Education level → Occupation → Employment situation;

[0171] S4. Create an independent generator G for each node of the causal relationship graph M i, where \(i = 1, 2, 3, 4, 5, 6, 7\), to obtain a generator set \(G=\{G_1, G_2, G_3, G_4, G_5, G_6, G_7\}\). Each node of the causal relationship graph \(M\) corresponds to a feature in the sample set \(C\), and a discriminator \(D\) is defined.

[0172] During this process, each generator \(G\) i simulates the distribution corresponding to the \(i\)-th feature by generating data. A discriminator \(D\) is defined to evaluate the generation quality of the generator set \(G\) and to determine whether the data generated by the generator \(G\) i conforms to the feature distribution in the causal relationship graph \(M\).

[0173] S5. Input into the generator set \(G\), and generate variables in sequence according to the topological order of the causal relationship graph \(M\). First, generate the root nodes, and then generate their child nodes until all variables are generated, obtaining the generated sample set where, are the observable features, sensitive features, latent features, and label values simulated and generated according to the causal relationship graph \(M\) and the generator set \(G\);

[0174] According to the directed acyclic graph structure of the causal relationship graph \(M\), generate features layer by layer according to the following steps:

[0175] First, perform a topological sort on the causal relationship graph \(M\) to determine the node generation order. Start generating from the root nodes and generate dependent features layer by layer backward.

[0176] Generation of root nodes:

[0177] Race and gender are root nodes and do not depend on other nodes, so they are directly generated:

[0178]

[0179] Generation of non-root nodes:

[0180] Generation of marital status:

[0181] Generation of education level:

[0182] Generation of weekly working hours:

[0183] Generation of occupation:

[0184] Generation of employment status:

[0185] According to the topological order, generate all nodes in sequence until all features in the causal relationship graph \(M\) are generated, and finally output the generated sample set

[0186] S6. Input the sample set C and the generated sample set into the discriminator D. The generator set G and the discriminator D are jointly trained in an adversarial manner to obtain the optimized generator set G opt ={G 1,opt , G 2,opt , G 3,opt , G 4,opt , G 5,opt , G 6,opt , G 7,opt}, and the discriminator D Opt ;

[0187] Take the sample set C = (X, S, U, Y) and the generated sample set as the training input of the discriminator D. The discriminator D is used to distinguish real samples from generated samples and simultaneously evaluate the distribution consistency of the generated samples;

[0188] The loss function L D of the discriminator D is expressed as:

[0189] where E real represents the expectation of the real sample set C, E fake represents the expectation of the generated sample set , D(C) represents the probability of the discriminator's prediction of the authenticity of the real sample set C, represents the probability of the discriminator's prediction of the authenticity of the generated sample set . This loss function L D optimizes the discriminator D by maximizing the prediction probability of real samples and minimizing the prediction probability of generated samples;

[0190] The adversarial training process of the discriminator D is as follows:

[0191] Fix the generator set G and update the discriminator D: Update the parameters using the gradient descent method to improve the discriminator D's ability to distinguish between the real sample set C and the generated sample set ;

[0192] Fix the discriminator D and update the generator set G: Minimize the following generator set loss function L G to maximize the probability of the generated sample set deceiving the discriminator D:

[0193] According to the adversarial training results, select the optimal solutions of the generator set G and the discriminator D: The generator set G opt ={G 1,opt , G 2,opt , G 3,opt , G 4,opt , G5,opt} and discriminator D Opt 。

[0194] S7. Input the sample set C into the generator set G opt ={G 1,opt ,G 2,opt ,...,G i,opt ,...,G p,opt} and discriminator D Opt ,for different values of each sensitive feature S, generate corresponding counterfactual fair samples, and aggregate all counterfactual fair samples to obtain a counterfactual fair dataset Data that ensures the fairness requirements of each sensitive feature fair 。

[0195] Input the sample set C=(X, S, U, Y) into the optimized generator set G opt ={G 1,opt ,G 2,opt ,G 3,opt ,G 4,opt ,G 5,opt} and discriminator D opt ,for different value sets {0, 1, 2} of each sensitive feature S, generate corresponding counterfactual sample sets one by one:

[0196]

[0197]

[0198] Aggregate the counterfactual samples corresponding to different values of all sensitive features S, and calculate the final fairness dataset by taking the mean:

[0199]

[0200] The present invention is compared with other algorithms, and the comparison metrics include classification accuracy ACC, maximum mean discrepancy MMD, and Wasserstein-1 distance. The accuracy metric ACC represents the accuracy of whether the obtained classification results match the actual classification results. The higher the ACC, the better the classification effect. The maximum mean discrepancy MMD is a metric for measuring the distribution difference between the generated data and the real data. It evaluates whether the generated data is similar to the real data in distribution by comparing different feature distributions. The smaller the MMD, the closer the generated data is to the real data in terms of the overall feature distribution, thus ensuring the authenticity of the data. The Wasserstein-1 distance (also known as the Earth Mover’s Distance) is a metric for measuring the distance between two probability distributions. It represents the "cost" required to "transport" one distribution to another. The smaller the WASS, the more similar the generated data is to the real data in terms of the feature distribution. This similarity ensures that the generated data retains the structural characteristics of the real data, which helps to generate high-quality data samples with fairness and authenticity. Table 2 is used to test the classification performance of the present invention. The second column, Original Data, is the original data. The third column, TabFairGAN, is a generative adversarial network proposed in the paper "Tabfairgan: Fair tabular data generation with generative adversarial networks." that generates a synthetic data set that meets both authenticity and fairness criteria. The fourth column, DECAF, is a generative adversarial network proposed in the paper "Decaf: Generating fair synthetic data using causally-aware generative networks." that achieves unbiased data reconstruction by focusing on causal parent variables using conditional causal reweighting. The last column is the method of the present invention. The comparison results are shown in Table 2. From the results in the table, it can be seen that the method proposed by the present invention maintains the highest accuracy rate, and the fairness of the model has been greatly improved, which proves the effectiveness of the present invention.

[0201] Table 2. Comparison results of classification performance and fairness metrics between the disclosed method of the present invention and existing methods

[0202]

[0203] Example 3

[0204] Such as Figure 4As shown in the figure, this embodiment provides a counterfactual fairness synthetic data generation device based on causal inference. The counterfactual fairness synthetic data generation device based on causal inference includes: a sample set generation module 401, a sample expansion module 402, a causal relationship graph establishment module 403, a generator discriminator creation module 404, a generated sample set acquisition module 405, a generator discriminator joint training module 406, and a counterfactual fairness data set generation module 407. The specific functions of each module are as follows:

[0205] The sample set generation module 401 is used to input the law school admission training data Data train , identify the observable features X, sensitive features S, and label values Y in the training data Data train , and combine them into a sample form and form a sample set C;

[0206] The sample expansion module 402 is used to input the sample set C into a variational autoencoder for training and learning, extract the latent features U from the sample set C. The latent features U represent latent factors that cannot be directly observed but may affect the label value Y. Finally, the sample form is expanded to

[0207] The causal relationship graph establishment module 403 is based on the dependencies between the observable features X, sensitive features S, latent features U, and label values Y, and establishes a causal relationship graph M. The causal relationship graph M is a directed acyclic graph, which is used to clarify the causal paths and dependencies between each feature. Each node in the graph represents a feature in the sample set C, and the directed edges in the graph represent the causal relationships between two features in the sample set C;

[0208] The generator discriminator creation module 404 is used to create an independent generator G for each node of the causal relationship graph M i , i = 1, 2,..., p, to obtain a generator set G = {G1, G2,..., G i ,..., G p}, where p is the total number of all features in the causal relationship graph M. Each node of the causal relationship graph M corresponds to a feature in the sample set C, and a discriminator D is defined;

[0209] The generated sample set acquisition module 405 is used to input the sample set C into the generator set G, generate variables in turn according to the topological order of the causal relationship graph M, first generate the root node, and then generate its child nodes until all variables are generated, to obtain a generated sample set where is the observable feature simulated according to the causal relationship graph M and the generator set G, where is a sensitive feature simulated and generated according to the causal relationship diagram M and the generator set G, where is a potential feature simulated and generated according to the causal relationship diagram M and the generator set G, where is a label value simulated and generated according to the causal relationship diagram M and the generator set G;

[0210] The generator discriminator joint training module 406 is used to input the sample set C and the generated sample set into the discriminator D, and the generator set G and the discriminator D are jointly trained in an adversarial manner to obtain the optimized generator set G opt ={G 1,opt , G 2,opt ,..., G i,opt ,..., G p,opt} and the discriminator D Opt , where G i,opt represents the final optimized generator for the i-th feature node;

[0211] The counterfactual fair dataset generation module 407 inputs the sample set C into the generator set G opt ={G 1,opt , G 2,opt ,..., G i,opt ,..., G p,opt} and the discriminator D Opt , and for different values of each sensitive feature S, generates corresponding counterfactual fair samples, and aggregates all the counterfactual fair samples to obtain the counterfactual fair dataset Data fair .

[0212] Example 4

[0213] This embodiment provides an electronic device, which can be a computer. As Figure 5 shown, it includes a processor 502, a memory, an input device 503, a display 504, and a network interface 505 connected by a system bus 501. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium 506 and an internal memory 507. The non-volatile storage medium 506 stores an operating system, a computer program, and a database. The internal memory 507 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the processor 502 executes the computer program stored in the memory, it implements the method for generating counterfactual fair synthetic data based on causal reasoning proposed in the above-mentioned Embodiment 1. The method for generating counterfactual fair synthetic data based on causal reasoning includes the following steps:

[0214] S1. Input the recruitment screening class training data Data train, identify training data Data train The observable features X, sensitive features S and label values Y in are combined and expressed as samples And form a sample set C;

[0215] S2, input the sample set C into the variational autoencoder for training and learning, extract the potential feature U from the sample set C, the potential feature U represents the potential factor that cannot be directly observed but may affect the label value Y, and finally, the sample form Expand to

[0216] S3. Based on the dependencies among observable features X, sensitive features S, potential features U and label values Y, a causal relationship graph M is established. The causal relationship graph M is a directed acyclic graph, which is used to clarify the causal path and dependency relationship between each feature. Each node in the graph represents a feature in the sample set C, and a directed edge in the graph represents the causal relationship between two features in the sample set C.

[0217] S4. Create an independent generator G for each node of the causal graph M. i , i=1,2,…,p, get the generator set G oppt = {G1, G2, ..., G i , ..., G p}, where p is the total number of all features in the causal graph M, each node of the causal graph M corresponds to a feature in the sample set C, and a discriminator D is defined;

[0218] S5. Input the sample set C into the generator set G, and generate variables in sequence according to the topological order of the causal graph M. First, generate the root node, then generate its child nodes, until all variables are generated, and obtain the generated sample set in is the observable feature generated by simulation based on the causal graph M and the generator set G, where is the sensitive feature generated by simulation based on the causal relationship graph M and the generator set G, where is the potential feature generated by simulating the causal graph M and the generator set G, where is the label value generated by simulation based on the causal relationship graph M and the generator set G;

[0219] S6. Combine sample set C and generate sample set Input to the discriminator D, the generator set G and the discriminator D are jointly trained in an adversarial manner to obtain the optimized generator set G opt = {G 1,opt , G 2,opt , ..., G i,opt , ..., G p,opt}, where G i,opt represents the final optimized generator for the i-th feature node, and the discriminator D Opt ;

[0220] S7. Input the sample set C into the generator set G opt ={G 1,opt , G 2,opt ,..., G i,opt ,..., G p,opt} and the discriminator D Opt , and for different values of each sensitive feature S, generate corresponding counterfactual fair samples, and aggregate all the counterfactual fair samples to obtain a counterfactual fair dataset Data that ensures the fairness requirements of each sensitive feature fair .

[0221] Example 5

[0222] This example provides a storage medium, which is a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements a method for generating counterfactual fair synthetic data based on causal inference in Example 1 above. The method for generating counterfactual fair synthetic data based on causal inference includes the following steps:

[0223] S1. Input the recruitment screening class training data Data train , identify the observable features X, sensitive features S, and label values Y in the training data Data train , and merge them into a sample form and form a sample set C;

[0224] S2. Input the sample set C into a variational autoencoder for training and learning, extract the latent feature U from the sample set C. The latent feature U represents a latent factor that cannot be directly observed but may affect the label value Y. Finally, the sample form is extended to

[0225] S3. Based on the dependencies among the observable features X, sensitive features S, latent features U, and label values Y, establish a causal relationship graph M. The causal relationship graph M is a directed acyclic graph used to clarify the causal paths and dependencies among each feature. Each node in the graph represents a feature in the sample set C, and the directed edges in the graph represent the causal relationships between two features in the sample set C;

[0226] S4. Create an independent generator G i for each node of the causal relationship graph M, i = 1, 2,..., p, to obtain the generator set G = {G1, G2,..., G i ,..., G p}, where p is the total number of all features in the causal relationship graph M. Each node of the causal relationship graph M corresponds to a feature in the sample set C, and a discriminator D is defined;

[0227] S5. Input the sample set C into the generator set G, and generate variables in sequence according to the topological order of the causal relationship graph M. First, generate the root node, and then generate its child nodes until all variables are generated, obtaining the generated sample set where is the observable feature simulated according to the causal relationship graph M and the generator set G, where is the sensitive feature simulated according to the causal relationship graph M and the generator set G, where is the latent feature simulated according to the causal relationship graph M and the generator set G, where is the label value simulated according to the causal relationship graph M and the generator set G;

[0228] S6. Input the sample set C and the generated sample set into the discriminator D. The generator set G and the discriminator D are jointly trained in an adversarial manner to obtain the optimized generator set G opt ={G 1,opt , G 2,opt ,..., G i,opt ,..., G p,opt}, where G i,opt represents the final optimized generator for the i-th feature node, and the discriminator D Opt ;

[0229] S7. Input the sample set C into the generator set G opt ={G 1,opt , G 2,opt ,..., G i,opt ,..., G p,opt} and the discriminator D Opt . For different values of each sensitive feature S, generate corresponding counterfactual fair samples, and aggregate all the counterfactual fair samples to obtain the counterfactual fair dataset Data that ensures the fairness requirements of each sensitive feature fair .

[0230] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0231] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0232] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A counterfactual fairness synthetic data generation method based on causal inference, characterized in that The counterfactual fairness synthetic data generation method includes the following steps: S1. Input the recruitment screening type training data Data train , identify the observable feature X, sensitive feature S and label value Y in the training data Data train , and merge them into the sample form to form the sample set C; S2. Input the sample set C into the variational autoencoder for training and learning, extract the latent feature U from the sample set C, where the latent feature U represents latent factors that cannot be directly observed but may affect the label value Y. Finally, the sample form is extended to S3. Based on the dependencies among the observable feature X, the sensitive feature S, the latent feature U, and the label value Y, establish a causal relationship graph M. The causal relationship graph M is a directed acyclic graph, which is used to clarify the causal paths and dependencies between each feature. Each node in the graph represents a feature in the sample set C, and the directed edge in the graph represents the causal relationship between two features in the sample set C; S4. Create an independent generator G for each node of the causal relationship graph M i , where i = 1, 2, …, p, to obtain the generator set G = {G1, G2, ..., G i , ..., G p}, where p is the total number of all features in the causal relationship graph M, each node of the causal relationship graph M corresponds to a feature in the sample set C, and a discriminator D is defined; S5. Input the sample set C into the generator set G, and generate variables in sequence according to the topological order of the causal relationship graph M. First, generate the root nodes, and then generate their child nodes until all variables are generated, obtaining the generated sample set wherein are respectively the observable features, sensitive features, latent features, and label values simulated according to the causal relationship graph M and the generator set G; S6. Input the sample set C and the generated sample set into the discriminator D. The generator set G and the discriminator D are jointly trained in an adversarial manner to obtain the optimized generator set G opt ={G 1,opt , G 2,opt ,..., G i,opt ,..., G p,opt}, where G i,opt represents the final optimized generator for the i-th feature node, and the discriminator D Opt ; S7. Input the sample set C into the generator set G opt ={G 1,opt , G 2,opt ,..., G i,opt ,..., G p,opt} and the discriminator D Opt . For different values of each sensitive feature S, generate corresponding counterfactual fair samples, and aggregate all the counterfactual fair samples to obtain the counterfactual fair dataset Data that ensures the fairness requirements of each sensitive feature fair .

2. The counterfactual fairness synthetic data generation method based on causal reasoning according to claim 1, wherein The process of step S1 is as follows: Express the observable feature X in the training dataset Data train as a numerical matrix of dimension n×m where n is the number of samples and m is the feature dimension of each sample; The training dataset Data train The sensitive feature S in it can be represented as a numerical matrix of dimension n×k where k is the dimension of the sensitive feature S; The training dataset Data train The label value Y in it is represented as a vector of dimension n×1 where the values of the vector elements are 0 or 1. A value of 0 indicates that the data has a risk of unfairness, while a value of 1 indicates that there is no such potential unfairness risk; Merge the observable feature X, the sensitive feature S, and the corresponding label value Y to form a sample set C. A sample C in the sample set C i =(X i , S i , Y i ) represents the complete information on the observable feature X i , the sensitive feature S i , and the label value Y i of the i-th data record.

3. The counterfactual fairness synthetic data generation method based on causal reasoning according to claim 1, wherein The process of step S2 is as follows: Input the sample set C into the variational autoencoder for training and learning. The variational autoencoder consists of a decoder and an encoder; Among them, the encoder takes the observable feature X and the label value Y in the sample set C as inputs, maps the inputs to the latent space containing the latent feature U through a neural network, and outputs the parameters describing the distribution q(U|X, Y) of the latent feature U, including the mean μ and the log variance logσ of the latent feature U 2 , Construct the latent feature U according to the following formula: U = μ + δ·ε where δ = exp(0.5·logσ 2 ), and ε is a random noise sampled from the standard normal distribution N(0, 1); The decoder takes the latent feature U and the sensitive feature S as inputs, reconstructs the observable feature X and the label value Y, and outputs the reconstructed observable feature X reconst and the label value Y reconst ; Define the preliminary optimization loss function as follows: L r = E q(U|X,Y) [-log(p(X,Y|U,S))] + KL[q(U|X,Y)||p(U)] Among them, p(X, Y|U, S) = p(Y|X, U, S)·p(X|U, S). The first term p(Y|X, U, S) describes the generation probability of the label value Y under the conditions of the observable feature X, the latent feature U, and the sensitive feature S, reflecting the influence of the observable feature X and the sensitive feature S on the label value Y. The second term p(X|U, S) describes the generation process of the observable feature X under the conditions of the latent feature U and the sensitive feature S, ensuring that the latent space has interpretability for the observable feature X, F q(U|X,Y) [-log(p(X, Y|U, S))] measures the similarity between the observable feature X and the label value Y and the reconstructed observable feature X generated by the decoder through the negative value of the log-likelihood reconst and the label value Y reconst The similarity of, KL[q(U|X, Y)|p(U)] measures the difference between the distribution q(U|X, Y) of the latent feature U generated by the encoder and the prior distribution p(U) of the latent feature U. p(U) is the standard normal distribution N(0, 1), and KL[·||·] is the KL divergence; The maximum mean discrepancy (MMD) is introduced as a regularization term, and MMD is used to measure the difference between two distributions and The difference is defined as: Among them, is a feature mapping function for projecting data into a replicated nuclear Hilbert space; Finally, the expression of the loss function L for training and optimizing the variational autoencoder is as follows: where α ≥ 0 is an important hyperparameter that controls the importance of the distribution averaging term, denotes the number of different pairs of sensitive features (s, s′), |S| denotes the number of different values of the sensitive feature S, p(U|s) denotes the probability distribution of the latent feature U given that the value of the sensitive feature is s, and p(U|s′) denotes the probability distribution of the latent feature U given that the value of the sensitive feature is s′.

4. The counterfactual fairness synthetic data generation method based on causal reasoning according to claim 1, wherein In the step S4, each generator G i simulates the distribution corresponding to the i-th feature by generating data, and defines a discriminator D for evaluating the generation quality of the generator set G and discriminating whether the data generated by the generator G i conforms to the feature distribution in the causal relationship diagram M.

5. The counterfactual fairness synthetic data generation method based on causal reasoning according to claim 1, characterized in that, The process of step S5 is as follows: According to the directed acyclic graph structure of the causal relationship graph M, generate features layer by layer according to the following steps: First, perform a topological sort on the causal relationship graph M to determine the node generation order. Start generating from the root node and generate the dependent features layer by layer backward; The root node has no parent node and is generated by a random noise variable: Among them, represents the generation eigenvalue of the root node r in the causal relationship graph M; G r () represents a generator for generating the root node features; Z r represents a random noise independently sampled from the standard normal distribution N(0, 1); The generation of non-root nodes depends on the generation results of the parent nodes and random noise. The generation process is as follows: Among them, represents the sample value generated by the i-th feature node in the causal relationship diagram M, represents the set generated by the parent node of the i-th node; Z i is random noise independently sampled from the standard normal distribution N(0, 1); the generator set G = {G1, G2,..., G i ,..., G p} performs combined modeling according to the node relationship; Generate all nodes in sequence according to the topological order until all features in the causal relationship graph M are generated, and finally output the generated sample set 6. The counterfactual fairness synthetic data generation method based on causal reasoning according to claim 1, wherein The process of step S6 is as follows: The sample set C = (X, S, U, Y) and the generated sample set are used as the training inputs of the discriminator D, which is used to distinguish real samples from generated samples and evaluate the distribution consistency of the generated samples; The loss function \(L\) of the discriminator \(D\) D is expressed as: Among them, E real represents the expectation of the real sample set C, and E fake represents the expectation of the generated sample set . D(C) represents the probability of the discriminator predicting the authenticity of the real sample set C, represents the probability of the discriminator predicting the authenticity of the generated sample set . This loss function L D optimizes the discriminator D by maximizing the prediction probability of real samples and minimizing the prediction probability of generated samples; The adversarial training process of the discriminator D is as follows: Fix the generator set G and update the discriminator D: Use gradient descent to update the parameters to improve the discriminator D's ability to distinguish between the real sample set C and the generated sample set ; Fix the discriminator D and update the generator set G: Minimize the following generator set loss function L G to maximize the probability of the generated sample set fooling the discriminator D: According to the adversarial training results, select the optimal solutions of the generator set G and the discriminator D: The generator set G opt ={G 1,opt , G 2,opt ,..., G i,opt ,..., G p,opt}, where G i,opt represents the final optimized generator and discriminator D for the i-th feature node Opt .

7. The counterfactual fairness synthetic data generation method based on causal reasoning according to claim 1, wherein The process of step S7 is as follows: Input the sample set C = (X, S, U, Y) into the optimized generator set G opt = {G 1,opt , G 2,opt ,..., G p,opt} and the discriminator D opt , for each different value set {s1, s2,..., s k ,..., s K} of each sensitive feature S, generate the corresponding counterfactual sample sets one by one: Among them, indicates the sample set s1, s2,..., s k generated when the sensitive feature S takes the value S k ,..., s K are K different values of the sensitive feature S. During the generation process, other features are kept unchanged, and only the sensitive feature S is adjusted to evaluate the fairness impact; Aggregate the counterfactual samples corresponding to different values of all sensitive features S, and calculate the final fairness dataset by taking the mean:

8. A counterfactual fairness synthetic data generation device based on causal reasoning, which is used to execute the counterfactual fairness synthetic data generation method based on causal reasoning according to any one of claims 1 to 7 above, and is characterized in that The counterfactual fairness synthetic data generation device includes: A sample set generation module for inputting recruitment screening type training data Data train , identifying observable features X, sensitive features S, and label values Y in the training data Data train , combining and representing them in the form of samples and forming a sample set C; A sample expansion module is used to input a sample set C into a variational autoencoder for training and learning, extract latent features U from the sample set C, where the latent features U represent latent factors that cannot be directly observed but may affect the label value Y. Finally, the sample form is expanded to A causal relationship graph establishment module, which based on the dependencies among the observable feature X, the sensitive feature S, the latent feature U, and the label value Y, establishes a causal relationship graph M. The causal relationship graph M is a directed acyclic graph, which is used to clarify the causal paths and dependencies between each feature. Each node in the graph represents a feature in the sample set C, and the directed edge in the graph represents the causal relationship between two features in the sample set C; A generator discriminator creation module for creating an independent generator G for each node of the causal relationship graph M i where \(i = 1, 2, \ldots, p\), to obtain a set of generators \(G=\{G_1, G_2, \ldots, G i , \ldots, G p \}\), where \(p\) is the total number of all features in the causal relationship graph M, each node of the causal relationship graph M corresponds to a feature in the sample set C, and a discriminator D is defined; The generated sample set acquisition module is used to input the sample set C into the generator set G, generate variables in sequence according to the topological order of the causal relationship graph M, first generate the root nodes, and then generate their child nodes until all variables are generated, obtaining the generated sample set Where is the observable feature simulated according to the causal relationship graph M and the generator set G, where are respectively the observable feature, sensitive feature, latent feature, and label value simulated according to the causal relationship graph M and the generator set G; The generator-discriminator joint training module is used to input the sample set C and the generated sample set into the discriminator D. The generator set G and the discriminator D are jointly trained in an adversarial manner to obtain the optimized generator set G opt ={G 1,opt ,G 2,opt ,...,G i,opt ,...,G p,opt} and the discriminator D Opt ,where G i,opt represents the final optimized generator for the i-th feature node; The counterfactual fairness dataset generation module inputs the sample set C into the generator set G opt ={G 1,opt , G 2,opt ,..., G i,opt ,..., G p,opt} and the discriminator D Opt . For different values of each sensitive feature S, corresponding counterfactual fairness samples are generated, and all the counterfactual fairness samples are aggregated to obtain the counterfactual fairness dataset Data that ensures the fairness requirements of each sensitive feature fair .

9. An electronic device, comprising a processor and a memory for storing processor-executable programs, characterized in that, When the processor executes the program stored in the memory, it implements the counterfactual fairness synthetic data generation method based on causal inference according to any one of claims 1 to 7.

10. A storage medium stores a program, characterized in that, When the program is executed by the processor, it implements the counterfactual fairness synthetic data generation method based on causal inference according to any one of claims 1 to 7.

Citation Information

Cited By

  • Medical synthetic data analysis method and device based on causal reasoning, and medium

    CN121075691A

  • Data set construction method oriented to special reasoning model

    CN121257759A

  • A data set construction method for a special reasoning model

    CN121257759B

  • Causal decoupling method and device based on multi-scale noise and adversarial supervision

    CN121415085A