Boundary-guided black box fairness testing method
By constructing a decision boundary exploration direction vector in the latent space of a generative adversarial network and generating and detecting individual discriminatory samples, the problem of balancing the efficiency and naturalness of discriminatory sample generation in existing technologies is solved, and the efficiency and effectiveness of black-box fairness testing are improved.
Patent Information
- Application Number
- CN202510829679.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-26
AI Technical Summary
Existing black-box fairness testing methods cannot effectively balance the naturalness, efficiency, and effectiveness of discriminatory sample generation. Some methods ignore the naturalness of generated samples, resulting in reduced efficiency and effectiveness, while other methods reduce search efficiency for the sake of naturalness.
By constructing an exploration direction vector pointing to the decision boundary in the latent space of the generative adversarial network, the generative adversarial network model maps the latent vector samples to instance samples. Taking advantage of the model's instability near the decision boundary, a boundary-guided latent vector sample exploration strategy is designed to generate test samples distributed near the decision boundary, and detect individual discrimination samples through small perturbations.
It achieves high efficiency in discriminatory sample search and naturalness in generated samples, improves the efficiency and effectiveness of black-box fairness testing, ensures that the generated samples conform to the real data distribution, and reduces the waste of high-quality samples.
Smart Images

Figure CN120705050A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and software testing, and in particular to a boundary-guided black box fairness testing method. Background Art
[0002] Neural network models are applied to numerous fields, including credit scoring, medical diagnosis, and fraud detection, to assist human decision-making. While neural network models demonstrate excellent performance across a variety of decision-making tasks, they can learn biases inherent in the data and produce unfair decisions. Furthermore, neural network models suffer from poor interpretability, making it difficult to trace and correct the unfair decisions they produce. Recent research and practice have shown that some neural network models may exhibit discriminatory behavior based on protected attributes such as gender and occupation. For example, Amazon's resume screening algorithm was exposed for scoring female applicants significantly lower than male applicants with equal qualifications, resulting in a lower approval rate for female resumes than male applicants. Therefore, systematically testing and correcting discriminatory behavior in neural network models is crucial.
[0003] To test model fairness, many studies have proposed generating test samples that violate fairness requirements, thereby revealing discriminatory behavior in the model. Fairness is generally categorized into two types: individual fairness and group fairness. Individual fairness requires that similar individuals who differ only in protected attributes receive similar decision outcomes, while group fairness requires that groups divided by protected attributes be treated equally. Assessing a model's individual fairness can reveal more discriminatory behavior that is often overlooked by group fairness.
[0004] Existing methods for testing individual fairness of models mainly include:
[0005] (1) A method for testing the fairness of a deep neural network model, publication number: CN116662153A: This method obtains several samples and queries the target model to obtain the prediction results of each sample, then clusters the samples to obtain clusters and trains the replacement models respectively. Based on the corresponding replacement models, a new set of seed samples is generated for each sample cluster, and perturbations are applied to the extracted seeds, thereby making them become samples that violate the fairness conditions with a high probability. Based on multiple perturbations applied to the currently discovered samples that violate the fairness conditions, more samples that violate the fairness conditions are discovered to fully verify the fairness of the deep neural network model.
[0006] (2) A natural fairness test case generation method for machine learning models, Publication No.: CN112433952A:
[0007] This method constructs a dataset to simulate the decision-making process of a machine learning model in the latent space, deriving an approximate proxy decision boundary. Furthermore, by exploiting the model's non-robust nature on the decision boundary, a latent vector candidate detection strategy is designed to generate fairness test cases that conform to natural laws near the proxy decision boundary.
[0008] Existing new fair testing methods cannot balance the naturalness, efficiency, and effectiveness of discriminatory sample generation. Related work either completely ignores the naturalness of generated samples and uses simple random bootstrapping strategies to maximize the efficiency and effectiveness of sample generation, or adopts rough sample search strategies for the sake of naturalness, which reduces the efficiency and effectiveness of discriminatory sample search. Summary of the Invention
[0009] To address the problem that existing black-box fairness testing methods fail to balance the naturalness, effectiveness, and efficiency of discriminatory sample generation, this paper provides a boundary-guided black-box fairness testing method. This method guides test samples by constructing an exploration direction vector pointing toward the decision boundary in the latent space of a generative adversarial network. This ensures the efficiency and effectiveness of discriminatory sample search and the naturalness of the generated discriminatory samples.
[0010] To achieve the above objectives, the present invention provides the following solutions:
[0011] A new boundary-guided black-box fair testing method, the method comprising:
[0012] Based on the decision-making characteristics of the neural network model, latent vector samples are randomly generated in the latent space. The trained generative adversarial network model is used to map them into instance samples in the input space of the real dataset, and the predicted labels are obtained by inputting the neural network model. An auxiliary dataset consisting of latent vector samples and corresponding predicted labels is constructed, and the centroid of data with different predicted labels is calculated to obtain the direction vector pointing to the decision boundary.
[0013] Taking advantage of the instability of neural network models near the decision boundary, a boundary-guided latent vector sample exploration strategy is designed. First, a random unit vector is generated and its cosine value with the direction vector is calculated to see if it exceeds a threshold. If so, it is used as the sample search direction; otherwise, a new one is generated. Then, an exploration step is given by the average fitness of the current latent vector sample set, and new latent vector samples closer to the decision boundary are explored along the sample exploration direction through vector calculation.
[0014] Merge the new latent vector sample set with the current latent vector sample set, calculate the probability distribution of the samples based on the fitness of all latent vector samples in the merged sample set, and then perform random sampling with unequal probabilities to select the candidate latent vector sample set for this round;
[0015] Use the generative adversarial network model to map the candidate latent vector sample set to the candidate test sample set, and test the performance of the target model on the generated test samples; if a test sample is judged to be an individual discrimination sample, it is added to the discrimination sample set;
[0016] Iterate the above exploration process and randomly perturb the test sample set obtained by the final exploration. Select a small number of non-sensitive attribute columns of a given sample and randomly perturb the attribute values of the selected perturbation columns within a very small range to obtain perturbed test samples. Test the performance of the target model on the perturbed test samples. If a perturbed test sample is judged to be an individual discrimination sample, it is added to the discrimination sample set. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 This is a flow chart of a boundary-guided black-box fairness testing method in one embodiment of the present invention.
[0019] Figure 2 301 and 302 in an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0021] Most existing fairness testing methods focus only on the efficiency and effectiveness of generating individual discriminatory samples in neural network models (abbreviated as models, hereinafter the same), ignoring whether these samples conform to the distribution of real data. Some samples may even violate objective constraints in the real world. Methods that focus on the naturalness of test sample generation use rough exploration strategies to search for individual discriminatory samples, resulting in a large number of high-quality samples being wasted, seriously reducing the efficiency and effectiveness of testing. To this end, an embodiment of the present invention is dedicated to providing a boundary-guided black-box fairness testing framework. By constructing an auxiliary dataset, the decision process of the model is mapped into the latent space. Based on this auxiliary dataset, the relative spatial position of the decision boundary in the latent space is determined. Utilizing the model's weak robustness near the decision boundary, a boundary-guided latent vector sample exploration strategy is designed to generate test samples distributed near the decision boundary. Small random perturbations are added to these test samples to generate a large number of similar test samples also distributed near the decision boundary, and samples that violate individual fairness are detected from them.
[0022] like Figure 1 As shown, an embodiment of the present invention provides a boundary-guided black-box fairness testing method, which is used by the tester to search for individual discrimination samples distributed near the decision boundary of the model. The method includes the following steps:
[0023] Step 101: Use a real-world dataset to train a generative adversarial network model so that it learns the data distribution characteristics of the training set data.
[0024] Taking structured tabular data as an example, after removing the labels in the training set, the data is input into the table generative adversarial network for learning, and a generator G is obtained, which can map any latent vector in the latent space to the distribution of real data. The generation process g can be expressed as:
[0025] g(z):Z→X
[0026] Where X represents the input space of the data sample, Z represents the latent space, and g(z) maps the latent vector z to the input space of the data sample to obtain the corresponding data sample.
[0027] Step 102: Latent vector samples are randomly sampled in the latent space. The generator maps them into samples in the data input space. The samples are input into the model to obtain the predicted output. High-quality latent vector samples are selected based on the predicted output to construct an auxiliary dataset.
[0028] Randomly sample 1,000,000 latent vector samples from the latent space according to Gaussian distribution to form a set Z rand . Use the generator to convert Z randThe latent vector samples in are mapped to the data input space and input into the target neural network model for prediction. The prediction output and the corresponding latent vector samples together construct an auxiliary dataset D au , the auxiliary dataset can be expressed as:
[0029] D au =(z,f θ (g(z))):z∈Z rand
[0030] Among them, f θ Represents the target neural network model. The auxiliary dataset maps the decision space of the target model into the latent space, thereby approximating the decision process of the target model.
[0031] Furthermore, the auxiliary dataset D is selected au The predicted output is greater than the threshold ε to improve the quality of the sample data. au The threshold ε is empirically set to 0.7.
[0032] Step 103: derive a direction vector pointing to the decision boundary.
[0033] The auxiliary dataset D au The latent vector samples in are divided according to their corresponding predicted labels, and the centroids of the two types of latent vector sample sets are calculated. The two centroids are subtracted to obtain the direction vector of the latent vector samples in the corresponding class pointing to the decision boundary.
[0034] Step 201: Obtain an initial latent vector sample set.
[0035] From the auxiliary dataset D au Select a batch of latent vectors as the initial latent vector sample set V init =[z1,z2,…,z n ] and assign it to the current latent vector sample set V cur , generally V init The size of is set to 1000, that is, n is 1000.
[0036] Steps 202-203: Generate a random unit vector as the exploration direction, and determine whether the cosine value between the unit vector and the direction vector exceeds a set threshold.
[0037] Generate a random unit vector as the exploration direction, calculate the cosine value between it and the direction vector pointing to the decision boundary, and determine whether this cosine value is greater than the set threshold c thr If it is greater than, jump to step 203, otherwise jump to step 201. In order to ensure the efficiency and coverage of the discrimination sample search, the threshold c thr It is empirically set to 0.7.
[0038] In step 204, an exploration step is given according to the average fitness score of the current latent vector sample set, so as to adjust the distance moved by the latent vector each time it is explored.
[0039] Specifically, the generator generates the current latent vector sample set V cur Perform deduction and output the current data sample set X on the data input domain cur =[x1,x2,…,x n ], then all samples are input into the target neural network model to obtain the predicted output and calculate the fitness score FC of the sample cur =[fc1,fc2,…,fc n ], where the calculation process of the fitness score fit can be expressed as:
[0040] fit(x)=1-(Prob(x,l)-Prob(x,l'))
[0041] Among them, Prob(x,l) represents the target model f θ The probability of predicting sample x as label l, where l and l' represent the two labels of the binary classification. The fitness score can measure the relative distance between the sample and the decision boundary. The higher the fitness score, the closer it is to the decision boundary.
[0042] The average fitness score of the sample set describes the distance between the sample set as a whole and the decision boundary, and a value is randomly given in a value interval as the exploration step d according to its range. step , the higher the average fitness score, the smaller the exploration step size given.
[0043] Step 205 : Generate a new test latent vector sample set by moving the latent vector samples toward the decision boundary through vector calculation.
[0044] For the current latent vector sample set V cur , move it towards the decision boundary through vector calculation to obtain a new latent vector sample set V new ,The movement of the latent vector can be expressed as:
[0045] z new =z cur +uv·dis
[0046] Among them, uv represents the unit exploration direction vector of this exploration, and dis represents the distance moved by the exploration latent vector. step The distance between the two centroids controls the size of dis.
[0047] Step 206: Determine whether the number of explored paths reaches a set value.
[0048] By setting the number of exploration paths n path Control V new In order to ensure efficiency and breadth of exploration, n is usually path Set to 2. When V new When the number of latent vectors reaches 2*n, jump to step 206; otherwise, jump to step 201.
[0049] Step 207: Calculate the fitness score and probability distribution of the obtained latent vector samples.
[0050] The current latent vector sample set V cur Merge into the new latent vector sample set V new , that is, V new =[z1,z2,…,z n ,z'1,…,z' 2n ]. V is converted to new Map to the data input domain to obtain a new sample set X new =[x1,x2,…,x n ,x'1,…,x' 2n ], calculate X new The fitness score FC of all samples in new =[fc1,fc2,…,fc n ,fc'1,…,fc' 2n ], and calculate the probability distribution of the sample P = [p1, p2, ..., p n ].
[0051] Step 208: Obtain a candidate test sample set for this round of exploration.
[0052] According to the probability distribution P of the sample, the new latent vector sample set V is selected by the unequal probability random sampling strategy. new and the new sample set X new Select the corresponding latent vector v expl and sample x expl , and add them to the candidate test latent vector sample set V of this round of exploration expl and candidate test sample set X expl and V expl Assign to V cur .
[0053] Step 209: In the candidate discrimination sample set X expl Identify individual discrimination samples.
[0054] Modify the candidate discrimination sample set X explGiven the protected attribute value of sample x in the model, we obtain the dual sample x', input x and x' into the target model, if the predicted outputs of the two are different, then x is determined to be an individual discrimination sample, and x is added to the discrimination sample set G discr .
[0055] Step 210: Determine whether the exploration depth reaches the set maximum global iteration number.
[0056] Record the number of iterations of sample exploration and determine whether it has reached the maximum global iteration number n global , usually n global Set to 2, if reached, jump to step 301, otherwise jump to step 201.
[0057] Step 301: randomly select a small number of non-protected attribute columns.
[0058] For the candidate test sample set X obtained in the last global iteration expl For a given sample in , the ratio c of the perturbed column to the total number of attributes is pert Randomly select non-protected attribute columns as perturbed columns. In order to take into account the similarity between the generated perturbed samples and the original samples and the coverage of the perturbed samples, c peret Set to 0.2.
[0059] Step 302: Obtain a disturbed test sample.
[0060] For all selected disturbed columns, according to the value range of its attribute column and the attribute value disturbance ratio v pert Perturb the instance sample to obtain the perturbed sample x pert In this embodiment, v pert Set to 0.1.
[0061] Step 303: Determine individual discrimination samples.
[0062] Given a perturbed sample x pert The protected attribute value of the dual sample x' pert , change x pert and x' pert Input the target model, if the predicted outputs of the two are different, then determine x pert is an individual discrimination sample, and x pert Add to the discrimination sample set G discr .
[0063] Step 304: Determine whether the set maximum number of perturbation iterations has been reached.
[0064] Record whether the number of times a given sample is perturbed reaches the maximum number of local iterations n local , usually n localSet to 50 to take into account both efficiency and coverage. local Then the local perturbation of a sample is completed, otherwise jump to step 301.
[0065] Step 305: Determine whether to traverse the candidate test sample set X expl All instance samples in .
[0066] Determine whether to traverse the candidate test sample set X expl If yes, then jump to step 306, otherwise jump to step 301.
[0067] Step 306: Determine the batch traversal of the initial latent vector sample set V init Whether the total number of latent vector samples reaches the set threshold.
[0068] Traverse the initial latent vector sample set V in batches init Whether the total number of latent vector samples reaches the set threshold α, if yes, the test is completed, otherwise jump to step 201. In this embodiment, α is set to 30000.
[0069] See also Figure 2 In one embodiment of the present invention, the Census dataset is used as an example to illustrate the specific execution process of steps 301 to 302. The input of the Census dataset is a 13-dimensional vector, and the eighth dimension represents gender (represented by 0 / 1).
[0070] For the candidate test sample set X obtained in the last global iteration expl A given sample x exlpore =[4,0,1,10,0,5,2,0,1,0,0,49,0], based on the ratio of perturbed columns to the total number of attributes c pert = 0.2 Randomly select two non-protected attribute columns as the perturbed columns, such as columns 2 and 12. Perturb the ratio v according to the attribute value pert = 0.1 random perturbation x exlpore The attribute values of the 2nd and 12th columns, x exlpore The attribute value ranges of the second and 12th columns are [0,7] and [0,99] respectively, so the perturbation ranges of the attribute values of the two columns are [0,1] and [40,58] respectively. A value is randomly selected in these two intervals to generate the perturbed sample x. pert =[4,1,1,10,0,5,2,0,1,0,0,44,0].
Claims
1. A boundary-guided black-box fairness testing method, characterized in that: The method comprises: The tester randomly samples latent vectors from the latent space of the generative adversarial network model and maps them to the data input space through the generator to obtain corresponding instance samples. The target neural network model is used to predict these instance samples, constructing an auxiliary dataset consisting of latent vectors and predicted labels. The centroid of the latent vector data for different label classes is calculated, and the direction vector pointing to the decision boundary is derived. Taking advantage of the fact that neural network models often produce unstable predictions near the decision boundary, a boundary-guided exploration strategy is designed. First, some latent vectors are selected from the auxiliary dataset as initial samples. Based on the direction of the direction vector, multiple exploration paths are constructed from the initial samples toward the decision boundary. The exploration direction is determined by a generated random unit vector, and the exploration distance is the distance between the two centroids multiplied by an exploration step. Samples closer to the decision boundary are selected from the explored samples, and global test samples are iteratively generated. Individual discriminatory samples are identified by modifying the protected attribute values of the global test samples. A small random noise is added to the global test sample finally obtained to generate a large number of local test samples distributed in the sample space near the given global test sample, and the individual discrimination samples in the identification are modified by modifying the protected attribute values of the local test samples.
2. A boundary-guided black-box fairness testing method according to claim 1, characterized in that: The direction vector and auxiliary data set construction process requires a specific table to generate an adversarial network model, the target test model is various neural network models, and the data type of the model input data is tabular data. The table-generating adversarial network learns the feature distribution of real data and randomly samples a large number of latent vector samples from its latent space. It uses a generator to map the latent vector samples to the data input domain to obtain the corresponding sample set. Use the target neural network model to predict the corresponding labels of the samples and construct an auxiliary dataset consisting of latent vector samples and corresponding predicted labels; The latent vector is divided according to its predicted label and the centroid of the latent vector sample data of different predicted labels is calculated. The centroids are connected to obtain the direction vector pointing to the decision boundary and the distance between the two centroids.
3. A boundary-guided black-box fairness testing method according to claim 1, characterized in that: The construction of multiple exploration paths toward the decision boundary specifically includes: Determine the exploration direction, that is, for a random unit vector, when the cosine value between it and the direction vector pointing to the decision boundary is greater than a set threshold, it is used as the exploration direction of this round; Calculate the fitness score of the current latent vector sample. The fitness calculation process fit can be expressed as: fit(x)=1-(Prob(x,l)-Prob(x,l')) Among them, Prob(x,l) represents the probability that the target neural network model predicts that sample x is label l, and l and l' represent the two labels of binary classification; That is, for all current samples, their fitness scores can be expressed as a measure of how close any sample is to the decision boundary in the sample space of the target model; According to the fitness score, the exploration step length of the next round of exploration is given and the sampling probability of the current sample being selected for the next round of exploration is constructed; By calculating the moving latent vector through vector calculation, a new latent vector sample set is constructed. The movement of the latent vector can be expressed as: z new =z cur +uv·dis Wherein, uv represents the unit direction vector of this exploration, and dis represents the moving distance of this exploration, whose size is adjusted by the distance between the two centers of mass and the exploration step size; Calculate the fitness score and sampling probability of the new latent vector sample set. According to the sampling probability of the latent vector sample, select the latent vector sample and the instance sample through the unequal probability random sampling strategy to construct the latent vector sample set and the corresponding sample set for this round of exploration. Iterate the exploration to form an exploration path.
4. A boundary-guided black-box fairness testing method according to claim 1 or 3, characterized in that: The construction of the exploration path requires the fastest search for samples close to the decision boundary, that is, selecting multiple paths, which will find a set of samples that are generally close to the decision boundary.
5. The boundary-guided black box fairness testing method according to claim 1, characterized in that: The process of generating local test samples requires generating a large number of similar samples close to the decision boundary, that is, randomly selecting a small number of non-protected attributes in the global test samples and perturbing their attribute values within a very small range.
Citation Information
Patent Citations
Deep neural network model fairness test method, system and device and medium
CN112433952A
Natural fairness test case generation method for machine learning model
CN116662153A