General adversarial disturbance generation method for table data based on Gumbel-Softmax
By using the Gumbel-Softmax mechanism in the adversarial sample generation of tabular data, the problem of difficulty in processing discrete features in tabular data is solved. The generated adversarial samples have high success rate and good migration, and the optimization process is efficient.
Patent Information
- Application Number
- CN202510279609.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
AI Technical Summary
In the adversarial sample generation task facing table data, gradient optimization is difficult to directly apply because table data contains discrete features and traditional gradient methods are difficult to deal with.
A general tabular data counter-perturbation generation method based on the Gumbel-Softmax mechanism is adopted. By feature encoding and mapping the table data into the dictionary space, Gumbel-Softmax is used to generate adversarial samples that meet the problem domain constraints.
The generated adversarial samples have high success rate and good mobility, can induce prediction errors on multiple target models, and the optimization process is always carried out in the problem domain space, with small feature modification and high time efficiency.
Smart Images

Figure CN120106031A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of adversarial attacks on deep learning, and specifically relates to a universal adversarial perturbation generation method for tabular data. Background Art
[0002] Adversarial sample generation is an important area of research in machine learning security and robustness, revealing potential threats and vulnerabilities in existing machine learning models. By adding specific noise to normal samples, researchers can create adversarial samples that appear almost indistinguishable from the original samples to humans, but can mislead machine learning models into making incorrect predictions or classifications.
[0003] In image recognition, a common machine learning field, attackers usually use gradient descent methods to add small and carefully designed perturbations to images to generate adversarial samples. This process involves minimizing or maximizing the loss function and updating these perturbations during backpropagation while ensuring that the generated adversarial samples are reasonable. For example, the maximum value of the perturbation is limited to no more than a certain threshold, or the image pixel value is kept in a reasonable range of 0 to 255.
[0004] However, in the task of generating adversarial samples for discrete data such as tabular data, gradient optimization-based methods face significant challenges. Since optimization algorithms such as gradient descent rely on continuous differentiable mathematical properties to guide parameter updates, and tabular data often contain discrete features (such as categorical variables, Boolean values, etc.), traditional gradient methods are difficult to apply directly. To overcome this problem, existing methods usually treat tabular data as floating-point numbers after feature encoding, and try to use gradient descent to generate adversarial samples until the attack goal is achieved. However, this strategy often results in the generated adversarial samples being outside the original problem space or requiring a large number of feature modifications, which limits its effectiveness in practical applications. Summary of the invention
[0005] In order to solve the above problems, a general tabular data adversarial perturbation generation method based on the Gumbel-Softmax mechanism is proposed. This method aims to generate general adversarial perturbations for multiple tabular datasets that can mislead prediction models with a very high success rate. The generated adversarial samples have excellent transferability and can induce prediction errors on multiple target models.
[0006] The first aspect of the present invention discloses a method for generating general table data adversarial perturbations based on a Gumbel-Softmax mechanism, the method comprising:
[0007] S1: Clean and feature encode the public table dataset, and optimize the dictionary space size according to the number of possible values for each feature to ensure that the original sample features and the dictionary space features can be converted into each other.
[0008] S2: Map the data features in the dataset to the dictionary space.
[0009] S3: Generate an efficient perturbation generation model by training the dataset in the dictionary space through multiple rounds of iteration.
[0010] S4: Divide the value range of each feature into intervals according to a predefined ratio to precisely control the feature modification amplitude.
[0011] S5: Apply Gumbel-Softmax to generate adversarial samples that conform to the problem domain constraints in the dictionary space, and aggregate and statistics to generate a universal adversarial perturbation.
[0012] S6: Add the universal adversarial perturbation to the samples with unflipped labels, repeat step S5 and iteratively optimize the loss term until the training is completed to obtain the final universal adversarial perturbation.
[0013] Furthermore, the specific steps of encoding the tabular dataset into the dictionary space in step S1 are as follows:
[0014] S11: Given a dataset X containing n features, for each feature X i (0 ≤ i < n), statistically calculate the global maximum value of its number of values by traversing the dataset and the minimum value and use these boundary values as the benchmark for feature encoding.
[0015] S12: Define the discrete number of values V i of feature X i :
[0016]
[0017] where the legal values of X i cover all integer values within the interval.
[0018] S13: Determine the maximum number of values V max of all features:
[0019]
[0020] This parameter is used to unify the one-hot encoding dimensions of each feature to ensure the structural consistency of the dictionary space.
[0021] S14: Map each feature X i to a one-hot encoding vector R max with a dimension of 1 × V i , and construct a dictionary feature matrix by vertically stacking all feature encoding vectors Its formal representation is:
[0022] R=[R1;R2;…;Rn] (3)
[0023] Each row of the matrix corresponds to a feature of the original dataset, and each column represents the possible value of the feature in the unified encoding space. Initialize to The zero matrix of , the subscript indicates the optimization round.
[0024] Furthermore, step S2 of mapping the table data features to the dictionary space is specifically as follows:
[0025] S21: In dataset X, feature X i The global range boundary of and For the i-th feature original value x of sample x i , converted to a non-negative integer offset o by normalization i :
[0026]
[0027] This operation converts any discrete eigenvalue Linear mapping to dictionary space. Normalized offset o i With the original eigenvalue x i A bijective relationship is formed to ensure the equivalence between the dictionary space and the original feature space.
[0028] S22: Based on the calculated offset o i , construct a unique hot encoding vector with dimension Vmax Its mathematical definition is:
[0029]
[0030] Furthermore, the disturbance generation model structure of step S3 is specifically as follows:
[0031] The model defines multiple fully connected layers to process input data and output prediction results. The input layer converts the sample features n×V max The dimension is flattened to a smaller dimension and nonlinearity is introduced through the ReLU activation function. Subsequently, a series of intermediate layers are used to further reduce the dimension and extract more abstract features. Each layer also uses the ReLU activation function to maintain nonlinear characteristics.
[0032] Furthermore, the step S4 of dividing the intervals according to the number of values of each feature is specifically as follows:
[0033] Current sample feature x i The offset value in the dictionary space is o i, and define the allowed modification range ∈∈[0,1]. The minimum extreme value boundary of the feature that can be modified and the maximum extreme value boundary It can be calculated by the following formula:
[0034]
[0035] This constraint ensures that the eigenvalue does not exceed the bounds during the optimization process and that the difference with the original sample is within a reasonable range.
[0036] Furthermore, the general adversarial perturbation step S5 is generated by Gumbel-Softmax as follows:
[0037] S51: According to the i-th feature x i The minimum extreme boundary of and the maximum extreme value boundary Gumbel noise g is added to the matrix distribution Θ, and the temperature coefficient τ is set to a value close to 0.
[0038] S52: Add Gumbel noise g to the possible value range of each feature and limit the noise addition area:
[0039]
[0040] where g i,f ~Gumbel(0,1), τ>0 is the temperature parameter that controls the smoothness of the Gumbel-Softmax distribution. When τ approaches 0, the generated samples will asymptotically approach the one-hot vector:
[0041]
[0042] Indicator function Returns 1 if the condition is met, otherwise returns 0.
[0043] S53: Generate an adversarial sample for each original sample through the total loss function in formula (9), where α and β are the weight coefficients of each loss term.
[0044]
[0045] The loss function includes classification loss and L1 regularization term. The classification loss is shown in formula (10): It is used to measure the difference between the model prediction value and the actual label to ensure that the adversarial sample misleads the model:
[0046]
[0047] in, is the adversarial objective function, T represents the sample feasible domain, which refers to the part that needs to satisfy the problem domain constraints in the feature space, and D represents the cost metric between the adversarial sample generated by the adversary and the original sample, and δ is the added perturbation. M is the target model, and M(x+δ) predicts the category of the adversarial sample so that it is different from its own category y.
[0048] The L1 regularization term, such as formula (11), limits the number of feature modifications and prevents excessive perturbations by adding the sum of the absolute values of the weights to the loss function as a penalty:
[0049]
[0050] x is the original sample, x adv is the adversarial sample it generates, and n is the number of features of the sample.
[0051] S54: Generate an adversarial sample that can maximize the model error rate for each sample in the training set through formula (9).
[0052] S55: After generating adversarial samples for each sample, summarize and count the feature encodings of the adversarial samples generated in each round. In the adversarial samples generated in each round, the encoding vector of the i-th feature of the k-th adversarial sample in the dictionary space is For each round of adversarial samples, the universal adversarial perturbation generated The update rules are:
[0053]
[0054] in It is a binary indicator function, which takes the value 1 only when the condition is met. and They represent the unique hot encoding of the kth original sample and the adversarial sample in the i-th feature dimension respectively.
[0055] S56: Use the newly generated universal adversarial perturbation in each round to balance the contribution of historical perturbations and newly generated perturbations in proportion α∈(0,1), and merge the new perturbations into the global perturbation matrix:
[0056]
[0057] The adversarial samples generated by the model are within the value range of their respective features. The more times the adversarial samples appear on a certain feature value, the greater the impact of this feature will have on the model's prediction.
[0058] Furthermore, step S6 of adding the universal adversarial perturbation to the sample without flipped labels is specifically as follows:
[0059] S61: Add the universal perturbation to the dictionary space sample of the unflipped sticky note and offset it by the sample's feature value o i Neighborhood Perform a maximum activation frequency search to determine the optimal perturbation position and the optimal index Calculated by:
[0060]
[0061] If there are multiple maximum values, the distance o is preferred. i The nearest index. The final generated adversarial encoding:
[0062]
[0063] S62: During the training process, if the sample with added noise successfully misleads the model, its subsequent optimization is skipped; otherwise, the optimization is continued and the overall noise is adjusted according to the exponential decay rate until the training is completed and the final universal adversarial perturbation is obtained.
[0064] Beneficial effects of the present invention:
[0065] 1. Compared with the previous method for generating adversarial samples for tabular data, this paper constructs a general adversarial perturbation generation method for tabular data based on Gumbel-Softmax by softening the discrete features. This solves the problem that there is no general adversarial perturbation generation method for tabular data in the current research community.
[0066] 2. The present invention combines the dictionary space of constructed features with Gumbel-Softmax, so that adversarial samples can always access the target model with real-world data during the generation process. Adversarial samples can achieve better attack success rate and transferability with standards that are more in line with the real world.
[0067] 3. The optimization of samples by the present invention is always in the problem domain space, and can generate universal adversarial perturbations with a smaller modification feature amplitude and less time. This method can find loopholes in the prediction model based on tabular data in a relatively simple way, which is of great significance to improving the robustness and security of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 It is the overall flow chart of the present invention;
[0069] Figure 2 It is a conceptual diagram of the problem space and feature space of the present invention;
[0070] Figure 3 A schematic diagram of the process of constructing a dictionary space for the present invention;
[0071] Figure 4 A schematic diagram of the process of dividing the optimization interval of features of the present invention; DETAILED DESCRIPTION
[0072] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0073] Figure 1 This example demonstrates a general tabular data adversarial perturbation generation method based on the Gumbel-Softmax mechanism, which can effectively attack the prediction model based on tabular data training and cause it to misjudge the sample category.
[0074] Figure 2 The optimization differences between adversarial samples in the problem space and feature space are demonstrated. The problem space (irregular circular area) is composed of real samples and strictly follows the distribution of real data; the feature space covers a wider range and includes all potential samples that may affect model predictions, but some samples may deviate from the actual distribution. Traditional methods (such as class2→class3) approach the target category boundary by adjusting the eigenvalues in the feature space, but the samples need to be forcibly pulled back to the problem space through post-processing such as truncation or rounding, resulting in a significant attenuation of adversarial properties. This method innovatively adopts a direct optimization strategy in the problem space (class1→class3), which simultaneously satisfies the model aggressiveness and data distribution constraints during the iteration process. It can generate samples that are both in line with the laws of reality and have strong adversarial properties without post-processing, fundamentally solving the contradiction between adversarial properties and data legitimacy. The specific implementation steps of a general tabular data adversarial perturbation generation method based on the Gumbel-Softmax mechanism are as follows:
[0075] S1: Clean and feature encode the public table dataset, and optimize the dictionary space size according to the number of possible values for each feature to ensure that the original sample features and the dictionary space features can be converted into each other.
[0076] S2: Map the data features in the dataset into the dictionary space.
[0077] S3: Generate an efficient perturbation generation model by training the dataset in the dictionary space through multiple rounds of iterations.
[0078] S4: Divide the value range of each feature into intervals according to a predefined ratio to accurately control the feature modification range.
[0079] S5: Apply Gumbel-Softmax to generate adversarial samples in the dictionary space that meet the constraints of the problem domain, and summarize the statistics to generate universal adversarial perturbations.
[0080] S6: Add the universal adversarial perturbation to the samples without flipped labels, repeat step S5 and iteratively optimize the loss term until the training is completed to obtain the final universal adversarial perturbation.
[0081] Specifically, the method comprises the following steps:
[0082] (1) Download the Adult dataset X from the Internet and clean the data with default values. After cleaning, the eight features X i (0≤i<8) By traversing the data set, the global maximum value of its value is counted With minimum And these boundary values are used as the basis for feature encoding.
[0083] (2) Define feature X i The number of discrete values V i :
[0084]
[0085] Among them, X i The legal value coverage All integer values in the interval (the legal values of the age feature in the Adult dataset are integers from 17 to 90).
[0086] (3) Determine the maximum number of values V for all features max :
[0087]
[0088] This parameter is used to unify the unique-hot encoding dimensions of each feature to ensure the structural consistency of the dictionary space to support global optimization. max =99.
[0089] (4) Each feature X i Mapped to dimension 1×V max (1×99) one-hot encoded vector R i , and construct the dictionary feature matrix by vertically stacking all feature encoding vectors Its formal expression is:
[0090] R=[R1;R2;…;R8] (3)
[0091] Each row of the matrix corresponds to a feature of the original dataset, and each column represents the possible value of the feature in the unified encoding space. Initialize to The zero matrix of , the subscript indicates the optimization round.
[0092] Figure 3 The process of constructing the dictionary space in this example is shown. By counting the minimum and maximum values of each feature in the table data set, the number of values of each feature and the number of possible values after discretization are determined. When the data set contains n features, the maximum number of feature values V max , then we can construct a The dimensional dictionary space is used to realize the conversion of samples between feature space and problem space. This encoding mechanism enables each sample feature to be mapped to the corresponding vector index through the dictionary, and these indexes are used as model input for training or reasoning. At the same time, the reverse mapping mechanism allows the index sequence generated by the model to be decoded into the corresponding original feature value, realizing the effective conversion from model output to actual feature value.
[0093] (5) In the Adult dataset X, feature X i The global range boundary of and The original value x of the i-th feature of one of the samples x i can be converted to a non-negative integer offset o by normalization i :
[0094]
[0095] Make any discrete eigenvalue Linear mapping to the interval [0,V max -1], preserving the ordinal relationship of the features. At the same time, the offset o i With the original eigenvalue x i The bijective relationship is formed to ensure the equivalence between the dictionary space and the original feature space. For the feature whose age value range is [17,90], it can be mapped to [0,73], and the offset o i No optimization errors will occur due to insufficient dictionary dimensions.
[0096] (6) Based on the calculated offset o i , construct a unique hot encoding vector R with a dimension of 1×99 i ∈{0,1} 99 , which is mathematically defined as:
[0097]
[0098] This ensures that in a one-hot encoding, only one value is 1, and the rest are all 0. This ensures that each feature will not have multiple values at the same time.
[0099] (7) Multiple fully connected layers are defined to process input data and use the data to train the general perturbation generation model. In this embodiment, the activation function used by the general perturbation generation model is ReLU:
[0100]
[0101] The ReLU function is simple to calculate and only involves a threshold operation. This allows it to complete tasks with less time and resources during forward and backward propagation, while effectively alleviating the gradient vanishing problem.
[0102] (8) Divide the interval according to the number of values of each feature. The current sample feature x i The offset value in the dictionary space is o i , and define the allowed modification range ∈∈[0,1]. The minimum extreme value boundary of the feature that can be modified and the maximum extreme value boundary It can be calculated by the following formula:
[0103]
[0104] This ensures that the sample does not cross the boundary during the optimization process and that the feature difference between the original sample and the sample is small. For the age feature of the Adult dataset, it can be expressed as:
[0105]
[0106] Figure 4 The optimization interval division process of each feature in this example is demonstrated. First, the categorical features in the real table sample are feature encoded and then mapped to the designed dictionary space. For example, for sample 1, the age is 20 years old and the occupation is a student. After the occupation is numerically encoded, the occupation is encoded as 1. According to the characteristic extreme values of the statistical data set, it can be concluded that the minimum extreme value boundary of the age of this sample is 17 and the maximum extreme value boundary is 23; the minimum extreme value boundary of the occupation is 0 and the maximum extreme value boundary is 2. In this way, the search area can be limited during the optimization process to prevent the optimization algorithm from entering unreasonable areas. This not only ensures that the generated samples meet the requirements of the problem domain space, but also saves time and computing resources.
[0107] (9) In this embodiment, the process of applying Gumbel-Softmax in the dictionary space and generating a universal adversarial perturbation includes: In the dictionary space Θ, each row represents a feature, and each column represents a feature from 0 to V max-1(99-1). We can then use formula (8) to independently sample each feature to get the sample.
[0108] z i ~Categor i cal(π i ) (8)
[0109] where π i =Softmax(Θ i ) is the probability vector of the i-th feature.
[0110] Since sample z is highly discrete, it needs to be relaxed and differentiated. In order to solve the highly discrete nature of sample z, Gumbel-Softmax is used to approximate the gradient of the discrete distribution. For the probability vector π of the i-th feature i =softmax(Θ i ), in which Gumbel noise is added to achieve continuous relaxation, allowing the use of gradient-based methods for optimization. By optimizing the parameter matrix Θ, the sample z~P Θ Become an adversarial example for the model.
[0111] (10) Input the original data into the extended dictionary space and define the left-changeable interval value as The right interval value can be changed to Gumbel noise is added to the dictionary distribution Θ and each feature is sampled, as shown in formula (7):
[0112]
[0113] In formula (9), when the temperature coefficient τ approaches 0, the forward propagation can approximately sample a one-hot vector (indicating category selection):
[0114]
[0115] If and only if j is a set log(π i,k )+g i,k The largest index is Returns 1 if the condition is met, otherwise returns 0.
[0116] The role of formula (11) is to ensure that the correct gradient is calculated and transmitted during back propagation without changing the forward propagation result.
[0117] ret=y hard -y soft .detach()+y soft (11)
[0118] The .datach() function is used to separate the tensor from the computational graph, stopping the flow of gradients so that the gradient calculation is not affected during the back propagation process. hard The approximate one-hot vector y is soft It is obtained by the argmax operation, that is, it only retains the position with the highest probability in each sample as 1, and the rest of the positions are 0. So that in the forward propagation process, -y in formula (11) soft .detach() and y soft The values of hard Forward propagation ensures that the generated samples are discrete values that satisfy the problem domain space. At the same time, when the gradient is obtained by back propagation, since y hard It is obtained through the nonlinear argmax operation and is not differentiable, so its gradient is 0, and only y soft Participates in the back propagation. Therefore, it is ensured that the forward propagation uses discrete choices (y hard ), while back propagation uses a continuous distribution (y soft ), so that it is possible to make clear choices in the forward process and ensure the correct transfer of gradients during back propagation.
[0119] (11) Then, in the process of back propagation, in order to prevent the model from optimizing the undesirable areas in the process of optimizing matrix parameters, when the model back propagates the gradient, it also only optimizes the optimizable areas and truncates the gradients of the non-optimizable areas. This ensures that the generated adversarial samples are always within the problem domain space.
[0120] (12) Through the total loss function of formula (12), an adversarial sample is generated for each sample. The features of the adversarial sample in the dictionary space are summarized and statistically analyzed to form a universal adversarial perturbation.
[0121]
[0122] The loss function includes classification loss and L1 regularization term, and α and β are the weights of each loss term. The classification loss is shown in formula (13): It is used to measure the difference between the model prediction value and the actual label to ensure that the model can accurately distinguish different categories. The adversarial loss function is used to measure the difference between the prediction of the model M for the sample x+δ after adding the perturbation δ and the true label y. The perturbation δ that maximizes the loss function is found through the argmax function, that is, an adversarial sample that maximizes the model error rate is generated. At the same time, the constraints are met, D(x+δ,x)<τ ensures that the size of the perturbation δ is within a certain threshold τ, and at the same time ensures that the generated adversarial sample always meets certain specific constraints.
[0123]
[0124] The L1 regularization term, such as formula (14), adds the sum of the absolute values of the weights to the loss function as a penalty to prevent the features from being modified too much. x is the generated adversarial sample, x adv is the generated adversarial sample, and n is the number of features of the sample.
[0125]
[0126] (13) Generate an adversarial sample x that can maximize the model error rate for each sample x in the training set through formula (12) adv .
[0127] (14) After generating an adversarial sample for each sample, the features of the adversarial sample in the dictionary space are summarized and counted. In each round of adversarial samples generated, the encoding vector of the i-th feature of the k-th adversarial sample in the dictionary space is For each round of adversarial samples, the universal adversarial perturbation generated The update rules are:
[0128]
[0129] in It is a binary indicator function, which takes the value 1 only when the condition is met. and They represent the unique hot encoding of the kth original sample and the adversarial sample in the i-th feature dimension respectively.
[0130] (15) The new universal adversarial perturbation generated in each round is used to balance the contribution of historical perturbations and newly generated perturbations in proportion α∈(0,1), and the new perturbations are integrated into the global perturbation matrix:
[0131]
[0132] In the value range of each feature of the adversarial samples generated by the model, the more times the adversarial samples appear in a certain feature value, the greater the impact of this feature on the model's prediction.
[0133] (16) Add the universal perturbation to the dictionary space sample of the unflipped sticky note and offset it according to the characteristic value of the sample i Neighborhood Perform a maximum activation frequency search to determine the optimal perturbation position and the optimal index Calculated by:
[0134]
[0135] If there are multiple maxima, the distance o is preferred. i The nearest index. The final generated adversarial encoding:
[0136]
[0137] (17) In the subsequent training process, if the sample with added noise can mislead the model, further optimization of the sample is skipped. On the contrary, if the sample fails to successfully mislead the model, it will continue to be optimized and the overall noise will be gradually adjusted at an exponential decrease rate until the entire training process is completed, thereby obtaining the final universal adversarial perturbation.
[0138] A universal adversarial perturbation is constructed by quantifying the statistical frequency of key features. This perturbation is then applied to a new set of original samples to generate adversarial samples. The validity and rationality of the newly generated samples are ensured based on the legal value range of each feature. Finally, these carefully constructed adversarial samples are fed into the target model for testing to verify whether they can cause the model to output incorrect prediction results.
[0139] The system implementation of the present invention is the same as the specific implementation of the method.
[0140] A person skilled in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the relevant hardware through a program, and the program may be stored in a computer-readable storage medium, which may include: ROM, RAM, disk or CD, etc.
[0141] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation modes of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A universal adversarial perturbation generation method for tabular data based on Gumbel-Softmax, characterized in that: The steps include: S1: Clean and feature encode the public table dataset, and optimize the dictionary space size according to the number of possible values for each feature to ensure that the original sample features and the dictionary space features can be converted into each other. S2: Map the data features in the dataset into the dictionary space. S3: Generate an efficient perturbation generation model by training the dataset in the dictionary space through multiple rounds of iterations. S4: Divide the value range of each feature into intervals according to a predefined ratio to accurately control the feature modification range. S5: Apply Gumbel-Softmax to generate adversarial samples in the dictionary space that meet the constraints of the problem domain, and summarize the statistics to generate universal adversarial perturbations. S6: Add the universal adversarial perturbation to the samples without flipped labels, repeat step S5 and iteratively optimize the loss term until the training is completed to obtain the final universal adversarial perturbation.
2. The method for generating universal adversarial perturbations for tabular data based on Gumbel-Softmax according to claim 1, characterized in that: In S1, the dictionary space construction includes a feature extraction module and a feature encoding module. The feature extraction module extracts and cleans features from the original table data, removes outliers and default values, and ensures data quality; the feature encoding module counts the possible values of categorical data and discrete data, determines the maximum number of possible values as the size of the dictionary space, and provides a consistent numerical representation for each feature.
3. The method for generating universal adversarial perturbations for tabular data based on Gumbel-Softmax according to claim 2, characterized in that: The feature extraction module extracts and cleans features from the original table data, removes abnormal values and default values, and ensures data quality. The process includes: For public table data sets with different characteristics, specific strategies are used to handle default values and outliers, and strict data cleaning procedures are performed on data that does not conform to the predefined problem space to ensure the quality and consistency of the data set.
4. The method for generating universal adversarial perturbations for tabular data based on Gumbel-Softmax according to claim 2, characterized in that: The feature encoding module counts the possible values of categorical data and discrete data, determines the maximum possible number of values as the size of the dictionary space, and provides a consistent numerical representation for each feature. The process includes: The data that has completed the data cleaning process is feature coded and a detailed statistical analysis of the number of possible values is performed. Based on this analysis, the maximum number of possible values is determined as the standard for constructing the dictionary space size, thereby providing an accurate data representation for subsequent processing.
5. The method for generating universal adversarial perturbations for tabular data based on Gumbel-Softmax according to claim 1, characterized in that: In S2, the original eigenvalues are normalized to non-negative integer offsets relative to the minimum value to ensure the bijective relationship with the dictionary space. A unique hot encoding vector of the corresponding dimension is generated according to the offset, and all feature encodings are stacked vertically to form a dictionary feature matrix.
6. The method for generating universal adversarial perturbations for tabular data based on Gumbel-Softmax according to claim 1, characterized in that: In S3, the input layer flattens the dictionary space features into low-dimensional vectors and introduces nonlinearity through the ReLU activation function. The multi-layer fully connected network gradually compresses the dimensions and extracts abstract features, and the output layer generates prediction results.
7. The method for generating universal adversarial perturbations for tabular data based on Gumbel-Softmax according to claim 1, characterized in that: In S4, according to the dictionary space offset and modification amplitude parameter of the current eigenvalue, the minimum extreme value boundary and the maximum extreme value boundary allowed for modification are calculated to ensure that the eigenvalue does not exceed the global value range of the original data set during the optimization process.
8. The method for generating universal adversarial perturbations for tabular data based on Gumbel-Softmax according to claim 1, characterized in that: In S5, the process of generating universal adversarial perturbations includes the following steps: S5-1: Add Gumbel noise in the feature modifiable interval and set the temperature coefficient close to 0 to approximate one-hot vector sampling; S5-2: Maximize the model prediction error through the classification loss function and combine L1 regularization to limit the number of feature modifications; S5-3: During the back-propagation process, only the gradient of the optimizable area is updated, and the gradient of the non-optimizable area is truncated to ensure that the generated adversarial examples are always within the problem domain space. S5-4: Count the feature encoding frequency of each round of adversarial samples and generate a universal perturbation matrix.
9. The method for generating universal adversarial perturbations for tabular data based on Gumbel-Softmax according to claim 1, characterized in that: In S6, after adding the universal perturbation to the sample, the position with the highest activation frequency in the feature offset neighborhood is searched as the optimal perturbation. If the adversarial sample fails to mislead the model, the noise is adjusted according to the exponential decay rate and the optimization step is repeated.
10. The method for generating universal adversarial perturbations for tabular data based on Gumbel-Softmax according to claim 9, characterized in that: Adjust the noise with an exponential decay rate and repeat the optimization steps including: According to the feature encoding of the adversarial samples generated in each round, the modification frequency of each feature value is counted, and the historical perturbations and the newly generated perturbations are fused in proportion. The high-frequency modification direction dominates the final perturbation matrix.