Transverse federated learning-oriented adversarial learning-based data value protection method
By constructing a feature extraction network and an adversarial task classifier in lateral federated learning, and utilizing gradient inversion layers and adversarial learning mechanisms, the problem of model theft is solved, and the effectiveness and security of the model on specific tasks are achieved, preventing the leakage of data value.
Patent Information
- Application Number
- CN202510948895.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-28
AI Technical Summary
In existing horizontal federated learning systems, models are vulnerable to being stolen by malicious actors using transfer learning techniques for unauthorized tasks, leading to data leaks and damage to client interests. Meanwhile, traditional privacy protection methods also affect model performance.
A three-tiered protection system is adopted, which involves constructing a feature extraction network, a main task classifier, and multiple adversarial task classifiers. By utilizing gradient reversal layers and adversarial learning mechanisms, the model is optimized to maximize adversarial loss, weaken the identifiability of intermediate shared features, and construct a total loss function for training.
It effectively prevents the model from generalizing in unauthorized tasks, improves the robustness and security of the model across different tasks, and protects the value of the model and data from being misappropriated.
Smart Images

Figure CN120850340A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep learning model security technology, and in particular to a model protection method combining adversarial training, which is a data value protection method based on adversarial learning for horizontal federated learning to prevent pre-trained model parameters from being maliciously stolen for unauthorized tasks. Background Technology
[0002] With increasing awareness of data silos and privacy protection, federated learning, as a decentralized machine learning method, enables multiple clients to collaboratively train models without sharing data. In horizontal federated learning, due to the overlap of feature spaces and the heterogeneity of sample spaces, model training can help multiple clients share useful knowledge. However, this also introduces potential risks to model security.
[0003] While federated learning can protect data privacy during model training, the global model still faces the risk of being stolen by malicious actors through reverse engineering. Attackers can use transfer learning techniques to apply the shared model to unauthorized domains, thereby leaking sensitive client data. Such data breaches not only threaten privacy and security but may also harm the client's business interests.
[0004] Current horizontal federated learning systems mainly suffer from the following data value protection vulnerabilities:
[0005] 1) By fine-tuning the global model to perform cross-domain knowledge transfer, unexpected data leakage is caused, harming the interests of the client;
[0006] 2) Traditional privacy protection methods limit model utility. For example, differential privacy can significantly reduce the performance of the main task. The process of constructing adversarial tasks in TIPRDC is both time-consuming and computationally expensive, and is not suitable for filtering out the data value of all non-main tasks.
[0007] Therefore, how to protect the privacy and value of data without significantly affecting model performance has become a pressing technical challenge. Summary of the Invention
[0008] To address the aforementioned issues, the present invention proposes a data value protection method based on adversarial learning for horizontal federated learning. This method protects the security of horizontal federated learning models by constructing a three-tiered protection system.
[0009] The specific technical solution for achieving the objective of this invention is as follows:
[0010] A data value protection method based on adversarial learning for horizontal federated learning, the method comprising the following steps:
[0011] Step 1: Receive the raw input data locally on the client and input it into the feature extraction network, which is a multi-layer convolutional neural network used to extract intermediate shared feature representations;
[0012] Step 2: Input the intermediate shared feature representation into the main task classifier and multiple adversarial task classifiers respectively, wherein:
[0013] The main task classifier is used to execute the authorized main task and calculate the main task prediction loss;
[0014] The adversarial task classifier is connected to the feature extraction network through a gradient inversion layer to perform unauthorized tasks and calculate the adversarial task prediction loss.
[0015] Step 3: Backpropagate the feature extraction network and the main task classifier using the main task loss to optimize the performance of the main task;
[0016] Step 4: For each adversarial task classifier, its classification loss is backpropagated through the gradient reversal layer and applied to the feature extraction network to maximize the adversarial loss, thereby weakening the recognizability of intermediate shared features for unauthorized tasks.
[0017] Step 5: Construct the total loss function as a weighted sum of the main task loss and the losses of each adversarial task. The total loss function is as follows:
[0018]
[0019] L c Main task loss, Loss of the kth adversarial mission
[0020] Step 6: During the horizontal federated learning process, multiple clients execute steps 1-5 respectively and upload their local model parameters to the federated server. The server performs parameter aggregation to form a global model and distributes the global model to each client for continued training. Training automatically terminates after reaching the preset maximum number of rounds (100 rounds), or when the main task performance index converges and the adversarial task performance reaches the suppression target.
[0021] Furthermore, the output of the main task classifier uses the softmax activation function, and the output is the probability distribution of each category.
[0022] Furthermore, the hidden layer of the adversarial task classifier uses the LeakyReLU activation function, and the output layer uses the sigmoid activation function.
[0023] Furthermore, the loss function of the adversarial task is related to the generator parameters θ. G The gradient is based on the adversarial coefficient λ k Adjustments were made to update the parameters.
[0024] Furthermore, the federated server uses a weighted average strategy to aggregate local model parameters from the client.
[0025] Furthermore, the primary task is gender recognition, while the adversarial task is facial expression recognition, including smile detection, glasses wearing detection, or age estimation.
[0026] Furthermore, the client only performs model training and gradient calculation locally, without uploading any raw data, thus ensuring data privacy.
[0027] Features of this invention:
[0028] Dual-path adversarial architecture: The main task and adversarial task are designed in parallel, and the adversarial learning mechanism is used to prevent the model from generalizing in unauthorized tasks.
[0029] Adversarial perturbation mechanism: Limit the model's performance on unauthorized tasks to prevent knowledge transfer attacks.
[0030] Feature Confusion Network: Employs an adversarial learning architecture based on gradient inversion to desensitize intermediate features of the model, preventing the model from leaking sensitive information.
[0031] Beneficial effects of the present invention
[0032] 1. Preventing model theft: Since the model is only effective on features relevant to specific tasks, even if an attacker steals the model, its performance on other tasks will be greatly reduced, thus preventing the model's value from being stolen.
[0033] 2. Enhanced model security: Through adversarial loss, the model actively reduces its dependence on irrelevant features during training, thereby improving the model's robustness and security across different tasks.
[0034] 3. Task specificity: The model's performance on a specific task will not be affected by changes in the external environment or tampering by attackers. Attached Figure Description
[0035] Figure 1 This is a flowchart of the present invention. Detailed Implementation
[0036] First, in the application context of this invention, the main task is set as gender recognition. This task is a client-authorized target task, and its prediction results are used in practical application scenarios, such as demographic analysis or user profile modeling. Meanwhile, the training data may contain information that can be used to infer other sensitive attributes, such as facial expressions (smiling), eyeglass wearing status, or age. Although these unauthorized tasks can be learned from shared intermediate features, they are privacy-related or sensitive business attributes and could be misused by attackers using transfer learning after the model is stolen, leading to the leakage of data value.
[0037] See Figure 1 The network consists of a feature extraction network (red), a main task classifier (purple), and multiple adversarial task classifiers (green). The adversarial task classifiers are connected to the feature extraction network via a gradient inversion layer. This gradient inversion layer, during backpropagation-based training, multiplies the gradient by a negative constant, thus reversing its direction. Otherwise, the training process is the same as the standard approach, minimizing the prediction loss of the main task classifier. The gradient inversion mechanism ensures that the adversarial task classifiers cannot correctly identify their corresponding tasks.
[0038] This invention provides a data value protection method based on adversarial learning for horizontal federated learning, which combines a deep model training structure and adversarial mechanisms, and is implemented through the following steps:
[0039] Step 1: Local processing and feature extraction of input data
[0040] The client receives local raw data samples and inputs them into the feature extraction network. The feature extraction network employs a multi-layer convolutional neural network structure, including several convolutional layers, pooling layers, and fully connected layers, to map the raw data into an intermediate shared feature representation. This feature representation will serve as the common input basis for both the main task and the adversarial task.
[0041] The feature extraction network selects the ReLU activation function, z l =ReLU(W l z l-1 +b l ), l=1,2,3,4,z l It is the output of the l-th layer, W l It is the weight matrix, b l It is a bias term.
[0042] Feature layers are selected using ReLU activation, h c =ReLU(W c1 z+b c1 z is the input passed from the previous layer, which is then activated by ReLU to obtain h. c As the output of the feature layer.
[0043] Step 2: Construct a parallel task path
[0044] The intermediate shared features are then input into the main task classifier and multiple adversarial task classifiers, respectively.
[0045] The main task classifier is used to predict the core task label authorized by the user (such as gender recognition). Its structure includes one or more fully connected layers with softmax as the activation function.
[0046] p = softmax(W c2 h c +b c2 ), h c It is the output of the feature layer, which, after linear transformation, is passed through the softmax function to obtain the predicted probability of each category.
[0047] Each adversarial task classifier is used for unauthorized tasks (such as smile recognition and glasses recognition) and consists of several hidden layers using the LeakyReLU activation function. The output layer uses the Sigmoid activation function. This represents the predicted probability for each adversarial task.
[0048] Each adversarial classifier is connected to the feature extractor via a gradient inversion layer (GRL) to generate adversarial gradients on the features during backpropagation.
[0049] Step 3: Calculate the losses of the main task and the adversarial task
[0050] The main task uses the cross-entropy loss function to calculate the difference between the model's predicted output and the true label, with the goal of minimizing this loss to improve task performance.
[0051]
[0052] y i c is the true label of the i-th sample, p i c is the predicted probability of the model for that category.
[0053] Each adversarial task also uses the cross-entropy loss function to measure its ability to discriminate intermediate features. For unauthorized tasks, each classifier needs to minimize its loss function.
[0054]
[0055] However, to counteract this, the feature extraction network needs to maximize the loss of these classifiers so that they cannot make correct judgments, which is achieved through gradient inversion layers (GRL).
[0056] The gradient reversal layer needs to be applied to the loss of each adversarial classifier so that during backpropagation, the feature extractor will be trained to maximize the loss of the adversarial classifier, i.e., to make the features indistinguishable from unauthorized tasks.
[0057]
[0058] The loss function for adversarial tasks depends on the generator parameters θ G The gradient is based on the adversarial coefficient λ k Adjustments were made to update the parameters.
[0059] Step 4: Construct a composite loss function
[0060] The main task loss and all adversarial task losses are combined in a weighted manner to form the total loss function:
[0061]
[0062] Step 5: Federated Aggregation and Global Model Update
[0063] After each client completes the above training locally, it only uploads the local model parameters (without uploading the original data). The federated server then performs weighted aggregation to generate a global model, which is distributed to each client for the next round of local training, thus realizing the iteration of horizontal federated learning.
[0064] The training process terminates when any of the following conditions are met:
[0065] (1) Reach the preset maximum number of iterations; set the maximum number of iterations (e.g., 100 iterations), and the training will automatically terminate after reaching this number of iterations.
[0066] (2) Performance convergence of the main task validation set: Based on the performance indicators of the main task on the validation set (such as accuracy and loss), when the change is less than a certain set threshold for several consecutive rounds (such as loss change <0.001), it is considered to have converged and training is stopped.
[0067] (3) The discrimination performance of the unauthorized task is lower than the set threshold; if the discrimination performance of the unauthorized task is always lower than a certain preset threshold (e.g., AUC<0.55), it is considered as convergence and training stops.
[0068] (4) The overall loss function converges or stabilizes. The process stops when the overall loss function reaches stability or optimality, combining the objective functions of the main task and the adversarial task.
[0069] Through the above steps, this invention significantly reduces the performance of the model on unauthorized tasks while ensuring the performance of the main task, thereby protecting the value and security of the model and the data it contains.
[0070] This invention combines privacy protection mechanisms with adversarial training to construct a secure model with targeted capability suppression properties, effectively preventing model parameters from being maliciously used for tasks other than their intended purpose. Examples demonstrate that, while maintaining performance on the primary task, the model's cross-task transferability is significantly reduced. This is of great significance for protecting the proprietary value of machine learning models.
Claims
1. A data value protection method based on adversarial learning for horizontal federated learning, characterized in that, The method includes the following steps: Step 1: Receive the raw input data locally on the client and input it into the feature extraction network, which is a multi-layer convolutional neural network used to extract intermediate shared feature representations; Step 2: Input the intermediate shared feature representation into the main task classifier and multiple adversarial task classifiers respectively, wherein: the main task classifier is used to perform the authorized main task and calculate the main task prediction loss; The adversarial task classifier is connected to the feature extraction network through a gradient inversion layer to perform unauthorized tasks and calculate the adversarial task prediction loss. Step 3: Backpropagate the feature extraction network and the main task classifier using the main task loss to optimize the performance of the main task; Step 4: For each adversarial task classifier, its classification loss is backpropagated through the gradient reversal layer and applied to the feature extraction network to maximize the adversarial loss, thereby weakening the recognizability of intermediate shared features for unauthorized tasks. Step 5: Construct the total loss function as a weighted sum of the main task loss and the losses of each adversarial task. The total loss function is as follows: L c Main task loss, Loss of the kth adversarial mission; Step 6: During the horizontal federated learning process, multiple clients execute steps 1-5 respectively and upload their local model parameters to the federated server. The server performs parameter aggregation to form a global model and distributes the global model to each client for continued training. Training automatically terminates after reaching the preset maximum number of rounds, or when the main task performance index converges and the adversarial task performance reaches the suppression target.
2. The data value protection method according to claim 1, characterized in that, The output of the main task classifier uses the softmax activation function, and the output is the probability distribution of each category.
3. The data value protection method according to claim 1, characterized in that, The hidden layer of the adversarial task classifier uses the LeakyReLU activation function, and the output layer uses the sigmoid activation function.
4. The data value protection method according to claim 1, characterized in that, The loss function of the adversarial task depends on the generator parameters θ. G The gradient is based on the adversarial coefficient λ k Adjustments were made to update the parameters.
5. The data value protection method according to claim 1, characterized in that, The federated server uses a weighted average strategy to aggregate local model parameters from clients.
6. The data value protection method according to claim 1, characterized in that, The primary task is gender recognition, while the adversarial task is facial expression recognition, including smile detection, glasses wearing detection, or age estimation.
7. The data value protection method according to claim 1, characterized in that, The client only performs model training and gradient calculation locally, without uploading any raw data, thus ensuring data privacy.
Citation Information
Cited By
Generalized time-frequency positioning method for distributed training scene
CN121692206A