A backdoor robust image classification method based on federated learning

By computing malicious client scores and geometric center-weighted aggregation in federated learning, combined with perturbation pruning, the vulnerability of federated learning image classification models to backdoor attacks is solved, achieving high-precision and robust image classification.

CN119229208BActive Publication Date: 2025-12-26HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411413543.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-11
Publication Date
2025-12-26
Estimated Expiration
2044-10-11

AI Technical Summary

Technical Problem

Federated learning is vulnerable to backdoor attacks in image classification tasks, posing a threat to model security, and existing technologies struggle to enhance robustness while maintaining model classification performance.

Method used

By calculating client-side malicious scores and geometric center weighted aggregation, combined with perturbation cropping, the robustness of the image recognition model is enhanced. A backdoor robust image classification method using federated learning is adopted. The central server calculates the similarity and malicious scores of the client models, performs weighted geometric center aggregation, and then crops the global model after aggregation to resist backdoor attacks.

Benefits of technology

This improves the robustness of the image recognition model against backdoor attacks, ensures high-precision image classification performance, and enhances the model's security and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229208B_ABST
    Figure CN119229208B_ABST
Patent Text Reader

Abstract

The application discloses a backdoor robust image classification method based on federated learning, comprising the following steps: 1, a server initializes a global model and distributes it to a client; 2, the client initializes a local model based on the received global model, and trains a local neural network based on a local data set; 3, the client calculates a local model update and sends a local update vector to the server; 4, the server records the update vector submitted by the client; 5, the server aggregates the model according to an improved federated aggregation algorithm; 6, the server distributes the trained global neural network to each client; 7, the client classifies images by using the trained global neural network. The application realizes cooperative image classification with stronger robustness against backdoor attacks by using federated learning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of anomaly detection, in particular to a backdoor robust image classification method based on federated learning. BACKGROUND

[0002] Image classification is a key task of computer vision, which aims to assign a class label from a predefined set to each image. Although humans can easily recognize the content of an image, machines must rely on complex algorithms to achieve similar results. In recent years, although deep learning has developed in the field of image classification, as the amount of data grows and the awareness of personal privacy protection strengthens, it becomes increasingly challenging to obtain and centrally train the required high-quality data.

[0003] Federated learning (FL) is an emerging distributed machine learning framework that collaboratively trains a global model without sharing raw data by allowing devices to train the model on their local data and then uploading only the model updates to a central server. However, federated learning faces challenges in model security when applied to downstream tasks such as image classification. Due to its distributed nature, federated learning is vulnerable to backdoor attacks, which can inject adversarial triggers into the model to make it perform malicious tasks without affecting the main task performance, thus posing a threat to system security. Therefore, how to enhance the robustness of federated learning models to such backdoor attacks in image classification problems still needs to be explored, especially while ensuring the classification performance of the model. SUMMARY

[0004] The present application proposes a backdoor robust image classification method based on federated learning to enhance the robustness of federated learning image recognition models when applied to image classification task scenarios, thereby achieving high-precision image classification.

[0005] To achieve the above-mentioned application purposes, the following technical solutions are adopted:

[0006] The backdoor robust image classification method based on federated learning is characterized by being applied to a federated system composed of a central server and a client, and the following steps are performed:

[0007] Step 1. Define the total number of global training rounds of the federated system as and the total number of local training rounds of the client as , define of the clients as malicious clients, and initialize the current global training round number , the center server initializes a global image recognition model in the t-1th global training round and sends it to each client;

[0008] Step 2. The t-1th global training round is completed, and the tth global training round is started. Step 3. The t-1th global training round is completed, and the tth global training round is started. Step 4. The t-1th global training round is completed, and the tth global training round is started. Step 5. The t-1th global training round is completed, and the tth global training round is started.

[0009] Step 3. The t-1th global training round is completed, and the tth global training round is started. Step 4. The t-1th global training round is completed, and the tth global training round is started. , wherein represents the jth image sample in , represents the true label of , represents the total number of image samples in the local data set of the tth client

[0010] Step 4. The t-1th global training round is completed, and the tth global training round is started. Step 5. The t-1th global training round is completed, and the tth global training round is started. Step 6. The t-1th global training round is completed, and the tth global training round is started. Step 7. The t-1th global training round is completed, and the tth global training round is started.

[0011] Step 5. The t-1th global training round is completed, and the tth global training round is started. Step 6. The t-1th global training round is completed, and the tth global training round is started. Step 7. The t-1th global training round is completed, and the tth global training round is started.

[0012] Step 6. The t-1th global training round is completed, and the tth global training round is started. Step 7. The t-1th global training round is completed, and the tth global training round is started.

[0013] Step 7. The t-1th global training round is completed, and the tth global training round is started. Step 8. The t-1th global training round is completed, and the tth global training round is started. Step 9. The t-1th global training round is completed, and the tth global training round is started. ​​​​​​The training proceeds normally for one round, then returns to step 4 and executes sequentially until t > T, thus obtaining the final trained global image recognition model. This is used to classify and recognize input image samples; among them, This represents the global training rounds for R malicious clients to launch backdoor attacks;

[0014] Step 8. A malicious client from The random selection ratio is The image sample is used to insert a black pixel block at a specified coordinate as a backdoor trigger. Indicates the first The number of local image samples selected by each malicious client;

[0015] Step 9. Initialize the first The malicious client was in the first Number of local training rounds under each global training round =1;

[0016] Initialize the first In the nth global training round The first local training round A malicious client image recognition model = ;

[0017] Step 10. Place the first The malicious client was in the first Local dataset for each global training round enter Training and calculation Local gradient Thus, the first In the nth global training round The first local training round A malicious client image recognition model ;

[0018] Step 11. Assign to Then, return to step 10 and execute sequentially until... > Until then, thus obtaining the first After training in the nth global training round, the th... A malicious client image recognition model and calculate global gradient Then, store in the first... The th global training round A set of historical global gradients of a malicious client image recognition model In the center, and sent to the center server, step 5 is executed; wherein, The total number of local training rounds of the malicious client.

[0019] The characteristics of the backdoor robust image classification method based on federated learning also lie in that the step 4 comprises:

[0020] Step 4.1. Initialize the local training round number of the i-th client in the t-th global training round =1; initialize the i-th local training round of the i-th client in the t-th global training round =1; initialize the i-th local training round of the i-th client in the t-th global training round =1; initialize the i-th local training round of the i-th client in the t-th global training round =1; initialize the i-th local training round of the i-th client in the t-th global training round = ;

[0021] Step 4.2. input into for training, and calculate the local gradient of , so as to obtain the i-th local training round of the i-th client image recognition model in the t-th global training round , wherein, denotes the learning rate of the i-th client model; Step 4.3. assign to , then return to step 4.2 for sequential execution until

[0022] . , so as to obtain the i-th local training round of the i-th client image recognition model in the t-th global training round , and calculate the global gradient of and store it in the set of historical global gradients of the i-th client image recognition model in the t-th global training round , and then send it to the center server. Further, the step 5 comprises: Step 5.1. calculate the similarity between the i-th client image recognition model after training in the t-th global training round and the i-th client image recognition model after training

[0023] using formula (1)

[0024] Step 5.2. if the similarity is greater than a preset threshold, the i-th client image recognition model after training in the t-th global training round is determined as a malicious client image recognition model. ​​​​​ , thereby obtaining the similarity between the i-th client image recognition model and all other client image recognition models:

[0025] (1)

[0026] In formula (1), is a non-zero value taken to avoid the denominator being zero; denotes the historical global gradient set of the i-th client image recognition model in the t-th global training round;

[0027] Step 5.2. Select the maximum value from the similarity between the i-th client image recognition model and all other client image recognition models as the malicious score of the i-th client image recognition model after training in the t-th global training round. ;

[0028] Step 5.3. When , the updated similarity is obtained using formula (2), and the similarity between remains unchanged:

[0029] (2)

[0030] Step 5.4. According to the process of step 5.3, the malicious score of the i-th client image recognition model is respectively compared with the malicious scores of all other client image recognition models, and is updated, thereby obtaining the updated similarity between the i-th client image recognition model and all other client image recognition models.

[0031] Step 5.5. Select the maximum value from the updated similarity between the i-th client image recognition model and all other client image recognition models as the updated malicious score of the i-th client image recognition model after training in the t-th global training round. ;

[0032] Step 5.6. Calculate the weight assigned to the i-th client image recognition model in the t-th global training round using formula (3):

[0033] (3) ​​​​​​​​​​​​​

[0034] In formula (3), represents the i-th client image recognition model after training in the t-th global training round updated malicious score

[0035] Step 5.7. Calculate the i-th client image recognition model after the t-th global training round updated weight :

[0036] (4)

[0037] Step 5.8. After normalizing , the i-th client image recognition model after the t-th global training round final weight .

[0038] Further, the step 6 comprises:

[0039] Step 6.1. Define the current aggregation iteration number as , and initialize = 1;

[0040] Step 6.2. The center server calculates the weight of the i-th client image recognition model in the t-th global training round th aggregation iteration using formula (6) and formula (7): :

[0041] (6)

[0042] (7)

[0043] In formula (6) and formula (7), is a smoothing parameter; is the geometric center in the t-th global training round th aggregation iteration, is the geometric center in the t-th global training round -1th aggregation iteration, when = 1, the formula (8) is used to obtain :

[0044]

[0045] Step 6.3. Determine whether formula (9) is true, if true, it means that the i-th client image recognition model after the t-th global training round ​​​and the optimal geometric center under the tth global training round , otherwise, assign +1 to After that, return to step 6.2 for sequential execution.

[0046] (9)

[0047] In formula (9), is a preset iteration threshold, denotes the objective function of the geometric center iteration optimization, and has:

[0048] (10)

[0049] Step 6.4. Obtain the preliminary global image recognition model under the tth global training round by using formula (11) :

[0050] (11)

[0051] Step 6.5. Update by using formula (12) to obtain the global image recognition model under the tth global training round :

[0052] (12)

[0053] In formula (12), is a clipping threshold under the tth global training round, is noise subject to a Gaussian distribution under the tth global training round.

[0054] The electronic device comprises a memory and a processor, and the memory is used to store a program supporting the processor to execute the backdoor robust image classification method, and the processor is configured to execute the program stored in the memory.

[0055] The computer readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the backdoor robust image classification method are executed.

[0056] Compared with the prior art, the present application has the following advantages:

[0057] 1. This invention proposes a novel weighted parameter clustering anomaly detection method in federated learning. This method enhances the robustness of federated learning against backdoor attacks in image recognition tasks by calculating malicious client scores before aggregation, calculating the updated geometric center of the client's image recognition model instead of simple weighted aggregation, and perturbing and pruning the global image recognition model after aggregation, thus ensuring high-precision image classification.

[0058] 2. This invention introduces a weighted geometric median estimation algorithm, which considers the differences in importance of image recognition models of each client when evaluating the geometric median of the model update parameters. Compared with traditional methods, the aggregation results obtained by this algorithm are more robust, and while ensuring the model classification accuracy, it effectively improves the ability of the image recognition model to resist malicious interference. Attached Figure Description

[0059] Figure 1 This is a flowchart illustrating the present invention. Detailed Implementation

[0060] In this embodiment, a backdoor robust image classification method based on federated learning is described, such as... Figure 1 As shown, it is applied to a central server and In a federated system consisting of individual clients, the following steps are followed:

[0061] Step 1. Define the total number of global training rounds for the federated system as... The total number of local training rounds with the client is ,definition In each client This client is a malicious client; initialize the current global training rounds. The central server initializes the global image recognition model for the (t-1)th global training round. The data is then distributed to each client. This embodiment uses the MNIST dataset for training and evaluation. The MNIST dataset consists of 70,000 grayscale images across 10 categories, with each category containing 6,000 training examples and 1,000 test examples. In this embodiment, 10,000 images are randomly selected from the 60,000 training examples as public data, and the remaining 50,000 images are used as local data on the client. A multi-class logistic regression model is used, consisting of a linear layer with a softmax function, and the loss is calculated using the cross-entropy function. In this embodiment, T is set to 100. Let 30 be the number of clients, N be 20, and R be 4.

[0062] Step 2. Each client receives a global image recognition model. After that, initialize the (t-1)th global training round. A client image recognition model ;

[0063] Step 3. Define the first The local dataset of each client in the t-th global training round is: ,in, express The j-th image sample in the image, express The true label, Indicates the first Local datasets for each client The total number of image samples in the dataset;

[0064] Step 4. enter Perform local training to obtain the result of the (t+1)th global training round. A client image recognition model ;

[0065] Step 4.1. Initialize the first The number of local training rounds for each client in the t-th global training round =1; initialize the t-th global training round. The first local training round A client image recognition model = ;

[0066] Step 4.2. enter Training is performed in the process, and calculations are performed. Local gradient Thus, the t-th global training round is obtained. The first local training round A client image recognition model ,in, Indicates the first The learning rate of each client model; in this embodiment, the learning rate Take 0.001.

[0067] Step 4.3. Assign to Then, return to step 4.2 and execute sequentially until... > So far, we can obtain the training result after the t-th global training round. A client image recognition model and calculate global gradient And store it in the t-th global training round. Historical global gradient set of a client image recognition model After processing, it is sent to the central server;

[0068] Step 5. After receiving the historical global gradient sets of the client models trained in the t-th global training round from N clients, the central server calculates the similarity and malicious score between the global gradients of each client model to obtain... weight ;

[0069] Step 5.1. Calculate the training result after the t-th global training round using equation (1). A client image recognition model Compared with the training day A client image recognition model similarity between Thus obtain Similarity between the model and all other client-side image recognition models:

[0070] (1)

[0071] In equation (1), This is to avoid using a non-zero value in the denominator; Represents the t-th global training round. The historical global gradient set of each client image recognition model; in this embodiment; Pick .

[0072] Step 5.2. From The maximum similarity among all other client image recognition models is selected as the result of training in the t-th global training round. A client image recognition model Malicious ratings ;

[0073] Step 5.3. When Then, the updated similarity is obtained using equation (2). and keep and similarity between constant:

[0074] (2)

[0075] Step 5.8. After normalization, we obtain the th global training round t. client image recognition model final weight ;

[0076] Step 6. The center server calculates the weighted geometric center of the global gradient of each client image recognition model in the tth global training round based on the weight of each client image recognition model, and uses it to aggregate each client image recognition model in the tth global training round to obtain the global image recognition model in the tth global training round ;

[0077] Step 6.1. Define the current aggregation iteration number as , and initialize = 1;

[0078] Step 6.2. The center server calculates the weight of the i th client image recognition model in the j th aggregation iteration in the t th global training round using formula (6) and formula (7) :

[0079] (6)

[0080] (7)

[0081] In formula (6) and formula (7), is a smoothing parameter; is the geometric center in the j th aggregation iteration in the t th global training round, is the geometric center in the j-1 th aggregation iteration in the t th global training round, and when = 1, is obtained using formula (8) : In this embodiment, takes .

[0082]

[0083] Step 6.3. Determine whether formula (9) is established, if it is established, it means that the optimal geometric center is obtained and is taken as the optimal geometric center in the t th global training round, otherwise, assign + 1 to , then return to step 6.2 for sequential execution; in this embodiment, takes .

[0084] (9) ​​​

[0085] In equation (9), The preset iteration threshold, Let represent the objective function for iterative optimization of the geometric center, and we have:

[0086] (10)

[0087] Step 6.4. Use equation (11) to obtain the preliminary global image recognition model under the t-th global training round. :

[0088] (11)

[0089] Step 6.5. Use equation (12) to... The model is updated to obtain the global image recognition model in the t-th global training round. :

[0090] (12)

[0091] In equation (12), Let be the pruning threshold in the t-th global training round. The noise follows a Gaussian distribution during the t-th global training round; in this embodiment, Take 0.1t+2 (t is the current global iteration), Pick .

[0092] Step 7. After assigning t+1 to t, determine... If the condition is met, it means that R malicious clients launched an attack in this round, and the malicious clients proceed to step 8, while the remaining NR honest clients return to step 4 and execute sequentially. Otherwise, it means that all clients performed training normally in this round, and return to step 4 and execute sequentially until t>T, thus obtaining the final trained global image recognition model. This is used to classify and recognize input image samples; among them, This represents the global training rounds for R malicious clients to launch backdoor attacks; in this embodiment, Take 10.

[0093] Step 8. A malicious client from The random selection ratio is The image sample is used to insert a black pixel block at a specified coordinate as a backdoor trigger. Indicates the first The number of local image samples selected by each malicious client. In this embodiment, The coordinates of the malicious client-injected pixel block are set to 0.05, and the coordinates are [23,25], [24,24], [25,23], [25,25].

[0094] Step 9. Initialize the first The malicious client was in the first Number of local training rounds under each global training round =1;

[0095] Initialize the first In the nth global training round The first local training round A malicious client image recognition model = ;

[0096] Step 10. Place the first The malicious client was in the first Local dataset for each global training round enter Training and calculation Local gradient Thus, the first In the nth global training round The first local training round A malicious client image recognition model ;

[0097] Step 11. Assign to Then, return to step 10 and execute sequentially until... > Until then, thus obtaining the first After training in the nth global training round, the th... A malicious client image recognition model and calculate global gradient Then, store in the first... The th global training round Historical global gradient set of a malicious client image recognition model In the middle, and sent to the central server, to execute step 5; where, This represents the total number of local training rounds for the malicious client; in this embodiment, Take 30.

[0098] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.

[0099] In this embodiment, a computer readable storage medium stores a computer program, and the computer program is run by a processor to execute the steps of the above method.

Claims

1. A backdoor robust image classification method based on federated learning, characterized in that, Applicable to a federated system composed of a central server and a client, and proceeds as follows: Step 1. Define the total number of global training rounds of the federal system as and the total number of local training rounds of the client as , define of the clients as malicious clients, initialize the current global training round number , the center server initializes the global image recognition model under the t-1th global training round and distributes it to each client; Step 2. The tth client receives the global image recognition model after the (t-1)th global training round , initializes the tth client image recognition model based on the global image recognition model ; Step 3. Define the local data set of the t-th global training round for the j-th client as where, denotes the j-th image sample in denotes the true label of denotes the total number of image samples in the local data set of the j-th client ​​​​​ Step 4. enter Perform local training to obtain the result of the (t+1)th global training round. A client image recognition model ; Step 5. After the center server receives the historical global gradient set of the trained client model of the t-th global training round sent by N clients, the similarity between the global gradients of each client model and the malicious score are calculated to obtain the weight of ; ; Step 6. The center server calculates the weighted geometric center of the global gradient of each client image recognition model in the t-th global training round based on the weight of each client image recognition model, and uses it to aggregate each client image recognition model in the t-th global training round to obtain a global image recognition model in the t-th global training round ; Step 6.

1. Define the current aggregation iteration number as and initialize = 1; Step 6.

2. The center server calculates the weight of the i-th client image recognition model in the t-th global training round using formula (6) and formula (7) the weight of the i-th client image recognition model in the t-th global training round : (6) (7) In formula (6) and formula (7), is a smoothing parameter; is the geometric center in the t-th global training round under the -th aggregation iteration, is the geometric center in the t-th global training round under the -1-th aggregation iteration, when = 1, formula (8) is used to obtain : (8) Step 6.

3. Determine whether formula (9) is established, if yes, it indicates that the optimal geometric center under the t-th global training round is obtained and as the optimal geometric center under the t-th global training round , otherwise, assign +1 to After that, return to step 6.2 for sequential execution; (9) In formula (9), is a preset iteration threshold value, denotes a target function of geometric center iteration optimization, and has: (10) Step 6.

4. Obtain a preliminary global image recognition model at the t-th global training round using formula (11) : (11) Step 6.

5. Update the global image recognition model under the t-th global training round by using formula (12) : (12) : (12) In formula (12), is a clipping threshold under the t-th global training round, is noise subject to a Gaussian distribution under the t-th global training round; Step 7. After assigning t+1 to t, determine... Does this hold true? If so, it means that R malicious clients are in the [number]th [phase]. The attack proceeds in round 1, and step 8 is executed. The remaining NR honest clients return to step 4 and execute sequentially. Otherwise, it indicates that all clients are in the NR round. The training proceeds normally for one round, then returns to step 4 and executes sequentially until t > T, thus obtaining the final trained global image recognition model. This is used to classify and recognize input image samples; among them, This represents the global training rounds for R malicious clients to launch backdoor attacks; Step 8. The malicious client selects a random proportion of picture samples and implants a black pixel block at the specified coordinates of the picture sample as a backdoor trigger; represents the number of local picture samples selected by the malicious client.​​ Step 9. Initialize the number of local training epochs for the malicious client at the first global training epoch = 1; Step 10. For each global training epoch, do the following: Initialize the first global training round Initialize the first local training round of the first malicious client image recognition model Initialize the first malicious client image recognition model Initialize the first malicious client image recognition model Initialize the first malicious client image recognition model ; Step 10. Train the first malicious client in the local dataset of the first global training round in the input of the first local training round of the first global training round, calculate the local gradient of the first malicious client image recognition model, thereby obtaining the first malicious client image recognition model of the first local training round of the first global training round. ;​​​​​​​​​ Step 11. Assign to Then, return to step 10 and execute sequentially until... > Until then, thus obtaining the first After training in the nth global training round, the th... A malicious client image recognition model and calculate global gradient Then, store in the first... The th global training round Historical global gradient set of a malicious client image recognition model In, and send it to the central server to execute step 5; wherein, This represents the total number of local training rounds for the malicious client.

2. The backdoor-robust image classification method based on federated learning according to claim 1, characterized in that, The step 4 comprises: Step 4.

1. Initialize the local training round number of the tth global training round of the i th client = 1; Initialize the i th local training round of the tth global training round of the i th client image recognition model ;​​​​ Step 4.

2. enter Training is performed in the process, and calculations are performed. Local gradient Thus, the t-th global training round is obtained. The first local training round A client image recognition model ,in, Indicates the first The learning rate of each client-side model; Step 4.

3. Assign to Then, return to step 4.2 and execute sequentially until... > So far, we can obtain the training result after the t-th global training round. A client image recognition model and calculate global gradient And store it in the t-th global training round. Historical global gradient set of a client image recognition model After that, it is sent to the central server.

3. The backdoor-robust image classification method based on federated learning according to claim 2, characterized in that, The step 5 comprises: Step 5.

1. Calculate the similarity between the trained t-th global image recognition model and the trained j-th client image recognition model using formula (1) ​​​​​​ (1) In equation (1), It is a non-zero value chosen to avoid the denominator being zero; Represents the t-th global training round. The historical global gradient set of each client-side image recognition model; Step 5.

2. From the maximum value among the similarities with all other client image recognition models, the malicious score of the trained t-th global round trained t-th global round of the client image recognition model ; Step 5.

3. When the updated similarity is obtained using formula (2) and the similarity between and is kept constant: ​ (2) Step 5.

4. Follow the procedure in Step 5.

3. Malicious ratings The system is evaluated and updated based on the malicious scores of all other client-side image recognition models to obtain the results. Updated similarity between the image recognition models of all other clients; Step 5.

5. From the maximum value of the updated similarities between the trained image recognition model and all other client image recognition models is selected as the trained image recognition model of the t-th global training round the updated malicious score ;​ Step 5.

6. Compute the weight assigned to the t-th global training round under the i-th client image recognition model using formula (3) :​​ (3) In equation (3), This represents the training result after the t-th global training round. A client image recognition model Updated malicious ratings; Step 5.

7. Calculate the t-th global training round's 1-th client image recognition model using formula (4) updated weight updated weight : (4) Step 5.

8. Normalizing the global training round t-th client image recognition model After normalization, the final weight of the t-th global training round t-th client image recognition model is obtained . ​ 4. An electronic device comprising a memory and a processor, characterized in that The memory is configured to store a program supporting the processor to execute the backdoor-robust image classification method of any one of claims 1-3.

5. A computer-readable storage medium having stored thereon a computer program, characterized in that The computer program, when executed by the processor, performs the steps of the backdoor-robust image classification method of any one of claims 1-3.

Citation Information

Patent Citations

  • Robustness federated learning model aggregation method based on truth value discovery

    CN114186237A

  • Federal learning data poisoning defense method based on data enhancement

    CN117875455A