A federated learning attack and defense security test method based on multi-agent cooperation
Patent Information
- Application Number
- CN202610701005.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-09-15
AI Technical Summary
[0005]鉴于现有技术的上述缺点、不足,本发明提供一种多智能体协同的联邦学习攻防安全测试方法,其解决了测试流程自动化程度低的技术问题
[0023] This application provides a multi-agent collaborative federated learning attack and defense security testing method. The initial test task is automatically broken down and distributed to various functional agents by a master orchestration agent. Attack and defense strategy agents independently generate local policy patches. After automatic conflict arbitration by the master orchestration agent, an executable test plan is generated. Finally, the execution evaluation agent uniformly schedules federated training experiments and automatically collects metrics to calculate a comprehensive score. This architecture completely establishes an automated link between attack configuration, defense deployment, training execution, and result analysis, replacing the existing testing mode that requires manually writing scripts to connect each stage. This significantly improves testing efficiency and reproducibility, enabling large-scale, multi-round dynamic attack and defense game testing. It achieves a fully automated closed-loop process for federated learning security testing, significantly reducing the degree of human intervention.
Smart Images

Figure CN122764545A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a multi-agent collaborative federated learning attack and defense security testing method. Background Technology
[0002] With the widespread application of federated learning in privacy-preserving joint modeling scenarios, the security threats it faces, such as poisoning attacks and backdoor attacks, are becoming increasingly prominent. To evaluate the security defense capabilities of federated learning systems, academia and industry have proposed various security testing schemes and verification frameworks. For example, underlying federated learning frameworks such as Flower, FedML, and NVFlare all provide basic security testing support.
[0003] However, existing federated learning security testing solutions suffer from low automation in complex attack and defense scenarios. Specifically, the configuration of attack parameters, the deployment of defense strategies, the execution of federated training tasks, and the collection and analysis of test results are all handled by different scripts, configuration files, or command-line terminals. This requires testers to manually write numerous scripts for connection and intervention. This manual testing model is not only inefficient but also struggles to support large-scale, multi-round dynamic attack and defense game testing, severely hindering comprehensive and in-depth verification of the security of federated learning systems. Summary of the Invention
[0004] (a) Technical problems to be solved
[0005] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a multi-agent collaborative federated learning attack and defense security testing method, which solves the technical problem of low automation in the testing process.
[0006] (II) Technical Solution
[0007] To achieve the above objectives, the main technical solutions adopted by the present invention include:
[0008] In a first aspect, embodiments of the present invention provide a multi-agent collaborative federated learning attack and defense security testing method, comprising: receiving an initial test task through a master orchestration agent and decomposing it into multiple subtasks, including at least attack exploration subtasks and defense exploration subtasks; generating local attack strategy patches in the candidate attack space through an attack strategy agent in response to the attack exploration subtasks, and generating local defense strategy patches in the candidate defense space through a defense strategy agent in response to the defense exploration subtasks; performing conflict detection and arbitration on the attack strategy patches and defense strategy patches through the master orchestration agent to merge them into an executable test plan for the current round; and executing an evaluation agent. The system schedules training experiments in a federated learning environment based on an executable test plan to collect multi-dimensional evaluation metrics and calculate a comprehensive score. A memory knowledge module stores the attack strategy patch, defense strategy patch, executable test plan, multi-dimensional evaluation metrics, and comprehensive score of the current round as a test memory tuple, generating an experience feedback vector. The master orchestration agent determines whether the preset iteration termination condition is met based on the comprehensive score and experience feedback vector. If not, the attack strategy agent and defense strategy agent proceed to the next iteration. If the condition is met, the iteration ends, and the security test result is determined based on the test memory tuples stored from previous iterations.
[0009] Optionally, the attack strategy patch includes attack type, attack scope, attack parameter strength, and attack timing arrangement; wherein, the attack type includes at least one of data poisoning, tag flipping, model poisoning, backdoor trigger injection, and client collusion attack.
[0010] Optionally, the defense strategy patch includes robust aggregation rules, abnormal client filtering rules, and defense sensitivity parameters; wherein, the robust aggregation rules include at least one of the following: Krum algorithm, multi-Krum algorithm, coordinate median algorithm, truncated mean algorithm, or weighted average mechanism based on confidence weight.
[0011] Optionally, the main control orchestration agent performs conflict detection and arbitration on the attack strategy patch and the defense strategy patch to merge and generate an executable test plan for the current round. This includes: adopting domain isolation priority, determining the robust aggregation rules in the defense strategy patch for the server-side aggregation logic, and determining the attack parameter strength in the attack strategy patch for the malicious client's local training logic; when the expected values proposed by the attack strategy patch and the defense strategy patch for the same global hyperparameter are inconsistent, calculating the weighted average of the two as the final value of the global hyperparameter, or selecting the expected value of one of them as the final value of the global hyperparameter according to the test objective in the initial test task; and generating an executable test plan for the current round based on the robust aggregation rules, attack parameter strength, and final values of each global hyperparameter determined by the arbitration.
[0012] Optionally, the evaluation agent collects multi-dimensional evaluation indicators and calculates a comprehensive score based on the multi-dimensional evaluation indicators, including: obtaining multi-dimensional evaluation indicators; wherein the multi-dimensional evaluation indicators include at least two of global accuracy, attack success rate, and communication and resource overhead; performing extreme value normalization processing on the multi-dimensional evaluation indicators to obtain normalized indicators; and performing weighted fusion on the normalized indicators to obtain a comprehensive score.
[0013] Optionally, generating an experience feedback vector includes: determining the trend of the combined effect of attack strategy patches and defense strategy patches changing with each round based on test memory tuples stored in consecutive rounds; and generating an experience feedback vector to adjust the strategy search direction for the next round based on the changing trend.
[0014] Optionally, the preset iteration termination conditions include: the current iteration round reaches the preset maximum iteration number threshold; or within a preset number of consecutive rounds, the absolute value of the difference between the comprehensive scores of adjacent rounds is less than the preset convergence threshold, and the attack and defense game is determined to have entered a relatively balanced state.
[0015] Optionally, the security test results are determined based on the test memory tuples stored in each iteration, including: traversing all test memory tuples stored in the memory knowledge module; extracting the attack strategy patch that results in the highest overall score as the strongest attack path, and extracting the defense strategy patch that results in the lowest overall score as the optimal defense strategy; generating security test results that include the risk level of the federated learning system, the strongest attack path, and the optimal defense strategy; wherein, the risk level of the federated learning system is determined based on the highest attack success rate and the maximum global accuracy decrease recorded in each iteration.
[0016] In a second aspect, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multi-agent collaborative federated learning attack and defense security testing method as described in any one of the first aspects.
[0017] Thirdly, embodiments of the present invention provide an electronic device, comprising:
[0018] processor;
[0019] A memory for storing computer programs that can be executed by the processor;
[0020] Wherein, when the processor executes the computer program, it implements the steps of the multi-agent collaborative federated learning attack and defense security testing method as described in any one of the first aspects.
[0021] (III) Beneficial Effects
[0022] The beneficial effects of this invention are:
[0023] This application provides a multi-agent collaborative federated learning attack and defense security testing method. The initial test task is automatically broken down and distributed to various functional agents by a master orchestration agent. Attack and defense strategy agents independently generate local policy patches. After automatic conflict arbitration by the master orchestration agent, an executable test plan is generated. Finally, the execution evaluation agent uniformly schedules federated training experiments and automatically collects metrics to calculate a comprehensive score. This architecture completely establishes an automated link between attack configuration, defense deployment, training execution, and result analysis, replacing the existing testing mode that requires manually writing scripts to connect each stage. This significantly improves testing efficiency and reproducibility, enabling large-scale, multi-round dynamic attack and defense game testing. It achieves a fully automated closed-loop process for federated learning security testing, significantly reducing the degree of human intervention. Attached Figure Description
[0024] Figure 1 The flowchart shown is a security testing method for multi-agent collaborative federated learning attack and defense provided in an embodiment of this application;
[0025] Figure 2 The diagram shows a structural diagram of an empirical rule extraction model provided in an embodiment of this application;
[0026] Figure 3 The diagram illustrates a detailed flowchart of a multi-agent collaborative federated learning attack and defense security testing method provided in an embodiment of this application. Detailed Implementation
[0027] To better explain and facilitate understanding of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] With the widespread application of federated learning in data-sensitive scenarios such as privacy computing, collaborative training on smart terminals, healthcare, and finance, its ability to perform joint modeling without aggregating raw data has significantly improved data privacy protection. However, it has also broken the security boundaries of traditional centralized training, introducing new security risks. For example, malicious actors may launch data poisoning, model poisoning, backdoor attacks, client-side collusion attacks, or data heterogeneity may lead to robust aggregation failures and privacy leaks.
[0029] To assess and address these security risks, various security testing schemes and attack / defense verification frameworks for federated learning have been proposed. Existing standard procedures for federated learning security testing typically include: manually setting fixed attack scenarios, configuring specific aggregation rules or defense strategies, executing multiple rounds of federated training, and finally outputting the model's performance metrics after being attacked. Currently, industry and academia have launched many excellent underlying frameworks for federated learning and privacy computing, such as Flower, FedML, and NVFlare.
[0030] However, most of the aforementioned systems are positioned at the level of building basic communication architecture or general federated training processes. For security defense assessment, existing frameworks mainly rely on static security operator calls, or only support offline testing and verification for single nodes and single attack types. When facing real federated adversarial networks, existing technologies lack a central scheduling mechanism that can coordinate a global perspective and automatically break down attack and defense tasks. With the continuous evolution of attack and defense methods, existing federated learning security testing solutions still face the following pressing technical problems when dealing with complex, dynamic, and distributed federated scenarios:
[0031] The testing process is fragmented and lacks automation: In existing solutions, attack configuration, defense configuration, training execution, and result analysis are often independent modules that lack automated linkage and orchestration, requiring a large amount of manual script writing to connect them.
[0032] Highly reliant on manual parameter tuning, making it difficult to cover complex attack and defense combinations: The testing process often relies on manual repeated adjustment of experimental parameters (such as attack ratio, poisoning intensity, defense threshold, etc.). Faced with a huge parameter space, manual enumeration is difficult to cover targeted complex attack and defense combination scenarios.
[0033] Lack of multi-role collaboration mechanism: Most existing testing systems are based on a single controller, which can only serially complete the distribution of fixed configurations and result evaluation, and cannot simulate the complex, multi-node game of multiple attackers and multiple defense mechanisms in the real world.
[0034] Lack of dynamic attack and defense game and continuous iteration capabilities: Existing testing systems usually only support "single verification" of fixed attacks or fixed defenses, and cannot continuously iterate strategies based on historical experimental states, experience memory and evaluation feedback, making it difficult to form a dynamic attack and defense game chain to deeply explore the system's vulnerabilities.
[0035] The distributed execution and state synchronization mechanisms are imperfect: In a distributed federation environment, test execution, state synchronization, result aggregation and policy correction lack a unified closed-loop coordination mechanism, making it difficult to form a stable, reusable and scalable security testing system.
[0036] Therefore, there is an urgent need to propose a multi-agent collaborative federated learning attack and defense security testing method to achieve full-process automation and collaboration from test target understanding, attack and defense strategy generation, experiment orchestration and execution, indicator collection and evaluation to strategy iterative optimization.
[0037] Based on this, this application provides a multi-agent collaborative federated learning attack and defense security testing method. The initial test task is automatically broken down and distributed to various functional agents by a master orchestration agent. Attack and defense strategy agents independently generate local policy patches. After automatic conflict arbitration by the master orchestration agent, an executable test plan is generated. Finally, the execution evaluation agent uniformly schedules federated training experiments and automatically collects metrics to calculate a comprehensive score. This architecture completely establishes an automated link between attack configuration, defense deployment, training execution, and result analysis, replacing the existing testing mode that requires manually writing scripts to connect each stage. This significantly improves testing efficiency and reproducibility, enabling large-scale, multi-round dynamic attack and defense game testing. It achieves a fully automated closed-loop process for federated learning security testing, significantly reducing the degree of human intervention.
[0038] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present invention can be understood more clearly and thoroughly, and that the scope of the present invention can be fully conveyed to those skilled in the art.
[0039] Please see Figure 1 , Figure 1 A flowchart of a multi-agent collaborative federated learning attack and defense security testing method provided in an embodiment of this application is shown. It should be understood that this federated learning attack and defense security testing method can be executed by an electronic device, and the specific device can be configured according to actual needs; this embodiment is not limited to this. For example, the electronic device can be a computer or a server, etc. Specifically, the federated learning attack and defense security testing method includes:
[0040] Step S110: The master control orchestration agent receives the initial test task and breaks it down into multiple subtasks, including at least an attack exploration subtask and a defense exploration subtask.
[0041] It should be understood that the specific subtasks included in these multiple subtasks can be set according to actual needs, and the embodiments of this application are not limited thereto.
[0042] To facilitate understanding of step S110, it will be described below through specific embodiments.
[0043] Specifically, the user's test requirements are obtained by requesting the access module and formalized into a tuple. This tuple defines the complete context constraints and available resource boundaries for this test task, and the meanings of each element in the tuple are as follows:
[0044] This represents a set of test targets, which may contain one or more test targets. For example, the test targets could be maximizing the attack success rate (i.e., finding the system's most vulnerable scenario), verifying the robustness of a specified defense algorithm in non-independent and identically distributed data (i.e., non-IID data) scenarios, or evaluating the lower bound of the system's accuracy loss under specific attack types.
[0045] This represents the set of environmental parameters of the federated learning system under test. For example, the... The configuration items may include, but are not limited to: dataset type (e.g., CIFAR-10, EMNIST, or custom medical image data), model network structure (e.g., ResNet-18, VGG-16, or Transformer structure), total number of clients / participants in co-training, and data distribution type of local data for each client (e.g., independent and identically distributed, or non-independent and identically distributed based on Dirichlet distribution and the values of its distribution parameters).
[0046] This represents the set of test constraints, which sets resource limits and termination boundaries during test execution to ensure that the test can be completed within a controllable computational overhead. For example, The constraints include, but are not limited to: the maximum communication overhead allowed by the system (i.e., the total number of rounds of global model aggregation or the total amount of data transmitted), the maximum number of training rounds in a single federated training experiment, the maximum number of iterations in multi-agent games, and the maximum computing resources allowed to be used (e.g., GPU memory limits or CPU core count limits).
[0047] This serves as a candidate capability space for the platform, and This includes a list of candidate attack libraries (e.g., Gaussian noise poisoning supporting gradient perturbation, label flipping attacks, model replacement attacks, backdoor trigger injection attacks, etc.) and a list of candidate defense libraries (e.g., supporting Krum aggregation, multi-Krum aggregation, median aggregation, and anomalous gradient filtering based on cosine similarity, etc.). It should be noted that subsequent attack and defense strategy agents will be defined within this candidate space. Within the constraints, local policy patches are searched and generated.
[0048] And, the master orchestration agent received Subsequently, based on internal intent parsing rules, the global test task is decomposed into a set of parallel or serial subtasks for each functional agent, and this set of subtasks can be represented as: The meanings of each element in this subtask set are as follows:
[0049] This indicates an attack exploration task, which is issued to the attack strategy agent, instructing it to explore the platform's candidate capability space. Within the candidate attack library, to achieve the test target set. Guided by this, search for and generate attack strategy patch configurations;
[0050] This indicates a defense exploration mission, which is issued to the defense strategy agent, instructing it to explore the platform's candidate capability space. Within the candidate defense library, search and generate corresponding robust aggregation rules and anomaly detection rule configurations;
[0051] and This indicates the execution and evaluation task, which is distributed to the execution and evaluation agent. It includes the scheduling constraints of the underlying federated execution engine (corresponding to the set of federated environment configurations). ), and based on the test target set Determined weighting rules for multidimensional evaluation indicators;
[0052] and This indicates the memory and reporting tasks, which are distributed to the memory knowledge module and the report output module, respectively clarifying the key state nodes that need to be recorded in this round of testing, as well as the risk profile dimensions that need to be presented in the final safety assessment report.
[0053] Therefore, through the above-mentioned decomposition and distribution mechanism, the master orchestration agent transforms the initial test intent into parallel or serial subtasks that can be executed independently by each functional agent, providing a structured task-driven foundation for subsequent multi-agent collaborative games.
[0054] In step S120, the attack strategy agent responds to the attack exploration subtask and generates a local attack strategy patch in the candidate attack space, and the defense strategy agent responds to the defense exploration subtask and generates a local defense strategy patch in the candidate defense space.
[0055] It should be understood that the content of the attack strategy patch can be configured according to actual needs, and the embodiments of this application are not limited thereto.
[0056] Optionally, the attack strategy patch includes attack type, attack scope, attack parameter strength, and attack timing arrangement; wherein, the attack type includes at least one of data poisoning, tag flipping, model poisoning, backdoor trigger injection, and client collusion attack.
[0057] It should also be understood that the content of the defense strategy patch can be configured according to actual needs, and the embodiments of this application are not limited thereto.
[0058] Optionally, the defense strategy patch includes robust aggregation rules, abnormal client filtering rules, and defense sensitivity parameters; wherein, the robust aggregation rules include at least one of the following: Krum algorithm, multiple Krum algorithm, coordinate median algorithm, truncated mean algorithm, or weighted average mechanism based on confidence weight.
[0059] To facilitate understanding of step S120, a specific embodiment will be described below.
[0060] Specifically, for the complex parameter space of federated learning, this application abandons the traditional global static configuration mode and adopts a "Local Strategy Patch" mechanism. The attack strategy agent and the defense strategy agent work in parallel, each searching and generating local strategy patches in an independent policy space. Neither knows the other's specific configuration, thereby improving the efficiency and coverage of policy search.
[0061] In this process, after receiving the attack exploration task from the master orchestration agent, the attack strategy agent, combined with experience feedback information from the historical memory pool, generates a structured local attack strategy patch within the given candidate attack space. This can be formally represented as:
[0062] .
[0063] The meanings of each element are as follows:
[0064] The attack type is indicated, and it includes at least one of the following: data poisoning (injecting incorrect or biased samples into training data to disrupt model training), label flipping (maliciously altering data labels to cause the model to learn incorrect mapping relationships), model poisoning / model replacement (directly uploading maliciously constructed model updates to replace or pollute the global model), backdoor trigger injection (implanting hidden trigger conditions into the model so that specific inputs trigger incorrect predictions), and client collusion attack (multiple malicious clients colluding to carry out attacks to amplify the destructive effect);
[0065] This indicates the attack scope, which is defined as the set of clients selected as malicious nodes and their proportion. For example, randomly selecting 10% of a total of 100 clients as malicious nodes, or launching a targeted attack against clients holding a specific data distribution;
[0066] This indicates the strength of attack parameters, which defines the specific intensity of the attack behavior, such as the scaling factor of the poison data, the step size of the gradient perturbation, and the pixel size and transparency of the backdoor trigger pattern.
[0067] This indicates the timing of the attack, defining the intervals within which the attack occurs during the federated training process. For example, it could involve continuous poisoning at the beginning of federated training, or launching a sudden backdoor injection attack in the last few communication rounds before the model converges.
[0068] Furthermore, after receiving the defense exploration task from the master orchestration agent, the defense strategy agent generates a structured local defense strategy patch in the candidate defense space based on the defense targets that enhance the model's aggregation robustness and anomaly detection capabilities. This defense strategy patch... This can be formally represented as:
[0069] .
[0070] The meanings of each element are as follows:
[0071] This refers to robust aggregation rules, which define the aggregation algorithm used by the server to resist the influence of malicious gradients, including but not limited to at least one of the following: Krum algorithm, multiple Krum algorithm, coordinate median algorithm, truncated mean algorithm, or weighted average mechanism based on client confidence weight.
[0072] This refers to the abnormal client filtering rules, which define the detection strategies for removing or downweighting suspicious client updates before global aggregation. Examples include gradient pruning based on cosine similarity, dynamic reputation scoring based on local validation set loss values, or outlier detection rules based on cluster analysis.
[0073] This represents the defense sensitivity parameter, which defines the specific execution threshold of the aforementioned defense algorithms and filtering rules, such as the norm threshold of gradient pruning, the upper limit of the number of client updates allowed to participate in aggregation, and the variance of Laplacian noise or Gaussian noise injected when introducing differential privacy protection mechanisms.
[0074] Therefore, using the above method, the attack strategy agent and the defense strategy agent can independently complete the optimal strategy search in their own dimension without needing to know the other's complete configuration. The local policy patches generated by both agents... and All data is submitted to the master orchestration agent in the form of reports. This decoupling design significantly reduces the state and action space of a single agent, enabling the system to efficiently cope with highly complex federated adversarial environments.
[0075] Step S130: The main control orchestration agent performs conflict detection and arbitration on the attack strategy patch and the defense strategy patch, so as to merge them to generate an executable test plan for the current round.
[0076] Specifically, a domain isolation priority is adopted. For server-side aggregation logic, robust aggregation rules in the defense strategy patch are determined, and for malicious client local training logic, attack parameter strength in the attack strategy patch is determined. When the expected values proposed by the attack strategy patch and the defense strategy patch for the same global hyperparameter are inconsistent, the weighted average of the two is calculated as the final value of the global hyperparameter, or the expected value of one of them is selected as the final value of the global hyperparameter according to the test objective in the initial test task. Based on the robust aggregation rules, attack parameter strength, and final values of each global hyperparameter determined by the above arbitration, an executable test plan for the current round is generated.
[0077] For example, since the two patches mentioned above are generated independently by different agents in a decoupled parallel mode, they may propose conflicting desired configuration values for shared global hyperparameters (such as global learning rate, batch size, and the proportion of clients participating in aggregation in each round). Therefore, the master orchestration agent needs to perform unified conflict resolution and configuration merging. Specifically, the master orchestration agent uses domain isolation priority rules to divide and determine the configuration domains that are exclusive to each of the attack strategy patch and the defense strategy patch. Specifically:
[0078] Regarding server-side aggregation logic: the selection of the server-side aggregation algorithm belongs to the core decision domain of the defender, and the attacker should not interfere. Therefore, the master orchestration agent determines to adopt the robust aggregation rule in the defense strategy patch, that is, according to the domain isolation priority rule, the robust aggregation rule in the defense strategy patch has absolute priority in the configuration of the server-side aggregation logic. For example, if the robust aggregation rule specified in the defense strategy patch is the "Multi-Krum" algorithm, then regardless of whether the attack strategy patch contains any implicit expectations for the aggregation rule, the final executable test plan will force the use of "Multi-Krum" as the server-side aggregation rule;
[0079] Regarding the local training logic of malicious clients: The configuration of attack behaviors during the local training process of malicious clients belongs to the core decision domain of the attacker, and the defender should not interfere. Therefore, the master orchestration agent determines the attack parameter strength in the attack strategy patch, that is, the attack parameter strength in the attack strategy patch has absolute priority in the configuration of the local training logic of malicious clients. For example, if the attack parameter strength specified in the attack strategy patch is "the scaling factor of the poisoned data is 5.0", then regardless of whether the defense strategy patch involves suggestions on hyperparameters for local training of the client, the final executable test plan will force the application of this scaling factor (5.0) to the local training process of the client marked as malicious.
[0080] By isolating these domains, it is ensured that the defender has priority in choosing the server-side aggregation algorithm, while the attacker has priority in configuring the specific intensity of local malicious behavior, thereby avoiding configuration conflicts between the two parties in their respective core functional domains.
[0081] Furthermore, when the attack strategy patch and the defense strategy patch propose inconsistent configuration values for the same global hyperparameter, the master orchestration agent arbitrates the conflicting expected values. Here, the expected value refers to the specific configuration value proposed by the attack strategy patch or the defense strategy patch for a given global hyperparameter. For example, the attack strategy patch might expect the global learning rate to be set to 0.01 so that malicious gradient updates can more significantly affect the global model; while the defense strategy patch might expect the global learning rate to be set to 0.05 to maintain the stability of model training. When these two proposed values are inconsistent, a global hyperparameter conflict occurs. The master orchestration agent arbitrates the conflicting expected values using one of the following methods:
[0082] Calculate the weighted average of the expected values of the attack strategy patch and the defense strategy patch, and use this as the final value of the global hyperparameter; or
[0083] Based on the test objectives in the initial test task, the expected value of the attack strategy patch or the expected value of the defense strategy patch is selected as the final value of this global hyperparameter.
[0084] For example, if the initial test objective is to find the vulnerability of the federated learning system under the worst attack, the master orchestration agent can give higher weight to the expected value of the attack strategy patch or directly adopt the attacker's expected value; if the test objective is to verify the effectiveness of a specific defense strategy, the defender's expected value can be adopted first; if the test objective is to evaluate the adversarial effect of both the attacker and the defender in a balanced way, the master orchestration agent can take the weighted average of the two expected values.
[0085] Subsequently, based on the configuration results determined by the aforementioned arbitration process, the master orchestration agent integrates the parameters that have resolved conflicts, specifically including:
[0086] Robust aggregation rules determined by domain isolation priority;
[0087] Attack parameter strength determined by domain isolation priority;
[0088] The final values of each global hyperparameter determined by global hyperparameter conflict arbitration.
[0089] The master orchestration agent merges the above configuration items to generate a globally executable test plan for the current round t. The executable test plan will be sent to the execution evaluation agent, which will then schedule the underlying federated learning environment to execute the federated training experiment for this round.
[0090] Step S140: The evaluation agent schedules the federated learning environment to perform training experiments according to the executable test plan to collect multi-dimensional evaluation indicators and calculate a comprehensive score based on the multi-dimensional evaluation indicators.
[0091] Specifically, the evaluation agent receives the executable test plan for the current round generated by the master orchestration agent. Then, it is converted into a configuration file format recognizable by the underlying federated learning execution engine. In a distributed computing environment or simulation environment, the execution evaluation agent is responsible for scheduling the participating nodes and executing the federated training experiments in this round in the following order: according to the attack scope defined by the attack strategy patch in the executable test scheme. The clients participating in the federated training are divided into a set of normal clients and a set of malicious clients; the federated training process is initiated, and attacks are scheduled at the times specified in the attack strategy patch. The corresponding communication round triggers a malicious client to execute a specified poisoning attack or backdoor trigger injection behavior; after receiving local model gradients or model weight updates uploaded by each client, the server calls the abnormal client filtering rules defined in the defense strategy patch. With robust aggregation rules The collected updates are filtered and aggregated to update the global model.
[0092] After the experiment was completed, the evaluation agent quantitatively evaluated the offensive and defensive game performance in the current round t. Specifically, the evaluation agent collected the following multi-dimensional raw evaluation metrics:
[0093] Global precision It refers to the prediction accuracy of the aggregated global model obtained after this round of federated training on a clean test sample set, which is used to measure the ability of the defense mechanism to maintain the performance of the main task under attack conditions.
[0094] Attack success rate This refers to the probability that a malicious sample will be identified as a target label specified by the attacker in a global model for backdoor or targeted poisoning attacks, which is used to measure the actual effect of the attack.
[0095] Communication and resource overhead This refers to the additional communication bandwidth consumption and computational latency introduced by executing the anomaly detection algorithm and the complex robust aggregation algorithm in this round of testing, and is used to measure the operational cost of the defense mechanism.
[0096] To transform the aforementioned multi-dimensional evaluation metrics into a single quantitative signal that can directly guide the multi-agent system in the next round of policy iteration, the evaluation agent employs an extreme value normalization and weighted fusion algorithm to calculate the comprehensive score for this round. The calculation process is as follows:
[0097] First, extreme value normalization is performed on each original evaluation index. The normalization formula is as follows:
[0098] ;
[0099] In the formula, The value is the normalized value; This represents the raw indicator value collected in the current round; This indicates the minimum value of the indicator across all historical rounds executed. This indicates the maximum value of the indicator across all historical rounds executed.
[0100] Subsequently, based on the test objective preference weights set in the initial test task, the normalized indicators are weighted and fused to obtain the comprehensive score for this round. For example, if the test objective in the initial test task is to evaluate the vulnerability of the federated learning system under the worst-case attack conditions, the comprehensive score function can be defined as:
[0101] ;
[0102] In the formula, This represents the overall score for the current round t, and this score directly quantifies the system risk level under the current "attack-defense" combination; , and These are the weight coefficients corresponding to the attack effect, model accuracy loss, and resource cost, respectively, and satisfy the following conditions: + + =1; This represents the normalized attack success rate. This represents the global precision after normalization. This represents the normalized communication and resource overhead.
[0103] It should be noted here that the test target preference weight (i.e. , and The request access module determines the test target set from the initial test task. The weighting can be determined by the user or specified directly when submitting test requirements. For example, if the test objective is to evaluate the system's vulnerability to extreme attacks, the attack success rate metric will be assigned a higher weight; if the test objective is to verify the effectiveness of the defense strategy, the global accuracy metric will be assigned a higher weight. Furthermore, this preference weight is distributed to the evaluation agent during the task initialization phase along with the evaluation sub-task, and is used to calculate the overall score.
[0104] It should also be noted that the above comprehensive score directly quantifies the system risk level under the current "attack-defense" configuration combination: the higher the score, the more advantageous the attacker is and the higher the system vulnerability under this configuration; the lower the score, the better the defender's defense effect and the stronger the system security.
[0105] Step S150: The attack strategy patch, defense strategy patch, executable test plan, multi-dimensional evaluation index and comprehensive score of the current round are stored as test memory tuples of the current round through the memory knowledge module, and an experience feedback vector is generated.
[0106] Specifically, considering that most existing testing frameworks are single-validation processes and cannot accumulate testing experience, this invention introduces a knowledge memory module to achieve experience reuse, strategy optimization, and automated convergence determination for multi-agent games in multiple rounds.
[0107] After the t-th round of federated training experiment and evaluation, the memory knowledge module receives the comprehensive score and detailed evaluation metrics from the evaluating agent. The memory knowledge module stores the complete context information of this round of testing as a structured test memory tuple in the test memory pool for persistent storage. This test memory tuple can be represented as:
[0108] ;
[0109] In the formula, This represents the test memory tuple for the current round t; This indicates the attack strategy patch for the current round t; This indicates the defense strategy patch for the current round t; This represents the executable test plan for the current round t; This represents the multidimensional evaluation index for the current round t; This represents the overall score for the current round t, including global accuracy, attack success rate, and communication and resource overhead.
[0110] By storing the attack and defense configurations, execution plans, and evaluation results of each round in the form of structured tuples, the system has accumulated historical test experience data covering various attack and defense combination scenarios, providing a data foundation for subsequent experience rule extraction and strategy iteration optimization.
[0111] After storing the test memory tuples for this round, the memory knowledge module generates an experience feedback vector to guide the generation of the next round's strategy, based on the historical tuple data stored in the test memory pool over multiple consecutive rounds. Specifically, this generation process includes:
[0112] The memory knowledge module analyzes memory tuples from multiple consecutive rounds of testing to determine the changing trend of the combined effect of the attack strategy patch and the defense strategy patch across rounds. This combined effect can be reflected by changes in the overall score, attack success rate, and global accuracy across each round. For example, analysis reveals patterns such as "when the poisoning ratio is below a certain threshold, the global accuracy loss under a specific robust aggregation algorithm approaches zero."
[0113] Furthermore, the memory knowledge module generates an experience feedback vector based on the aforementioned trend to adjust the strategy search direction for the next round. In a preferred embodiment, the memory knowledge module can invoke a pre-set experience rule extraction model to analyze data from multiple consecutive rounds, extract stage-specific experience rules, and output a structured experience feedback vector. For example, after identifying the correlation between the aforementioned poisoning ratio and accuracy loss, an experience feedback vector containing indications such as "suggest shrinking the attack parameter space" or "suggest switching the attack type" is generated. .
[0114] It should be understood that the model structure of this empirical rule extraction model can be set according to actual needs, and the embodiments of this application are not limited thereto.
[0115] Optionally, such as Figure 2 As shown, the empirical rule extraction model may include an input layer, a configuration parsing layer, a multidimensional trend correlation analysis layer, and an output layer connected in sequence.
[0116] The input layer receives a sequence of test memory tuples consisting of k consecutive rounds of test memory tuples. Each test memory tuple in the sequence contains an attack strategy patch, a defense strategy patch, an executable test plan, multi-dimensional evaluation metrics, and a comprehensive score for the corresponding round. The specific value of k can be set according to actual needs, and this embodiment is not limited to this.
[0117] Additionally, the configuration parsing layer transforms the discrete configuration items and continuous parameter values in each round's local attack strategy patch and local defense strategy patch into standardized configuration feature vectors. This layer contains the following functional units:
[0118] Attack Configuration Parsing Unit: This unit parses the local attack strategy patches for each round and generates attack configuration feature vectors based on the combination of attack type, attack scope, and attack timing. These vectors characterize the attacker's strategy orientation for that round. For example, typical configuration patterns such as "early full-scale poisoning," "late-stage targeted backdoor injection," and "continuous low-intensity tag flipping" are mapped to corresponding feature representations.
[0119] Defense Configuration Parsing Unit: This unit parses the local defense strategy patches for each round and generates a defense configuration feature vector based on a combination of robust aggregation rules, abnormal client filtering rules, and defense sensitivity parameters. This vector characterizes the defense's strategy orientation for that round. For example, typical configuration patterns such as "Krum aggregation combined with strict cosine similarity filtering," "multiple Krum aggregation combined with lenient thresholding," and "median aggregation combined with differential privacy noise" are mapped to corresponding feature representations.
[0120] The attack-defense confrontation intensity calculation unit calculates the attack-defense confrontation intensity index for this round by comparing the numerical relationship between the attack parameter intensity in the attack configuration feature vector and the defense sensitivity parameter in the defense configuration feature vector. For example, assuming the attack parameter intensity specified in the attack strategy patch is the "scaling factor of the poison data," with a value of 5; and the defense sensitivity parameter specified in the defense strategy patch is the "norm threshold of gradient clipping," with a value of 2. The attack-defense confrontation intensity index can be obtained by calculating the ratio of the attack parameter intensity to the defense sensitivity parameter, i.e., dividing 5 by 2, resulting in a ratio of 2.5. This ratio is the attack-defense confrontation intensity index for this round.
[0121] It's important to note that a higher attack-defense intensity index indicates a greater advantage of the attacker's disturbance strength relative to the defender's tolerance limit, leading to more intense confrontation. Conversely, if the attack parameter strength is low while the defense sensitivity parameter is high, the ratio is less than 1, indicating that the defender has a large tolerance margin under the current configuration, and the attack behavior is easily filtered by the defense mechanism, resulting in a lower attack-defense intensity. When the attack parameter strength and defense sensitivity parameter values are similar, the ratio approaches 1, indicating a relatively balanced confrontation under the current configuration. Therefore, through the above calculation method, the attack-defense intensity index transforms the numerical comparison between the attack parameter strength and the defense sensitivity parameter into a continuous quantitative indicator, facilitating subsequent multi-dimensional trend correlation analysis to accurately capture the numerical variation patterns of the combined effects of attack strategy patches and defense strategy patches under different confrontation intensities.
[0122] Therefore, the configuration parsing layer ultimately outputs the configuration feature vector for each round, which integrates three types of information: attack configuration feature vector, defense configuration feature vector, and attack-defense confrontation intensity index.
[0123] Additionally, a multi-dimensional trend correlation analysis layer is used to analyze the temporal correlation between configuration feature vectors, multi-dimensional evaluation indicators, and comprehensive scores across rounds, in order to determine the changing trend of the combined effect of attack strategy patches and defense strategy patches across rounds. This layer includes:
[0124] Configuration time-series evolution analysis unit: Performs evolution analysis on the configuration feature vector sequence of consecutive k rounds to identify whether the attack configuration has changed (e.g., the attack type has changed from data poisoning to backdoor injection), whether the defense configuration has been adjusted (e.g., the robust aggregation rule has changed from Krum to multi-Krum), and whether the attack and defense confrontation intensity index shows an increasing or decreasing trend.
[0125] Combined Effect Trend Extraction Unit: Extracts trends from the comprehensive score sequence and multi-dimensional evaluation index sequence (such as attack success rate sequence and global accuracy sequence) over k consecutive rounds to determine the changing trend of the combined effect. For example, it identifies trend types such as a continuous increase in the comprehensive score (the attacker gradually gains the upper hand), a continuous decrease (the defender gradually adapts), or a trend towards stabilization (both sides reach a stalemate).
[0126] The configuration-effect correlation mapping unit maps the temporal evolution of configuration feature vectors with the changing trends of combined effects across rounds, identifying configuration change factors that significantly impact the overall score. For example, this unit can identify that after the defense configuration in round m is switched from "multiple Krum combined with a loose threshold" to "Krum combined with strict cosine similarity filtering," the upward trend of attack success rate in subsequent rounds is significantly suppressed, and the overall score enters a downward trend; or it can identify correlation patterns such as "when the attack parameter strength is below a certain threshold, the global accuracy loss under a specific robust aggregation rule is close to zero."
[0127] Therefore, the final output of this multidimensional trend correlation analysis layer is a series of structured records of attack-defense correlation patterns. Each record contains the following information: the type and round of configuration change, the type of trend of change in the combined effect, and a description of the causal relationship between the configuration change and the effect change. For example, a typical attack-defense correlation pattern record can be expressed as: "In round m, the defense configuration was switched from A to B, which caused the upward trend of the subsequent attack success rate to stop, and the overall score changed from rising to falling."
[0128] Furthermore, this output layer receives structured attack-defense correlation records from the multi-dimensional trend correlation analysis layer and generates an empirical feedback vector to adjust the direction of the next round of strategy search. This output layer includes:
[0129] The strategy adjustment direction decision unit receives the attack-defense correlation record as input, analyzes the causal relationships in each record, and generates specific strategy adjustment instructions for the attack strategy agent or the defense strategy agent. For example, upon receiving the attack-defense correlation record "In round m, the defense configuration was switched from multiple Krum combined with a relaxed threshold to Krum combined with strict cosine similarity filtering, causing the subsequent attack success rate to stop its upward trend and the overall score to change from rising to falling," the unit determines that the defense configuration switch has a significant positive effect on suppressing attacks, and therefore generates a strategy adjustment instruction: "It is recommended that the defense strategy agent maintain or further strengthen the current defense configuration in the next iteration," or "It is recommended that the attack strategy agent switch the attack type to try to bypass the current defense configuration." Similarly, upon receiving the attack-defense correlation record "When the attack parameter strength is below 2.0, the global accuracy loss under a specific robust aggregation rule is close to zero," the unit determines that the current attack strength is insufficient to pose an effective threat to the defense, and therefore generates a strategy adjustment instruction: "It is recommended that the attack strategy agent raise the search lower bound of the attack parameter strength to above 2.0 in the next iteration."
[0130] Feedback Vector Encoding Unit: Encodes the policy adjustment instructions generated by the policy adjustment direction decision unit into a structured empirical feedback vector. Each dimension of this empirical feedback vector corresponds to a different policy adjustment operation type and adjustment magnitude, which can be directly parsed by the attack policy agent and the defense policy agent, and used to narrow down the candidate policy space, perform targeted mutation, or explore in the next iteration.
[0131] The structural description of the above-described empirical rule extraction model is merely exemplary. In actual implementation, this model can be replaced by other analytical models or rule extraction modules with equivalent functions, such as a policy evaluation model based on reinforcement learning, a trend analysis model based on Bayesian optimization, or an association rule mining module based on statistical hypothesis testing. Any computational model capable of extracting attack-defense correlation patterns from historical test data and generating policy adjustment instructions falls within the protection scope of this invention.
[0132] In step S160, the master orchestration agent determines whether the preset iteration termination condition is met based on the comprehensive score and experience feedback vector. If the condition is not met, the attack strategy agent and the defense strategy agent are driven to proceed to the next iteration. If the condition is met, the iteration ends, and the security test result is determined based on the test memory tuples stored in previous iterations.
[0133] It should be understood that the specific conditions for the preset iteration termination condition can be set according to actual needs, and the embodiments of this application are not limited thereto.
[0134] For example, preset iteration termination conditions include: the current iteration round reaches a preset maximum iteration count threshold, for details please refer to Figure 3 The relevant process; or within a preset number of consecutive rounds, if the absolute value of the difference between the comprehensive scores of adjacent rounds is less than a preset convergence threshold, it is determined that the attack and defense game has entered a relatively balanced state.
[0135] For example, the master orchestration agent receives a comprehensive score. With experience feedback vector The system then executes the iteration termination logic. The system's termination conditions mainly include the following two items: a maximum iteration count threshold constraint, referring to the current round... ( (Maximum number of game rounds set at the beginning of the task); game convergence constraints (local Nash equilibrium approximation), calculate the change in score over k consecutive rounds:
[0136] ;
[0137] in, This represents the difference in overall scores between adjacent rounds; This represents the preset minimum convergence threshold, and its specific value can be set according to actual needs. If satisfied, it indicates that both the attack and defense strategies have reached their bottlenecks within the current parameter space, and the game has entered a relatively balanced state.
[0138] If the master orchestration agent determines that the iteration termination condition is not met, the experience feedback vector generated in this round is used as context information to drive the attack strategy agent and the defense strategy agent into the (t+1)th iteration. In this iteration, each agent no longer performs blind random search, but instead narrows its candidate strategy space based on the experience feedback vector, and performs targeted mutation or exploration, thereby accelerating the search for the most vulnerable attack and defense links in the federated learning system.
[0139] When the master orchestration agent determines that any iteration termination condition is met, it triggers a test termination command and calls the report output module. The report output module iterates through the historical test memory tuples stored in the memory knowledge module, i.e., the history of all test memory tuples from round 1 to round t, and generates a structured security assessment test report accordingly. The security assessment test report includes:
[0140] (1) Risk level determination: Based on the highest attack success rate and the maximum global accuracy decline recorded in each iteration, the security vulnerability of the current federated learning system is classified as high-risk, medium-risk or low-risk.
[0141] (2) Extraction of the optimal countermeasure path: Extract the attack strategy patch that leads to the highest comprehensive score from the memory tuples of previous tests as the strongest attack path, and the defense strategy patch with the best defense effect as the optimal defense strategy.
[0142] (3) Vulnerability profile and reinforcement suggestions: The performance of the federated learning system in resisting attacks across various evaluation metrics is displayed in a visual manner (such as radar charts), and parameter tuning suggestions and algorithm replacement suggestions are automatically generated based on test results.
[0143] Through the aforementioned multi-round game and automated report generation mechanism, this application achieves a fully automated closed loop for federated learning security testing, from task parsing, policy generation, experiment execution to result evaluation and report output, and has the adaptive capability to iterate policies based on historical experience.
[0144] In other words, the security test results are determined based on the test memory tuples stored in each iteration, including: traversing all test memory tuples stored in the memory knowledge module; extracting the attack strategy patch that results in the highest overall score as the strongest attack path, and extracting the defense strategy patch that results in the lowest overall score as the optimal defense strategy; generating security test results that include the risk level of the federated learning system, the strongest attack path, and the optimal defense strategy; wherein, the risk level of the federated learning system is determined based on the highest attack success rate and the maximum global accuracy decrease recorded in each iteration.
[0145] In summary, by employing the above technical solutions, this application achieves the following technical effects:
[0146] This application, by setting up a master orchestration agent and a test request access module, can automatically transform user test intentions into a set of tasks readable by the system. It establishes a complete link from target parsing, policy generation, experiment execution to result evaluation, effectively solving the problems of fragmented test processes and the need for extensive manual script writing in existing technologies, thus improving the automation and efficiency of security testing. Furthermore, the multi-agent role division and local policy patch generation mechanism proposed in this application allow attack and defense policy agents to independently search and mutate policies, with the master orchestration agent arbitrating and merging conflicts. This effectively overcomes the search difficulties caused by the large parameter space under global static configuration, enabling the coverage of more complex attack and defense combination scenarios. Moreover, by introducing a collaborative interaction mechanism among the master orchestration agent, attack policy agent, defense policy agent, and execution evaluation agent, this application simulates an adversarial environment with multiple attackers and multiple defense mechanisms operating in parallel, filling the gap in existing testing systems' lack of multi-role collaborative capabilities. Furthermore, by setting up a knowledge memory module and a comprehensive scoring feedback loop, the system can record experimental states and configuration differences in multiple rounds of gameplay, and extract empirical rules to guide the next iteration. This enables the system to possess self-learning and adaptive capabilities, continuously uncovering security vulnerabilities in federated learning systems and approaching a relative equilibrium state in the attack-defense game. Additionally, this application establishes a multi-dimensional weighted evaluation index that integrates global accuracy, attack success rate, and communication and resource overhead. Based on this quantitative system, the report output module can automatically generate a structured security assessment report containing risk level, optimal attack-defense path, and system hardening recommendations.
[0147] To facilitate understanding of the embodiments of this application, specific embodiments are described below.
[0148] Specifically, let's define the scenario for the implementation example:
[0149] Application scenario: Jointly training image classification models in medical or financial institutions (using ResNet-18 network architecture).
[0150] Dataset and Distribution: CIFAR-10 dataset, split using non-independent identically distributed (Non-IID) partitioning, Dirichlet distribution parameters. .
[0151] Federation node size: 1 central server and 10 participating parties (clients).
[0152] The initial testing intent was to "evaluate the robustness of the federated model against model poisoning attacks and to find the optimal defense strategy."
[0153] Based on the above scenario, the steps for implementing a multi-agent collaborative federated learning attack and defense security testing method are as follows:
[0154] First, task access and initialization: The system receives the above test intent, and the test request access module converts it into a structured task. The constraints are set as follows: the maximum number of communication rounds is 50, the maximum number of multi-agent iterative games is 10, and the upper limit for the proportion of malicious clients is set at 20%.
[0155] Secondly, task parsing and decomposition of the master orchestration agent parsing The task is broken down into an attack exploration subtask (limited to the model poisoning category), a defense exploration subtask (finding a robust aggregation algorithm), and an execution evaluation subtask, and then distributed to the corresponding agents.
[0156] And, attack strategy patch generation (round 1):
[0157] The attack strategy agent generates initial local attack strategy patches from the candidate attack library based on the test target. :
[0158] Gaussian noise model poisoning attack;
[0159] Client2 and Client5 were randomly selected as malicious nodes (accounting for 20%).
[0160] The noise scaling factor is set to 5.0;
[0161] Poisoning continued from round 1 to round 50.
[0162] And, defense strategy patch generation (round 1):
[0163] In order to defend against poisoning, the defense strategy agent generates an initial local defense strategy patch. :
[0164] It employs the basic Krum aggregation algorithm;
[0165] No additional client-side removal rules;
[0166] No additional threshold.
[0167] And, conflict arbitration and experimental enforcement
[0168] Master orchestration intelligent agent and The merge was performed, and no hyperparameter conflicts were found. The first round of global test plan was then generated. The evaluation agent invokes the underlying federated execution engine to initiate and complete 50 rounds of federated training in a distributed environment consisting of 10 nodes.
[0169] And, multi-dimensional indicator collection and comprehensive scoring:
[0170] Experimental results from evaluating the agent's data collection revealed that, due to the highly non-independent and identically distributed (Non-IID) data distribution, the Krum algorithm misjudged and discarded long-tail updates from normal clients as malicious updates, instead selecting malicious gradients with masquerading noise, thus affecting the global model accuracy. The percentage plummeted from 82% to 45% when there was no attack baseline.
[0171] The system calculates the overall score for this round. (A high score indicates that the attacker has an absolute advantage in the current offensive and defensive combination, and the system is extremely vulnerable.)
[0172] And, the accumulation of knowledge and feedback from experience:
[0173] The memory knowledge module stores the above configurations, indicators, and scores into the memory pool. and extract experience feedback "In scenarios with a high degree of Non-IID ( The single Krum algorithm cannot effectively defend against Gaussian model poisoning with a scaling factor of 5.0. It is recommended that the defender broaden the aggregation selection range, and the attacker can try to reduce the noise concealment.
[0174] And, multiple rounds of iteration and collaborative optimization:
[0175] The master orchestration agent determines that the maximum number of iterations has not been reached (round 1 < 10) and has not converged, triggering round 2 iteration: the defense agent (adjusted based on feedback) will... The aggregation algorithm in the code has been modified to "Multi-Krum" and an "anomaly filtering rule based on cosine similarity" has been added. The attack agent (adjusted based on feedback) will... The poisoning method has been upgraded to "adaptive scaling targeted poisoning".
[0176] The system repeats the steps of conflict arbitration and experimental execution until the knowledge is stored in memory and feedback is received. As the game progresses to the 8th round, the strategies of both sides reach a local Nash equilibrium (after 3 consecutive rounds). (with minimal fluctuations), the main control agent determines that the system has converged and the termination condition is met.
[0177] Finally, the safety assessment report is output:
[0178] The report output module summarizes the data from 8 rounds of gameplay and outputs the final security assessment report:
[0179] 1. Risk rating: Medium to high risk.
[0180] 2. Optimal attack path: Alternating model poisoning with an adaptive scaling factor of 2.5 (which can bypass similarity detection to the greatest extent).
[0181] 3. Optimal Defense Recommendation: For the current distribution of Non-IID data, it is recommended to adopt a combined defense strategy of "Multi-Krum combined with Coordinate-wise Median", which can maintain global accuracy above 76% under extreme attacks.
[0182] It should be understood that the above-described multi-agent collaborative federated learning attack and defense security testing method is merely exemplary. Those skilled in the art can make various modifications based on the above method, and the modified solutions also fall within the protection scope of this application.
[0183] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the multi-agent collaborative federated learning attack and defense security testing method described above.
[0184] This application also provides an electronic device, including:
[0185] processor;
[0186] Memory is used to store computer programs that can be executed by a processor;
[0187] Among them, the steps of implementing the multi-agent collaborative federated learning attack and defense security testing method as described above are included when the processor executes the computer program.
[0188] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0189] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions.
[0190] It should be noted that any reference numerals placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In claims that enumerate several means, several of these means may be embodied by the same hardware. The use of the terms first, second, third, etc., is merely for convenience of expression and does not indicate any order. These terms can be understood as part of the component names.
[0191] Furthermore, it should be noted that in the description of this specification, the terms "one embodiment," "some embodiments," "embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0192] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the claims should be interpreted to include both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0193] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, then this invention should also include these modifications and variations.
Claims
1. A multi-agent collaborative federated learning attack and defense security testing method, characterized in that, include: The master orchestration agent receives the initial test task and breaks it down into multiple sub-tasks, including at least attack exploration sub-tasks and defense exploration sub-tasks. The attack strategy agent responds to the attack exploration subtask and generates local attack strategy patches in the candidate attack space; the defense strategy agent responds to the defense exploration subtask and generates local defense strategy patches in the candidate defense space. The main control orchestration agent performs conflict detection and arbitration on the attack strategy patch and the defense strategy patch, so as to merge them to generate an executable test plan for the current round. The evaluation agent schedules the federated learning environment to perform training experiments according to the executable test scheme, so as to collect multi-dimensional evaluation indicators and calculate a comprehensive score based on the multi-dimensional evaluation indicators. The attack strategy patch, defense strategy patch, executable test plan, multi-dimensional evaluation index and comprehensive score of the current round are stored as test memory tuples of the current round through the memory knowledge module, and an experience feedback vector is generated. The master orchestration agent determines whether the preset iteration termination condition is met based on the comprehensive score and the experience feedback vector. If the condition is not met, the attack strategy agent and the defense strategy agent are driven to perform the next round of iteration. If the condition is met, the iteration ends, and the security test result is determined based on the test memory tuples stored in previous iterations.
2. The federated learning attack and defense security testing method according to claim 1, characterized in that, The attack strategy patch includes attack type, attack scope, attack parameter strength, and attack timing arrangement; wherein, the attack type includes at least one of data poisoning, tag flipping, model poisoning, backdoor trigger injection, and client collusion attack.
3. The federated learning attack and defense security testing method according to claim 1, characterized in that, The defense strategy patch includes robust aggregation rules, abnormal client filtering rules, and defense sensitivity parameters; wherein, the robust aggregation rules include at least one of the following: Krum algorithm, multiple Krum algorithm, coordinate median algorithm, truncated mean algorithm, or weighted average mechanism based on confidence weight.
4. The federated learning attack and defense security testing method according to claim 3, characterized in that, The step of using the master orchestration agent to perform conflict detection and arbitration between the attack strategy patch and the defense strategy patch, in order to merge them to generate an executable test plan for the current round, includes: The domain isolation priority is adopted. For the server-side aggregation logic, the robust aggregation rules in the defense strategy patch are determined. For the malicious client local training logic, the attack parameter strength in the attack strategy patch is determined. When the expected values proposed by the attack strategy patch and the defense strategy patch for the same global hyperparameter are inconsistent, the weighted average of the two is calculated as the final value of the global hyperparameter, or the expected value of one of them is selected as the final value of the global hyperparameter according to the test objective in the initial test task. Based on the robust aggregation rules determined by the arbitration, the attack parameter strength, and the final values of each of the global hyperparameters, an executable test plan for the current round is generated.
5. The federated learning attack and defense security testing method according to claim 1, characterized in that, The process of collecting multi-dimensional evaluation indicators by executing an evaluation agent and calculating a comprehensive score based on the multi-dimensional evaluation indicators includes: Obtain the multidimensional evaluation metrics; wherein the multidimensional evaluation metrics include at least two of global accuracy, attack success rate, and communication and resource overhead; The multidimensional evaluation index is subjected to extreme value normalization to obtain the normalized index; The normalized indicators are weighted and fused to obtain the comprehensive score.
6. The federated learning attack and defense security testing method according to claim 1, characterized in that, The generation of the experience feedback vector includes: Based on the test memory tuples stored in consecutive rounds, the trend of the combined effect of the attack strategy patch and the defense strategy patch changing with the rounds is determined; Based on the changing trend, an empirical feedback vector is generated to adjust the search direction for the next round of strategies.
7. The federated learning attack and defense security testing method according to claim 1, characterized in that, The preset iteration termination conditions include: The current iteration round has reached the preset maximum iteration count threshold; or If, within a predetermined number of consecutive rounds, the absolute value of the difference between the comprehensive scores of adjacent rounds is less than a predetermined convergence threshold, the game is considered to have entered a relatively balanced state.
8. The federated learning attack and defense security testing method according to claim 1, characterized in that, Determining the security test result based on the test memory tuples stored in each iteration includes: Iterate through all test memory tuples stored in the memory knowledge module; The attack strategy patch that results in the highest overall score is extracted as the strongest attack path, and the defense strategy patch that results in the lowest overall score is extracted as the optimal defense strategy. Generate security test results that include the risk level of the federated learning system, the strongest attack path, and the optimal defense strategy; wherein, the risk level of the federated learning system is determined based on the highest attack success rate and the maximum global accuracy decrease recorded in each iteration.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the multi-agent collaborative federated learning attack and defense security testing method as described in any one of claims 1 to 8.
10. An electronic device, characterized in that, include: processor; A memory for storing computer programs that can be executed by the processor; The processor, when executing the computer program, implements the steps of the multi-agent collaborative federated learning attack and defense security testing method as described in any one of claims 1 to 8.