Algorithm robustness test method for multi-agent system and related equipment

By screening the target agent in the multi-agent reinforcement learning algorithm and generating adversarial perturbation values, the problem of limited robustness test performance and difficult deployment in the existing technology is solved, and effective robustness testing and performance evaluation of the agent team is achieved.

CN120012864APending Publication Date: 2025-05-16HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510177554.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When testing the robustness of multi-agent reinforcement learning algorithms, the prior art has problems such as limited performance, low attack immediacy and high deployment difficulty, especially in complex environments.

Method used

The robustness of the agent team is tested by filtering and determining the target agent in the agent model, generating anti-perturbation values ​​and adding them to the target agent's observations. This method includes two scenarios: ally observation attack and enemy observation attack, using action score values ​​and adversarial perturbation values ​​to mislead the agent's actions.

Benefits of technology

The robustness test of multi-agent systems is realized, which can effectively evaluate the performance and attack resistance of the agent team in complex environments, and the method is more realistic and concealed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012864A_ABST
    Figure CN120012864A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-agent system-oriented algorithm robustness test method and related equipment, which are used for realizing robustness test of agents. The method provided by the embodiment of the invention comprises the following steps: obtaining and forming an agent team set in a trained agent model; determining a target observation attack scene, performing observation attack on the to-be-screened agents in the agent team set based on the target observation attack scene, and screening and determining a target agent from all the to-be-screened agents; in the trained agent model, determining an action score value of any to-be-screened agent at the current moment, and generating a confrontation disturbance value for the trained agent model through the action score value; adding the confrontation disturbance value to the target agent so as to test the robustness of the target agent team set after the target agent team set is attacked; wherein the target agent team set is an agent team set comprising the target agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of information security technology, and in particular to an algorithm robustness testing method for a multi-agent system and related equipment. Background Art

[0002] There have been many studies on robustness testing of reinforcement learning algorithms, but only a few studies have considered multi-agent reinforcement learning environments. For example, early studies on robustness testing of multi-agent reinforcement learning applied adversarial attacks to multi-agent reinforcement learning algorithms. Some studies only attacked a fixed target agent in the agent team, but this approach may be limited in performance when the environment becomes more complex. Other studies attacked multiple fixed target agents. In such attacks, when the target agent changes, the adversary model needs to be retrained.

[0003] Recently, some works use multi-agent reinforcement learning algorithms or differential evolution algorithms to dynamically select target agents from the agent team, which means that the target agents can be plural and the targets can change over time. These methods can better test the robustness of multi-agent reinforcement learning algorithms because they may cause great harm to the team. However, their methods of selecting target agents and corresponding actions may affect the immediacy of attack implementation. In addition, some methods use adversarial samples to attack the observations of specific target agents. However, the limitations of using agent observations and the difficulty of deploying such white-box attacks have not been carefully considered. Generally speaking, the adversary should not have the observations of all agents at the beginning, and should not have too much information about the agent model before the game or task begins.

[0004] Therefore, there is an urgent need for a more reasonable reinforcement learning algorithm for intelligent agents to achieve robustness testing of intelligent agents. Summary of the invention

[0005] The embodiments of the present application provide an algorithm robustness testing method and related equipment for a multi-agent system, which are used to implement robustness testing of agents.

[0006] The first aspect of the embodiment of the present application provides an algorithm robustness testing method for a multi-agent system, comprising:

[0007] Acquire and form an agent team set in the trained agent model; wherein the agent team set includes a plurality of agents to be screened, and there is an association relationship between any two agents to be screened;

[0008] Determine a target observation attack scenario, and perform observation attacks on the to-be-screened agents in the agent team based on the target observation attack scenario, and screen and determine the target agent from all the to-be-screened agents;

[0009] In the trained agent model, determining the action score value of any agent to be screened at the current moment, and generating an adversarial disturbance value for the trained agent model through the action score value; wherein the action score value is obtained by observing the selected action of any agent to be screened in the agent team set;

[0010] The adversarial disturbance value is added to the target agent to obtain a trained agent model, so as to test the robustness of the target agent team set when the target agent team set is attacked; wherein the target agent team set is an agent team set including the target agent.

[0011] Optionally, if the target observation attack scenario is an ally observation attack scenario, the observation attack is performed on the to-be-screened intelligent agents in the intelligent agent team set based on the target observation attack scenario, and the target intelligent agent is screened and determined from all the to-be-screened intelligent agents, including:

[0012] Determine a first agent to be screened and a first local observation value corresponding to the first agent to be screened in the agent team set; wherein the first agent to be screened is the agent to be screened selected by the attacker when attacking the agent team set;

[0013] Determine other agents to be screened in the agent team set, excluding the first agent to be screened, and obtain other local observation values ​​of the other agents to be screened based on the first local observation value; wherein the other agents to be screened are located in the first local observation value;

[0014] Obtaining the first action score to be disturbed of the other to-be-screened intelligent agent through the other local observation values; wherein the first action score to be disturbed is used to represent the action score of the other to-be-screened intelligent agent when it is not disturbed;

[0015] Based on the ally observation attack algorithm, the attacker perturbs the other agents to be screened to obtain first perturbation action scores of all other agents to be screened after being perturbed, so as to determine first action score differences of all other agents to be screened according to the first perturbation action score and the first perturbation action score;

[0016] The target intelligent agent is determined from all other intelligent agents to be screened according to the first action score difference.

[0017] Optionally, if the target observation attack scenario is an enemy observation attack scenario, the observation attack is performed on the to-be-screened intelligent agents in the intelligent agent team set based on the target observation attack scenario, and the target intelligent agent is screened and determined from all the to-be-screened intelligent agents, including:

[0018] Obtaining a second local observation value when the attacker attacks the agent team set; wherein the second local observation value is an observation value observed by the attacker when observing the agent team set;

[0019] If there is an agent to be screened in the second local observation value, obtaining the second to-be-perturbed action score values ​​of all agents to be screened in the second local observation value;

[0020] Based on the enemy observation attack algorithm, the attacker perturbs all the agents to be screened to obtain second perturbation action scores of all the agents to be screened after being perturbed, so as to determine the second action score difference of all the agents to be screened according to the second perturbation action score and the second perturbation action score;

[0021] The target intelligent agent is determined from among all the intelligent agents to be screened according to the second action score difference.

[0022] Optionally, the method further comprises:

[0023] If the attacker fails to observe the agent to be screened during the enemy observation attack algorithm, or if the attacker fails, the enemy observation attack algorithm is stopped and the disturbance is stopped.

[0024] Optionally, determining the action score value of any agent to be screened at the current moment includes:

[0025] At the current moment, obtaining the current moment observation value of the trained agent model;

[0026] Input the current moment observation value into the trained agent model to obtain the current moment to-be-perturbed action score values ​​of all the agents to be screened in the agent team set; wherein any of the current moment to-be-perturbed action score values ​​is used to represent the action score of any agent to be screened under observation at the current moment;

[0027] The action score value having the maximum value among all the action score values ​​to be disturbed at the current moment is used as the action score value.

[0028] Optionally, generating an adversarial disturbance value for the trained agent model through the action score value includes:

[0029] Obtaining an input gradient value of the action score value for the trained agent model and a perturbation step length for the trained agent model;

[0030] The input gradient value is attacked using a white-box attack algorithm to obtain the adversarial disturbance value associated with the disturbance step size.

[0031] Optionally, the obtaining of an input gradient value of the action score value for the trained agent model includes:

[0032] Obtaining an observation value of any to-be-screened agent in the trained agent model; wherein the observation value of any to-be-screened agent corresponds to a feature value of multiple features;

[0033] Using the fast gradient sign algorithm in the white-box attack algorithm, a disturbance is added to any feature in the observation value of any of the agents to be screened, and the observed change value output by the trained agent model is determined;

[0034] Estimating based on the observed change value to obtain the input gradient value;

[0035] The fast gradient sign algorithm is: The X is the observed value of any agent to be screened, and the X is a multidimensional vector, the n is the vector dimension, and the X i is the i-th dimension of X, used to characterize the eigenvalue index of the observation, is the observed change value.

[0036] A second aspect of the embodiment of the present application provides an algorithm robustness testing system for a multi-agent system, comprising:

[0037] An acquisition unit, used for acquiring and forming an agent team set from the trained agent model; wherein the agent team set includes a plurality of agents to be screened, and there is an association relationship between any two agents to be screened;

[0038] A determination unit, used to determine a target observation attack scenario, and based on the target observation attack scenario, perform an observation attack on the to-be-screened intelligent agents in the intelligent agent team set, and screen and determine a target intelligent agent from all the to-be-screened intelligent agents;

[0039] A generating unit, configured to determine, in the trained agent model, an action score value of any agent to be screened at a current moment, and generate an adversarial perturbation value for the trained agent model through the action score value; wherein the action score value is obtained by observing a selected action of any agent to be screened in the agent team set;

[0040] The acquisition unit is also used to add the adversarial disturbance value to the target intelligent agent to obtain a trained intelligent agent model, so as to test the robustness of the target intelligent agent team set when the target intelligent agent team set is attacked; wherein, the target intelligent agent team set is an intelligent agent team set including the target intelligent agent.

[0041] The algorithm robustness testing system for multi-agent systems provided in the second aspect of an embodiment of the present application is used to execute the algorithm robustness testing method for multi-agent systems described in the first aspect.

[0042] A third aspect of the embodiment of the present application provides an algorithm robustness testing device for a multi-agent system, comprising:

[0043] CPU, memory, input and output interfaces, wired or wireless network interfaces, and power supply;

[0044] The memory is a short-term storage memory or a persistent storage memory;

[0045] The central processing unit is configured to communicate with the memory and execute instruction operations in the memory to perform the algorithm robustness testing method for multi-agent systems described in the first aspect.

[0046] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, which includes instructions. When the instructions are executed on a computer, the computer executes the algorithm robustness testing method for a multi-agent system described in the first aspect.

[0047] A fifth aspect of an embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed on a computer, the computer executes the algorithm robustness testing method for a multi-agent system described in the first aspect.

[0048] It can be seen from the above technical scheme that the embodiments of the present application have the following advantages: through an algorithm robustness testing method for a multi-agent system disclosed in the embodiments of the present application, by selecting an agent in the agent team as the target agent, and then generating adversarial disturbances through the previous and next action scores of any agent in the agent team in the agent model, and adding them to the observation of the target agent, the action of the agent is misled, thereby causing damage to the agent team, thereby completing the robustness test of the agent in the agent model, making the testing method more valuable. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0050] Figure 1 A flowchart of an algorithm robustness testing method for a multi-agent system disclosed in an embodiment of the present application;

[0051] Figure 2 A flowchart of another algorithm robustness testing method for a multi-agent system disclosed in an embodiment of the present application;

[0052] Figure 3 A flowchart of another algorithm robustness testing method for a multi-agent system disclosed in an embodiment of the present application;

[0053] Figure 4 A schematic diagram of an attack framework for a robustness test disclosed in an embodiment of the present application;

[0054] Figure 5 A schematic diagram of an intelligent agent selection for an ally observing attack disclosed in an embodiment of the present application;

[0055] Figure 6 A schematic diagram of an intelligent agent selection for enemy observation and attack disclosed in an embodiment of the present application;

[0056] Figure 7 A graph showing the change in team winning rate and reward of an intelligent agent team whose allies observe and attack, disclosed in an embodiment of the present application;

[0057] Figure 8 A graph showing the team winning rate and reward changes of an intelligent agent team observed and attacked by an enemy disclosed in an embodiment of the present application;

[0058] Fig. 9 A schematic diagram of the structure of an algorithm robustness testing system for a multi-agent system disclosed in an embodiment of the present application;

[0059] Fig.10 A schematic diagram of the structure of an algorithm robustness testing device for a multi-agent system disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0060] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0061] It should be noted that the descriptions involving "first", "second", etc. in this application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in this field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0062] It should be noted in advance that Centralized Training with Decentralized Execution (CTDE) is a common framework in the design of multi-agent reinforcement learning algorithms. In CTDE, agents can share their local information during training so that the algorithm can update the agent's strategy, and these strategies only need to use local observations to guide the corresponding agent to choose actions during execution. CTDE combines the advantages of centralized training and decentralized execution in this form. The QMIX algorithm is a successful implementation of the CTDE framework. It successfully solves some problems in the Value Decomposition Networks (VDN) algorithm. QMIX is a value-based multi-agent reinforcement learning algorithm (MARL) algorithm. QMIX uses the CTDE framework and uses a centralized hybrid network to receive global observations to guide the agents to update their model parameters. The hybrid network uses the Q value of the action selected by each agent to estimate the joint action value, which can then be used to update the network of each agent to help the team achieve better performance. Adversarial samples are designed to add tiny perturbations to the model input, causing the model to misclassify the input with high confidence. Adversarial attacks can be divided into two categories according to the implementation environment: white-box attacks and black-box attacks. In a white-box attack, the adversary needs to have enough knowledge of the model parameters and the gradient of the loss function to launch an attack. The Fast Gradient Sign Attack (FGSM) is a classic white-box attack that uses the gradient of the loss function with respect to the model input to quickly generate adversarial perturbations. For ease of understanding and description, the above-mentioned English abbreviations and related Chinese descriptions will not be repeated in the following.

[0063] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0064] To understand the robustness testing method of the intelligent agent proposed in the embodiment of the present application, please refer to Figure 4 , Figure 4 A schematic diagram of an attack framework for a robustness test disclosed in an embodiment of the present application. Figure 4 It can be seen that this robustness testing method mainly proposes a two-step attack framework.

[0065] Specifically, in the first attack, at a specific time, an agent is selected from the agent team (including agent 1, agent 2, ..., agent N) through a screening method as the target agent. Figure 4 In the example above, agent 1 is selected from agents 1-N as the target agent (i.e., the victim agent). The second step is to generate adversarial perturbations to add to the observations of the target agent (e.g. Figure 4 Specifically, we can use the local observation value of the agent and use the black box method to generate adversarial perturbations, which will eventually mislead the target agent to choose actions that deviate from the original action strategy. It can be understood that, among them, o n That is, the observation value of the corresponding agent N, and the corresponding o v is the added adversarial perturbation.

[0066] It is further understood that the purpose of the first attack is to select the target agent, so as to add adversarial disturbances in the next step. In order to make the attack concealed, destructive and feasible, the algorithm needs to have the following properties: the number of attacks launched is as small as possible; the reduction in team rewards is as large as possible; the implementation conditions are as realistic as possible. Therefore, the embodiment of the present application mainly uses the properties of multi-agent reinforcement learning itself to select the optimal attack time and victim agent with as few conditions as possible. For this purpose, two scenarios are designed: ally observation attack and enemy observation attack. For further understanding and explanation, please refer to Figure 1 , Figure 2 and Figure 3 The embodiment shown.

[0067] Figure 1 The flowchart of a method for testing the robustness of an algorithm for a multi-agent system disclosed in an embodiment of the present application includes steps 101 to 104.

[0068] 101. Obtain and form an agent team set from the trained agent model.

[0069] Since in reinforcement learning, the input of the agent model is the observation value, and the output is the score of the agent for each action, then in testing the robustness of the agent, it is necessary to first define the agent model, and thereby obtain and form the agent team set. It should be noted that the agent team set includes multiple agents to be screened, and there is an association relationship between any two agents to be screened.

[0070] Furthermore, in this embodiment, it is not necessary to consider information such as model parameters of the agent model. At the same time, the agent team set can also be described as an agent team, where the agent team is Figure 4 As shown, it consists of agents 1-N.

[0071] 102. Determine a target observation attack scenario, and conduct observation attacks on the intelligent agents to be screened in the intelligent agent team based on the target observation attack scenario, and screen and determine the target intelligent agent from all the intelligent agents to be screened.

[0072] Then, after determining the agent team, it is necessary to train the method of screening the target agent under different observation and attack scenarios. Specifically, it is necessary to first determine the target observation and attack scenario at this time, which includes at least the ally observation and attack scenario (AOA, Ally ObservationAttack) or the enemy observation and attack scenario (EOA, EnemyObservationAttack). Then, combined with the above-mentioned target observation and attack scenarios, the target observation and attack scenarios can be used to concentrate the observation and attack of different agents to be screened on the agent team, so as to screen and determine the target agent among all the agents to be screened.

[0073] In one specific embodiment, in the AOA scenario, the attacker can obtain the observation value of an agent in the agent team, and then obtain the observation of other agents to be screened in the observation of the agent, and then conduct observation attacks on other agents to be screened, so as to select one of the agents to be screened as the target agent. Figure 2 Steps 202 to 205 are examples of the embodiment shown.

[0074] In other feasible technical solutions, in the EOA scenario, the information of the agent can only be obtained through the enemy's observation. Then, the attacker can obtain the observation value of the attacker. It should be noted that the observation value includes the agent team (including the agent to be screened). Then, the attacker observes and attacks the agent team, and then selects a candidate agent from the agent team as the target agent. For specific selection methods, please refer to Figure 2 Steps 206 to 209 are examples of the embodiment.

[0075] 103. In the trained intelligent agent model, determine the action score value of any intelligent agent to be screened at the current moment, and generate an adversarial disturbance value for the trained intelligent agent model through the action score value.

[0076] Therefore, after determining the target agent, we can select any agent to be screened in the agent model, thereby defining the action score value of the action selected by the agent to be screened at the current moment, and then, we can generate the adversarial perturbation value corresponding to the agent model through the action score value. It should be noted that the action score value is obtained by observing the selected action of any agent to be screened in the agent team.

[0077] In one of the specific embodiments, it can be understood that since the traditional method of generating adversarial perturbations includes the FGSM algorithm, etc., the algorithm requires the gradient of the model loss function. In multi-agent reinforcement learning, the loss function related to the team reward is only used in training and cannot be obtained during the execution process. In order to generate adversarial samples more reasonably, a black-box adversarial interference method designed for the action score of the agent can be used. Specifically, in reinforcement learning, the input of the agent model is the observation value at the current moment, and the output is the score of the Q function for each action, that is, the Q value (action score value). The value (that is, the maximum value) of the action selected by the kth agent is defined as Q k , and then use the FGSM method to attack the gradient of the agent's Q value with respect to the observation to obtain the adversarial perturbation value.

[0078] 104. Add the adversarial perturbation value to the target agent to obtain a trained agent model, so as to test the robustness of the target agent team set when the target agent team set is attacked.

[0079] Thus, combined with the above step 103, the adversarial disturbance value can be added to the observation of the target agent, so that a trained agent model can be obtained, so that when the target agent team set is attacked, the robustness of the agent team at this time can be tested. It should be noted that the target agent team set is the agent team set including the target agent.

[0080] In one specific embodiment, the generated adversarial perturbation value can be added to the observation value of the target intelligent agent, so that the action of the target intelligent agent is misled. Furthermore, since the observation values ​​of other to-be-screened intelligent agents in the intelligent agent team set are included in the observation value of the target intelligent agent, the observation values ​​of other to-be-screened intelligent agents will also be disturbed, thereby causing damage to the team, thereby forming the target intelligent agent team set described above.

[0081] Furthermore, after executing step 104, experiments are conducted in a multi-agent test platform environment (SMAC, StarCraft Multi-Agent Challenge). By attacking the agent team trained using the QMIX algorithm on different maps, the changes in the team winning rate and rewards of the agent team after being attacked are obtained, thereby further verifying the performance of the method.

[0082] The present embodiment discloses an algorithm robustness testing method for a multi-agent system. An agent in an agent team is selected as the target agent. Then, in the agent model, adversarial perturbations are generated based on the previous and next action scores of any agent in the agent team and added to the observations of the target agent. This misleads the actions of the agent, thereby damaging the agent team. This completes the robustness test of the agent in the agent model, making the testing method more valuable.

[0083] For the above Figure 1 For a detailed description of step 102, please refer to Figure 2 Figure 2 This is a flow chart of another algorithm robustness testing method for a multi-agent system disclosed in an embodiment of the present application, including steps 201 to 209.

[0084] 201. Determine the target observation attack scenario.

[0085] As shown in the above step 102, the target observed attack scenario includes at least AOA and EOA scenarios. Furthermore, in this embodiment, it is necessary to preferentially determine which implementation scenario the observed attack scenario belongs to, AOA or EOA.

[0086] 202. Determine a first agent to be screened and a first local observation value corresponding to the first agent to be screened in the agent team.

[0087] It should be noted that, in this embodiment, steps 202 to 205 are a method for screening target agents in the AOA scenario. Specifically, in this embodiment, an agent to be screened can be selected from the agent team set, and at the same time, the local observation value of the agent to be screened can be obtained. Among them, the agent to be screened at this time is named the first agent to be screened for easy distinction. It should be noted that the first agent to be screened is the agent to be screened selected by the attacker when attacking the agent team set.

[0088] In one specific embodiment, see Figure 5 , Figure 5 This is a schematic diagram of an intelligent agent selection for an ally observing attack disclosed in an embodiment of the present application. Figure 5 As can be seen in , the attacker can obtain the local observation value of at least one agent to be screened in the agent team set by attacking the agent team set in the agent model. Figure 5 In the example, the agent team set includes at least agents 1-4, among which the agent attacked by the attacker is listed as agent 1, i.e., the first agent to be screened.

[0089] Furthermore, this can also be described as the attacker (or adversary) being able to control an agent in the agent team. t The observed value o t It can be understood that at this time t=1.

[0090] 203. Determine the agent team concentration, exclude other agents to be screened except the first agent to be screened, and obtain other local observation values ​​of other agents to be screened based on the first local observation value.

[0091] Then, it is possible to determine other agents to be screened in the agent team set except the first agent to be screened, and then, through the first local observation value, it is possible to obtain the local observation values ​​of other agents to be screened, that is, other local observation values. It should be noted that the other agents to be screened are located in the first local observation value.

[0092] In one specific embodiment, in combination Figure 5 It can be seen that Agent 2 and Agent 3 are other agents in Agent 1's observation, but they are not in Agent 1's observation. Furthermore, it can be known that at this time, Agent 2 and Agent 3 are allied agents with Agent 1. Therefore, it can be known that after the attacker (or adversary) has mastered Agent 1 and its observation, he can then obtain the other allied agents (agents) in Agent 1's observation. i ,iin o t ) observations. Furthermore, the local observation values ​​of other ally agents at this time can also be obtained.

[0093] 204. Obtain the first disturbed action scores of other to-be-screened intelligent agents through other local observation values, and based on the ally observation attack algorithm, disturb the other to-be-screened intelligent agents through the attacking party to obtain the first disturbed action scores of all other to-be-screened intelligent agents after being disturbed, so as to determine the first action score difference of all other to-be-screened intelligent agents according to the first to-be-disturbed action score and the first disturbed action score.

[0094] Then, by obtaining other local observation values, the first action score to be disturbed of other intelligent agents to be screened can be obtained, and then based on the ally observation attack algorithm, the other intelligent agents to be screened can be disturbed by the attacker, so as to obtain the first disturbed action score of all other intelligent agents to be screened after being disturbed. Finally, the difference in the first action score of all other intelligent agents to be screened is determined according to the first action score to be disturbed and the first disturbed action score. It should be noted that the first action score to be disturbed is used to characterize the action score of other intelligent agents to be screened when they are not disturbed.

[0095] In one specific embodiment, the action scores of other agents to be screened in the preliminary observation (not yet disturbed) are first obtained, that is, the first action score to be disturbed. This action score can also be understood as the Q value in the QMIX algorithm, that is, the action score of a certain action under the current observation. Then, the attacker interferes with other agents to be screened (the specific interference value can be found in Figure 3 The embodiment shown in the figure) is used to obtain the action scores of the other agents to be screened after being disturbed (already disturbed), that is, the first disturbed action score. Then, by comparing the difference between the action scores before and after, the first action score to be disturbed and the first disturbed action score with the largest difference before and after are determined, and the first action score difference between the two is determined.

[0096] Combined with the above QMIX algorithm, define Q i It is the Q value of the action selected by agent i before being disturbed (actually the maximum Q value), and Q is defined as i ′ is the Q value of the action selected by agent i after being disturbed, ΔQ i =Q i -Q i ′ is the difference in Q value before and after the agent i is disturbed, that is, the difference in the score of the first action.

[0097] 205. Determine a target intelligent agent from all other intelligent agents to be screened based on the first action score difference.

[0098] Furthermore, based on the first action score difference, the to-be-screened intelligent agent that satisfies the first action score difference among all other to-be-screened intelligent agents can be selected as the target intelligent agent.

[0099] In one specific embodiment, the agent with the largest Q value difference can be selected as the target agent.

[0100] Furthermore, in the AOA method, the selected agent stops attacking when there are no other agents in the observation or when the selected agent dies, so it is a sparse attack method. Figure 5 As shown, the attacker has the observation of agent 1, so the attack stops when there are no other agents in the observation of agent 1; or when agent 1 dies, the attack also stops.

[0101] 206. Obtain a second local observation value when the attacker attacks the team set of intelligent agents.

[0102] It should be noted that, in this embodiment, steps 203 to 209 are a method for screening target agents in the EOA scenario. Specifically, in this embodiment, due to the enemy's observation attack, it is initially impossible to obtain the observation value of any agent, and the agent's information can only be obtained through the enemy's observation. Therefore, it is necessary to obtain the second local observation value of the attacker when attacking the agent team set. It should be noted that, in this embodiment, the second local observation value is the observation value observed by the attacker when observing the agent team set. The second local observation value is not directly related to the first local observation value, but is only used to distinguish the observation value. For the convenience of understanding and description, it will not be repeated later.

[0103] In one specific embodiment, see Figure 6 , Figure 6 This is a schematic diagram of an intelligent agent selection for enemy observation and attack disclosed in an embodiment of the present application. Figure 6 As can be seen in , the attacker can observe the local observation value of at least one to-be-screened agent in the agent team set by attacking the agent team set in the agent model. Figure 6 In the example, the agent team set includes at least agents 1-N, where the agents observed by the attacker are listed as agent 1 or agent 2, etc., which are not described in detail here. It should also be noted that at this time, the attacker (or adversary) can also have a local observation value of attacker 1, that is, the observed agent 1 or agent 2, etc. It should also be noted that since there may be multiple attackers, for example, it can also include attacker 2, ..., attacker N, which is not limited here. For the convenience of description, attacker 1 is used for detailed description later.

[0104] Furthermore, this can also be described as obtaining the attacker's (or adversary's) observation value o e .

[0105] 207. If there are agents to be screened in the second local observation value, obtain the second action to be disturbed score values ​​of all agents to be screened in the second local observation value.

[0106] Furthermore, based on step 206, when there are agents to be screened in the second local observation value, the second to-be-perturbed action scores of all agents to be screened in the second local observation value can be obtained.

[0107] In one specific embodiment, if an agent (in this embodiment, agent 1 and agent 2) appears in the second local observation value of the attacker, that is, in o e Agents appear in e), the second action score to be disturbed (undisturbed) of agent 1 and agent 2 in the second local observation value can be calculated, that is, the undisturbed Q value at this time.

[0108] 208. Based on the enemy observation attack algorithm, all the intelligent agents to be screened are disturbed by the attacking party to obtain the second disturbance action score values ​​of all the intelligent agents to be screened after being disturbed, so as to determine the second action score difference of all the intelligent agents to be screened according to the second action score value to be disturbed and the second disturbance action score value.

[0109] Furthermore, based on the enemy observation attack algorithm, all the intelligent entities to be screened can be disturbed by the attacker, and the second disturbance action score value of all the intelligent entities to be screened after being disturbed can be obtained. Then, the difference in the second action score of all the intelligent entities to be screened can be determined according to the second action score value to be disturbed and the second disturbance action score value.

[0110] In one specific embodiment, using Figure 3 The disturbance generated in the manner shown in the embodiment disturbs agent 1 and agent 2, and then calculates the second disturbance action score value after the disturbance of agent 1 and agent 2. Then, by comparing the difference between the action score values ​​before and after, the second action score value to be disturbed and the second disturbance action score value with the largest difference before and after are determined, and the difference between the second action scores of the two is determined. It should be noted that the method of calculating the action score difference in this embodiment is similar to step 204, and will not be repeated here.

[0111] It should also be noted that, in this embodiment, it can also be understood that at this time the attacker does not have any observation value of any intelligent agent, but has observation value of attacker 1. At this time, in the intelligent agent team composed of agents 1-N, agents 1 and 2 are under the observation of attacker 1. Therefore, through the selection method, the final target intelligent agent is selected from agents 1 and 2.

[0112] 209. Determine a target intelligent agent from among all intelligent agents to be screened according to the second action score difference.

[0113] Therefore, the target agent can be determined from all agents to be screened through the difference in the second action score.

[0114] In one specific embodiment, the agent with the largest Q value difference can be selected as the target agent.

[0115] Furthermore, in the EOA method, the selected agent stops attacking when there are no other agents in the observation, or when the selected agent dies, so it is a sparse attack method. Figure 6As shown in the figure, when there is no agent in the attacker's observation or the enemy dies, the EOA method will not launch an attack, so it is a sparse attack. That is, the attacker has the observation value of attacker 1 (which can also be distinguished as enemy 1), so when there are no other agents in the observation value of enemy 1, the attack will stop; or when enemy 1 dies, the attack will also stop.

[0116] Through the algorithm robustness testing method for a multi-agent system disclosed in this embodiment, the attack time can be screened through the local observation of the agent in the agent team or the local observation of the enemy, and the target agent can be effectively selected. In this way, the requirement of minimizing the disturbance to the target agent observation can be met. At the same time, the use of the agent observation value is strictly limited, making the method more practical, and the robustness test is completed under realistic conditions, making the test method more valuable.

[0117] For the above Figure 1 For a detailed description of step 103, please refer to Figure 3 , Figure 3 This is a flow chart of another algorithm robustness testing method for a multi-agent system disclosed in an embodiment of the present application, including steps 301 to 303.

[0118] 301. In the trained intelligent agent model, at the current moment, obtain the current moment observation value of the trained intelligent agent model, and input the current moment observation value into the trained intelligent agent model to obtain the current moment disturbed action score value of all the intelligent agents to be screened in the intelligent agent team.

[0119] This embodiment mainly describes the method of generating adversarial disturbances. Specifically, in the agent model, it is necessary to obtain the current moment observation value of the agent model at the current moment, and input the current moment observation value into the agent model, so as to obtain the current moment to be disturbed action score value of all the agents to be screened in the agent team. It should be noted that the score value of the action to be disturbed at any current moment is used to characterize the action score of any agent to be screened under the observation at the current moment. It should also be noted in advance that traditional methods for generating adversarial disturbances include FGSM, etc., but most of them require the gradient of the model loss function. In multi-agent reinforcement learning, the loss function of the team reward is only used in training and cannot be obtained during the execution process. In order to generate adversarial samples more reasonably, this embodiment can also be understood as a black box adversarial disturbance method.

[0120] In one of the specific embodiments, in an actual black box scenario, since the gradient of the action score (Q value) cannot be obtained through the parameters of the agent model, it is necessary to design an estimate of the gradient. Therefore, it is necessary to obtain the observation value at the current moment in the agent model, that is, the observation value at the current moment, and then use the current moment observation value as the input of the agent model to output the score of the Q function (QMIX function) for each action. Thus, the score value of the action to be disturbed at the current moment of all the agents to be screened is obtained. It should be noted that the score value of the action to be disturbed can also be understood as the Q value (that is, the score) of the action selected by the agent.

[0121] 302. The action score value having the maximum value among all the action score values ​​to be disturbed at the current moment is taken as the action score value.

[0122] Therefore, based on step 301, the action score value with the maximum value among all the action score values ​​to be disturbed at the current moment can be used as the action score value.

[0123] In one specific embodiment, the Q value (that is, the maximum Q value) of the action selected by the kth agent can be defined as Q k .

[0124] 303. Obtain the input gradient value of the action score value for the trained intelligent agent model and the perturbation step size for the trained intelligent agent model, and use a white box attack algorithm to attack the input gradient value to obtain an adversarial perturbation value associated with the perturbation step size.

[0125] Furthermore, the input gradient value of the action score value for the intelligent agent model and the perturbation step size of the intelligent agent model can be obtained, so that the input gradient value can be attacked using the white box attack algorithm to obtain the adversarial perturbation value associated with the perturbation step size.

[0126] In one specific embodiment, in the process of obtaining the input gradient value of the agent model in the action score value, the observation value of any agent to be screened in the agent model can also be obtained. Among them, the observation value of any agent to be screened corresponds to the feature values ​​of multiple features. Then, using the fast gradient sign algorithm in the white box attack algorithm, a disturbance is added to any feature in the observation value of any agent to be screened to determine the observation change value output by the agent model; finally, an estimation is performed based on the observation change value to obtain the input gradient value. Among them, the fast gradient sign algorithm is X is the observation value of any agent to be screened, and X is a multidimensional vector, n is the vector dimension, X i is the i-th dimension of X, used to characterize the eigenvalue index of the observation, is the observed change value.

[0127] Furthermore, combined with the above description, it can be understood that the subscripts 1, 2, ..., n of X refer to the observed feature value index. It can be simply understood that X is the observed value of the agent, X is a multidimensional vector, assuming there is an n-dimensional vector, X i represents the i-th dimension of the observation X, where i is the feature index. Then, by adding a small perturbation to the i-th feature of the observation and observing the change in the Q value output at this time, we can get To estimate the gradient. For example, in combination with the above step 302, define the Q value of agent k as Q k , then the i-th feature of the observation value X of agent k, that is, X i If we add a disturbance to the Q value of agent k, the Q value of agent k will change, and the change value is ΔQ k , using ΔQ k Divide by X i The disturbance added to the target agent is Δx. Then the FGSM formula can be used to generate adversarial disturbances and added to the target agent.

[0128] Specifically, there are It should be noted that ε is the size of the perturbation step, ▽ X Q k is the gradient of the Q value with respect to the model input. sgn represents the sign function, d pertubation (X) is the calculated adversarial perturbation. It should also be noted that the perturbation step size is the size of the adversarial perturbation, which is determined by the attacker.

[0129] Thus, after generating the adversarial perturbation, we can perform Figure 1 Step 104 shown. Furthermore, after completing steps 101 to 104, experiments can also be conducted in the SMAC environment. By attacking the agent team trained using the QMIX algorithm on different maps, the changes in the team win rate and rewards of the agent team after being attacked are obtained to verify the performance of the method. We also designed ablation experiments for each component of the method to verify the effectiveness of the method. Specifically, QMIX generates a loss function for team rewards during training, but after the training is completed, the loss function for team rewards does not exist during actual execution, so the loss function for team rewards cannot be used during the attack.

[0130] Therefore, for ease of understanding and description, please refer to Figure 7 and Figure 8 ,in, Figure 7 A graph showing the change in team winning rate and reward of an intelligent agent team whose allies observe and attack, disclosed in an embodiment of the present application; Figure 8This is a graph of the team winning rate and reward changes of an intelligent agent team under enemy observation and attack disclosed in an embodiment of the present application. Figure 7 include Figure 7 (a) and Figure 7 (b) Figure 7 (a) is a graph showing the change in team win rate under ally observation attack. Figure 7 (b) is a graph showing the changes in team rewards under ally observation attacks. Figure 8 include Figure 8 (a) and Figure 8 (b) Figure 8 (a) is a graph showing the change in team win rate under enemy observation attack. Figure 8 (b) is a graph showing the changes in team rewards under enemy observation attacks.

[0131] Among them, Figure 7 In the example, we select two maps, 2s3z and MMM, under the SMAC environment to attack the agent team trained with the QMIX algorithm. Among them, only one agent is selected as the target agent. Figure 7 In (a), the AOA method can reduce the winning rate of the agent team from 100% to 4% in the 2s3z map, and from 99% to 0% in the MMM map. Figure 7 In (b), the team reward drops from nearly 20 to 11.6 in the 2s3z map and from nearly 20 to 6.1 in the MMM map. These results show that the AOA method can cause great damage to the agent team using partial observations.

[0132] In the EOA method, an experiment was also conducted to select all agents within the agent range as targets for attack. Figure 8 In (a), the winning rate of the agent team in the 2s3z map decreases from 100% to 0%, and the winning rate of the agent team in the MMM map changes from 99% to 10%. Figure 8 In (b), the team reward drops from nearly 20 to 8.3 in the 2s3z map and from nearly 20 to 10.3 in the MMM map. This shows that the attack effect can be greatly improved by sacrificing the stealth of the attack.

[0133] The algorithm robustness testing method for a multi-agent system disclosed in this embodiment, in addition to minimizing the requirement of disturbance to the observation of the target agent, also considers query-based attacks and alternative model attacks to generate black-box adversarial disturbances, which is the first work to conduct black-box adversarial attacks in a multi-agent environment. Furthermore, various experiments were conducted in the embodiments of the present application, showing that the technical solution of the present application can greatly reduce the winning rate and team rewards of the agent team trained by QMIX with a lower disturbance step size. The effectiveness of each component of the proposed method was also verified through ablation studies. At the same time, it also combines concealment and destructiveness, which is more realistic in implementation, making it suitable for robustness testing of models trained by MARL algorithms. In short, the query-based attack method proposed in this embodiment is a black-box adversarial attack method for a multi-agent environment, which is both innovative and effective.

[0134] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0135] See also Fig. 9 , Fig. 9 A schematic diagram of the structure of an algorithm robustness testing system for a multi-agent system disclosed in an embodiment of the present application.

[0136] The acquisition unit 901 is used to acquire and form an agent team set from the trained agent model; wherein the agent team set includes a plurality of agents to be screened, and there is an association relationship between any two agents to be screened;

[0137] A determination unit 902 is used to determine a target observation attack scenario, and based on the target observation attack scenario, perform observation attacks on the to-be-screened agents in the agent team, and screen and determine the target agent from all the to-be-screened agents;

[0138] The generating unit 903 is used to determine the action score value of any agent to be screened at the current moment in the trained agent model, and generate an adversarial perturbation value for the trained agent model through the action score value; wherein the action score value is obtained by observing the selected action of any agent to be screened in the agent team set;

[0139] The acquisition unit 901 is also used to add the adversarial disturbance value to the target intelligent agent to obtain a trained intelligent agent model, so as to test the robustness of the target intelligent agent team set when the target intelligent agent team set is attacked; wherein, the target intelligent agent team set is an intelligent agent team set including the target intelligent agent.

[0140] Exemplarily, when the target observed attack scenario is an ally observed attack scenario, the system includes:

[0141] The determining unit 902 is specifically used to determine the first agent to be screened and the first local observation value corresponding to the first agent to be screened in the agent team set; wherein the first agent to be screened is the agent to be screened selected by the attacker when attacking the agent team set;

[0142] The determination unit 902 is further used to determine other agents to be screened in the agent team set except the first agent to be screened, and obtain other local observation values ​​of other agents to be screened based on the first local observation value; wherein the other agents to be screened are located in the first local observation value;

[0143] The acquisition unit 901 is specifically used to acquire the first action score to be disturbed of other agents to be screened through other local observation values; wherein the first action score to be disturbed is used to represent the action score of other agents to be screened when they are not disturbed;

[0144] The acquisition unit 901 is further used to perturb other agents to be screened through the attacking party based on the ally observation attack algorithm to obtain the first perturbation action score values ​​of all other agents to be screened after being perturbed, so as to determine the first action score difference of all other agents to be screened according to the first perturbation action score value and the first perturbation action score value;

[0145] The determination unit 902 is further configured to determine a target agent from among all other agents to be screened according to the first action score difference.

[0146] Exemplarily, when the target observation attack scenario is an enemy observation attack scenario, the system includes:

[0147] The acquisition unit 901 is specifically used to acquire a second local observation value when the attacker attacks the agent team set; wherein the second local observation value is an observation value observed by the attacker when observing the agent team set;

[0148] The acquisition unit 901 is further configured to acquire the second to-be-perturbed action score values ​​of all the to-be-screened agents in the second local observation value when there are to-be-screened agents in the second local observation value;

[0149] The acquisition unit 901 is further used to perturb all the agents to be screened through the attacking party based on the enemy observation attack algorithm to obtain the second perturbation action score values ​​of all the agents to be screened after being perturbed, so as to determine the second action score difference of all the agents to be screened according to the second perturbation action score value and the second perturbation action score value;

[0150] The determination unit 902 is specifically configured to determine a target agent from among all agents to be screened according to the second action score difference.

[0151] Exemplarily, the system further includes: a stopping unit 904;

[0152] The stopping unit 904 is used to stop the enemy observation attack algorithm and stop the disturbance when the attacker fails to observe the agent to be screened or the attacker fails in the enemy observation attack algorithm.

[0153] Exemplarily, the system further includes: a setting unit 905;

[0154] The acquisition unit 901 is specifically used to acquire the current moment observation value of the trained agent model at the current moment;

[0155] The acquisition unit 901 is also used to input the current moment observation value into the trained agent model to obtain the current moment to-be-perturbed action score value of all the agents to be screened in the agent team set; wherein any current moment to-be-perturbed action score value is used to represent the action score of any agent to be screened under observation at the current moment;

[0156] The setting unit 905 is specifically configured to take the action score value with the maximum value among all the action score values ​​to be disturbed at the current moment as the action score value.

[0157] Exemplarily, the system includes:

[0158] An acquisition unit 901 is specifically used to acquire an input gradient value of the action score value for the trained agent model and a perturbation step length of the trained agent model;

[0159] The acquisition unit 901 is further used to attack the input gradient value using a white box attack algorithm to obtain an adversarial disturbance value associated with the disturbance step size.

[0160] Exemplarily, the system includes:

[0161] The acquisition unit 901 is specifically used to acquire the observation value of any agent to be screened in the trained agent model; wherein the observation value of any agent to be screened corresponds to the feature values ​​of multiple features;

[0162] The determination unit 902 is specifically used to use the fast gradient sign algorithm in the white box attack algorithm to add a disturbance to any feature in the observation value of any agent to be screened, and determine the observation change value output by the trained agent model;

[0163] The acquisition unit 901 is further used to estimate based on the observed change value to obtain the input gradient value;

[0164] The fast gradient sign algorithm is X is the observation value of any agent to be screened, and X is a multidimensional vector, n is the vector dimension, X i is the i-th dimension of X, used to characterize the eigenvalue index of the observation, is the observed change value.

[0165] See below Fig.10 The structural diagram of an algorithm robustness testing device for a multi-agent system disclosed in an embodiment of the present application includes:

[0166] CPU 1001, memory 1005, input / output interface 1004, wired or wireless network interface 1003 and power supply 1002;

[0167] The memory 1005 is a temporary storage memory or a permanent storage memory;

[0168] The CPU 1001 is configured to communicate with the memory 1005 and execute the instructions in the memory 1005 to perform the aforementioned Figures 1 to 3 An algorithm robustness testing method for a multi-agent system in any of the illustrated embodiments.

[0169] The embodiment of the present application also provides a chip system, the chip system includes at least one processor and a communication interface, the communication interface and the at least one processor are interconnected through a line, and the at least one processor is used to run a computer program or instruction to execute the aforementioned Figures 1 to 3 An algorithm robustness testing method for a multi-agent system in any of the illustrated embodiments.

[0170] The embodiment of the present application also provides a computer-readable storage medium, the computer-readable storage medium includes instructions, when the instructions are executed on a computer, the computer executes the aforementioned Figures 1 to 3 An algorithm robustness testing method for a multi-agent system in any of the illustrated embodiments.

[0171] The present application also provides a computer program product including instructions, which, when executed on a computer, enables the computer to execute the aforementioned Figure 1 An algorithm robustness testing method for a multi-agent system in the illustrated embodiment.

[0172] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0173] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0174] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0175] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0176] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk and other media that can store program code.

Claims

1. A method for testing algorithm robustness of multi-agent systems, characterized in that: The method comprises: Acquire and form an agent team set in the trained agent model; wherein the agent team set includes a plurality of agents to be screened, and there is an association relationship between any two agents to be screened; Determine a target observation attack scenario, and perform observation attacks on the to-be-screened agents in the agent team based on the target observation attack scenario, and screen and determine the target agent from all the to-be-screened agents; In the trained agent model, determining the action score value of any agent to be screened at the current moment, and generating an adversarial disturbance value for the trained agent model through the action score value; wherein the action score value is obtained by observing the selected action of any agent to be screened in the agent team set; The adversarial disturbance value is added to the target agent to test the robustness of the target agent team set when the target agent team set is attacked; wherein the target agent team set is an agent team set including the target agent.

2. The algorithm robustness testing method for multi-agent systems according to claim 1, characterized in that: If the target observation attack scenario is an ally observation attack scenario, the observation attack is performed on the to-be-screened intelligent agents in the intelligent agent team based on the target observation attack scenario, and the target intelligent agent is screened and determined from all the to-be-screened intelligent agents, including: Determine a first agent to be screened and a first local observation value corresponding to the first agent to be screened in the agent team set; wherein the first agent to be screened is the agent to be screened selected by the attacker when attacking the agent team set; Determine other agents to be screened in the agent team set, excluding the first agent to be screened, and obtain other local observation values ​​of the other agents to be screened based on the first local observation value; wherein the other agents to be screened are located in the first local observation value; Obtaining the first action score to be disturbed of the other to-be-screened intelligent agent through the other local observation values; wherein the first action score to be disturbed is used to represent the action score of the other to-be-screened intelligent agent when it is not disturbed; Based on the ally observation attack algorithm, the attacker perturbs the other agents to be screened to obtain first perturbation action scores of all other agents to be screened after being perturbed, so as to determine first action score differences of all other agents to be screened according to the first perturbation action score and the first perturbation action score; The target intelligent agent is determined from all other intelligent agents to be screened according to the first action score difference.

3. The algorithm robustness testing method for multi-agent systems according to claim 1, characterized in that: If the target observation attack scenario is an enemy observation attack scenario, the observation attack is performed on the to-be-screened intelligent agents in the intelligent agent team set based on the target observation attack scenario, and the target intelligent agent is screened and determined from all the to-be-screened intelligent agents, including: Obtaining a second local observation value when the attacker attacks the agent team set; wherein the second local observation value is an observation value observed by the attacker when observing the agent team set; If there is an agent to be screened in the second local observation value, obtaining the second to-be-perturbed action score values ​​of all agents to be screened in the second local observation value; Based on the enemy observation attack algorithm, the attacker perturbs all the agents to be screened to obtain second perturbation action scores of all the agents to be screened after being perturbed, so as to determine the second action score difference of all the agents to be screened according to the second perturbation action score and the second perturbation action score; The target intelligent agent is determined from among all the intelligent agents to be screened according to the second action score difference.

4. The algorithm robustness testing method for multi-agent systems according to claim 3 is characterized in that: The method further comprises: If the attacker fails to observe the agent to be screened during the enemy observation attack algorithm, or if the attacker fails, the enemy observation attack algorithm is stopped and the disturbance is stopped.

5. The algorithm robustness testing method for multi-agent systems according to claim 1, characterized in that: The step of determining the action score of any agent to be screened at the current moment includes: At the current moment, obtaining the current moment observation value of the trained agent model; Input the current moment observation value into the trained agent model to obtain the current moment to-be-perturbed action score values ​​of all the agents to be screened in the agent team set; wherein any of the current moment to-be-perturbed action score values ​​is used to represent the action score of any agent to be screened under observation at the current moment; The action score value having the maximum value among all the action score values ​​to be disturbed at the current moment is used as the action score value.

6. The algorithm robustness testing method for multi-agent systems according to claim 1, characterized in that: The step of generating an adversarial disturbance value for the trained agent model through the action score value includes: Obtaining an input gradient value of the action score value for the trained agent model and a perturbation step length for the trained agent model; The input gradient value is attacked using a white-box attack algorithm to obtain the adversarial disturbance value associated with the disturbance step size.

7. The algorithm robustness testing method for multi-agent systems according to claim 6, characterized in that: The obtaining of the input gradient value of the action score value for the trained agent model includes: Obtaining an observation value of any to-be-screened agent in the trained agent model; wherein the observation value of any to-be-screened agent corresponds to a feature value of multiple features; Using the fast gradient sign algorithm in the white-box attack algorithm, a disturbance is added to any feature in the observation value of any of the agents to be screened, and the observed change value output by the trained agent model is determined; Estimating based on the observed change value to obtain the input gradient value; The fast gradient sign algorithm is: The X is the observed value of any agent to be screened, and the X is a multidimensional vector, the n is the vector dimension, and the X i is the i-th dimension of X, used to characterize the eigenvalue index of the observation, is the observed change value.

8. An algorithm robustness testing system for multi-agent systems, characterized in that: The algorithm robustness testing system comprises: An acquisition unit, used for acquiring and forming an agent team set from the trained agent model; wherein the agent team set includes a plurality of agents to be screened, and there is an association relationship between any two agents to be screened; A determination unit, used to determine a target observation attack scenario, and based on the target observation attack scenario, perform an observation attack on the to-be-screened intelligent agents in the intelligent agent team set, and screen and determine a target intelligent agent from all the to-be-screened intelligent agents; A generating unit, configured to determine, in the trained agent model, an action score value of any agent to be screened at a current moment, and generate an adversarial perturbation value for the trained agent model through the action score value; wherein the action score value is obtained by observing a selected action of any agent to be screened in the agent team set; The acquisition unit is also used to add the adversarial disturbance value to the target agent to test the robustness of the target agent team set when the target agent team set is attacked; wherein the target agent team set is an agent team set including the target agent.

9. An algorithm robustness testing device for a multi-agent system, characterized in that: The device comprises: CPU, memory, input and output interfaces, wired or wireless network interfaces, and power supply; The memory is a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instruction operations in the memory to perform the algorithm robustness testing method for a multi-agent system as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes instructions, which, when executed on a computer, enable the computer to execute the algorithm robustness testing method for a multi-agent system as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent agent training method, data processing method and data processing system

    CN120996077A