Multi-agent black box test method for unmanned aerial vehicle countermeasure, electronic equipment and medium

Through the multi-agent black box testing method, zero-order optimization and random coordinate descent are used to generate adversarial samples, solving the complex black box testing problem in drone confrontation scenarios, and achieving efficient and low-cost security testing and decision-making performance evaluation.

CN120278218APending Publication Date: 2025-07-08ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510224990.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The drone confrontation scenario is complex, and it is difficult to conduct white box testing. There are many input parameters and complex influencing factors. It requires a highly robust decision model and can make effective decisions in unknown situations. Black box testing is important to judge the security of decision-making.

Method used

The multi-agent black box testing method is adopted, by configuring the drone adversarial scenario, building a loss function and performing zero-order optimization to solve the approximate gradient, using the approximate gradient to generate adversarial samples, comparing the difference in reward value to determine decision performance, and combining zero-order optimization and random coordinate descent methods for safety testing.

Benefits of technology

It realizes efficient and low-computing cost black box testing, which can discover potential problems and debug the model decision process, reduce platform costs and time costs, and ensure the security of the model decision process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278218A_ABST
    Figure CN120278218A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent black-box test method for unmanned aerial vehicle confrontation, electronic equipment and a medium, and the method comprises the steps: configuring a multi-agent unmanned aerial vehicle confrontation scene, including setting the number of agents, a state space, an action set of each agent, a reward function and a state transfer function; a loss function is constructed according to input and output of the intelligent agent, and the intelligent agent takes the t-moment state as output and takes the action and the (t + 1)-moment state obtained by action execution as output; performing zero-order optimization on the loss function to solve an approximate gradient; carrying out random coordinate descent on the input value of the to-be-tested sample by utilizing the approximate gradient to obtain an adversarial sample; and respectively inputting the to-be-tested sample and the confrontation sample into the intelligent agent, comparing reward values output by the intelligent agent under the to-be-tested sample or the confrontation sample, and if a difference value of the reward values exceeds a threshold value, judging that the performance of a decision made by the intelligent agent under the confrontation sample becomes poor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of reinforcement learning, and particularly relates to a multi-agent black-box testing method, an electronic device, and a medium for unmanned aerial vehicle (UAV) confrontation. Background Art

[0002] With the progress of artificial intelligence and autonomous system technologies, UAVs and other aerial agents have been widely used in military, civilian, and commercial fields.

[0003] Currently, the UAV confrontation scenario is relatively complex, making it difficult to conduct white-box testing; moreover, there are numerous input parameters and complex influencing factors. The decision-making model of the UAV agent needs to have high robustness, be able to exclude the interference of input perturbations or input anomalies, and be able to make effective decisions even in the face of unknown situations, and be able to adapt to various emergency situations and changes.

[0004] Therefore, it is particularly important to judge the decision-making safety of the UAV agent decision-making model through black-box testing. Summary of the Invention

[0005] In view of this, the present invention provides a multi-agent black-box testing method, an electronic device, and a medium for UAV confrontation.

[0006] In a first aspect, an embodiment of this aspect provides a multi-agent black-box testing method for UAV confrontation, and the method includes:

[0007] Configure a multi-agent UAV confrontation scenario, including setting the number of agents, the state space, the action set of each agent, the reward function, and the state transition function;

[0008] Construct a loss function based on the input and output of the agent. The agent takes the state at time t as the input and the state at time t + 1 obtained by the action and the execution of the action as the output; perform zero-order optimization on the loss function to solve for the approximate gradient;

[0009] Use the approximate gradient for random coordinate descent of the input value of the sample to be tested to obtain an adversarial sample; input the sample to be tested and the adversarial sample into the agent respectively, and compare the reward values output by the agent under the sample to be tested or the adversarial sample. If the difference in the reward values exceeds the threshold, it is determined that the decision-making performance of the agent under the adversarial sample deteriorates. The beneficial effects of the present invention are:

[0010] In a second aspect, an embodiment of this aspect provides an electronic device, including a memory and a processor, and the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned multi-agent black-box testing method for UAV confrontation.

[0011] In a third aspect, embodiments of this aspect provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-mentioned multi-agent black-box testing method for UAV countermeasure is implemented.

[0012] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0013] The present invention provides a multi-agent black-box testing method for UAV countermeasure, which is based on zero-order optimization for gradient calculation. The zero-order optimization method can perform black-box testing without using a surrogate model, and has the characteristics of high efficiency and low computing power requirements, which can reduce the platform cost and time cost. Through gradient optimization, potential problems are found and debugged in black-box testing, which is conducive to familiarizing with the decision-making process of the agent model, and then tracing the root cause to take corresponding solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0015] Figure 1 It is a schematic diagram of the UAV countermeasure air combat scenario provided by the embodiments of the present invention;

[0016] Figure 2 It is a schematic diagram of the multi-agent black-box testing method for UAV countermeasure provided by the embodiments of the present invention;

[0017] Figure 3 It is a flowchart of the multi-agent black-box testing method for UAV countermeasure provided by the embodiments of the present invention;

[0018] Figure 4 It is a schematic diagram of black-box testing based on zero-order optimization provided by the embodiments of the present invention;

[0019] Figure 5 It is a schematic diagram of an electronic device provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0021] It should be noted that, without conflict, the features in the following embodiments and implementation manners can be combined with each other.

[0022] As Figure 1 shown, an embodiment of the present invention provides a multi-agent black box testing method for UAV confrontation, and the method includes:

[0023] Step S1, configure a multi-agent UAV confrontation scenario, including setting the number of agents, state space, action sets of each agent, reward function, and state transition function.

[0024] Specifically, in this example, the multi-agent UAV confrontation scenario is a real-time decision-making problem in continuous action and state space, and can be modeled according to the Markov decision process. The following tuple is constructed in this example:

[0025] (N, S, A1, A2, A3,..., A N , γ, R1, R2, R3,..., R N , F)

[0026] In the formula, N represents the number of agents, and the number of agents is greater than or equal to two; S is all possible state spaces in UAV confrontation; the joint state is defined by A = A1 × A2 × A3 ×... × A N ; A1, A2, A3,..., A N represent the action sets of each agent; γ is the discount factor, γ ∈ [0, 1]; R1, R2, R3,..., R N represent the reward values of each agent after executing the specified action to reach the next state under the existing state S; the state transition function F: S × A → F(S) defines the probability distribution of the actions of the multi-agent in all possible states according to the current state and joint action. The ultimate goal of the agent decision-making model is to obtain the maximum cumulative reward value.

[0027] Among them, the state space O i of the i-th agent is expressed as follows:

[0028]

[0029] In the formula, L i represents the survival state of the i-th agent, D i represents the distance between the i-th agent and its friendly UAV, α i represents the angle of the i-th agent, represents the distance between the i-th agent and the enemy UAV, represents the angle between the i-th agent and the enemy UAV, H i represents the flight altitude of the i-th agent

[0030] Among them, the available actions of the i-th agent are to fire or only move without firing; when the i-th agent chooses to fire, the speed and heading angle change in the current state are the same as those in the previous state. When an enemy aircraft enters the circular area with a radius of 100 of the i-th agent, the i-th agent will have a probability of p fire to choose to fire, and after firing, there is a probability of p broke to shoot down the enemy aircraft.

[0031] Among them, the setting process of the reward function includes: setting corresponding reward values for each situation including completely destroying enemy agents, destroying one enemy agent, completely destroying friendly agents, destroying one friendly agent, staying within a specified distance from friendly UAVs for a certain period of time, and exceeding the limited flight range; among them, the reward values for completely destroying enemy agents and destroying one enemy agent are positive, and the reward value corresponding to completely destroying enemy agents is greater than the reward value corresponding to destroying one enemy agent; the reward values for completely destroying friendly agents and destroying one friendly agent are negative, and the reward value corresponding to completely destroying friendly agents is greater than the reward value corresponding to destroying one friendly agent. At the same time, since the multi-agent UAV confrontation scenario involves the cooperation between multiple agents, considering the communication problem between agents, we set a distance-based sparse reward to limit the agents in the cluster to maintain a suitable distance and punish the agents that exceed the restricted flight area.

[0032] In this example, our agents are trained using the reinforcement learning method, and the training method of the enemy is fixed. Therefore, to achieve the optimal decision, that is, to ensure that as many of our agents as possible remain, while as many enemy agents as possible are eliminated. Therefore, when our side destroys an enemy agent, the reward value is positive, and when our agent is destroyed, the reward value is negative. When all our (enemy) agents are destroyed, the global minimum (maximum) value is obtained. The expression of the reward function set in this example is as follows:

[0033]

[0034] Step S2, construct a loss function according to the input and output of the agent. The agent takes the state at time t as the output, and takes the action and the state at time t+1 obtained by executing the action as the output; perform zero-order optimization on the loss function to solve the approximate gradient.

[0035] Specifically, the step S2 includes the following sub-steps:

[0036] Step S201, construct a loss function according to the input and output of the agent, and the expression is as follows:

[0037]

[0038] In the formula denotes the expected output sequence, and y denotes the actual output sequence. denotes the overall error, and log 0 is defined as -∞.

[0039] It should be noted that since the black-box attack cannot observe the internal structure of the model and can only observe the input and output of the model, the loss function is only related to the output F of the model and the class label l expected by the model.

[0040] Step S202, solve the first-order gradient based on the loss function is defined as The expression is as follows:

[0041]

[0042] In the formula, h = 0.0001, e i is the standard unit vector, and only the i-th component is 1.

[0043] Furthermore, the gradient calculated in this way will have a slight decrease in accuracy, but the final experiment shows that the probability of a successful attack is still extremely high.

[0044] Step S203, solve the second-order gradient based on the first-order gradient is defined as The expression is as follows:

[0045]

[0046] Step S3, use the approximate gradient to perform random coordinate descent on the input value of the sample to be tested to obtain an adversarial sample; input the sample to be tested and the adversarial sample into the agent respectively, and compare the reward values output by the agent under the sample to be tested or the adversarial sample. If the difference in the reward values exceeds the threshold, it is determined that the decision-making performance of the agent under the adversarial sample deteriorates.

[0047] Furthermore, as Figure 3 shown, the process of using the approximate gradient to perform random coordinate descent on the input value of the sample to be tested to obtain an adversarial sample includes:

[0048] In the random coordinate descent method, a dimension is randomly selected for value update in each iteration, and the update method is to make argmin δ f(x + δe i ) minimized by any first-order or second-order method to approximately find an optimal δ value.

[0049] The update rule used in the present invention is the ADAM update rule, and this method can make the solution converge faster. In the update rule, the step size η and the state of ADAM need to be clarified first. and some hyperparameters β1 = 0.9, β2 = 0.999, ∈ = 10 -8 。

[0050] In each iteration, through

[0051]

[0052] where T i represents the i-th iteration, M i represents the state of M in ADAM after the i-th iteration, v i represents the state quantity of v in ADAM after the i-th iteration, represents the value of M that participates in influencing the perturbation size after the i-th iteration, represents the value of v that participates in influencing the perturbation size after the i-th iteration.

[0053] Update these parameters, and after the update, through

[0054]

[0055] calculate δ, and finally through

[0056] x i ←x i +δ

[0057] realize the coordinate update of the i-th dimension;

[0058] Concatenate each dimension to obtain the adversarial example.

[0059] Furthermore, the specific UAV game confrontation scenario is as Figure 1 shown. The goal is to use the strategy obtained by the optimization algorithm to enable the red UAV to autonomously complete effective strikes on hostile units such as the blue UAV. Train the red UAV through the multi-agent proximal policy optimization method. Finally, the red UAV completes tactical coordination, offensive and defensive games by continuously adjusting the altitude, azimuth attitude, and speed thrust, and finally shoots down all enemy UAVs at the smallest possible loss cost.

[0060] In summary, the present invention passes a multi-agent black-box testing method for UAV confrontation, which is used to perform security testing on the multi-agent game reinforcement learning model. In order to discover the problem of security vulnerabilities in the model strategy, by means of zero-order optimization to solve the approximate gradient, the gradient solution is achieved at a relatively low computing power cost without using the learning method to clone and obtain a shadow model with the same strategy as the original model, and the adversarial samples are quickly obtained by using the method of stochastic coordinate descent to achieve targeted attacks to discover the security vulnerabilities in its strategy. This method will not suffer losses caused by the lack of transferability of attacks generated by attempting to clone alternative models. Discovering and fixing the security vulnerabilities of the model in this way can help ensure that this sensitive information is not obtained by malicious attackers.

[0061] Correspondingly, the present application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the multi-agent black-box testing method for UAV confrontation as described above. As Figure 5 shown, it is a hardware structure diagram of any device with data processing capabilities where the multi-agent black-box testing method for UAV confrontation provided by the embodiment of the present invention is located. In addition to Figure 5 the processors, memory, and network interfaces shown, any device with data processing capabilities where the device in the embodiment is located usually includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.

[0062] Correspondingly, the present application also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the multi-agent black-box testing method for UAV confrontation as described above is implemented. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both the internal storage unit of any device with data processing capabilities and the external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store the data that has been output or will be output.

[0063] The above embodiments are only used to illustrate the design concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A multi-agent black-box testing method for UAV countermeasure, characterized in that The method includes: Configuring a multi-agent UAV confrontation scenario, including setting the number of agents, state space, action sets of each agent, reward function, and state transition function; Constructing a loss function based on the input and output of the agent, where the agent takes the state at time t as the output and the action and the state obtained by executing the action at time t+1 as the output; performing zero-order optimization on the loss function to solve for the approximate gradient; Using the approximate gradient to perform stochastic coordinate descent on the input value of the test sample to obtain an adversarial sample; inputting the test sample and the adversarial sample into the agent respectively, and comparing the reward values output by the agent under the test sample or the adversarial sample. If the difference in the reward values exceeds the threshold, it is determined that the decision-making performance of the agent under the adversarial sample deteriorates.

2. The multi-agent black-box testing method for UAV countermeasure according to claim 1, characterized in that, Constructing a multi-agent UAV confrontation scenario, constructing the following tuple: (N, S, A1, A2, A3,..., A N , γ, R1, R2, R3,..., R N , F) Wherein, N represents the number of agents, and the number of agents is greater than or equal to two; S is all possible state spaces in the UAV confrontation; the joint state is defined by A = A1×A2×A3×...×A N ; A1, A2, A3,..., A N represent the action sets of each agent; γ is the discount factor, γ ∈ [0, 1]; R1, R2, R3,..., R N represent the reward values of each agent after executing the specified action to reach the next state under the existing state S; the state transition function F: S×A→F(S) defines the probability distribution of the actions of multiple agents in all possible states according to the current state and the joint action.

3. The multi-agent black box testing method for UAV countermeasure according to claim 1 or 2, characterized in that The state space O of the i-th agent i has the following expression: where, L i represents the survival state of the i-th agent, D i represents the distance between the i-th agent and its friendly aircraft, α i represents the angle of the i-th agent, represents the distance between the i-th agent and the enemy aircraft, represents the angle between the i-th agent and the enemy aircraft, H i represents the flight altitude of the i-th agent.

4. The multi-agent black box testing method for UAV countermeasure according to claim 1 or 2, characterized in that The optional actions of the i-th agent are to fire or just move without firing; when the i-th agent chooses to fire, the changes in speed and heading angle of the current state are the same as those of the previous state. When the enemy aircraft enters the effective attack range of the i-th agent, the i-th agent will have a probability of p fire to choose to fire. After firing, there is a probability of p broke to shoot down the enemy aircraft.

5. A multi-agent black-box testing method for UAV countermeasure according to claim 1 or 2, characterized in that The process of setting the reward function includes: setting corresponding reward values for each situation including completely destroying enemy agents, destroying one enemy agent, completely destroying friendly agents, destroying one friendly agent, staying within a specified distance from friendly UAVs for a certain period of time, and exceeding the specified flight range. Among them, the reward values for completely destroying enemy agents and destroying one enemy agent are positive, and the reward value corresponding to completely destroying enemy agents is greater than the reward value corresponding to destroying one enemy agent; The reward values for completely destroying friendly agents and destroying one friendly agent are negative, and the reward value corresponding to completely destroying friendly agents is greater than the reward value corresponding to destroying one friendly agent.

6. The multi-agent black-box testing method for UAV countermeasure according to claim 1, characterized in that The expression of the loss function is as follows: In the formula, x represents the input of the agent, that is, the state at time t; F(x) represents the output of the agent, that is, the action and the state obtained by executing the action at time t+1; i represents the i-th agent, l represents the expected class label l of the agent, and k is a tuning parameter for attack transferability, and k≥0.

7. A multi-agent black-box testing method for UAV countermeasure according to claim 1, characterized in that The process of performing zero-order optimization on the loss function to solve for the approximate gradient includes: Solve the first-order gradient based on the loss function Is defined as The expression is as follows: where h = 0.0001, e i is a standard unit vector and only the i-th component is 1; Solving the second-order gradient based on the first-order gradient is defined as The expression is as follows:

8. A multi-agent black-box testing method for UAV countermeasure according to claim 1, characterized in that The process of using the approximate gradient to perform stochastic coordinate descent on the input value of the test sample to obtain an adversarial sample includes: In the random coordinate descent method, at each iteration, a dimension is randomly selected for value update. The update method is to approximately find an optimal δ value by making argmin δ f(x + δe i ) minimized through any first-order or second-order method; Based on the ADAM update rule, define the step size η, the state of ADAM and hyperparameters β1 = 0.9, β2 = 0.999, ∈ = 10 -8 ; In each iteration, through where T i represents the i-th iteration, M i represents the state of M in ADAM after the i-th iteration, v i represents the state quantity of v in ADAM after the i-th iteration, represents the M value that participates in influencing the perturbation magnitude after the i-th iteration, represents the v value that participates in influencing the perturbation magnitude after the i-th iteration; Updating these parameters, and after updating, through Realize the coordinate update of the i-th dimension; Concatenating each dimension to obtain an adversarial sample.

9. An electronic device, comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the multi-agent black-box testing method for UAV confrontation described in any one of claims 1-8 above.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the multi-agent black-box testing method for UAV confrontation described in any one of claims 1-8.