Method and device for determining continuous action iteration dilemma CAID noise

By establishing an agent dynamic model and using an adaptive filtering mechanism in the context of CAID, we can determine whether the noise is positive excitation noise, which solves the problem of how to find positive excitation noise that promotes cooperation in the agent system, and achieves the optimization of system stability and cooperative behavior.

CN120145833APending Publication Date: 2025-06-13NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510217284.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the context of CAID, how to find positive excitation noise that can promote cooperation in intelligent systems such as drones and unmanned vehicles, taking into account the impact of communication noise on the game behavior and results of the agent.

Method used

By determining the agent's fitness based on the game gain matrix, a CAID agent dynamics model is established, and combining the adaptive filtering mechanism and the optimal gain factor, it is determined whether the noise is positive excitation noise. Noise that meets specific conditions is determined as positive excitation noise.

Benefits of technology

Effectively reduce the impact of negative excitation noise, enhance system stability, and promote the convergence and optimization of cooperative behavior of agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145833A_ABST
    Figure CN120145833A_ABST
Patent Text Reader

Abstract

The invention discloses a continuous action iteration dilemma CAID noise determination method and device, and relates to the technical field of artificial intelligence. The method is used for analyzing the influence of noise on intelligent agent cooperation behaviors under the CAID background, and searching positive excitation noise capable of promoting cooperation for intelligent agents such as unmanned aerial vehicles and unmanned vehicles. The method comprises the steps of obtaining a CAID-based agent dynamic model according to an execution strategy evolution law and fitness difference, and obtaining an optimal gain factor based on characteristics of a system, a first global consensus cost function and a connected graph; when communication noise exists when the agent executes a task, obtaining a CAID-based agent dynamic model influenced by the communication noise; and according to the CAID-based agent dynamic model influenced by the communication noise, the adaptive filtering mechanism and the optimal gain factor, obtaining a judgment condition of the positive excitation noise, and determining the noise meeting the judgment condition in the game environment as the positive excitation noise.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more particularly to a method and device for determining continuous action iterated dilemma (CAID) noise. Background Art

[0002] In the real world, from intelligent transportation systems to drone swarm collaboration, cooperation is a key element for the normal operation of these complex systems. For a system environment composed of self-interested and rational agents such as drones and unmanned vehicles, the emergence mechanism of their cooperative behavior has attracted much attention. To deeply study this mechanism, researchers rely on the powerful framework of evolutionary game theory.

[0003] The research and application of evolutionary games are extensive. In scenarios such as drone formation coordination and unmanned vehicle traffic flow management, existing research has explored the consistency among agents, including aspects such as payoff consistency and behavior consistency. However, the binary game models adopted by many studies have limitations and are difficult to accurately depict the dynamic changes of agents' cooperative behavior. In contrast, the CAID (continuous action iterated dilemma) model can more accurately describe the strategy evolution of agents such as drones and unmanned vehicles during the cooperation process. It is found that under specific conditions, multiple iterations will make the strategies selected by agents show significant convergence.

[0004] However, interference factors in real scenarios, such as communication noise, will have a significant impact on the game behavior and results of agents. Taking drone formation and unmanned vehicle driving as examples, communication noise will interfere with agents' perception of the surrounding environment and the strategies of other agents, thereby affecting their learning rate and cooperative behavior. Usually, noise is regarded as an obstacle to efficiency and optimization, but the stochastic resonance effect shows that appropriate noise may also bring positive effects. Therefore, in the context of CAID, how to find positive incentive noise that can promote cooperation for agents such as drones and unmanned vehicles has become an urgent problem to be solved. Summary of the Invention

[0005] Embodiments of the present invention provide a method and device for determining continuous action iterated dilemma (CAID) noise, which are used to analyze the impact of noise on the cooperative behavior of agents in the context of CAID, and to find positive incentive noise that can promote cooperation for agents such as drones and unmanned vehicles.

[0006] Embodiments of the present invention provide a method for determining continuous action iterated dilemma (CAID) noise, including:

[0007] Determining the fitness of agents under different execution strategies based on the game payoff matrix; obtaining an agent dynamics model based on CAID according to the execution strategy evolution law and the fitness difference among different agents;

[0008] Based on the relationship between the agent dynamics model and the agent control input, a system with CAID as the background is obtained; based on the characteristics of the system, the first global consensus cost function, and the connected graph, a third global consensus cost function and an optimal gain factor are obtained.

[0009] When there is communication noise in the agent's task execution, the agent dynamics model affected by the communication noise is obtained according to the agent dynamics model and the communication noise; according to the agent dynamics model affected by the communication noise, the adaptive filtering mechanism, and the optimal gain factor, a judgment condition for positive excitation noise is obtained, and the noise that satisfies the judgment condition in the game environment is determined as positive excitation noise.

[0010] Preferably, the fitness of the agent under different execution strategies is as follows:

[0011]

[0012] The fitness difference between different agents is as follows:

[0013] ΔP ji = P(x j ) - P(x i )

[0014] where P(x i ) represents the fitness of agent i under the corresponding execution strategy, b ij represents the communication situation between agent i and agent j, P(x j ) represents the fitness of agent j under the corresponding execution strategy, r 0 , r 1 , r 2 and r 3 represent the payoffs of the agent under different execution strategies, x i represents the execution strategy of agent i, x j represents the execution strategy of agent j, and ΔP ji represents the fitness difference between agent j and agent i, and N represents the number of agents in the undirected graph.

[0015] Preferably, the system with CAID as the background is as follows:

[0016]

[0017] The first global consensus cost function is as follows:

[0018]

[0019] The third global consensus cost function is as follows:

[0020] J c3 = ∫ 0 ∞ x T (0)[2e -φLt Le -φLt + φ 2 e -φLt L 2 e -φLt x(0)dt

[0021] where represents the dynamics model of agent i, u i (t) represents the control input of intelligent machine i, φ represents the gain factor, φ ≥ 0, b ij represents the communication situation between agent i and agent j, b ij = 1, x j (t) represents the execution strategy of agent j at time t, x i (t) represents the execution strategy of agent i at time t, L represents the Laplacian matrix, x(t) represents the vector form of the agent execution strategy set, J c1 represents the first global consensus cost function, represents the cumulative consensus error of the partial state within the infinite time range [0, ∞], represents the error from the maximum value of the execution strategy, J u represents the agent control energy cost, J c3 represents the third global consensus cost function, x(0) represents the vector form of the initial execution strategy set of the agent, T represents the transpose, and e represents the base of the natural logarithm.

[0022] Preferably, after obtaining the optimal gain factor, it further includes:

[0023] According to the optimal gain factor and the agent dynamics model, obtain the optimal learning rate of the agent:

[0024] The optimal gain factor is as follows:

[0025]

[0026] The optimal learning rate of the agent is as follows:

[0027]

[0028] where φ * represents the optimal gain factor, ψ 1 represents the normalized eigenvector corresponding to the eigenvalue λ 1 [L] of the Laplacian matrix, T represents the transpose, and L represents the Laplacian matrix, represents the optimal learning rate of the agent, θ i represents the degree of agent i.

[0029] Preferably, the agent dynamics model is as follows:

[0030]

[0031] The agent dynamics model affected by communication noise is as follows:

[0032]

[0033] where represents the dynamics model of agent i, x i (k + 1) represents the execution strategy of agent i in the (k + 1)-th iteration, x i (k) represents the execution strategy of agent i in the (k + 1)-th iteration, A ij represents the element of the weighted adjacency matrix, θ i represents the degree of agent i, τ ij represents the learning rate of agent i, x j (t) represents the execution strategy of agent j at time t, x i (t) represents the execution strategy of agent i at time t, N represents the number of agents in the undirected graph, represents the agent dynamics model affected by communication noise, S ij represents the communication noise, k ij represents the communication noise coefficient.

[0034] Preferably, the judgment conditions for the positive incentive noise include:

[0035]

[0036] where τ ij (t) represents the learning rate, represents the optimal learning rate, represents the learning rate affected by noise, t 0 and t f represent the interval of the cumulative mean square error.

[0037] The embodiment of the present invention provides a device for determining the continuous action iteration dilemma CAID noise, including:

[0038] The first obtaining unit is configured to determine the fitness of the agent under different execution strategies based on the game payoff matrix; obtain the agent dynamics model based on CAID according to the execution strategy evolution law and the fitness difference between different agents;

[0039] A second obtaining unit, configured to obtain a system with CAID as the background according to the relationship between the agent dynamics model and the agent control input; and obtain a third global consensus cost function and an optimal gain factor based on the system, the first global consensus cost function, and the characteristics of the connected graph.

[0040] A determination unit, configured to, when there is communication noise in the agent's task execution, obtain an agent dynamics model affected by the communication noise according to the agent dynamics model and the communication noise; and obtain a judgment condition for positive excitation noise according to the agent dynamics model affected by the communication noise, the adaptive filtering mechanism, and the optimal gain factor, and determine the noise that satisfies the judgment condition in the game environment as positive excitation noise.

[0041] Preferably, the agent dynamics model is as follows:

[0042]

[0043] The agent dynamics model affected by the communication noise is as follows:

[0044]

[0045] The judgment condition for the positive excitation noise includes:

[0046]

[0047] Wherein, represents the dynamics model of agent i, x i (k + 1) represents the execution strategy of agent i in the (k + 1)-th iteration, x i (k) represents the execution strategy of agent i in the (k + 1)-th iteration, A ij represents the element of the weighted adjacency matrix, θ i represents the degree of agent i, τ ij represents the learning rate of agent i, x j (t) represents the execution strategy of agent j at time t, x i (t) represents the execution strategy of agent i at time t, N represents the number of agents in the undirected graph, represents the agent dynamics model affected by the communication noise, S ij represents the communication noise, k ij represents the communication noise coefficient, τ ij (t) represents the learning rate, represents the optimal learning rate, represents the learning rate affected by the noise, t 0 and t f represent the interval of the cumulative mean square error.

[0048] An embodiment of the present invention provides a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the method for determining the continuous action iterative dilemma CAID noise described in any one of the above.

[0049] An embodiment of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the method for determining the continuous action iterative dilemma CAID noise described in any one of the above.

[0050] An embodiment of the present invention provides a method and device for determining continuous action iterative dilemma CAID noise. The method includes: determining the fitness of an agent under different execution strategies based on a game payoff matrix; obtaining an agent dynamics model based on CAID according to the execution strategy evolution law and the fitness difference between different agents; obtaining a system with CAID as the background according to the relationship between the agent dynamics model and the agent control input; obtaining a third global consensus cost function and an optimal gain factor based on the system, the first global consensus cost function, and the characteristics of the connected graph; when there is communication noise in the agent's task execution, obtaining an agent dynamics model affected by communication noise according to the agent dynamics model and the communication noise; obtaining a judgment condition for positive incentive noise according to the agent dynamics model affected by communication noise, the adaptive filtering mechanism, and the optimal gain factor, and determining the noise that satisfies the judgment condition in the game environment as positive incentive noise. This method uses the consensus control theory to solve the asymptotic convergence problem in the CAID environment, avoids problems such as instability caused by parameter adjustment when using neural networks for fitting, and provides a determined optimal gain factor; further, considering the subtle role of noise in the CAID model, the optimal learning rate of the unmanned aerial vehicle is obtained, and the definition of the concept of positive incentive noise in the context of CAID is given here. Combining with the adaptive filtering mechanism, it dynamically ensures that only positive incentive noise affects the system, effectively reducing the impact of negative incentive noise and enhancing the system stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0052] Figure 1 It is a schematic flowchart of the method for determining continuous action iterative dilemma CAID noise provided by the embodiment of the present invention;

[0053] Figure 2 Schematic diagram of the complex network involved in the global consensus cost function provided by the embodiment of the present invention;

[0054] Figure 3A Schematic diagram of the iteration result curve of the general CAID provided by the embodiment of the present invention;

[0055] Figure 3B Schematic diagram of the iteration result curve of the optimal consensus control of CAID provided by the embodiment of the present invention;

[0056] Figure 3C Schematic diagram of the change curve of the global consensus cost with respect to the gain factor provided by the embodiment of the present invention;

[0057] Figure 4 Schematic diagram of the CAID for finding the positive incentive noise direction provided by the embodiment of the present invention;

[0058] Figure 5A Schematic diagram of the noise affecting the learning rate of CAID provided by the embodiment of the present invention;

[0059] Figure 5B Schematic diagram of the learning rate obtained by applying the adaptive filtering mechanism based on the definition of positive incentive noise provided by the embodiment of the present invention;

[0060] Figure 6 Schematic diagram of the structure of the device for determining the CAID noise of the continuous action iteration dilemma provided by the embodiment of the present invention. Detailed implementation manners

[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] Figure 1 Schematic diagram of the process for determining the CAID noise of the continuous action iteration dilemma provided by the embodiment of the present invention, as Figure 1 shown, the method includes the following steps:

[0063] Step 101, determine the fitness of the agent under different execution strategies based on the game payoff matrix; obtain the agent dynamics model based on CAID according to the execution strategy evolution law and the fitness difference between different agents;

[0064] Step 102: According to the relationship between the agent dynamics model and the agent control input, obtain a system with CAID as the background; based on the characteristics of the system, the first global consensus cost function, and the connected graph, obtain the third global consensus cost function and the optimal gain factor.

[0065] Step 103: When there is communication noise during the agent's task execution, obtain the agent dynamics model affected by communication noise according to the agent dynamics model and the communication noise; according to the agent dynamics model affected by communication noise, the adaptive filtering mechanism, and the optimal gain factor, obtain the judgment condition for positive excitation noise, and determine the noise that satisfies the judgment condition in the game environment as positive excitation noise.

[0066] In the following embodiments, taking the agent as a drone as an example, the method for determining the continuous action iterative dilemma CAID noise provided by the embodiments of the present invention is introduced in detail. However, in practical applications, the specific type of the agent is not limited, that is, the agent can also be an unmanned vehicle.

[0067] In step 101, the connection between drones can be represented by an undirected graph, where nodes represent drones and edges represent the relationship between drones. The undirected graph can be represented by the following formula:

[0068] G=(V,B) (1-1)

[0069] Where V represents the set of nodes (drones), and this set is a non-empty set, V={v 1 ,…,v N}}. B is the adjacency matrix, which represents how the nodes in an undirected graph are connected, B=(b ij )∈R N×N .

[0070] Furthermore, since drones will interact with each other, in the embodiments of the present invention, an undirected network is used to describe the relationship between drones. In this network, b ij is an element of the adjacency matrix, which is used to represent the communication situation between drone i and drone j. If b ij =1, it means that drone i and drone j can communicate with each other or sense each other's states; if b ij =0, it means that there is no communication between drone i and drone j or they cannot sense each other. And b ij =b ji indicates that this communication relationship is two-way, that is, if drone i can sense drone j, then drone j can also sense drone i, regardless of direction.

[0071] In this embodiment, Θ represents the degree matrix, as shown below:

[0072]

[0073] Among them, diag represents a diagonal matrix. This degree matrix is an N×N square matrix (related to the number of drones in the undirected graph G), and the elements θ 1 , …, θ N in the distribution correspond to the degree of each drone, and Θ ∈ R N×N .

[0074] Furthermore, the degree of the drone is represented by the following formula:

[0075]

[0076] Among them, θ i represents the degree of drone i, which represents the number of edges directly connected to drone i, that is, the number of other drones that can directly communicate or sense each other with drone i, and b ij represents the communication situation between drone i and drone j.

[0077] The Laplacian matrix associated with the undirected graph G provided by the embodiment of the present invention is as follows:

[0078] L = (l ij ) (2)

[0079] Among them, L represents the Laplacian matrix. L is an N×N real matrix, and the elements l ij in the matrix are all real numbers. N corresponds to the number of drones in the undirected graph G at the same time. l ij represents the element l ii in the i-th row and i-th column of the Laplacian matrix L (that is, the element on the diagonal). Its value is the sum of the values of b ij for all j except node i itself (j ≠ i). L ∈ R N×N .

[0080] In the example of drones, l ij represents the number of other drones that can directly communicate or sense each other with drone i. When b ij = 1 (that is, drones i and j can communicate or sense each other's states), l ij = -1; when b ij = 0 (drones i and j cannot communicate or sense each other's states), l ij = 0.

[0081] In this embodiment, considering N drones participating in the game, each drone i can adopt various continuous flight strategies x iare selected from these flight strategies x i represents the degree of cooperation of UAV i, from pure cooperative behavior x i = 1 to pure betrayal behavior x i = 0. In game theory, the payoff matrix can be used to describe the payoffs of UAVs under different combinations of flight strategies, and the payoff matrix can be expressed as

[0082]

[0083] where r 0 , r 1 , r 2 , r 3 represent the payoffs of the UAVs under the corresponding flight strategies. Specifically: r 0 represents the payoff obtained when both UAVs choose cooperation, r 3 represents the payoff obtained when both choose betrayal. When one of the two UAVs chooses cooperation and the other chooses betrayal, the payoff obtained by the UAV that chooses cooperation is r 1 , and the payoff obtained by the betraying UAV is r 2 .

[0084] Furthermore, in the embodiments of the present invention, the payoff matrix of the chicken game is also provided, as shown below:

[0085]

[0086] where g represents the payoff and s represents the cost. In this embodiment, the parameters are set as g = 5 and s = 0.5.

[0087] In the embodiments of the present invention, according to the game payoff matrix, the fitness of UAVs under different flight strategies and the fitness differences between different UAVs can be determined. Among them, the advantage of UAV i under a flight strategy can be represented by a fitness function, as shown below:

[0088]

[0089] Furthermore, the fitness difference between different UAVs (UAV i and UAV j) can be expressed by the following formula:

[0090] ΔP ji = P(x j ) - P(x i ) (5)

[0091] where P(x i ) represents the fitness of UAV i under the corresponding flight strategy, P(x j ) represents the fitness of UAV j under the corresponding flight strategy, b ijIndicates the communication situation between UAV i and UAV j, b ij = 1, r 0 , r 1 , r 2 and r 3 are the benefits of the UAV under the corresponding flight strategy, x i represents the flight strategy of UAV i, x j represents the flight strategy of UAV j, ΔP ji represents the fitness difference between UAV j and UAV i.

[0092] In the embodiment of the present invention, based on the imitation dynamics theory, the UAV can learn and transition to a neighboring flight strategy with a certain probability. In the embodiment of the present invention, the flight strategy evolution law can be expressed by the following formula:

[0093]

[0094] where, x i (k + 1) represents the flight strategy of UAV i in the (k + 1)-th iteration, θ i represents the degree of UAV i, τ ij represents the probability that UAV i learns the neighboring flight strategy in the k-th iteration, that is, the learning rate of UAV i, τ ij = sig(α|ΔP ji |) / θ j , θ j represents the degree of UAV j, ΔP ji represents the fitness difference between UAV i and UAV j, b ij represents the communication situation between UAV i and UAV j, b ij = 1, x i (k) represents the flight strategy of UAV i in the k-th iteration, x j (k) represents the flight strategy of UAV j in the k-th iteration, and α represents a scalar parameter for controlling the learning rate, α > 0.

[0095] In the embodiment of the present invention, the update rate of the UAV from the flight strategy in the k-th iteration to the flight strategy in the (k + 1)-th iteration is determined as the UAV dynamics model. In other words, according to the flight strategy evolution law of the UAV (the flight strategy of the UAV in the iteration), the learning rate of the UAV, the degree of the UAV, and the fitness difference between different UAVs, the UAV dynamics model is obtained through the following formula:

[0096]

[0097] where, represents the UAV i dynamics model (the update rate of the flight strategy of UAV i), xi (k + 1) represents the flight strategy of UAV i in the (k + 1)-th iteration, x i (k) represents the flight strategy of UAV i in the k-th iteration, θ i represents the degree of UAV i, τ ij represents the learning rate of UAV i, b ij represents the communication situation between UAV i and UAV j, b ij = 1, N represents the number of UAVs, x j (t) represents the flight strategy of UAV j at time t, x i (t) represents the flight strategy of UAV i at time t.

[0098] Furthermore, let the element A of the weighted adjacency matrix ij = τ ij b ij / θ i , then the UAV dynamics model shown in formula (7) can be reformulated as:

[0099]

[0100] where, represents the UAV i dynamics model, A ij represents the element of the weighted adjacency matrix, that is, it represents assigning weights to each connecting edge of the undirected graph τ ij represents the learning rate of UAV i, θ i represents the degree of UAV i, b ij represents the communication situation between UAV i and UAV j, b ij = 1.

[0101] In the embodiment of the present invention, after determining the UAV dynamics model, according to the chicken game provided above, the payoff matrix of the chicken game and the UAV dynamics model, a continuous action iterative chicken game dynamics model can be obtained.

[0102] According to the introduction of the undirected graph in step 101, it can be known that the undirected graph is composed of a node set and an edge set. The undirected graph means that the edges have no direction, and the Laplacian matrix has symmetric positive semi-definite characteristics.

[0103] Furthermore, the condition for the undirected graph G to be considered connected is that the Laplacian matrix has a simple zero eigenvalue, and there is a unique eigenvector 1 in the null space of the Laplacian matrix N . For the connected graph G, the N eigenvalues of the Laplacian matrix satisfy the following formula:

[0104] λ 1 [L]= 0 ≤ λ2 [L] ≤ … ≤ λ N [L] (8)

[0105] Among them, ψ in i is the normalized eigenvector corresponding to λ i [L], that is and the Laplacian matrix can be diagonalized, which can be expressed as: L = Ψ T ΛΨ, Λ = diag([λ 1 [L], …, λ N [L]]).

[0106] In step 102, a system with CAID as the background can be obtained according to the UAV dynamics model and the UAV control input, as follows:

[0107]

[0108] Among them, x i (t) represents the UAV i dynamics model, u i (t) represents the UAV i control input, u i (t) ∈ R. In practical applications, when the UAV i control input u i (t) is given, the closed-loop system of formula (9) can be obtained.

[0109] If then the closed-loop system reaches asymptotic consensus, that is, the flight strategy convergence under game theory is achieved.

[0110] In order to enable the UAVs to reach consensus, considering the UAV dynamics model provided in the embodiments of the present invention, proportional control can be adopted, that is, a representation of the UAV control input is obtained:

[0111]

[0112] Among them, u i (t) represents the UAV i control input, φ represents the gain factor, φ ≥ 0, b ij represents the communication situation between UAV i and UAV j, b ij = 1, x j (t) represents the flight strategy of UAV j at time t, x i (t) represents the flight strategy of UAV i at time t.

[0113] Based on formula (10), formula (9) can be written in a compact form, as follows:

[0114]

[0115] Among them, \(L\) represents the Laplacian matrix, and \(x(t)\) represents the vector form of the UAV flight strategy set.

[0116] In the system based on the CAID background, the embodiment of the present invention introduces the CAID optimal consensus control method, and the key part is the global consensus cost function. Since there are multiple different expressions for the global consensus cost function in the embodiment of the present invention, in order to clearly distinguish them, in the embodiment, according to the order of their appearance, they are respectively named the first global consensus cost function \(J\) c1 , the second global consensus cost function \(J\) c2 , and the third global consensus cost function \(J\) c3 . It should be emphasized here that the first, second, and third are only for the sake of more clear expression in the text, and they do not mean that these three different forms of global consensus cost functions have essentially different meanings.

[0117] Specifically, the first global consensus cost function is as follows:

[0118]

[0119] In this formula, \(J\) c1 represents the first global consensus cost function, represents the cumulative consensus error of partial states within the infinite time range \([0, \infty]\), represents the error from the maximum value of the flight strategy, and \(J\) u represents the control energy cost of the UAV.

[0120] Looking further, the cumulative consensus error of partial states within the infinite time range \([0, \infty]\) in formula (11) the error from the maximum value of the flight strategy and the control energy cost \(J\) of the UAV u are respectively represented by the following formulas:

[0121] First, for the cumulative consensus error of partial states within the infinite time range \([0, \infty]\) its expression is:

[0122]

[0123] Next, the error from the maximum value of the flight strategy its expression is:

[0124]

[0125] Finally, the control energy cost \(J\) of the UAV u , its expression is:

[0126]

[0127] Here it should be noted that represents the cumulative consensus error of the partial state within the infinite time range [0, ∞], represents a specific connection relationship, [x i (t) - x j (t)] 2 represents the square of the difference in the flight strategies of UAV i and UAV j at time t, and then the value obtained by integrating this squared value over the time range [0, ∞], which is used to measure the cumulative consensus error of the partial state of the system within infinite time, with the aim of ensuring that the system can reach a consensus. represents the error from the maximum value of the flight strategy, b im represents the connection coefficient with the maximum node, x m (t) represents the maximum value of the UAV strategy, and N represents the number of UAVs in the undirected graph.

[0128] In the embodiments of the present invention, to achieve this goal, an undirected graph G Figure 2 as shown can be constructed (1) =(V, B (1) ). In this undirected graph, as an element of B (1) , it represents the connection relationship between nodes other than the maximum node (UAV), while b im represents the connection coefficient with the maximum node. The purpose of setting b im is to improve the cooperation level of UAVs because, in the context of CAID, the states of UAVs ultimately tend to converge to the average of the initial values.

[0129] Furthermore, an undirected graph G Figure 2 as shown can be constructed (2) =(V, B (2) ). In this undirected graph, as an element of B (2) , when , it means there is an interaction between UAV i and UAV j. Subsequently, the undirected graph G Figure 2 in (1) =(V, B (1) ) and the undirected graph G (2) =(V, B (2) ) are iterated, and can be obtained, which represents the connection relationship of all nodes.

[0130] Based on the above, the first global consensus cost function can be reformulated. Combining Equation (12-1), Equation (12-2), and Equation (12-3), the first global consensus cost function can be expressed as:

[0131]

[0132] In this formula, J c1 represents the first global consensus cost function, represents the element of B (1) , x i (t) represents the strategy of UAV i at time t, x j (t) represents the strategy of UAV j at time t, x m (t) represents the maximum value of the UAV strategy, u i (t) represents the control input of UAV i, is the element of B (2) , which takes the value of 1 only when there is a connection relationship between UAV i and UAV j, represents the connection coefficient of all nodes,

[0133] Furthermore, the first global consensus cost function shown in Equation (13) is transformed into a more compact form, that is, the second global consensus cost function is obtained, and its expression is as follows:

[0134] J c2 = ∫ 0 ∞ [2x(t) T Lx(t)+φ 2 x(t) T L 2 x(t)]dt (14)

[0135] In this formula, J c2 represents the second global consensus cost function, x(t) represents the vector form of the UAV flight strategy set, T represents the transpose, L represents the Laplacian matrix, and φ represents the gain factor.

[0136] In the embodiment of the present invention, when a connected graph G with a symmetric Laplacian matrix L having a simple zero eigenvalue is given, based on the relevant knowledge of the matrix exponential, the third global consensus cost function can be further obtained, which is specifically as follows:

[0137] J c3 = ∫ 0 ∞ x T (0)[2e -φLt Le -φLt +φ 2 e-φLt L 2 e -φLt x(0)dt (15)

[0138] Among them, J c3 represents the third global consensus cost function, x(0) represents the vector form of the initial flight strategy set of the UAV, T represents the transpose, L represents the Laplacian matrix, φ represents the gain factor, t represents time, and e represents the base of the natural logarithm.

[0139] In the embodiment of the present invention, in order to find the optimal gain factor φ * , it is necessary to take the derivative of the third global consensus cost function with respect to φ, which is specifically as follows:

[0140]

[0141] Furthermore, let Through a series of mathematical transformations and derivations, the following formula can be obtained:

[0142] x T (0)[∫ 0 ∞ 2Lte -φLt Le -φLt dt]x(0)+φ 2 x T (0)[∫ 0 ∞ Lte -φLt L 2 e -φLt dt]x(0)-φx T (0)[∫ 0 ∞ e -φLt L 2 e -φLt dt]x(0)=0 (17)

[0143] Furthermore, since the Laplacian matrix has a diagonalized form, that is, L = Ψ T ΛΨ, based on this, the first term included in formula (17) can be transformed into:

[0144]

[0145] Among them, λ represents the eigenvalue of the Laplacian matrix.

[0146] Next, in a similar way, the second term included in formula (17) can be transformed into:

[0147]

[0148] Similarly, the third term included in formula (17) can be transformed into:

[0149]

[0150] In the embodiment of the present invention, substituting formulas (18)-(20) into formula (17), through a series of algebraic operations and simplifications, the optimal gain factor can be finally obtained, which can be expressed as:

[0151]

[0152] where φ * represents the optimal gain factor, x(0) represents the vector form of the initial flight strategy set of the UAV, and ψ 1 is the normalized eigenvector corresponding to the eigenvalue λ 1 of the Laplacian matrix L, that is, ψ 1 is the vector obtained through normalization,

[0153] Figure 3A is a schematic diagram of the iteration result curve of the general CAID, Figure 3B is a schematic diagram of the iteration result curve of the optimal consensus control of CAID, Figure 3A is a schematic diagram of the curve of the global consensus cost with respect to the change of the gain factor. According to Figure 3A and Figure 3B it can be seen that after the optimal consensus control method provided by the embodiment of the present invention, consensus can be quickly reached, effectively improving the convergence speed and cooperation level; furthermore, according to Figure 3C the extreme point of it can be seen that the optimal gain factor can indeed be obtained through the global consensus cost function, and this gain factor can minimize the cost function. That is, according to Figures 3A - 3C it can be determined that the method provided by the embodiment of the present invention can prompt the system to quickly reach consensus and improve the cooperation level, and at the same time, a deterministic optimal gain factor can be obtained.

[0154] In step 103, based on the optimal gain factor φ * determined in step 102 and the UAV dynamics model, the following optimal learning rate can be obtained, which is as follows:

[0155]

[0156] where represents the optimal learning rate, φ * represents the optimal gain factor, and θ i represents the degree of the UAV i.

[0157] In the embodiments of the present invention, considering the situation of communication noise when the UAV executes tasks, the communication noise will affect the perception and learning ability between UAVs, that is, the UAV will have observation errors when perceiving the flight strategies of other UAVs. This kind of error will inevitably cause problems in the process of copying the flight strategies, and then affect the quality of the task execution results. Assuming there is communication noise, according to the UAV dynamics model shown in formula (7-1), the UAV dynamics model affected by communication noise can be obtained as follows:

[0158]

[0159] Among them, represents the UAV dynamics model affected by communication noise, S ij represents communication noise, k ij represents the communication noise coefficient.

[0160] In practical applications, according to the above formula, the change of the UAV flight strategy over time can be obtained

[0161] Furthermore, when there is communication noise when the UAV executes tasks, the filtering parameters can be dynamically adjusted according to the UAV dynamics model affected by communication noise, the adaptive filtering mechanism, the communication situation between UAVs (i.e., the communication situation reflected by b ij ), and the optimal gain factor, to obtain the judgment condition of positive incentive noise, and then determine the noise that satisfies the judgment condition in the game environment as positive incentive noise, that is, to ensure that only positive incentive noise contributes to the system.

[0162] In the embodiments of the present invention, the judgment conditions of positive incentive noise include the following two formulas:

[0163]

[0164] Among them, τ ij (t) represents the learning rate, represents the optimal learning rate, represents the learning rate affected by noise.

[0165] In the embodiments of the present invention, if the noise in the game environment satisfies the above formula, it means that at any time t, the learning rate affected by noise is always within the interval defined by the normal learning rate τ ij (t) and the optimal learning rate ; at the same time, in the interval [t 0 , t fWithin it, the cumulative mean square error between the learning rate and the optimal learning rate in the presence of noise is less than that between the learning rate and the optimal learning rate without noise. These two judgment conditions ensure that noise will not cause the learning rate to deviate from the reasonable range, which reflects the rationality of the influence of noise on the learning rate from one aspect. At the same time, it also shows that the existence of noise makes the learning rate closer to the optimal learning rate within a certain period of time, reflecting the positive effect of noise on the optimization of the learning rate from the cumulative effect.

[0166] During the actual flight process, if the communication noise of the UAV causes the learning rate of the UAV to change towards the optimal learning rate and reduces the cumulative mean square error relative to it, then the noise in the environment where the UAV is located can be considered as positive incentive noise.

[0167] It should be noted that the above-mentioned "noise in the game environment" refers to various uncertain and interfering factors that affect the decision-making of UAVs and the game results during the game process. It may come from multiple aspects, such as the UAV's incomplete understanding of the strategies, preferences, etc. of other UAVs, and the random changes in the external environment; "communication noise" is mainly the interference that occurs during the information transmission process of UAVs, resulting in information distortion, inaccuracy or incompleteness. In the game environment, communication noise can be regarded as a specific manifestation of the noise in the game environment. If there is noise in the communication between UAVs, it will affect the UAV's understanding of information such as each other's intentions and strategies, and then interfere with the entire game process and results, becoming an uncertain factor affecting the game; "positive incentive noise" generally refers to some uncertain factors or interferences that occur during the implementation of the incentive mechanism, making the incentive effect deviate from the expectation, but this deviation is a fluctuation in the positive incentive direction. In the game environment, positive incentive noise can also be regarded as a part of the noise in the game environment. For example, in some cooperative games, positive incentive measures are to encourage UAVs to adopt cooperative strategies.

[0168] In the CAID environment, the existence of noise in the game will affect the UAV's ability to perceive the flight strategies adopted by its neighbors, affect the learning rate in the evolution of flight strategies, and then affect the cooperative behavior within the system, such as Figure 4 As shown, if the influence caused by noise causes the learning rate τ ij to change in the direction of the optimal learning rate shown by the solid arrow direction, then this noise can be characterized as positive incentive noise. On the contrary, if the influence on the learning rate τ at a certain moment ij changes in the direction of the dashed arrow, then this noise cannot be characterized as positive incentive noise.

[0169] Communication noise will affect the perception and learning ability between UAVs and affect its learning rate. Figure 5AShows the influence of general communication noise on the learning rate. Among them, the smooth curve is the learning rate without noise, and the other is the learning rate affected by communication noise, which will produce irregular fluctuations and deviate from the range defined by positive incentive noise; in the embodiment of the present invention, an adaptive filtering mechanism is combined to determine that the noise in the game environment is positive incentive, ensuring that only positive incentive noise affects the system and effectively reducing the influence of negative incentive noise. Figure 5B The learning rate change curve obtained by applying the adaptive filtering mechanism based on the definition of positive incentive noise is compared Figure 5A It can be seen that all the noise acting on the system can be ensured to be positive incentive. It can continuously affect the learning rate to promote cooperative behavior.

[0170] To more clearly introduce the method for determining the continuous action iteration dilemma CAID noise provided by the embodiment of the present invention, the following takes an autonomous vehicle as an example to further introduce this method. Specifically, this embodiment includes the following steps:

[0171] Step 201, assume that there are N = 4 autonomous vehicles driving on a road. The relationship between the vehicles is represented by an undirected graph G=(V, B), where V = {v 1 ,…,v N} is the vehicle set, Describes the connection relationship between vehicles (if vehicle i and vehicle j can communicate with each other and sense each other's states, then b ij = 1, otherwise b ij = 0), and the specific representation is as follows:

[0172]

[0173] Specifically, the payoffs of vehicles under different driving strategy combinations are described according to the payoff matrix of the chicken game: if vehicle i and vehicle j find that they are driving towards each other, when vehicle i and vehicle j meet, if both choose to decelerate and give way (i.e., choose the cooperative driving strategy), the payoffs are each g = 5, the speed decreases but relative safety is ensured; if one vehicle decelerates (cooperates) and the other vehicle accelerates (betrays), then the accelerating vehicle obtains a payoff The decelerating vehicle has a payoff of 0 due to the reduced speed but relative safety; if both vehicles accelerate recklessly (i.e., choose the betraying driving strategy), a car accident will occur and even cause deaths, and both vehicle i and vehicle j have to bear the cost s = 1. Therefore, in this case, the payoff matrix is defined as:

[0174]

[0175] For vehicle i, its fitness function P(x i ) can be expressed by the formula to be determined, where x i represents the driving strategy of vehicle i (where its x i ranges from a completely conservative driving strategy x i = 0 to an aggressive driving strategy x i = 1). The fitness difference between different vehicles can be determined by the formula ΔP ji = P(x j ) - P(x i ).

[0176] Furthermore, according to the imitation dynamics, the formal strategy of vehicle i in the (k + 1)-th iteration can be obtained, from which the dynamic model formula of vehicle i based on CAID can be obtained When A ij = τ ij b ij / θ i , the dynamic model of vehicle i can be expressed as:

[0177] Step 202, the initial state of the vehicle is taken as x(0) = [0.5, 0.6, 0.4, 0.8] T , according to the vehicle dynamic model and the control input u i (t) of vehicle i, the system with CAID as the background is obtained where x i (t) represents the driving state of vehicle i (such as vehicle speed, driving direction, etc.), that is, the degree of adopting a decelerating driving strategy, and u i (t) represents the control input for vehicle i. For example, according to the interaction situation between vehicle i and surrounding vehicles, driving strategies, and the overall traffic conditions, an instruction to adjust the vehicle speed of vehicle i or an instruction to change the driving path of vehicle i can be issued, so as to coordinate the driving behaviors of vehicle i and other vehicles.

[0178] To enable all vehicles to reach a consensus (for example, all vehicles reach a unified passing strategy at an intersection to avoid collisions), proportional control is adopted for the control input of vehicle i

[0179] Or In this embodiment, the control input of vehicle i comprehensively considers the driving state differences between vehicle i and other vehicles that can communicate and sense each other, and promotes the entire vehicle group to reach a consensus by adjusting the driving state of vehicle i itself.

[0180] Introduce the first global consensus cost function where

[0181] represents the cumulative consensus error of partial vehicle states within an infinite time range (e.g., the cumulative error of the speed differences among different vehicles);

[0182] represents the error from the optimal driving strategy (such as the strategy x m (t) for passing through an intersection most efficiently); represents the control energy cost (such as the control energy consumed by a driving strategy with frequent speed adjustments).

[0183] After a series of derivations, by taking the derivative of the global consensus cost function with respect to φ, the optimal gain factor can be calculated as

[0184] Step 203, according to the optimal gain factor φ * and the vehicle dynamics model, the optimal learning rate can be obtained:

[0185]

[0186] In practical applications, since there is communication noise when vehicles execute tasks, the communication noise will affect the perception and learning ability among vehicles. That is, communication noise needs to be added to the vehicle dynamics model, and then the vehicle dynamics model affected by communication noise can be obtained: where S ij represents Gaussian white noise, and k ij represents the coefficient of Gaussian white noise. That is, the change of the vehicle execution strategy over time can be obtained according to the above formula.

[0187] Furthermore, when there is sensor noise when the vehicle is executing tasks, the filtering parameters can be dynamically adjusted according to the vehicle dynamics model affected by sensor noise, the adaptive filtering mechanism, the communication situation among vehicles (i.e., the communication situation reflected by b ij ) and the optimal gain factor to obtain the judgment condition for positive excitation noise, and then the noise that meets the judgment condition in the game environment is determined as positive excitation noise, that is, ensuring that only positive excitation noise contributes to the system.

[0188] In summary, the embodiments of the present invention provide a method and device for determining continuous action iterative dilemma CAID noise. This method uses consensus control to solve the asymptotic convergence problem in the CAID environment, avoiding problems such as instability caused by parameter adjustment when using neural networks for fitting, and providing a determined optimal gain factor. Further, considering the subtle role of noise in the CAID model, an optimal learning rate is obtained, and the definition of the concept of positive incentive noise in the context of CAID is given here. Combining with the adaptive filtering mechanism, it dynamically ensures that only positive incentive noise affects the system, effectively reducing the impact of negative incentive noise and enhancing the system stability.

[0189] Based on the same inventive concept, the embodiments of the present invention provide a device for determining continuous action iterative dilemma CAID noise. Since the principle of this device for solving technical problems is similar to that of the method for determining continuous action iterative dilemma CAID noise, the implementation of this device can refer to the implementation of the method, and the repeated parts will not be elaborated here.

[0190] Figure 6 It is a schematic structural diagram of the device for determining continuous action iterative dilemma CAID noise provided by the embodiments of the present invention, as Figure 6 shown, the device includes a first obtaining unit 601, a second obtaining unit 602, and a determining unit 603.

[0191] The first obtaining unit 601 is configured to determine the fitness of the agent under different execution strategies based on the game payoff matrix; obtain the agent dynamics model based on CAID according to the execution strategy evolution law and the fitness difference between different agents.

[0192] The second obtaining unit 602 is configured to obtain a system in the context of CAID according to the relationship between the agent dynamics model and the agent control input; obtain a third global consensus cost function and an optimal gain factor based on the system, the first global consensus cost function, and the characteristics of the connected graph.

[0193] The determining unit 603 is configured to, when there is communication noise in the agent's task execution, obtain the agent dynamics model affected by the communication noise according to the agent dynamics model and the communication noise; obtain the judgment condition of positive incentive noise according to the agent dynamics model affected by the communication noise, the adaptive filtering mechanism, and the optimal gain factor, and determine the noise satisfying the judgment condition in the game environment as positive incentive noise.

[0194] Preferably, the agent dynamics model is as follows:

[0195]

[0196] The agent dynamics model affected by the communication noise is as follows:

[0197]

[0198] The judgment conditions for the positive excitation noise include:

[0199]

[0200] Among them, represents the dynamic model of agent i, and x i (k + 1) represents the execution strategy of agent i in the (k + 1)-th iteration, and x i (k) represents the execution strategy of agent i in the (k + 1)-th iteration, A ij represents the element of the weighted adjacency matrix, θ i represents the degree of agent i, τ ij represents the learning rate of agent i, x j (t) represents the execution strategy of agent j at time t, and x i (t) represents the execution strategy of agent i at time t, and N represents the number of agents in the undirected graph. represents the dynamic model of the agent affected by communication noise, S ij represents the communication noise, k ij represents the communication noise coefficient, τ ij (t) represents the learning rate, represents the optimal learning rate, represents the learning rate affected by noise, t 0 and t f represent the interval of the cumulative mean square error.

[0201] It should be understood that the units included in the above-mentioned determining device for the continuous action iteration dilemma CAID noise are only logical divisions based on the functions implemented by the device. In actual applications, the above units can be superimposed or split. And the functions implemented by the determining device for the continuous action iteration dilemma CAID noise provided in this embodiment correspond one by one to the determining method for the continuous action iteration dilemma CAID noise provided in the above embodiment. For the more detailed processing flow implemented by this device, it has been described in detail in the first method embodiment above, and will not be described in detail here.

[0202] Another embodiment of the present invention also provides a computer device, which includes: a processor and a memory; the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the electronic device executes each step of determining the continuous action iteration dilemma CAID noise in the method flow shown in the above method embodiment.

[0203] Another embodiment of the present invention further provides a computer-readable storage medium storing computer instructions, which, when running on a computer device, cause the computer device to execute each step of the method for determining the continuous action iteration dilemma CAID noise in the method flow shown in the above method embodiment.

[0204] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0205] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for determining the continuous action iterative dilemma CAID noise, characterized in that: include: Determine the fitness of the agent under different execution strategies based on the game payoff matrix; According to the evolution law of execution strategy and the fitness difference between different agents, the agent dynamics model based on CAID is obtained; According to the relationship between the agent dynamics model and the agent control input, a system with CAID as the background is obtained; Based on the system, the first global consensus cost function, and the characteristics of the connectivity graph, a third global consensus cost function and an optimal gain factor are obtained; When there is communication noise when the agent performs a task, the agent dynamics model affected by the communication noise is obtained according to the agent dynamics model and the communication noise; the judgment condition of the positive excitation noise is obtained according to the agent dynamics model affected by the communication noise, the adaptive filtering mechanism and the optimal gain factor, and the noise that meets the judgment condition in the game environment is determined as the positive excitation noise.

2. The method for determining the continuous action iterative dilemma CAID noise according to claim 1, characterized in that: The fitness of the agent under different execution strategies is as follows: The fitness differences between the different agents are as follows: ΔP ji =P(x j )-P(x i ) Among them, P(x i ) represents the fitness of agent i under the corresponding execution strategy, b ij represents the communication between agent i and agent j, P(x j ) represents the fitness of agent j under the corresponding execution strategy, r0, r1, r2 and r3 represent the benefits of the agent under different execution strategies, and x i represents the execution strategy of agent i, x j represents the execution strategy of agent j, ΔP ji represents the fitness difference between agent j and agent i, and N represents the number of agents in the undirected graph.

3. The method for determining the continuous action iterative dilemma CAID noise according to claim 1, characterized in that: The system based on CAID is as follows: The first global consensus cost function is as follows: The third global consensus cost function is as follows: in, represents the dynamic model of agent i, u i (t) represents the control input of smart machine i, φ represents the gain factor, φ≥0, b ij represents the communication between agent i and agent j, b ij =1,x j (t) represents the execution strategy of agent j at time t, x i (t) represents the execution strategy of agent i at time t, L represents the Laplacian matrix, x(t) represents the vector form of the agent execution strategy set, J c1 represents the first global consensus cost function, represents the cumulative consensus error of some states in the infinite time range [0,∞], represents the error with the maximum value of the execution strategy, J u represents the energy cost of agent control, J c3 represents the third global consensus cost function, x(0) represents the vector form of the agent’s initial execution strategy set, T represents transpose, and e represents the base of the natural logarithm.

4. The method for determining the continuous action iterative dilemma CAID noise according to claim 1, characterized in that: After obtaining the optimal gain factor, the method further includes: According to the optimal gain factor and the agent dynamics model, the optimal learning rate of the agent is obtained: The optimal gain factor is as follows: The optimal learning rate for the agent is as follows: Among them, φ * represents the optimal gain factor, ψ1 represents the normalized eigenvector of the Laplace matrix corresponding to the eigenvalue λ1[L], T represents the transpose, L represents the Laplace matrix, represents the optimal learning rate of the agent, θ i Represents the degree of agent i.

5. The method for determining the continuous action iterative dilemma CAID noise according to claim 1, characterized in that: The agent dynamics model is as follows: The agent dynamics model affected by communication noise is as follows: in, represents the dynamic model of agent i, x i (k+1) represents the execution strategy of agent i in the k+1th iteration, x i (k) represents the execution strategy of agent i in the k+1th iteration, A ij represents the elements of the weighted adjacency matrix, θ i represents the degree of agent i, τ ij represents the learning rate of agent i, x j (t) represents the execution strategy of agent j at time t, x i (t) represents the execution strategy of agent i at time t, N represents the number of agents in the undirected graph, represents the agent dynamics model affected by communication noise, S ij represents the communication noise, k ij Represents the communication noise factor.

6. The method for determining the continuous action iterative dilemma CAID noise according to claim 1, characterized in that: The judgment conditions of the positive excitation noise include: Among them, τ ij (t) represents the learning rate, represents the optimal learning rate, represents the learning rate affected by noise, t0 and t f An interval representing the cumulative mean square error.

7. A device for determining continuous action iterative dilemma CAID noise, characterized in that: include: The first obtaining unit is used to determine the fitness of the agent under different execution strategies based on the game payoff matrix; According to the evolution law of execution strategy and the fitness difference between different agents, the CAID-based agent dynamics model is obtained; A second obtaining unit is used to obtain a system with CAID as the background according to the relationship between the agent dynamics model and the agent control input; Based on the system, the first global consensus cost function, and the characteristics of the connectivity graph, a third global consensus cost function and an optimal gain factor are obtained; A determination unit is used to obtain, when there is communication noise when the agent performs a task, a dynamic model of the agent affected by the communication noise based on the dynamic model of the agent and the communication noise; obtain a judgment condition of positive excitation noise based on the dynamic model of the agent affected by the communication noise, the adaptive filtering mechanism and the optimal gain factor, and determine the noise that meets the judgment condition in the game environment as positive excitation noise.

8. The device for determining the continuous action iterative dilemma CAID noise according to claim 7, characterized in that: The agent dynamics model is as follows: The agent dynamics model affected by communication noise is as follows: The judgment conditions of the positive excitation noise include: in, represents the dynamic model of agent i, x i (k+1) represents the execution strategy of agent i in the k+1th iteration, x i (k) represents the execution strategy of agent i in the k+1th iteration, A ij represents the elements of the weighted adjacency matrix, θ i represents the degree of agent i, τ ij represents the learning rate of agent i, x j (t) represents the execution strategy of agent j at time t, x i (t) represents the execution strategy of agent i at time t, N represents the number of agents in the undirected graph, represents the agent dynamics model affected by communication noise, S ij represents the communication noise, k ij represents the communication noise coefficient, τ ij (t) represents the learning rate, represents the optimal learning rate, represents the learning rate affected by noise, t0 and t f An interval representing the cumulative mean square error.

9. A computer device, characterized in that: The computer device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the method for determining the continuous action iterative dilemma CAID noise according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor executes the method for determining continuous action iterative dilemma CAID noise according to any one of claims 1 to 6.