A Multi-Agent Pursuit-Evasion Game Method and Device Based on Reinforcement Learning
Through the combination of reinforcement learning and fuzzy learning, the problems of limited boundary value conditions and poor robustness in multi-agent pursuit and escape game are solved, and the optimal game strategy is generated, which is suitable for complex environmental tasks of drones, unmanned vehicles and other agents.
Patent Information
- Application Number
- CN202211552727.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-06
AI Technical Summary
The existing technology has problems such as limited boundary value conditions, poor robustness, and inability to deal with incomplete information and uncertain factors in solving the multi-agent pursuit and escape game. Especially in the application of drones, unmanned vehicles and other agents, it is difficult to generate optimal game strategies.
Using reinforcement learning-based methods, through fuzzy learning and Q learning, the optimal game strategy is generated independently, and the fuzzy processing and defuzzing algorithm are used, combined with local Q value tables and ε-greedy strategies, the dimensional disaster problem of continuous state space is solved, and the selection of optimal input state variables and the determination of control amount is realized.
It realizes the generation of autonomous strategy for multi-agent pursuit and escape game, which is globally optimal and robust, avoids dimensional disasters, and is suitable for the pursuit and escape tasks of agents in complex environments.
Smart Images

Figure CN115952729B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multi-agent reinforcement learning, and particularly to a multi-agent pursuit-evasion game method and device based on reinforcement learning. Background Art
[0002] The pursuit-evasion game problem is a typical differential game problem, which was first applied to military confrontation fields such as UAV combat, ship confrontation, satellite interception, and missile interception. With the popularization of intelligent agents such as UAVs and unmanned vehicles, the pursuit-evasion model also plays a very important role in the scheduling of industrial products, search and rescue, supervision and management, transportation management, etc. However, the numerical solution method based on the two-point boundary value problem is limited by the boundary conditions, has poor robustness, and cannot solve the pursuit-evasion game under incomplete information, the multi-person pursuit-evasion game under linear and planar dynamic models, and the pursuit-evasion differential game considering uncertain factors, etc. Summary of the Invention
[0003] The purpose of the present invention is to provide a multi-agent pursuit-evasion game method and device based on reinforcement learning. Through the game data of multi-agent pursuit-evasion, and by using fuzzy learning and Q learning for the exploration and exploitation of the environment, an optimal game strategy can be autonomously generated.
[0004] To achieve the above purpose, the present invention provides the following solutions:
[0005] A multi-agent pursuit-evasion game method based on reinforcement learning, comprising:
[0006] Performing fuzzy processing on the relative position state of the current pursuer and evader to determine the fuzzy state of the relative position state in the reinforcement learning device, obtaining the current fuzzy state variable;
[0007] According to the current fuzzy state variable and the trained association function, obtaining the maximum Q value function;
[0008] Based on the maximum Q value function, selecting the input state variable according to the optimal value under the current fuzzy state variable, obtaining the optimal input state variable strategy of the pursuit-evasion game training model in the current state;
[0009] Performing defuzzification processing on the optimal input state variable strategy by using a defuzzification algorithm to obtain the final actual control quantity.
[0010] Preferably, the training process of the association function includes:
[0011] Selecting the state variables of the pursuit-evasion game training model of the pursuer and evader, and storing the state variables of the pursuit-evasion game training model in the form of a fuzzy set;
[0012] Construct a local correlation function of the state variables of the pursuit-evasion game training model at the current moment and their adjacent state variables based on the state variables of the pursuit-evasion game training model at the current moment; the local correlation function is the local Q-value table.
[0013] Give the update rule of the correlation function in the fuzzy rule.
[0014] Determine the temporal difference error based on the update rule.
[0015] Update the local correlation function based on the temporal difference error to obtain the Q-value function at the next moment.
[0016] Use the Q-value function at the next moment as the output of the fuzzy inference device, and update the parameters of the fuzzy inference device using the gradient descent method.
[0017] Select the output variable result value according to the local Q-value table and the ε-greedy strategy.
[0018] Perform defuzzification on the input state variables using the weighted average method to obtain the action output at the next moment.
[0019] Input the action output into the pursuit-evasion game training model to obtain the model state variables at the next moment.
[0020] Obtain the reward at the next moment based on the update rule of the local correlation function in the given fuzzy rule.
[0021] Return to execute "Select the state variables of the pursuit-evasion game training models of both the pursuer and the evader, and store the state variables of the pursuit-evasion game training model in the form of a fuzzy set" until the local correlation function in the fuzzy rule converges.
[0022] Preferably, the pursuit-evasion game training model is:
[0023]
[0024] where t is the current moment, ξ(t) is the state variable at the current moment, is the differential of the state variable ξ(t) at the current moment, F(*) is the motion state dynamics model, G(*) is the input state dynamics model of the pursuer, K(*) is the input state dynamics model of the evader, U p is the input state variable of the pursuer, U e is the input state variable of the evader.
[0025] Preferably, the update rule of the local correlation function in the given fuzzy rule specifically includes:
[0026] Define the initial local correlation function.
[0027] The initial local association function is continuously updated using the Q - learning algorithm until convergence.
[0028] According to the specific embodiments provided by the present invention, the following technical effects are disclosed:
[0029] The present invention realizes the generation of strategies for multi - agent pursuit - evasion games through self - gaming. Based on the game data of multi - agent pursuit - evasion, by using fuzzy learning and Q - learning for the exploration and utilization of the environment, it can autonomously generate optimal game strategies. Moreover, the present invention reasonably divides the state - action space by using a fuzzy method. The Nash equilibrium solution generated according to the rules has global optimality and robustness. The local Q - value table composed of adjacent states of the current state avoids the curse of dimensionality problem caused by the continuous state space.
[0030] In addition, the present invention also provides a multi - agent pursuit - evasion game device based on reinforcement learning. The device includes:
[0031] A memory for storing computer control instructions; the computer control instructions are used to implement the above - provided multi - agent pursuit - evasion game method based on reinforcement learning;
[0032] A processor, connected to the memory, for retrieving and executing the computer control instructions.
[0033] Preferably, the processor includes:
[0034] A fuzzification processing module for fuzzifying the relative position state of the current pursuer - evader pair to determine the fuzzy state in which the relative position state is located in the reinforcement learning device to obtain the current fuzzy state variable;
[0035] A Q - value function determination module for obtaining the maximum Q - value function according to the current fuzzy state variable and the trained association function;
[0036] A variable strategy determination module for selecting the input state variable according to the optimal value under the current fuzzy state variable based on the maximum Q - value function to obtain the optimal input state variable strategy of the pursuit - evasion game training model in the current state;
[0037] A control quantity determination module for defuzzifying the optimal input state variable strategy using a defuzzification algorithm to obtain the final actual control quantity.
[0038] Preferably, the memory is a computer - readable storage medium.
[0039] Since the technical effects achieved by the above - provided device of the present invention are the same as those achieved by the multi - agent pursuit - evasion game method based on reinforcement learning provided by the present invention, they will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041] Figure 1 It is a flowchart of the multi-agent pursuit-evasion game method based on reinforcement learning provided by the present invention;
[0042] Figure 2 It is the overall framework diagram of reinforcement learning provided by the present invention;
[0043] Figure 3 It is the fuzzy device structure diagram provided by the present invention;
[0044] Figure 4 It is the fuzzy inference device structure diagram provided by the present invention;
[0045] Figure 5 It is the flowchart of the double-fuzzy device Q-learning architecture provided by the present invention;
[0046] Figure 6 It is the aircraft pursuit-evasion model diagram provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0048] Reinforcement learning is a learning method that interacts with the environment using a "trial and error" approach. It can be characterized by a Markov decision process. By calculating the size of the expected value of the cumulative return after performing an action in the current state, the rationality of the action selection can be judged. Based on this, through the Q-value function table generated by reinforcement learning, the long-term benefits of actions are considered, and it has a strong rate of return. Moreover, the learning process of the agent interacting with the environment does not depend on a good initial value, does not require solving the first-order necessary conditions, and only needs the return value of the environment to evaluate the executed action. Therefore, by establishing a multi-agent reinforcement learning model and allowing the agent to continuously explore and learn in the simulation environment and iterate repeatedly, an optimal fuzzy rule base can be generated to provide an optimal equilibrium solution for the agent's pursuit-evasion game.
[0049] The instantiation of specific practical problems within the framework of reinforcement learning is limited by the curse of dimensionality in the state space. To overcome the large-scale continuous state space, reasonable state space partitioning and description can reduce the scale of the problem and improve the efficiency and stability of reinforcement learning. Secondly, by introducing a local Q-value table, the dimensionality during training is reduced from the state-action space dimension to the action space dimension.
[0050] Based on this, the present invention provides a multi-agent pursuit-evasion game method and device based on reinforcement learning. Through the game data of multi-agent pursuit-evasion, and using fuzzy learning and Q-learning for the exploration and exploitation of the environment, it can autonomously generate optimal game strategies.
[0051] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0052] As Figure 1 shown, a multi-agent pursuit-evasion game method based on reinforcement learning according to the present invention includes:
[0053] Step 100: Perform fuzzy processing on the relative position state of the current pursuer and evader, and determine the fuzzy state in which the relative position state is located in the reinforcement learning device to obtain the current fuzzy state variable.
[0054] Step 101: According to the current fuzzy state variable and the trained correlation function, obtain the maximum Q-value function.
[0055] Step 102: Based on the maximum Q-value function, select the input state variable according to the optimal value under the current fuzzy state variable, and obtain the optimal input state variable strategy of the pursuit-evasion game training model in the current state.
[0056] Step 103: Use the defuzzification algorithm to perform defuzzification processing on the optimal input state variable strategy to obtain the final actual control quantity.
[0057] Furthermore, the training process of the above-mentioned correlation function includes:
[0058] Step 1: Select the state variables of the pursuit-evasion game training model for the pursuer and evader, and store the state variables of the pursuit-evasion game training model in the form of a fuzzy set. Specifically, select the state variables s=(x1,...,x N ) and action variables A={a1,a2,...,a M} of the pursuit-evasion game dynamics model, and divide the value space of each state into a superposition combination of multiple trigonometric functions through the triangular membership function to store the continuous variables in the form of a fuzzy set.
[0059] Among them, the pursuit-evasion game training model is:
[0060]
[0061] where \(t\) is the current time, \(\xi(t)\) is the state variable at the current time, is the differential of the state variable \(\xi(t)\) at the current time, \(F(*)\) is the dynamic model of the motion state, \(G(*)\) is the dynamic model of the input state of the pursuer, \(K(*)\) is the dynamic model of the input state of the escapee, \(U\) p is the input state variable of the pursuer, \(U\) e is the input state variable of the escapee.
[0062] Step 2: Construct the local correlation function of the state variable of the pursuit-evasion game training model at the current time and its adjacent state variables according to the state variable of the pursuit-evasion game training model at the current time The local correlation function is the local Q-value table. Specifically, through the Q-value function table of the fuzzy inference system, the adjacent local Q-value table is constructed based on the state-action value at the current time.
[0063] Step 3: Give the update rule of the correlation function in the fuzzy rule, specifically including: defining the initial local correlation function. Continuously update the initial local correlation function using the Q-learning algorithm until convergence.
[0064] Step 4: Determine the temporal difference error based on the update rule. The temporal difference error is where \(Q\) t (s t ) is the global Q-value function, is the maximum Q-value function, \(\gamma\in[0,1]\) is the discount factor, \(r\) t+1 is the reward at the next time. is the output variable result value at time \(t\), is the centroid of rule \(l\) at time \(t\) (i.e., the proportion value of rule \(l\) among all rules), \(q\) t is the local correlation function value at time \(t\), \(l = 1,2,\cdots,L\), and \(L\) is the total number of fuzzy rules.
[0065] Step 5: Update the local correlation function based on the temporal difference error to obtain the Q-value function \(Q\) t+1 (s t ) at the next time.
[0066] Step 6: Use the Q-value function \(Q\) t+1 (s t ) at the next time as the output of the fuzzy inference device, and update the parameters \(\Theta\) of the fuzzy inference device using the gradient descent method t+1。
[0067] Step 7: Select the output variable result value a according to the local Q-value table and the ε-greedy strategy l 。
[0068] Step 8: Perform defuzzification on the input state variables using the weighted average method to obtain the action output U at the next moment t+1 (S t )。
[0069] Step 9: Input the action output U t+1 (S t ) into the pursuit-evasion game training model to obtain the model state variable S at the next moment t+1 。
[0070] Step 10: Obtain the reward r at the next moment based on the update rule of the local correlation function in the given fuzzy rules t+1 。
[0071] Step 11: Return to execute Steps 1 - 10 until the local correlation function in the fuzzy rules converges, and the trained correlation function can be obtained
[0072] Embodiment 2
[0073] The reinforcement learning algorithm adopted in this embodiment has the basic architecture as Figure 2 shown. The agent interacts with the environment through states, actions, rewards. Assume the state of the environment at time t is denoted as st, and the agent executes an action a in the environment t . At this time, the action a t changes the original state of the environment and makes the agent reach a new state s at time t + 1 t+1 , and in the new state, the environment generates a feedback reward r t+1 for the agent. The agent based on the new state s t+1 and the feedback reward r t+1 executes a new action a t+1 , and so on, iteratively interacting with the environment through feedback signals. Reinforcement learning is an intelligent autonomous learning method. It does not require a tutor signal and does not require a strict mathematical model. Instead, it uses a "trial and error" method to interact with the environment for learning, continuously tries different behavioral strategies and improves them, adapts to the dynamic unknown environment, and effectively solves research hotspots such as incomplete information game problems
[0074] Based on this, the multi-agent pursuit-evasion game method based on reinforcement learning provided in this embodiment includes:
[0075] Step 1: Construct a pursuit-evasion game training model including a pursuer agent and an evader agent (hereinafter referred to as the pursuer and evader):
[0076]
[0077] Wherein, t represents the current moment of the system, and ξ represents the state variable of the dynamic model. represents the differential of the state variable of the dynamic model, generally including the position, velocity, Euler angles, etc. in the inertial coordinate system. F represents the dynamic model of the motion state, and G and K respectively represent the input state dynamic models of the pursuer and the evader. Generally, it is assumed that the agent is a rigid body model, and U p , U e are input state variables, where the subscripts p and e respectively represent the pursuer and the evader.
[0078] The objective functions of the pursuer and the evader are respectively established as follows:
[0079]
[0080] Wherein, J represents the objective function of the pursuer and the evader, d(t f ) represents the relative distance between the two parties, t0 and t f respectively represent the initial time and the termination time, and W p , W e respectively represent the energy weight coefficients. respectively represent the transpose of the input state variables, and d τ represents the integration factor.
[0081] Step 2: According to the fuzzy system model as shown in Figure 3 , construct a fuzzy Q-learning model for the multi-agent pursuit-evasion game. The fuzzy system model mainly includes modules such as input, fuzzification, fuzzy inference engine, fuzzy rules, defuzzification, and output.
[0082] Step 2-1: Select the state variables s = (s1,..., s N ) and action variables (i.e., input state variables) A = {a1, a2,..., a M} of the pursuit-evasion game training model for the pursuer and the evader. Wherein, s represents the set of model state variables, (s1,..., s N ) represents the N components of the model state variables, N is determined according to the actual pursuit-evasion game training model, A represents the set of model action variables, and {a1, a2,..., a M} represents the M components of the model action variables, and M is determined according to the actual pursuit-evasion game training model. And the value space of each state is divided into a superposition combination of multiple trigonometric functions through the triangular membership function, so as to store the continuous variables in the form of fuzzy sets. Among them, the triangular membership function is:
[0083]
[0084] Among them, x is the independent variable of the triangular membership function, a and c are used to represent the two lower vertices of the triangle, and b is used to represent the upper vertex of the triangle.
[0085] Step 2-2: Generate global continuous behavior according to the fuzzy rules, based on the ε-greedy strategy Select the specific value of the output variable result (output variable result value) a l
[0086] Among them, the fuzzy rule is: Indicates the l-th rule, s i , i = 1... N represents the state variables of the pursuit-evasion game training model of both the pursuer and the evader. Represents the corresponding fuzzy set of state variables, U l Represents the output variable result of the fuzzy rule, that is, the input state variable of the pursuit-evasion game training model, a l Represents the specific value of the output variable result, ε represents the probability value of the occurrence of a random event. Represents a given rule And the correlation function under the input state variable a.
[0087] Next, this embodiment gives the update rule of the correlation function in the fuzzy rule. First, define the Q-value function and the maximum Q-value function. Then, continuously update the correlation function through the Q-learning formula until it converges.
[0088] Among them, the Q-value function is:
[0089]
[0090] The maximum Q-value function is:
[0091] The Q-learning formula is:
[0092] In the formula, t represents the current time of the pursuit-evasion game training model, L represents the total number of fuzzy rules. Represents the maximum value of all input state variables a ∈ A in rule l, η is the learning rate. is the time difference error, γ ∈ [0, 1] is the discount factor, r t+1 is the reward at the next moment, obtained based on the differential of the objective function of both the pursuer and the evader established in Step 1.
[0093] Step 2-3: Through the Q-value fuzzy inference system in Step 2-1, construct the local correlation function of the current state and its adjacent states according to the current moment state variable value of the pursuit-evasion game training model in Step 1.
[0094] Step 2-4: Perform defuzzification according to the weighted average method, that is, convert a definite fuzzy quantity into a definite state variable or input variable. For example, the continuous input state variable at time t+1 is
[0095] Among them, the weighted average method is:
[0096] In the formula, represents the membership degree of the fuzzy set .
[0097] Step 3: According to the fuzzy inference system structure shown in Figure 4 , construct the Q-value fuzzy inference system of the pursuit-evasion game training model for both the pursuer and the evader. Figure 4 Among them, g represents the output of the fuzzy inference system, x i represents the input of the fuzzy inference system, N represents the dimension or number of input variables, Π represents the product, and ∑ represents the sum. To simplify the description of the hierarchical structure of the fuzzy inference system, this figure is illustrated with two inputs as an example. The actual number of input variables is determined by the specific model.
[0098] Step 3-1: Construct a fuzzy inference system. In the first layer, each state variable and input state variable of the pursuit-evasion game training model for both the pursuer and the evader correspond to three Gaussian membership functions Among them, x i represents the independent variable corresponding to the state variable, and σ and μ are the standard deviation and mean respectively; in the second layer, perform a product operation on each input; in the third layer, all nodes are fixed coefficients C l ; the fourth and fifth layers are used for defuzzification. The output of each layer can be expressed as H l :
[0099]
[0100] Among them, represents the l-th quantity in the j-th layer, L j represents the total number of quantities in the j-th layer, j = 1,..., 5, Π represents the product, and Σ represents the sum.
[0101] Step 3-2: According to the gradient descent method and the chain rule, we can obtain
[0102]
[0103] Among them, the gradient descent method is:
[0104] In the formula, Θ = [σ, μ, C] is the parameter list of the Q-value fuzzy inference system.
[0105] Step 4: The algorithm structure and process of multi-agent pursuit-evasion game reinforcement learning based on dynamic double fuzzy system Q-learning are as follows Figure 5 Let the current time be t, and the current model state variable be S t and the agent has executed the input state variable U t and has obtained the reinforcement learning reward r t , then the algorithm runs as follows:
[0106] Step 4-1: According to Step 2-1, store the continuous variable in the form of a fuzzy set as z t .
[0107] Step 4-2: Construct the local association function through Step 2-3 that is, the local Q-value table, and calculate the time difference error through Step 2-2 Update the association function and calculate the Q-value function Q at the next moment t+1 (s t ).
[0108] Step 4-3: Take the Q-value function Q t+1 (s t ) obtained in Step 4-2 as the output of the fuzzy inference system in Step 3, and update the parameters of the Q-value fuzzy inference system to Θ t+1 ;
[0109] Step 4-4: According to the local Q-value table obtained in Step 4-2, select the specific value a of the output variable result according to the ε-greedy strategy defined in Step 2-2 l ;
[0110] Step 4-5: Perform defuzzification operation on the input state variable according to Step 2-4 to generate the action output U at time t+1 t+1 (S t );
[0111] Step 4-6: Use U t+1 (S t ) to substitute into the pursuit-evasion game training model constructed in Step 1 to obtain the model state variable S at the next moment t+1 ;
[0112] Step 4-7: Obtain the reward r at the next moment according to Step 2-2 t+1
[0113] Step 4-8: The algorithm transfers to Step 4-1 to loop again until the association function in the fuzzy rule converges, that is, the training of the local association function is completed.
[0114] Step 5: Calculation of input state variables. The input quantity strategy in the pursuit-evasion game process is based on the finally converged correlation function Q t Make a selection and obtain a continuous control quantity through defuzzification. The calculation process is as follows:
[0115] Step 5-1: Fuzzify the relative position state of the current pursuer and evader, and judge the fuzzy state in which they are located in the reinforcement learning system;
[0116] Step 5-2: When the pursuit-evasion game motion state is S t According to the correlation function that has been trained in Step 4, obtain the maximum Q-value function
[0117] Step 5-3: Select the input state variables according to the optimal value in the current training model state S t to obtain the optimal input state variable strategy of the training model in the current state.
[0118] Step 5-4: Use the defuzzification formula to defuzzify the obtained optimal behavior to obtain the final actual control quantity.
[0119] Example Three
[0120] Learn and train the multi-agent pursuit-evasion game method based on dynamic fuzzy Q-learning provided in the above Example Two and Example Three. For the convenience of explanation, an agent dynamics model is used as Figure 6 shown Figure 6 where q2 is the target line-of-sight angle of the pursuer looking at the evader, α and β are the lead angles corresponding to the pursuer and evader respectively, P e is the evading party, abbreviated as the evader, and P p is the pursuing party, abbreviated as the pursuer. The main learning and training process is as follows:
[0121] Step 1: Build a specific training model, such as where x and y represent position variables, v represents the speed variable, Ψ represents the yaw angle, u represents the input variable, and the relative distance between the pursuer and evader is
[0122] Step 2: Initialize the system parameter values.
[0123] For example, use W e = W p = 30, v e = 0.2 km / s, v p = 0.3 km / s, η = γ = 0.8, ε = 0.3. Initialize the system state values. For example, x e = 1, y e = 3, Ψ e = 0; x p = 0, yp = 0, Ψ p = 0。
[0124] Step 3: Fuzzify the variables in the fuzzy Q - learning, which is implemented using the triangular membership function. Among them, the input relative - distance state S d The number of fuzzy sets is The fuzzy state description is S d = {very close, close, medium, far, very far}, and the center constants of the fuzzy sets are {1, 2, 3, 4, 5}. The input state S α , S β Similarly, use the trigonometric function for fuzzification. The number of fuzzy sets is The center constants of the fuzzy sets are kπ / 4 (-4 < k ≤ 4, k is a constant), and S α , S β is expressed as S a , S β = {left, left - back, back, right - back, right, right - front, front, left - front}. The fuzzy action sets of the pursuer and the evader are the same. The number of fuzzy sets is The action set uses the same triangular membership function, and the center constants of the fuzzy sets are {-0.2, -0.1, 0, 0.1, 0.2}. The fuzzy action set is expressed as A e = A p = {right - fast, right - slow, straight - ahead, left - slow, left - fast}.
[0125] Step 4: According to the algorithm flow in Step 4 of Embodiment 3, train the training model until the correlation function converges; that is, determine the size of the Q - matrix of the pursuer and the evader in the fuzzy Q - learning according to the fuzzy Q - learning parameter selection, and continuously update and iterate the fuzzy Q - learning matrix according to the reinforcement - learning iteration process. After multiple trainings, use the fuzzy rule base generated by the model as the decision basis.
[0126] Step 5: According to the calculation process in Step 5 of Embodiment 3, obtain the optimal input - state variable sequence of the training model in chronological order. That is, the input - quantity strategy in the pursuer - evader game is selected according to the finally - converged fuzzy rule base, and the continuous control quantity is obtained through defuzzification.
[0127] Embodiment 4
[0128] This embodiment provides a multi - agent pursuer - evader game device based on reinforcement learning. The device includes: a memory and a processor. The memory used can be a computer - readable storage medium.
[0129] The memory is used to store computer control instructions. The computer control instructions are used to implement the multi - agent pursuer - evader game method provided in Embodiment 1 or Embodiment 2 above.
[0130] The processor is connected to the memory. The processor is mainly used to retrieve and execute computer control instructions.
[0131] Furthermore, the processor provided in this embodiment may further include: a fuzzification processing module, a Q-value function determination module, a variable strategy determination module, and a control quantity determination module.
[0132] Among them, the fuzzification processing module is used to perform fuzzification processing on the relative position state of the current pursuer and evader, and determine the fuzzy state in which the relative position state is located in the reinforcement learning device to obtain the current fuzzy state variable.
[0133] The Q-value function determination module is used to obtain the maximum Q-value function according to the current fuzzy state variable and the trained association function.
[0134] The variable strategy determination module is used to select the input state variable according to the optimal value under the current fuzzy state variable based on the maximum Q-value function, and obtain the optimal input state variable strategy of the pursuit-evasion game training model under the current state.
[0135] The control quantity determination module is used to perform defuzzification processing on the optimal input state variable strategy by using the defuzzification algorithm to obtain the final actual control quantity.
[0136] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0137] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A multi-agent pursuit-evasion game method based on reinforcement learning, characterized in that, Including: Fuzzify the relative position state of the current pursuer and evader, and determine the fuzzy state of the relative position state in the reinforcement learning device to obtain the current fuzzy state variable; According to the current fuzzy state variable and the trained association function, obtain the maximum Q-value function; Based on the maximum Q-value function, select the input state variable according to the optimal value under the current fuzzy state variable, and obtain the optimal input state variable strategy of the pursuit-evasion game training model under the current state; Use the defuzzification algorithm to defuzzify the optimal input state variable strategy to obtain the final actual control quantity; The training process of the association function includes: Select the state variables of the pursuit-evasion game training model of the pursuer and evader, and store the state variables of the pursuit-evasion game training model in the form of a fuzzy set; among them, the value space of each state is divided into a superposition combination of multiple trigonometric functions through the triangular membership function, and the continuous variable is stored in the form of a fuzzy set; the pursuit-evasion game training model is: where \(t\) is the current time, \(\xi(t)\) is the state variable at the current time, is the differential of the state variable \(\xi(t)\) at the current time, \(F(*)\) is the dynamic model of the motion state, \(G(*)\) is the dynamic model of the input state of the pursuer, \(K(*)\) is the dynamic model of the input state of the escapee, \(U\) p is the input state variable of the pursuer, \(U\) e is the input state variable of the escapee; Construct the local association function of the state variable of the pursuit-evasion game training model and its adjacent state variables at the current moment according to the state variable of the pursuit-evasion game training model at the current moment; the local association function is the local Q-value table; Give the update rule of the association function in the fuzzy rule; Determine the temporal difference error based on the update rule; Update the local association function based on the temporal difference error to obtain the Q-value function at the next moment; Take the Q-value function at the next moment as the output of the fuzzy inference device, and use the gradient descent method to update the parameters of the fuzzy inference device; Select the output variable result value according to the local Q-value table and the ε-greedy strategy; Perform a defuzzification operation on the input state variable by using the weighted average method to obtain the action output at the next moment; Input the action output into the pursuit-evasion game training model to obtain the model state variable at the next moment; Obtain the reward at the next moment based on the update rule of the local association function in the given fuzzy rule; Return to execute "select the state variables of the pursuit-evasion game training model of the pursuer and evader, and store the state variables of the pursuit-evasion game training model in the form of a fuzzy set" until the local association function in the fuzzy rule converges.
2. The multi-agent pursuit-evasion game method based on reinforcement learning according to claim 1, wherein The update rule of the local association function given in the fuzzy rule specifically includes: Define the initial local association function; Use the Q-learning algorithm to continuously update the initial local association function until convergence.
3. A multi-agent pursuit-evasion game device based on reinforcement learning, characterized in that, Including: A memory for storing computer control instructions; the computer control instructions are used to implement the multi-agent pursuit-evasion game method based on reinforcement learning according to any one of claims 1-2; A processor connected to the memory for retrieving and executing the computer control instructions.
4. The multi-agent pursuit-evasion game device based on reinforcement learning according to claim 3, characterized in that, The processor includes: A fuzzification processing module for fuzzifying the relative position state of the current pursuer and evader, and determining the fuzzy state of the relative position state in the reinforcement learning device to obtain the current fuzzy state variable; A Q-value function determination module for obtaining the maximum Q-value function according to the current fuzzy state variable and the trained association function; A variable strategy determination module, configured to select an input state variable according to an optimal value under the current fuzzy state variable based on the maximum Q-value function, so as to obtain an optimal input state variable strategy of the pursuit-evasion game training model under the current state; A control quantity determination module, configured to perform defuzzification processing on the optimal input state variable strategy by using a defuzzification algorithm to obtain a final actual control quantity.
5. The multi-agent pursuit-evasion game device based on reinforcement learning according to claim 3, characterized in that, The memory is a computer-readable storage medium.