Multi-unmanned-boat pursuit-evasion game control method
By introducing sequential decision-making and collaborative design of observers and reinforcement learning into the pursuit party control algorithm of unmanned boats, the external interference and system uncertainty problems of unmanned boat clusters in pursuit and fugitive game control are solved, and a more efficient pursuit effect is achieved.
Patent Information
- Application Number
- CN202211507056.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-11-29
AI Technical Summary
Existing algorithms cannot effectively solve the control needs of unmanned boat clusters for pursuit and fugitive game, especially when facing external interference and system uncertainty.
Sequential decision-making is introduced into the chasing party control algorithm of unmanned boats, through the collaborative design of observers and reinforcement learning, self-game is carried out to approximate external interference and system uncertainty, and an optimal control strategy is formed through sequential game.
It improves the collaboration ability of unmanned boat clusters in pursuit and fugitive tasks, increases the success rate of pursuit, reduces the time required for pursuit, and the pursuit process is more stable and has better performance.
Smart Images

Figure CN115903820B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned boats, and particularly to a multi-unmanned boat pursuit-evasion game control method. Background Art
[0002] In recent years, with the exhaustion of land fuel resources, the strategic position of the ocean, which occupies about 71% of the earth's area, has been continuously improving. To fully explore and exploit marine resources, the development of marine equipment technology is indispensable. Marine intelligent equipment represented by unmanned boats (including underwater vehicles, underwater robots, surface unmanned boats, etc.) is the main carrier for maritime operations at the present stage.
[0003] Cluster unmanned boats refer to a group of multiple unmanned boats in formation. In recent years, the application of cluster unmanned boats has been increasing. Currently, cluster unmanned boats have played an important role in military fields such as encirclement and capture, expulsion, mine sweeping, and anti-submarine warfare, as well as in civilian fields such as material supply, terrain mapping, sea rescue, and unmanned search.
[0004] However, for the pursuit-evasion game control of unmanned boat clusters, there is currently no relatively effective algorithm that can be applied in practice. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-unmanned boat pursuit-evasion game control method to solve the problem that the existing algorithms cannot meet the pursuit-evasion game control requirements of unmanned boat clusters.
[0006] To solve the above technical problems, the present invention provides a multi-unmanned boat pursuit-evasion game control method, including:
[0007] When the unmanned boat is in pursuit-evasion, sequential decision-making is introduced into the control algorithm of the pursuer for self-play;
[0008] The observer calculates the optimal response of the observer according to the optimal control given by the controller to approximate the external interference and system uncertainty in the pursuer group; and
[0009] The controller receives the optimal response of the observer and recalculates the optimal control of the pursuer according to the optimal response, where the recalculation is performed one or more times to form sequential decision-making.
[0010] Optionally, in the multi-unmanned boat pursuit-evasion game control method, it further includes:
[0011] Cooperating the observer and the game of reinforcement learning, so that the observer acts as a follower to handle uncertainties, where the uncertainties include external interference and modeling errors;
[0012] Taking reinforcement learning as the controller to form a leader;
[0013] Through the sequential game between the observer and the controller, the reinforcement learning algorithm can cope with external disturbances and modeling errors, reach the Nash equilibrium, and achieve cooperative pursuit and capture in the game.
[0014] Optionally, in the multi-unmanned boat pursuit-evasion game control method described above, it further includes:
[0015] Set the reward function according to obstacle avoidance, tracking, surrounding, and control quantity consumption;
[0016] Use the reciprocal velocity obstacle method to set the obstacle avoidance reward and handle the obstacle avoidance problems of other static and dynamic obstacles at the same time;
[0017] Use the potential energy function to set the tracking reward and the surrounding reward. When the potential energy function is within the threshold distance and there is an obstacle avoidance requirement, the growth of the potential energy stops.
[0018] Optionally, in the multi-unmanned boat pursuit-evasion game control method described above, it further includes:
[0019] In the multi-unmanned boat pursuit-evasion game, each pursuer is set as a subsystem;
[0020] Use the observer of this subsystem as a follower to handle uncertainties including external disturbances and modeling errors;
[0021] Design the controller of this subsystem based on reinforcement learning to form a leader;
[0022] Enhance the interaction between the leader designed based on reinforcement learning and the environment through the observer to improve the control performance. Conversely, use the improvement of the control performance to improve the observation performance of the observer;
[0023] According to this process, establish a sequential game graph between the leader and the follower;
[0024] The escape unmanned boat strategy adopts the inherent model and strategy.
[0025] Optionally, in the multi-unmanned boat pursuit-evasion game control method described above, it further includes:
[0026] Step 1: According to the conventional unmanned boat swaying, yawing, and rolling motion equations:
[0027]
[0028] where v i (t), r i (t), ψ i (t), p i (t), φ i (t), u i (t) and f ψi (t), f φi(t) are respectively expressed as the swaying speed, yawing speed, yaw angle, rolling speed, roll angle, rudder angle of the \(i\)th following unmanned surface vehicle, and unknown uncertainties, \(\zeta,\omega\) n are expressed as the damping ratio and natural frequency, \(T\) v , \(T\) r are expressed as the time constant, \(K\) dv , \(K\) dr , \(K\) vr , \(K\) dp , \(K\) vp are expressed as the unmanned surface vehicle system gain;
[0029] Step 2: According to the swaying, yawing, and rolling motion equations of the \(i\)th following unmanned surface vehicle in the steps, define the system state \(x\) of the dynamic equation of the following unmanned surface vehicle i (t), and the output \(y\) measured by the angle sensor i (t), unknown uncertainties \(f\) caused by other factors such as waves and wind disturbances i (t) are respectively \(x\) i (t)=[v i (t)r i (t)\(\psi\) i (t)p i (t)\(\varphi\) i (t)] T , \(y\) i (t)=[\(\psi\) i (t)\(\varphi\) i (t)] T , \(f\) i (t)=[f ψi (t)f φi (t)] T , and the dynamic equation of the following unmanned surface vehicle is expressed as follows:
[0030]
[0031]
[0032] Abbreviate the dynamic equations of each unmanned surface vehicle subsystem as:
[0033]
[0034] Step 3: Design an observer for the \(i\)th unmanned surface vehicle subsystem, specifically:
[0035]
[0036] where \(L\) is the observer parameter matrix;
[0037] System error
[0038]
[0039] Optionally, in the above multi-unmanned-boat pursuit-evasion game control method, it further includes:
[0040] Step 4: To form a complete sequential game process, an auxiliary control law vi is introduced to make the observer and the controller form a non-cooperative game. The following observer is designed to improve the design of the observer in Step 3:
[0041]
[0042] The following performance index function is introduced to optimize the observer performance:
[0043]
[0044] where u i T Gv i One represents the influence of subsystem i on sub-observer i, and Q, R, G are symmetric positive definite matrices, which are used to adjust the weight ratio between the constraints in the performance index function.
[0045] Optionally, in the above multi-unmanned-boat pursuit-evasion game control method, it further includes:
[0046] Step 5: In the sequential game decision-making of the pursuer, the optimal response of sub-observer i of the pursuer i needs to be considered first. Assume that the control law u of subsystem i i is initialized as an admissible control at the beginning of the game, and the following sub-observer Hamiltonian function is introduced:
[0047]
[0048] where is the partial derivative of the performance index with respect to ;
[0049] The optimal performance index value satisfies the Hamilton-Jacobi (HJ) equation:
[0050]
[0051] The necessary condition for solving equation (10) is
[0052]
[0053] Ideally At this time Using adaptive dynamic programming to solve, the optimal auxiliary control law is obtained as:
[0054]
[0055] Substitute Equation (12) into the Hamilton-Jacobi (HJ) equation shown in Equation (10):
[0056]
[0057] Optionally, in the multi-unmanned-boat pursuit-evasion game control method described above, it further includes:
[0058] Step Six: On the basis of the fifth step, introduce the following control leader optimization objective function for the pursuer subsystem i based on reinforcement learning:
[0059]
[0060] Design the control law such that the following equation holds:
[0061]
[0062] where one term reduces the consumption of the control quantity, one term introduces an auxiliary control law to form a complete sequential non-cooperative game; δ i is the reward function, which consists of three parts: respectively, the tracking potential energy that plays a role in local obstacle avoidance the tracking potential energy that plays a role in tracking the evader and the surrounding potential energy that plays a role in circular navigation and surrounding 0 L and I are symmetric positive definite matrices, k 1 k 2 k 1 should be slightly greater than k 2 ;
[0063]
[0064]
[0065]
[0066]
[0067] In Equation (17), v t represents the current speed, and different rewards are given by judging whether v t belongs to the reciprocal speed obstacle method region;
[0068] a, b, c, d, e, f are constant values for adjusting the strategy performance, diff v represents the difference between the current speed of the unmanned boat and the desired speed, and ξ is the expected shortest time to collide with an obstacle at the current speed;
[0069] For equations (18) and (19), d ie , d e0 are the actual distance and the desired distance between the current pursuer and the evader respectively, and ε is an adjustable hyperparameter; if d ij <α, it is regarded that a collision occurs between the agents, and α is a small positive constant; so the situation of d ij = 0 will not occur, and similarly the situation of d ie = 0 will not occur; b ij is an indicator function, as shown in equation (20), where d range represents the action distance of the circumnavigation potential energy, and d range <d e0 / m, m>1 is a constant; which means: with the current pursuer as the center, a region with a radius of d range , the current pursuer only has the potential energy effect of circumnavigation around with other pursuers within this region;
[0070]
[0071] Optionally, in the multi-unmanned-boat pursuit-evasion game control method described above, it further includes:
[0072] Converting the unequal-length environmental state sequence into an equal-length state sequence, and using a BiGRU (Bidirectional Gated Recurrent Unit) to process the unequal-length environmental state sequence, where represents the state information of the i-th obstacle detected within the detection range of the pursuer, and o self represents the state information of the current unmanned boat itself, h ∈ R mx1 represents the environmental state information of the i-th pursuer detection range extracted by the BiGRU;
[0073] Then the environmental state information of the current pursuer can be expressed as:
[0074]
[0075] Optionally, in the multi-unmanned-boat pursuit-evasion game control method described above, it further includes:
[0076] Step eight: Designing a controller using reinforcement learning, the reinforcement learning algorithm framework adopts the approximate policy optimization algorithm, and according to the loss function in equation (14), define the following action value function:
[0077]
[0078] The local action reward function in the time interval (t, t + h] can be defined as:
[0079]
[0080] Let \(h\) be the time interval for each sampling. When \(h\rightarrow0\), the approximate equation (24) holds:
[0081]
[0082] Therefore, if the immediate reward function is defined as shown in equation (24), then the discounted return \(G\) t is as shown in equation (25), where \(\gamma\) is the discount factor:
[0083] \(G\) t \(=R\) t \(+\gamma R\) t+1 \(+\gamma\) 2 \(R\) t+2 \(+\gamma\) 3 \(R\) t+3 \(+\cdots\) (50)
[0084] Denote the current state of the unmanned boat as the action as \(u\), and the policy as \(\pi\). The action-value function \(Q -\) value and the state-value function \(V -\) value are:
[0085] \(Q\) π \((s\) t , \(u\) t ) \(=E\) π \((G\) t \(|S\) t \(=s\) t , \(U\) t \(=u\) t ) (51)
[0086] \(V\) π \((s\) t ) \(=E\) π \((G\) t \(|S\) t \(=s\) t ) (52)
[0087] Given the system state equation, the immediate reward function, and the state \(s\) t and the action \(u\) t information, according to the approximate policy optimization algorithm in reinforcement learning, fit the value function and the policy function, and give the optimal control \(u\) t ;
[0088] Step Nine: Based on the optimal control law of the subsystem generated in Step Eight introduce it in Step Five and minimize the performance index to obtain the optimal auxiliary control law and repeat this sequential game process, thereby completing the pursuit and encirclement of the escapee.
[0089] The inventors of the present invention have found through research that when using the reinforcement learning algorithm to complete the pursuit and evasion tasks of unmanned boats, the environmental uncertainties such as wind disturbance and wave disturbance, as well as the errors caused by modeling uncertainties, are rarely considered, making it difficult to apply the designed algorithm in practice.
[0090] Secondly, when using the pure control algorithm relying on the model to handle the pursuit and evasion tasks, difficulties such as underactuation and nonlinearity are often encountered, and the optimal control law is often difficult to solve.
[0091] In addition, the existing reinforcement learning control technologies often focus on the control objective (effect), rarely considering the issue of control energy consumption. In practical applications, the design and solution of the control strategy with the least energy consumption still need to be further studied.
[0092] Finally, in the current pursuit and evasion problems, the design of the reward function is mostly based on the absolute position of the unmanned boat, without considering the collision risks brought by relative speed, acceleration, etc., which restricts the further improvement of the control effect. At the same time, a large amount of offline and online training is also required.
[0093] Based on the above insights, the present invention provides a multi-unmanned-boat pursuit and evasion game control method. By introducing sequential decision-making into the control algorithm of the pursuer and performing "self-game": in the face of unknown factors such as external disturbances and system uncertainties, the observer is an effective method to solve the problem. An observer is introduced to solve this unknown factor, and then, the auxiliary controller improves the control effect, solving the problems that the uncertainty brought by the "trial-and-error - correction" characteristic of reinforcement learning is difficult to handle and the algorithm strategy is difficult to be practically applied when using the reinforcement learning algorithm for pursuit and evasion games. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] Figure 1 is a schematic diagram of the sequential game between the leader and the follower in an embodiment of the present invention;
[0095] Figure 2 is a schematic flowchart of the pursuer control algorithm based on reinforcement learning in an embodiment of the present invention;
[0096] Figure 3 is a schematic diagram of the potential energy function image of formula (18) in an embodiment of the present invention;
[0097] Figure 4 is a schematic diagram of the detection of obstacles by an unmanned boat in an embodiment of the present invention;
[0098] Figure 5 is a schematic diagram of extracting environmental features in an embodiment of the present invention;
[0099] Figure 6 is a schematic diagram of using the PPO algorithm to output the action u t in an embodiment of the present invention;
[0100] Figure 7 It is a schematic diagram of the sequential decision-making workflow of an embodiment of the present invention. Detailed implementation manners
[0101] The present invention will be further elaborated below in conjunction with the detailed implementation manners with reference to the accompanying drawings.
[0102] It should be noted that the components in the respective drawings may be exaggeratedly shown for illustration purposes and are not necessarily to scale correctly. In the respective drawings, the same or functionally identical components are provided with the same reference numerals.
[0103] In the present invention, unless otherwise specified, the expressions "arranged on...", "arranged above...", and "arranged over..." do not exclude the situation where there are intermediate objects between the two. In addition, "arranged on or above..." only represents the relative positional relationship between two components, and in certain situations, such as when the product direction is reversed, it can also be converted to "arranged under or below...", and vice versa.
[0104] In the present invention, the respective embodiments are only intended to illustrate the solutions of the present invention and should not be construed as restrictive.
[0105] In the present invention, unless otherwise specified, the quantifiers "a" and "one" do not exclude the scenario of multiple elements.
[0106] It should also be noted here that in the embodiments of the present invention, for the sake of clarity and simplicity, only a part of the components or assemblies may be shown, but those of ordinary skill in the art can understand that, under the teaching of the present invention, the required components or assemblies can be added according to the specific scenario requirements. In addition, unless otherwise stated, the features in different embodiments of the present invention can be combined with each other. For example, a certain feature in the second embodiment can be used to replace the corresponding or functionally identical or similar feature in the first embodiment, and the resulting embodiment also falls within the disclosure scope or the recorded scope of the present application.
[0107] It should also be noted here that within the scope of the present invention, the expressions "the same", "equal", "equivalent", etc. do not mean that the two values are absolutely equal, but allow a certain reasonable error, that is, the said expressions also cover "substantially the same", "substantially equal", "substantially equivalent". By analogy, in the present invention, the directional terms "perpendicular to", "parallel to", etc. also cover the meanings of "substantially perpendicular to" and "substantially parallel to".
[0108] In addition, the numbering of the steps of the respective methods of the present invention does not limit the execution order of the method steps. Unless otherwise specified, the method steps can be executed in different orders.
[0109] The following further elaborates on the multi-unmanned boat pursuit-evasion game control method proposed by the present invention in conjunction with the accompanying drawings and specific embodiments. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings are all in a very simplified form and use non-precise scales, only for conveniently and clearly assisting in explaining the purpose of the embodiments of the present invention.
[0110] The purpose of the present invention is to provide a multi-unmanned boat pursuit-evasion game control method to solve the problem that the existing algorithms cannot meet the pursuit-evasion game control requirements of unmanned boat clusters.
[0111] To achieve the above purpose, the present invention provides a multi-unmanned boat pursuit-evasion game control method, including: when the unmanned boat is in pursuit-evasion, sequential decision-making is introduced into the control algorithm of the pursuer to perform "self-game"; the observer calculates the optimal response of the observer according to the optimal control given by the controller to approximate the external interference and system uncertainty in the pursuer group; and the controller receives the optimal response of the observer and recalculates the optimal control of the pursuer according to the optimal response, and so on alternately to form sequential decision-making.
[0112] Figures 1 to 7 Embodiments of the present invention are provided, and the present invention solves the problems that the uncertainty brought by the "trial and error - correction" characteristic of reinforcement learning is difficult to handle and the algorithm strategy is difficult to be practically applied when using the reinforcement learning algorithm for pursuit-evasion games.
[0113] In the unmanned boat pursuit-evasion task, sequential decision-making is introduced into the control algorithm of the pursuer to perform "self-game": in the face of unknown factors such as external interference and system uncertainty, the observer is an effective method to solve the problem. In this patent, an observer is introduced to solve this unknown factor, and then, it assists the controller to improve the control effect. The observer first calculates the optimal response of the observer according to the optimal control given by the controller to approximate the uncertainty existing in the pursuer group. Then, the controller receives the optimal response information of the observer and recalculates the optimal control of the pursuer on this basis, and so on alternately to form sequential decision-making. When performing the pursuit-evasion task, the introduction of the sequential game method can improve the group cooperation ability of the pursuer, thereby increasing the success rate of the pursuit, reducing the time required for the pursuit and encirclement, and the pursuit process is more stable and has better performance.
[0114] Game collaborative design of the observer and reinforcement learning: The advantages of the observer and reinforcement learning control are comprehensively utilized. The observer is used as a follower to handle environmental uncertainty; reinforcement learning is used as the controller to form a leader. Through the sequential game between the observer and the controller, the designed algorithm can cope with uncertainty factors such as external interference and modeling errors and reach the Nash equilibrium to achieve the goal of game collaborative encirclement.
[0115] Better performance of the reinforcement learning algorithm: Due to the introduction of the observer, the agent can rely less on the observation data of the actual environment and can generate model-based data. While obtaining sufficient training data, it avoids the danger brought by direct interaction with the actual environment, thus ensuring the training and performance of the reinforcement learning algorithm.
[0116] The design of the reward function comprehensively considers obstacle avoidance, tracking, surrounding, and control effort consumption: The reciprocal velocity obstacle (RVO) method is used to design the obstacle avoidance reward. Based on the characteristics of RVO, the designed obstacle avoidance algorithm has better performance than the traditional potential field obstacle avoidance method, and it can not only avoid collisions with other unmanned boats but also handle the obstacle avoidance problems of other static and dynamic obstacles simultaneously; The potential function is used to design the rewards for tracking and surrounding, and the designed potential function can stop the growth of the potential at a relatively short distance when obstacle avoidance is required. Thus, the three parts of the rewards are reasonably distributed in their action areas; Since the control effort consumption problem is considered in the design of the reward function, the designed algorithm can take resource conservation into account.
[0117] In the pursuit-evasion game of multiple unmanned boats, each pursuer is set as a subsystem. The observer of this subsystem is used as a follower to handle environmental uncertainties; At the same time, the controller of this subsystem is designed based on reinforcement learning to form a leader. Through this follower, the observer enhances the interaction between the control leader designed based on reinforcement learning and the environment, improving the control performance. Conversely, the improvement of the control performance can enhance the observation performance of the observer follower. According to this process, the following sequential game graph between the leader (subsystem) and the follower (sub-observer) is established (as Figure 1 shown). Figure 2 It is the flowchart of the designed pursuer control algorithm based on reinforcement learning. At the same time, the evader unmanned boat strategy adopts an inherent model and strategy.
[0118] Step 1: According to the swaying, yawing, and rolling motion equations of a conventional unmanned boat:
[0119]
[0120] where v i (t), r i (t), ψ i (t), p i (t), φ i (t), u i (t) and f ψi (t), f φi (t) respectively represent the swaying speed, yawing speed, yaw angle, rolling speed, rolling angle, rudder angle, and unknown uncertainty of the i-th following unmanned boat, ζ, ω nDenoted as damping ratio and natural frequency, T v , T r Denoted as time constant, K dv , K dr , K vr , K dp , K vp Denoted as the gain of the unmanned boat system.
[0121] Step 2: According to the swaying, yawing, and rolling motion equations of the i following unmanned boats in the step, define the system state x of the dynamic equation of the following unmanned boats i (t), the measurable output y of the angle sensor i (t), the unknown uncertainties f caused by other factors such as waves and wind disturbances i (t) are x i (t) = [v i (t) r i (t) ψ i (t) p i (t) φ i (t)] T , y i (t) = [ψ i (t) φ i (t)] T , f i (t) = [f ψi (t) f φi (t)] T , and the following dynamic equation of the unmanned boat can be obtained as follows:
[0122]
[0123]
[0124] For the convenience of subsequent elaboration, the present invention simplifies the dynamic equations of each unmanned boat subsystem as:
[0125]
[0126] Step 3: Design an observer for the i-th unmanned boat subsystem, specifically:
[0127]
[0128] where L is the observer parameter matrix;
[0129] System error
[0130]
[0131] Step 4: To form a complete sequential game process, introduce an auxiliary control law vi , such that the observer and the controller form a non - cooperative game. Furthermore, design the following observer to improve the design of the observer in step three:
[0132]
[0133] Meanwhile, introduce the following performance index function to optimize the observer performance:
[0134]
[0135] where \(u\) i T \(G_v\) i One represents the influence of subsystem \(i\) on sub - observer \(i\), and \(Q\), \(R\), \(G\) are symmetric positive - definite matrices, which are used to adjust the weight ratio between various constraints in the performance index function.
[0136] Step Five: In the sequential game decision of the pursuer, it is necessary to first consider the optimal response (auxiliary control law) of sub - observer \(i\) of the pursuer \(i\). Assume that the control law \(u\) of subsystem \(i\) i is initialized as an admissible control at the beginning of the game, and introduce the following sub - observer Hamiltonian function:
[0137]
[0138] where is the partial derivative of the performance index with respect to .
[0139] The optimal performance index value satisfies the Hamilton - Jacobi (HJ) equation:
[0140]
[0141] The necessary condition for solving equation (10) is
[0142]
[0143] Ideally At this time Using ADP (Adaptive Dynamic Programming, a method for solving the optimal control law according to the performance index), the optimal auxiliary control law can be obtained as:
[0144]
[0145] Substitute equation (12) into the Hamilton - Jacobi (HJ) equation shown in equation (10):
[0146]
[0147] Step 6: Based on the fifth step, introduce the following pursuer subsystem i to optimize the objective function of the control leader based on reinforcement learning:
[0148]
[0149] Design control law So that the following is true:
[0150]
[0151] in One is to reduce the consumption of control volume, An auxiliary control law is introduced to form a complete sequential non-cooperative game. i is the reward function (defined as the smaller the better), which consists of three parts: Tracking potential energy to track the fugitive And the surrounding potential energy that plays a circumnavigation role (i.e. let the pursuers form a circle around the escapee), L and I are symmetric positive definite matrices, k 0 , k 1 , k 2 is a positive hyperparameter. Considering that the escapee should be captured after being caught, k 1 Should be slightly larger than k 2 .
[0152]
[0153]
[0154]
[0155]
[0156] In formula (17), v t Indicates the current speed, by judging v t Whether it belongs to the RVO area (speed domain), the reward is different. RVO (Reciprocal Speed Obstacle Method) is an obstacle avoidance algorithm that can simultaneously consider the position and relative speed of the current agent, so it has a good performance in the obstacle avoidance algorithm. The present invention designs the reward function of the obstacle avoidance part based on RVO. a, b, c, d, e, f are constant values that can adjust the performance of the strategy, diff v represents the difference between the current speed of the unmanned boat and the expected speed, and ξ is the expected shortest time to collide with the obstacle at the current speed. For equations (18) and (19), d ie ,d e0are the actual distance and the desired distance between the current pursuer (subsystem) and the evader respectively, and ε is an adjustable hyperparameter. If d ij < α, it is regarded that a collision occurs between the agents, and α is a small positive constant. So there will be no situation where d ij = 0, and similarly there will be no situation where d ie = 0. b ij is an indicator function, as shown in Equation (20), where d range represents the action distance of the circumferential navigation potential energy, and d range < d e0 / m, where m > 1 is a constant. It means that: with the current pursuer as the center, in the area with d range as the radius, the current pursuer only has the potential energy action of circumferential navigation around with other pursuers within this area.
[0157]
[0158] Next, specifically explain the three parts of the reward function:
[0159] The RVO reward function for obstacle avoidance The setting of the RVO to abandon the dangerous speed can largely ensure that there will be no collision between the agents, and this part will act crosswise with the other two potential energies, which can further constrain the obstacle avoidance problem at a relatively close distance. As the distance between the agents becomes farther, the role of the RVO becomes smaller, and at this time, mainly the other two potential energy functions are acting.
[0160] The tracking potential energy for tracking the evader Its role is to shorten the distance between the current pursuer and the evader until it remains at a desired distance d e0 (at this time the potential energy is 0). The function graph of Equation (18) is as Figure 3 shown (taking d e0= 5): In the part greater than 5: The potential energy almost linearly increases along y = x. In this way, normalization can be avoided, and training difficulties caused by excessive potential energy can also be avoided. Because when the two agents are far apart, such as at the initialization moment, there is a relatively large distance between the pursuer and the evader. At this time, the potential energy is large. After normalization, this part of the gap will become very small. If the pursuer and the evader shorten the distance at this time, the reward obtained will increase very little or hardly increase compared with the previous moment. Coupled with the overlapping effect of the three parts of the reward function, in fact, the reward can be regarded as not increasing, so the evader will not be encouraged to approach the pursuer. In the part less than 5: In the interval (ε, 5), it is jointly constrained by the potential energy function of this part and the RVO reward, so that the distance between the two agents is not too close. However, if the distance between the two agents approaches a certain extent, that is, in the interval (0, ε), the original potential energy stops increasing, and the RVO reward is used alone to avoid excessive overlap of the two parts of the reward. (At close range, multiple reward functions act together and overlap with each other, and the potential energy growth is relatively rapid at this time, which will bring unknown factors and training difficulties. Therefore, a working distance ε is artificially specified. When the distance between the two agents shrinks to ε, the potential energy no longer increases and remains a constant value. At this time, the RVO acts alone.)
[0161] The surrounding potential energy of the encircling Is generally consistent with the tracking potential energy in point 2, the difference is that an indicator function is introduced: Constraining the current pursuer only needs to cooperate with the left and right neighbors to complete the surrounding function.
[0162] Step 7: Considering the RVO reward item for obstacle avoidance: The number of obstacles (other unmanned boats) around each pursuer is not fixed, such as Figure 4 Shown: The left side is case 1, with two obstacles within the detection range, and the right side is case 2, with three obstacles within the detection range. The neural network framework in the reinforcement learning algorithm cannot process unequal-length environmental state sequences (sequences sent into the same neural network each time and with inconsistent lengths). It is necessary to first convert the unequal-length environmental state sequences into equal-length state sequences. This algorithm uses a BiGRU bidirectional recurrent gate unit to process unequal-length environmental state sequences, as Figure 5 Shown. Among them Represents the state information (speed information (v x , v y )) and position information (p x , p y )) of the i-th obstacle (other unmanned boats, excluding itself) detected within the detection range of the pursuer, and o self Represents the state information of the current unmanned boat itself (excluding other unmanned boats), h ∈ R mx1Represents the environmental state information within the detection range of the \(i\)-th pursuer extracted by the BiGRU (i.e., the overall characteristics of obstacles within a certain range of the current unmanned boat).
[0163] Then the environmental state information of the current pursuer can be expressed as:
[0164]
[0165] Step Eight: Design a controller using reinforcement learning. The reinforcement learning algorithm framework adopts PPO (Proximal Policy Optimization algorithm, a reinforcement learning algorithm with excellent performance and good stability. Through a period of training, the neural network adopted by the algorithm can have the following functions: according to the input state information, output the best action that can make the controlled object reach the expected goal under this state). According to the loss function in Equation (14), define the following action-value function:
[0166]
[0167] The local action reward function in the time interval \((t, t + h]\) can be defined as:
[0168]
[0169] where \(h\) is the time interval for each sampling. Let When \(h\rightarrow0\), Equation (24) can approximately hold:
[0170]
[0171] Therefore, the immediate reward function can be defined as shown in Equation (24), and the discounted return \(G\) t As shown in Equation (25), where \(\gamma\) is the discount factor:
[0172] \(G\) t \(= R\) t +\(\gamma R\) t+1 +\(\gamma\) 2 \(R\) t+2 +\(\gamma\) 3 \(R\) t+3 +… (77)
[0173] Denote the state of the current unmanned boat as the action as \(u\), and the policy as \(\pi\). The action-value function Q-value and the state-value function V-value are:
[0174] \(Q\) π \((s\) t ,u t ) = E π \((G\) t |S t = s t ,Ut = u t ) (78)
[0175] V π (s t ) = E π (G t |S t = s t ) (79)
[0176] As Figure 6 shown, given the system state equation (environment), the immediate reward function, and the state s t and the action u t information, according to the PPO algorithm in reinforcement learning, the value function and the policy function can be fitted, and the optimal control u can be given through the policy network t .
[0177] Step Nine: As Figure 7 shown, based on the optimal control law of the subsystem generated in Step Eight introduced in Step Five and minimizing the performance index to obtain the optimal auxiliary control law and repeating this sequential game process to complete the pursuit and capture of the escapee.
[0178] In summary, the above embodiments have described in detail different configurations of the multi-unmanned boat pursuit and escape game control method. Of course, the present invention includes but is not limited to the configurations listed in the above embodiments. Any content obtained by transformation based on the configurations provided in the above embodiments belongs to the scope protected by the present invention. Those skilled in the art can draw inferences from one instance to another based on the content of the above embodiments.
[0179] Each embodiment in this specification is described in a progressive manner. The key points of each embodiment are the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0180] The above description is only a description of the preferred embodiments of the present invention, and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the art according to the above disclosure belong to the scope protected by the claims.
Claims
1. A multi-unmanned boat pursuit-evasion game control method, characterized in that, it includes: When the unmanned boat is in pursuit-evasion, sequential decision-making is introduced into the control algorithm of the pursuer for self-play; The observer calculates the optimal response of the observer according to the optimal control given by the controller to approximate the external interference and system uncertainty in the pursuer group; and The controller receives the optimal response of the observer and recalculates the optimal control of the pursuer according to the optimal response, where the recalculation is performed one or more times to form sequential decision-making; Wherein the method further includes: Cooperating the game between the observer and reinforcement learning, so that the observer acts as a follower to handle uncertainties, and the uncertainties include external interference and modeling errors; Using reinforcement learning as the controller to form a leader; and Through the sequential game between the observer and the controller, the reinforcement learning algorithm can cope with external interference and modeling errors and reach Nash equilibrium to achieve game cooperative encirclement.
2. The multi-unmanned boat pursuit-evasion game control method according to claim 1, characterized in that, it further includes: Setting a reward function according to obstacle avoidance, tracking, surrounding, and control quantity consumption; Using the reciprocal speed obstacle method to set the obstacle avoidance reward and simultaneously handle the obstacle avoidance problems of other static and dynamic obstacles; Using the potential energy function to set the tracking reward and the surrounding reward, and the potential energy function stops the growth of potential energy when the obstacle avoidance requirement is met within the threshold distance.
3. The multi-unmanned boat pursuit-evasion game control method according to claim 1, characterized in that, it further includes: In the multi-unmanned boat pursuit-evasion game, each pursuer is set as a subsystem; Using the observer of the subsystem as a follower to handle uncertainties; Designing the controller of the subsystem based on reinforcement learning to form a leader; Enhancing the interaction between the leader designed based on reinforcement learning and the environment through the observer to improve the control performance, and vice versa using the improvement of the control performance to improve the observation performance of the observer; According to this process, establishing a sequential game graph between the leader and the follower; The escape unmanned boat strategy adopts an inherent model and strategy.
4. The multi-unmanned boat pursuit-evasion game control method according to claim 1, characterized in that, it further includes: Step 1: According to the conventional unmanned boat swaying, yawing, and rolling motion equations: (1) wherein and are respectively expressed as the swinging speed, yaw speed, yaw angle, roll speed, roll angle, rudder angle of deflection and unknown uncertainty following the unmanned boat, is expressed as the damping ratio and natural frequency, is expressed as the time constant, is expressed as the unmanned boat system gain; Step 2: According to the swing, yaw, and roll motion equations of the following unmanned boat, define the system state of the dynamic equation of the following unmanned boat , the angle sensor outputs , unknown uncertainties caused by waves, wind disturbances or other factors are respectively , and the following dynamic equation of the unmanned boat is expressed as follows: (2) (3) Abbreviating the dynamic equations of each unmanned boat subsystem as: (4) Step 3: Design the unmanned boat subsystem The observer is specifically as follows: (5) where L is the observer parameter matrix; System error : (6)。 5. The multi-unmanned boat pursuit-evasion game control method according to claim 4, characterized in that, it further includes: Step 4: To form a complete sequential game process, introduce an auxiliary control law , such that the observer and the controller form a non-cooperative game, and design the following observer to improve the observer design in Step 3: (7) Introducing the following performance index function to optimize the observer performance: (8) Among them A representative subsystem On the sub-observer The influence of, and Is a symmetric positive definite matrix, used to adjust the weight ratio between each constraint in the performance index function.
6. The multi-unmanned boat pursuit-evasion game control method according to claim 5, characterized in that, it further includes: Step 5: In the sequential game decision-making of the pursuer, it is necessary to first consider the optimal response of the pursuer's sub-observer. Assume that the control law of the subsystem is initially initialized as an admissible control at the beginning of the game, and the following sub-observer Hamiltonian function is introduced: (9) wherein is the partial derivative of the performance index pair ; Optimal performance metrics The value satisfies the Hamilton-Jacobi equation: (10) A necessary condition for solving equation (10) is (11) Ideally At this time Using adaptive dynamic programming to solve, the optimal auxiliary control law is obtained as follows: (12) Substituting Equation (12) into the Hamiltonian-Jacobi (HJ) equation shown in Equation (10): (13)。 7. The multi-unmanned boat pursuit-evasion game control method according to claim 6, characterized in that, it further includes: Step 6: On the basis of the fifth step, introduce the following pursuer subsystem Reinforcement learning-based control leader optimization objective function: (14) Design control law such that the following equation holds: (15) Among them One item reduces the consumption of the control quantity One item introduces an auxiliary control law to form a complete sequential non - cooperative game is the reward function, which consists of three parts: respectively, that plays a role in local obstacle avoidance , the tracking potential energy that plays a role in tracking the evader , and is a symmetric positive definite matrix is a positive hyperparameter. Considering that the pursuer will carry out the encirclement after catching up with the evader, should be slightly larger than ; (16) (17) (18) (19) In formula (17), represents the current speed, and different rewards are given by judging whether it belongs to the reciprocal speed obstacle method area. is a constant value for adjusting the performance of the strategy, represents the difference between the current speed of the unmanned boat and the desired speed, is the expected shortest time to collide with an obstacle at the current speed; For equations (18) and (19), are respectively the actual distance and the desired distance between the current pursuer and the evader, is an adjustable hyperparameter; if , it is regarded that a collision occurs between the agents, is a small positive constant; so the situation of will not occur, and similarly the situation of will not occur; is an indicator function, as shown in equation (20), where represents the action distance of the circumnavigation potential energy, and , m > 1 is a constant; it means that: with the current pursuer as the center, as the radius of the region, the current pursuer only has the potential energy action of circumnavigation around with other pursuers within this region; (20)。 8. The multi-unmanned boat pursuit-evasion game control method according to claim 7, characterized in that, it further includes: Convert the environment state sequence of unequal length into an equal-length state sequence, and use the BiGRU (Bidirectional Gated Recurrent Unit) to process the environment state sequence of unequal length, where represents the state information of the i-th obstacle detected within the detection range of the pursuer, and represents the state information of the current unmanned boat itself, represents the environment state information of the i-th pursuer detection range extracted by the BiGRU; Then the environmental state information of the current pursuer can be expressed as: (21)。 9. The multi-unmanned boat pursuit-evasion game control method according to claim 8, characterized in that, it further includes: Step 8: Design a controller using reinforcement learning. The reinforcement learning algorithm framework adopts the proximal policy optimization algorithm. According to the loss function in Equation (14), define the following action-value function: (22) In the time interval the local action reward function can be defined as: (23) is the time interval for each sampling. Let , when , approximately, equation (24) holds: (24) Therefore, define the immediate reward function as shown in Equation (24), and the discounted return is shown in Equation (25), where is the discount factor: (25) Record the current state of the unmanned boat as , the action is , and the strategy is ; The action value function Q-value and the state value function V-value are as follows: (26) (27) Given the system state equation, the immediate reward function, and the state and action information, according to the approximate policy optimization algorithm in reinforcement learning, fit the value function and the policy function, and give the optimal control through the policy network ; Step Nine: Based on the optimal control law of the subsystem generated in Step Eight , introduce it in Step Five and minimize the performance index to obtain the optimal auxiliary control law , and repeat this sequential game process to complete the pursuit and encirclement of the evader.
Citation Information
Patent Citations
Anti-interference method for Vehicular ad hoc networks based on power control and electronic equipment
CN108012248A
Multi-unmanned aerial vehicle intelligent collaborative decision-making method for hunting task
CN113467508A