Training method and device for roundup strategy network based on multi-agent reinforcement learning

Through the method based on multi-agent reinforcement learning, the roundup strategy network of the underwater bionic robot team was trained, which solved the problem of roundup and escape in the underwater environment, and achieved efficient, stable and adaptive cooperative roundup effect.

CN119476404BActive Publication Date: 2025-05-06INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411545735.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-05-06
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of underwater bionic robot rounding up and escape, especially in complex underwater environments, where high nonlinear impacts and environmental interference make designing effective strategies difficult.

Method used

A roundup strategy network based on multi-agent reinforcement learning is adopted to train the policy network of the roundup team in a liquid environment through training methods, and the evaluation network is used to perform global information processing and reward score calculation to realize centralized training and decentralized execution.

Benefits of technology

It improves the stability and scalability of the algorithm in complex environments, realizes cooperative roundup of bionic robot fish in complex environments, can meet the task requirements in various complex scenarios, and ensures the security and adaptability of the strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119476404B_ABST
    Figure CN119476404B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a training method and device for a capture strategy network based on multi-agent reinforcement learning, which is applied to a liquid environment containing at least two capturers and escapees. The capture strategy network corresponding to the at least two capturers can be used to obtain the action information of the at least two capturers through the state information of the at least two capturers at the current moment and the relative state information of the at least two capturers to other capturers and escapees; and the scoring information for the action information of the at least two capturers can be obtained by inputting the global information into the evaluation network, so as to train the capture strategy network corresponding to the at least two capturers. The paradigm of centralized training and decentralized execution is adopted to ensure the stability and scalability of the algorithm in complex environments, and further realize the efficient cooperative capture of bionic robot fish in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of swarm intelligent control, and specifically to a training method and device for a roundup strategy network based on multi-agent reinforcement learning, and a method and device, electronic equipment, storage medium and computer program product for rounding up using a roundup strategy network based on multi-agent reinforcement learning. Background Art

[0002] The capture-escape problem is a typical problem in the field of cooperation and adaptive decision-making of multi-robot systems, and has attracted researchers from multiple fields such as control theory, robotics, and game theory. In this task, the capture team must complete the capture task through close internal coordination and cooperation, and it also involves competition and confrontation between the two teams. Although there are some theoretical studies that can achieve the capture of escapees in complex situations with incomplete information and multiple task objectives, they are difficult to directly apply to actual robot scenarios due to their reliance on ideal assumptions, high modeling complexity, and high computational difficulty.

[0003] In recent years, reinforcement learning has gradually become a popular method for solving the capture-escape problem in practical applications because it can adapt to dynamic environments and does not require precise modeling. However, current research mainly focuses on unmanned ground vehicles (UGVs), unmanned aerial vehicles (UAVs) and unmanned surface vessels (USVs). In contrast, there are relatively few studies on capture-escape for underwater bionic robots. This is mainly because underwater bionic robots are underactuated systems with complex dynamic models, and the fluctuations of the underwater environment have a great impact on the movement of the robot. Therefore, this highly nonlinear effect and environmental interference pose a huge challenge to the design of effective capture-escape strategies for underwater bionic robots. Summary of the invention

[0004] The exemplary embodiments of the present disclosure provide a method and device for training a roundup strategy network based on multi-agent reinforcement learning, and a method and device, electronic device, storage medium, and computer program product for rounding up a roundup strategy network based on multi-agent reinforcement learning, which can at least solve the above-mentioned technical problems and other technical problems not mentioned above.

[0005] According to one aspect of the present disclosure, a training method for a capture strategy network based on multi-agent reinforcement learning is provided, which is applied to a liquid environment containing at least two capturers and an escapee. The training method for the capture strategy network based on multi-agent reinforcement learning includes: in a training round including multiple capture-escape action moments, performing the following operations for each moment: respectively obtaining the state information of the at least two capturers at the current moment and the relative state information of the at least two capturers to other capturers and the escapee; respectively inputting the state information of the at least two capturers and the relative state information of the at least two capturers to other capturers and the escapee into the capture strategy network corresponding to the at least two capturers, and obtaining the action information of the at least two capturers, wherein the action information of the at least two capturers refers to the current state information of the at least two capturers; The invention relates to a method for controlling the capture of at least two capturers by inputting the state information and action information of each of the at least two capturers into a state transition model to obtain the state information of each of the at least two capturers at the next moment and the corresponding reward score; and the global information is input into an evaluation network to obtain the scoring information for the action information of each of the at least two capturers, wherein the global information includes: the state information and action information of each of the at least two capturers at the current moment, and the state information and action information of the escaper at the current moment, wherein the action information of the escaper refers to the decision information of the escape action of the escaper at the current moment; in the training round, based on the reward score and the scoring information obtained by each of the at least two capturers in the training round, the capture strategy networks corresponding to the at least two capturers are trained respectively.

[0006] Optionally, the liquid environment also contains obstacles, wherein the respectively obtaining the status information of the at least two capturers at the current moment and the relative status information of the at least two capturers to other capturers and the escaper, further includes: respectively obtaining the relative status information of the at least two capturers to the obstacles at the current moment; wherein the respectively inputting the status information of the at least two capturers and the relative status information of the at least two capturers to other capturers and the escaper into the respectively corresponding capture strategy networks to obtain the respective action information of the at least two capturers, includes: respectively inputting the status information of the at least two capturers and the relative status information of the at least two capturers to other capturers, the escaper and the obstacle into the respectively corresponding capture strategy networks to obtain the respective action information of the at least two capturers; wherein the global information also includes: status information of the obstacle.

[0007] Optionally, the state information and action information of each of the at least two capturers are input into a state transition model to obtain the state information of each of the at least two capturers at the next moment, including: for each of the at least two capturers, the following operations are performed: based on the distance between the capturer and the escapee, determining the attraction force corresponding to the capturer; based on the distance between the capturer and the obstacle, determining the repulsion force corresponding to the capturer; based on the state information of the capturer at the current moment, determining an additional force corresponding to the capturer; based on the attraction force, the repulsion force, and the additional force, adjusting the action information of the capturer; inputting the state information of the capturer at the current moment and the adjusted action information of the capturer into the state transition model to obtain the state information of the capturer at the next moment.

[0008] Optionally, the training method includes multiple training stages, each training stage includes multiple training rounds, and the training of the capture strategy networks corresponding to the at least two capturers respectively includes: in a first training stage where no obstacles are set in the liquid environment, the capture strategy networks are trained based on a first preset scenario, wherein the first preset scenario is: the mobility of the capturer is less than that of the escapee, and the mobility difference between the capturer and the escapee is a first preset value, and the perception range of the capturer is a second preset value, and the perception range of the capturer is used to determine that the capture mission is successful when the escapee is within the perception range of the capturer; in a second training stage where obstacles are set in the liquid environment, the capture strategy networks are trained based on a second preset scenario, wherein the second preset scenario is obtained by gradually increasing the first preset value in the first preset scenario and gradually decreasing the second preset value.

[0009] Optionally, the bonus score is calculated based on a first bonus score, a second bonus score and a penalty score, the first bonus score being determined by the result of the capture mission, the second bonus score being determined by the distance between the capturer and the escapee, and the penalty score being determined by the distance between the at least two capturers and the obstacle.

[0010] Optionally, the motion information of the capturer includes heading change information; wherein, adjusting the motion information of the capturer based on the attraction, the repulsion, and the additional force comprises: adjusting the heading change information of the capturer based on the attraction, the repulsion, and the additional force, wherein the value of the additional force is determined by the heading change information.

[0011] Optionally, the action information of the capturer also includes a moving distance, and the minimum value of the moving distance is greater than the distance that the capturer swims in one swing cycle, and the distance that the capturer swims in one swing cycle is determined by the minimum turning radius of the capturer.

[0012] According to another aspect of the present disclosure, a method for encircling and capturing based on a multi-agent reinforcement learning capture strategy network is also provided, which is applied to a liquid environment containing at least two capturers and an escapee, and the method for encircling and capturing based on a multi-agent reinforcement learning capture strategy network comprises: respectively obtaining current state information of each of the at least two capturers at the current moment and relative state information of each of the at least two capturers to other capturers and the escapee; respectively inputting the state information of each of the at least two capturers and the relative state information of each of the at least two capturers to other capturers and the escapee into the capture strategy network corresponding to each of the at least two capturers to obtain action information of each of the at least two capturers; inputting the state information and action information of each of the at least two capturers into a state transition model to obtain the state information of each of the at least two capturers at the next moment; and executing the capture task of the at least two capturers on the escapee based on the state information of each of the at least two capturers at the next moment; wherein the capture strategy network is trained by any of the above-mentioned training methods for the capture strategy network based on multi-agent reinforcement learning.

[0013] According to another aspect of the present disclosure, there is also provided a training device for a capture strategy network based on multi-agent reinforcement learning, which is applied to a liquid environment containing at least two capturers and an escapee, and the training device for the capture strategy network based on multi-agent reinforcement learning comprises: an information acquisition unit; an action decision unit; a state transfer unit; an action evaluation unit; and a network training unit; wherein, in a training round including a plurality of capture-escape action moments, for each moment: the information acquisition unit respectively acquires the state information of the at least two capturers at the current moment and the relative state information of the at least two capturers to other capturers and the escapee; the action decision unit respectively inputs the state information of the at least two capturers and the relative state information of the at least two capturers to other capturers and the escapee into the capture strategy network corresponding to the at least two capturers, and obtains the action information of the at least two capturers, and the at least two capturers are respectively transferred to the capture strategy network corresponding to the at least two capturers. The action information of each capturer refers to the decision information of the capture action of each of the at least two capturers at the current moment; the state transfer unit inputs the state information and action information of each of the at least two capturers into the state transfer model to obtain the state information of each of the at least two capturers at the next moment and the corresponding reward score; and the action evaluation unit inputs the global information into the evaluation network to obtain the scoring information for the action information of each of the at least two capturers, wherein the global information includes: the state information and action information of each of the at least two capturers at the current moment, and the state information and action information of the escaper at the current moment, wherein the action information of the escaper refers to the decision information of the escape action of the escaper at the current moment; in the training round, the network training unit trains the capture strategy networks corresponding to each of the at least two capturers based on the reward scores and the scoring information obtained by each of the at least two capturers in the training round.

[0014] According to another aspect of the embodiment of the present disclosure, there is also provided a device for encircling based on a multi-agent reinforcement learning encirclement strategy network, which is applied to a liquid environment containing at least two encirclers and an escapee, and the device for encircling based on a multi-agent reinforcement learning encirclement strategy network comprises: an information acquisition unit, configured to: respectively acquire current state information of the at least two encirclers at the current moment and relative state information of the at least two encirclers with other encirclers and the escapee; an action prediction unit, configured to: respectively acquire the state information of the at least two encirclers and the relative state information of the at least two encirclers with other encirclers and the escapee; The information is respectively input into the capture strategy network corresponding to each of the at least two capturers to obtain the action information of each of the at least two capturers; the state transfer unit is configured to: input the state information and action information of each of the at least two capturers into the state transfer model to obtain the state information of each of the at least two capturers at the next moment; the capture execution unit is configured to: execute the capture task of the at least two capturers on the escapee based on the state information of each of the at least two capturers at the next moment; wherein the capture strategy network is trained by the training method of the capture strategy network based on multi-agent reinforcement learning described in any one of the above.

[0015] According to another aspect of an embodiment of the present disclosure, there is also provided an electronic device, comprising: at least one processor; and at least one memory storing computer executable instructions, wherein when the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to execute any of the training methods for the multi-agent reinforcement learning-based roundup strategy network and the method for roundup using the multi-agent reinforcement learning-based roundup strategy network as described above.

[0016] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium storing instructions is also provided. When the instructions are executed by at least one processor, the at least one processor is prompted to execute any of the training methods for the multi-agent reinforcement learning-based roundup strategy network and the method for roundup using the multi-agent reinforcement learning-based roundup strategy network as described above.

[0017] According to another aspect of an embodiment of the present disclosure, a computer program product is also provided, including a computer program / instruction, which, when executed by a processor, implements the training method of the roundup strategy network based on multi-agent reinforcement learning as described in any one of the above items and the method of rounding up the roundup strategy network based on multi-agent reinforcement learning as described above.

[0018] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects:

[0019] According to the training method and device of the multi-agent reinforcement learning-based roundup strategy network and the method and device, electronic device, storage medium and computer program product for roundup based on the multi-agent reinforcement learning disclosed in the present invention, centralized training is achieved through the evaluation network's processing of global information, and decentralized execution is achieved through the roundup strategy networks of the roundups. The paradigm of centralized training and decentralized execution ensures the stability and scalability of the algorithm in complex environments, and can further realize the cooperative roundup of bionic robot fish in complex environments. The present invention is not only efficient, but also can meet the task requirements in various complex scenarios.

[0020] In addition, an attractive-enhanced decision-making mechanism is adopted to predict the state of the captor at the next moment, which can ensure the security and adaptability of the strategy while ensuring efficient capture.

[0021] In addition, the curriculum learning method improves the practicality of training and effectively shortens the training cycle by gradually increasing the difficulty of training. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0023] Figure 1 A flowchart showing a training method for a hunting strategy network based on multi-agent reinforcement learning in an exemplary embodiment of the present disclosure is shown;

[0024] Figure 2 A schematic diagram showing an adaptive round-up strategy based on an attraction mechanism in an exemplary embodiment of the present disclosure;

[0025] Figure 3 A schematic diagram showing a capture task for a cluster of underwater bionic robotic fish in an exemplary embodiment of the present disclosure is shown;

[0026] Figure 4 A schematic diagram showing a multi-agent reinforcement learning algorithm framework for underwater scenarios in an exemplary embodiment of the present disclosure;

[0027] Figure 5 A schematic diagram showing the effect comparison between the phased training and the continuous training method using course learning in an exemplary embodiment of the present disclosure;

[0028] Figure 6 , Figure 7 , Figure 8A screenshot showing a video of an underwater multi-agent capture-escape experiment consisting of two capturers and one escaper in an exemplary embodiment of the present disclosure;

[0029] Fig. 9 A flowchart showing a method for performing roundups based on a roundup strategy network using multi-agent reinforcement learning in an exemplary embodiment of the present disclosure;

[0030] Fig.10 A block diagram showing a training device for a hunting strategy network based on multi-agent reinforcement learning in an exemplary embodiment of the present disclosure is shown;

[0031] Fig.11 A block diagram showing a device for performing roundups based on a roundup strategy network of multi-agent reinforcement learning in an exemplary embodiment of the present disclosure;

[0032] Fig.12 A block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0033] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation methods described in the following examples do not represent all implementation methods consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the attached claims.

[0035] It should be noted that the phrase "at least one of the items" in the present disclosure includes three types of parallel situations: "any one of the items", "a combination of any number of the items", and "all of the items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "executing at least one of step 1 and step 2" which means the following three parallel situations: (1) executing step 1; (2) executing step 2; (3) executing step 1 and step 2.

[0036] The multi-robot adaptive capture task involves a team of capturers collaborating to capture one or more escapees in a specific environment such as on the ground, in the air, or underwater. This requires not only effective coordination and cooperation within the capture team, but also adaptive adjustment strategies based on environmental information such as terrain. When formulating a suitable collaborative capture strategy, it is necessary to solve problems such as real-time path planning, coordination and control in the multi-robot system, which makes this problem highly challenging.

[0037] In recent years, bio-inspired underwater vehicles have attracted extensive attention from academia and industry due to their high maneuverability and low disturbance to the underwater environment. With the development of robotics, more and more researchers have begun to focus on multi-robot systems, which are highly regarded for their ability to cooperate efficiently and perform complex tasks. In underwater environments, studying the behavior of marine fish groups provides valuable insights into the development of multi-biomimetic robot swarm systems. These multi-robot systems are able to effectively complete complex tasks through cooperation between groups, which is crucial in fields such as marine exploration, security, and biological discovery.

[0038] In order to promote the development and improvement of underwater multi-robot systems and their corresponding technologies, a large number of studies have been conducted on underwater bionic robot swarms, such as formation, encirclement, formation control, and underwater football. Although these tasks demonstrate the collaborative capabilities of underwater bionic robot swarms, they have certain limitations. In addition, team competition in multi-robot systems is also a common phenomenon. In confrontational scenarios, these systems not only face complex interactions and confrontation strategies, but also need to deal with the high dynamics and unpredictability of the environment. This requires better coordination and cooperation capabilities and stronger adaptive decision-making capabilities among team members, making the research on cooperation and adaptive decision-making of multi-underwater robot systems a challenging field.

[0039] In order to solve the above problems, the present disclosure provides a training method and device for a hunting strategy network based on multi-agent reinforcement learning, and a method and device, electronic device, storage medium and computer program product for hunting based on a hunting strategy network based on multi-agent reinforcement learning. Centralized training is achieved through the processing of global information by the evaluation network, and decentralized execution is achieved through the hunting strategy networks of the hunters. The paradigm of centralized training and decentralized execution ensures the stability and scalability of the algorithm in complex environments, and can further realize the cooperative hunting of bionic robot fish in complex environments. The present disclosure is not only efficient, but also can meet the task requirements in various complex scenarios.

[0040] Next, we will refer to Figures 1 to 12The present invention specifically describes the training method and device of the roundup strategy network based on multi-agent reinforcement learning, and the method and device, electronic device, storage medium and computer program product for rounding up based on the roundup strategy network based on multi-agent reinforcement learning.

[0041] The training method of the trapping strategy network based on multi-agent reinforcement learning disclosed in the present invention can be applied to a liquid environment containing at least two trappers and escapers. Specifically, the liquid environment can be an environment below the water surface, and the intelligent trappers and escapers can both be bionic robotic fish. The goal of the trappers is to increase the teamwork performance so that they can use their numerical advantage to capture the escapers.

[0042] It is understandable that a training round can be determined by a preset step length. For example, a training round can be considered completed after the bionic robotic fish moves a preset step length. During this process, if the capture task is successful, the training round can be completed in advance.

[0043] In each training round, if the distance between any of the capturers in the capture team and the escapee is less than the preset value, it can be considered that a capture is successfully completed. Conversely, if the escapee successfully survives a round, it can be considered that the capture mission of that round has failed, and the capture mission of that round is also ended.

[0044] Figure 1 A flowchart showing a method for training a hunting strategy network based on multi-agent reinforcement learning in an exemplary embodiment of the present disclosure is shown.

[0045] Reference Figure 1 In step 101, in a training round including multiple capture-escape action moments, the following operations are performed for each moment, namely, steps 101-1 to 101-4:

[0046] In step 101 - 1 , the status information of at least two capturers at the current moment and the relative status information of at least two capturers to other capturers and the escaper are respectively obtained.

[0047] According to an exemplary embodiment of the present disclosure, state information may represent the environmental information perceived by the agent and its dynamic changes, and its quality directly determines the final performance of the strategy.

[0048] According to an exemplary embodiment of the present disclosure, the state information of the captor may include, but is not limited to, the position, yaw angle, etc. of the captor, and the relative state information of the captor and other captors and escapees may include the relative position, relative yaw angle, etc. between the intelligent bodies.

[0049] In step 101-2, the state information of at least two capturers and the relative state information of at least two capturers to other capturers and the escapee are respectively input into the capture strategy network corresponding to the at least two capturers to obtain the action information of at least two capturers. The action information of at least two capturers refers to the decision information of the capture action of at least two capturers at the current moment.

[0050] According to an exemplary embodiment of the present disclosure, the decentralized capture strategy network is responsible for making decisions and generating actions based on the local observations of the individual intelligent agents acting as the capturers (i.e., the state information of the capturers and the relative state information of the capturers, other capturers, and escapers), so as to capture the individual intelligent agents acting as the escapers. The intelligent agents acting as the escapers can have their strategy networks for escape trained separately to generate decision information for the escape action based on their current state information. In the execution phase, the intelligent agents i∈N acting as the capturers can generate decision information for the escape action based on the local observations s i Make decision action μ i , this decision action can be considered as the decision information of the round-up action.

[0051] According to an exemplary embodiment of the present disclosure, the liquid environment may also contain obstacles. Therefore, the local observation information of the intelligent agent individual as the hunter may also include the relative state information (eg, relative position information, etc.) of each hunter and the obstacle.

[0052] According to an exemplary embodiment of the present disclosure, the state information of at least two capturers and the relative state information of at least two capturers to other capturers, escapees and obstacles can be respectively input into the corresponding capture strategy networks to obtain the action information of at least two capturers.

[0053] According to an exemplary embodiment of the present disclosure, the relative state information of each captor in relation to other captors, escapers, and obstacles can be combined to design a multi-agent reinforcement learning state space for an underwater capture-escape scenario. This state space can be viewed as a collection of the respective and relative state information of all agents and obstacles in the environment.

[0054] According to an exemplary embodiment of the present disclosure, assuming that there is a capture and escape problem involving n intelligent agents, for the state space o of the capturer (pursuer) i i , which can be expressed as:

[0055] o i =[p i ,ψ i ,P,Ψ,P obs ]

[0056] Among them, p irepresents the position of agent i in the world coordinate system, ψ i represents the yaw angle of agent i, P = [s i,1 ,s i,2 ,…s i,j ,s i,n-1 ](i≠j) represents the relative position between agent i and other agents, Ψ=[ψ 1 , ψ 2 ,…,ψ j , ψ n-1 ](i≠j) represents the yaw angle with other agents, P obs A collection representing the locations of obstacles.

[0057] In step 101 - 3 , the state information and action information of at least two capturers are input into the state transition model to obtain the state information of at least two capturers at the next moment and the corresponding reward score.

[0058] According to an exemplary embodiment of the present disclosure, the reward mechanism of reinforcement learning can enable the intelligent agent to quickly learn an efficient capture strategy from a complex environment. The capturer intelligent agent in the capturer team can transition from the current state S to the subsequent state S′ according to the basic kinematic model, that is, the state transition model, and obtain the relevant reward r i In order to balance the cooperation within the team and the competition between teams, the rewards of the agents can be shared within the team. The learning goal of each agent is to maximize its cumulative reward in a training round.

[0059] Next, an adaptive round-up strategy based on the attraction mechanism is introduced.

[0060] Artificial potential field is a method of robot path planning. Its principle is to regard the target as the "gravitational source" that attracts the robot, and the obstacle as the "repulsive source" that generates repulsive force, so as to guide the robot to move from the starting position to the target position while avoiding obstacles. Therefore, the escaper can be regarded as the target of the captor. According to the motion characteristics of the bionic robot fish, by analyzing the target attraction, obstacle repulsion, and additional force in the environment, an adaptive capture strategy based on the attraction mechanism is designed.

[0061] According to an exemplary embodiment of the present disclosure, for each of the at least two captors, the following operations may be performed: based on the distance between the captor and the escapee, the attraction force corresponding to the captor is determined; based on the distance between the captor and the obstacle, the repulsion force corresponding to the captor is determined; based on the state information of the captor at the current moment, the additional force corresponding to the captor is determined; based on the attraction force, the repulsion force, and the additional force, the motion information of the captor is adjusted; the state information of the captor at the current moment and the adjusted motion information of the captor are input into the state transition model to obtain the state information of the captor at the next moment.

[0062] According to an exemplary embodiment of the present disclosure, the distance between the captor and the escapee can be abstracted as attraction, which can be expressed as follows:

[0063] F att (p) = -k att ·(pp goal )

[0064] Where p is the position of the pursuer, p goal is the escapee position, k att is the gravitational coefficient.

[0065] The repulsive force of the obstacle on the robot fish is related to the distance between the robot fish and the obstacle, and when the robot fish moves away from the obstacle beyond a certain distance, the repulsive force can be reduced to zero, which can be expressed as follows:

[0066]

[0067] Among them, d i =||pp obs || is the distance from the robot fish i to the obstacle obs, k rep is the repulsion coefficient, d 0 is the distance threshold for the repulsive force.

[0068] In order to enable the agent to adaptively cope with complex environments, the magnitude of the gravitational coefficient and the repulsive coefficient can be dynamically adjusted according to the current state of the agent:

[0069]

[0070] Among them, d goal =||p goal -p|| is the Euclidean distance from the robot's current position to the target position, d obs =||p obs -p|| is the Euclidean distance from the robot's current position to the obstacle, and α and β are the sensitivity coefficients for distance perception.

[0071] According to an exemplary embodiment of the present disclosure, the action information of the hunter may include heading change information.

[0072] According to an exemplary embodiment of the present disclosure, the heading change information of the captor may be adjusted based on the attractive force, the repulsive force, and the additional force, wherein the value of the additional force may be determined by the heading change information.

[0073] According to an exemplary embodiment of the present disclosure, in order to increase the adaptive adjustment capability of the intelligent agent to complex environmental changes, additional force may be added so that the intelligent agent can effectively explore a broad solution space, which may be expressed as follows:

[0074] F add =T·μ

[0075] Among them, T is the scaling factor, and the value of the exploration factor μ can be determined by the reinforcement learning strategy, that is, it can be determined by the roundup strategy network according to the current state information. Specifically, it can be determined by the output of the roundup strategy network and the heading change information in the action information.

[0076] The total force on the agent can be composed of attraction, repulsion, and an additional force based on the current state, which can be expressed as follows:

[0077] F total (p) = F att (p)+F rep (p)+F add

[0078] According to an exemplary embodiment of the present disclosure, the multi-agent reinforcement learning action space, i.e., action information, of the underwater capture-escape scenario can be designed by combining the relative state information of each capturer with other capturers, escapers, and obstacles.

[0079] In practical applications, the action space of the robot fish is mainly represented by moving a certain distance along a specific heading. According to an exemplary embodiment of the present disclosure, the action information of the hunter may also include the moving distance. Therefore, the action space may be defined by a parameter pair (μ, l), and the strategy may be expressed as π θ (μ, l|S). In it, μ represents the change of the heading of the agent, and l represents the movement distance of the agent.

[0080] It is worth noting that complex and variable speeds can lead to large differences in the control cycle of the intelligent agent. Therefore, in order to ensure the transferability of the strategy, in the exemplary embodiment of the present disclosure, the range of the heading angle can be set to [-70°, +70°], and l can be set to a fixed value. The value of l is related to the motion performance of the robot fish, and its size depends on the steering ability of the robot fish. The law of cosines shows that the relationship between the step length l, the heading angle ψ and the turning radius r is as follows:

[0081]

[0082] Obviously, the value of l is proportional to the change in heading angle. Therefore, when the robot fish swims with the minimum turning radius, the minimum value of l can be calculated to be 0.12m. In order to reflect the difference in mobility between intelligent agents, in the exemplary embodiments of the present disclosure, the minimum l of the capturer and escaper can be set to 0.18m and 0.2m respectively, both of which are greater than the distance they swim in one swing cycle.

[0083] According to an exemplary embodiment of the present disclosure, the minimum value of the moving distance is greater than the distance that the hunter swims in one swing cycle, and the distance that the hunter swims in one swing cycle is determined by the minimum turning radius of the hunter.

[0084] Next, we introduce an example of an intelligent robot fish acting as a hunter in an adaptive hunting strategy based on the attraction mechanism, in which the attraction mechanism adjusts its hunting movement strategy based on local observation information.

[0085] Figure 2 A schematic diagram showing an adaptive capture strategy based on an attraction mechanism in an exemplary embodiment of the present disclosure.

[0086] Reference Figure 2 , for the roundup P t , when the obstacle and the escaper are currently in the position shown in the figure, assuming that their original yaw angle is ψ t , under the action of the attraction mechanism, it can generate attraction F att , repulsive force F rep , and the additional force F add , the resultant force will be F total , this combined force will make the capturer P t The yaw angle is changed from the original yaw angle ψ t Modified to △ψ t On this basis, when the moving distance is l, it will move to the target point in the diagram at the next moment.

[0087] Next, we introduce the process of processing global information through a centralized evaluation network.

[0088] The centralized evaluation network can evaluate the value of the individual capture actions of the capturers through global information, namely global state information and action information.

[0089] In addition, the evaluation network is also continuously updated during the training process, and it can be updated according to its own gradient.

[0090] Return to reference Figure 1 In step 101-4, the global information is input into the evaluation network to obtain the scoring information for the action information of at least two capturers, wherein the global information includes: the state information and action information of at least two capturers at the current moment, and the state information and action information of the escaper at the current moment, wherein the action information of the escaper refers to the decision information of the escape action of the escaper at the current moment.

[0091] According to an exemplary embodiment of the present disclosure, the global information at the current moment can be processed through the evaluation network to score the capture strategy generated by each strategy network based on the current local observation information of the corresponding capturer agent (the state information of the agent and the relative state information between the agent and other agents). The parameters of the capture strategy network can be adjusted based on the scoring information; and the evaluation network is updated according to its own gradient.

[0092] According to an exemplary embodiment of the present disclosure, the global information may further include: state information of obstacles.

[0093] According to an exemplary embodiment of the present disclosure, the state information of the obstacle may include but is not limited to the position information of the obstacle, and the position information of the obstacle may be used to obtain the relative distance between the agent and the obstacle.

[0094] Next, the process of executing the encirclement and capture mission scenario for an underwater multi-bionic robotic fish cluster based on the formation configuration of the multi-bionic robotic fish system is introduced.

[0095] Figure 3 A schematic diagram of a capture mission for a cluster of underwater bionic robotic fish in an exemplary embodiment of the present disclosure is shown.

[0096] According to an exemplary embodiment of the present disclosure, referring to Figure 3 , we can design a capture-escape mission consisting of a capture team P and an escapee E, Figure 3 The formation configuration of the multi-bionic robotic fish system is demonstrated, and obstacles are also included in the scene.

[0097] Among them, the capture team can be composed of 3 people with a perception radius R c The robot fish agent is composed of i , P j , P k, and all individuals are isomorphic. The escapee E can be a person with a perception radius of R e A robotic fish agent.

[0098] Based on the current state information of each agent in the scene and the current position information of the obstacle, the local observation information of a single agent can be determined. Based on the local observation information and the global observation information (global information), the encirclement and capture movement strategy of the agent can be generated based on the attraction mechanism.

[0099] For example, at the current moment, for the hunter P i , whose local observation information includes the agent's yaw angle ψ i State information such as the relative yaw angle Ψ between the agent and the obstacle i,obs , the relative distance d between the agent and the obstacle i,obs , the relative distance d between the agent and the escaper i,e etc. relative status information; for the capturer P j , whose local observation information includes the agent's yaw angle ψ j State information such as the relative distance d between the agent and the escaper j,e etc. relative status information; for the capturer P k , whose local observation information includes the agent's yaw angle ψ k Based on the global observation information including the state information and relative state information of each capturer, the capture movement strategy for the current formation state can be generated by processing the global observation information through the evaluation network and processing the local observation information of each capturer by each capture strategy network, that is, the movement information: i , where the agent's relative yaw angle △ψ relative to the escaper i,e Next, move l i Distance; for the capturer P j , where the agent's relative yaw angle △ψ relative to the escaper j,e Next, move l j Distance; for the capturer P k , where the agent's relative yaw angle △ψ relative to the escaper k,e Next, move l k distance.

[0100] In addition, for the escaper E, its escape strategy network, evaluation network, etc. can be trained separately. Similarly, the escaper E can also determine the local observation information of the escaper E based on the current state information of each intelligent agent in the scene and the current position information of the obstacle, and generate the escape motion strategy of the intelligent agent based on the attraction mechanism based on the local observation information and the global observation information (global information).

[0101] For example, at the current moment, for the escaper E, its local observation information may include the yaw angle ψ of the agent e Based on the global observation information including the state information of each capturer and escapee and the relative state information, the escape motion strategy for the current formation state can be generated by processing the global observation information through the escapee's evaluation network and processing the local observation information of the escapee E through the escape strategy network, that is, the motion information for the escape: at the relative yaw angle △ψ of the agent e Next, move l e distance.

[0102] It can be understood that the above-mentioned state information, relative state information, action information, etc. can all be vectors including direction information.

[0103] Although the number of escapers is less than that of the capturers, in the exemplary embodiment of the present disclosure, the escapers can have higher mobility and swimming speed. The capturers' goal is to increase the teamwork performance so that they can use their numerical advantage to capture the more agile escapers. Each individual can sense R c In each training round, if the distance between any capturer in the capture team and the escaper is less than R c , then the capture is considered to be successfully completed. On the contrary, if the escapee successfully survives in a round, the capture mission of that round is considered to have failed.

[0104] Just as an example, in this task, the training can be arranged in a volume of 4×4m 2 In a pool environment, all entities can be confined to the boundaries with non-collisionable walls around them. The maximum speed of the captor can be set to 0.6m / s, the maximum angular velocity can be set to 0.9rad / s, and the maximum speed of the escaper can be set to 0.9m / s, the maximum angular velocity can be set to 1.2rad / s.

[0105] Return to reference Figure 1 In step 102, in a training round, based on the reward scores and scoring information obtained by the at least two hunters in the training round, the hunting strategy networks corresponding to the at least two hunters are trained respectively.

[0106] Next, a multi-agent reinforcement learning algorithm framework for underwater scenarios in an exemplary embodiment of the present disclosure is introduced.

[0107] Figure 4 A schematic diagram showing a multi-agent reinforcement learning algorithm framework for underwater scenarios in an exemplary embodiment of the present disclosure.

[0108] Reference Figure 4 The multi-agent reinforcement learning algorithm framework for underwater scenes can adopt the MADDPG network (Multi-Agent Deep Deterministic Policy Gradient, based on multi-agent deep deterministic policy gradient network), and can obtain a real-time decision network, namely the capture strategy network, through the centralized training with decentralized execution (CTDE) training paradigm. The input is the state information such as the relative position and yaw angle of the bionic robot fish, and the output is an independent control signal for each agent, thereby realizing the adaptive cooperative capture strategy of the capturer to the escapee.

[0109] In the underwater pursuit and escape training stage, we can first build a Multi-Agent Attraction-enhanced Adaptive Underwater Pursuit algorithm (MA3UP) based on MADDPG. Figure 4 As shown in (a) in the figure, it is a multi-agent reinforcement learning training framework for underwater scenes, which consists of a decentralized strategy network (i.e., the capture strategy network), a centralized evaluation network (i.e., the evaluation network), and a state transfer model based on the attraction mechanism (i.e., the state transfer model).

[0110] Among them, [μ 1 , μ 2 ,…], [l 1 , l 2 , …] represents the action set of all agents, μ can represent the relative yaw angle, l can represent the moving distance, and r represents the reward score.

[0111] The status information S of each hunter at the current moment 1 ,…,S 3 , input them into their corresponding roundup strategy networks respectively, and obtain the action information μ 1 +l 1 ,…,μ 3 +l 3 , that is, the strategic information of the round-up action, where θ 1 ,…,θ 3 Represents the parameters of each roundup strategy network.

[0112] The status information and action information S of each capturer can be 1 +μ 1 +l 1 ,…,S 3+μ 3 +l 3 Input into the evaluation network to obtain the scoring information Q for the round-up strategy (action information) 1 ,Q 2 ,Q 3 , where w represents the parameters of the evaluation network, which can be updated according to its own gradient; in addition, the score information Q 1 ,Q 2 ,Q 3 Can be used to adjust θ 1 ,…,θ 3 , to train the round-up strategy network.

[0113] For each hunter agent, after inputting the current state information S into the hunting strategy network θ to obtain the action information μ and l, the attraction F can be added based on the action information μ. att and the repulsive force F rep , so as to obtain the relative yaw angle △Ψ by summing ∑. On this basis, based on the moving distance l, the relative yaw angle △Ψ and the current state information S, based on the current state transition environment Env, the state information S′ and the reward score r at the next moment can be obtained.

[0114] Reference Figure 4 (b) shows the training process of centralized training and decentralized execution. In the execution phase, each agent can make decisions based on the observed local state information and obtain corresponding rewards through the environment state transition model. 1 , r 2 , ...] represents the reward set, and the learning goal of each agent is to maximize its cumulative reward Among them, γ is the attenuation factor.

[0115] It can be understood that the noise during the state transfer process is the additional force.

[0116] During the training phase, the centralized evaluation network can guide the update of the policy network. In the process of centralized training and decentralized execution, global information can ensure the stability of the network during training, while decisions made based on local information can ensure the scalability of the network.

[0117] During the training process, the agents can observe the relative positions of all predators and evaders. Each agent has a critic network that fully observes the environment and estimates the value of state-action pairs. In addition, each agent takes completely independent actions based on its own observations through an independent policy network.

[0118] Understandably, Figure 4The state information S shown in includes the state information of the current capturer and the relative state information between the current capturer and other capturers, escapees, and obstacles.

[0119] At this point, for underwater simulation environments and capture-and-escape mission scenarios, the MADDPG algorithm can be used to obtain a capture-and-escape strategy for underwater scenarios using a centralized training and distributed execution training paradigm.

[0120] According to an exemplary embodiment of the present disclosure, a multi-agent reinforcement learning reward function for an underwater capture-escape scenario may be designed by combining the relative state information of each capturer with other capturers, escapers, and obstacles.

[0121] According to an exemplary embodiment of the present disclosure, a bonus score can be calculated based on a first bonus score, a second bonus score and a penalty score, the first bonus score being determined by the result of the roundup mission, the second bonus score being determined by the distance between the roundup and the escapee, and the penalty score being determined by the distance between at least two roundups and an obstacle.

[0122] According to an exemplary embodiment of the present disclosure, each agent i is rewarded r in the process of interacting with the environment. i (i.e. r) can include three parts, expressed as follows:

[0123]

[0124] The exemplary embodiment of the present disclosure focuses on the reward for the capturer. The main task of the capturer is to capture the escapee. Therefore, when the capturer i meets the capture condition, he can obtain the first reward score, which is the main reward r as shown below: main,i :

[0125]

[0126] Among them, d i =‖p i -p evd ‖ represents the distance between the capturer i and the escaper evd.

[0127] In addition, since the positive samples in the early stage of training are sparse, in order to avoid the sparse reward problem, an auxiliary reward based on the attraction enhancement mechanism, that is, the second reward score, can be designed to guide the captor to move in the direction of approaching the escapee. Its specific form can be expressed as follows:

[0128]

[0129] Among them, ||p i -p j|| represents the distance between the captor i and the escaper j.

[0130] In addition to the main task of catching the escapee, the agent also needs to avoid obstacles. Therefore, a corresponding sub-goal reward can be designed. When the distance d between the agent and the obstacle is 0 Less than the threshold d min , a corresponding penalty can be given, namely the penalty score:

[0131] r sub = -δ·max(d min -d 0 ,0)

[0132] Among them, δ represents the weight coefficient of the sub-goal reward.

[0133] According to an exemplary embodiment of the present disclosure, the mission goal may be that multiple hunters, through an adaptive cooperative hunting strategy, surround a fugitive with higher mobility in a map with complex terrain. According to the mission requirements, a MA3UP network consisting of two hunters and one invader, i.e., three agents, may be designed in an exemplary embodiment of the present disclosure. The structure of the policy network and the evaluation network of each hunter agent are the same, and its hidden layer may have 64 hidden nodes. Among them, the input dimensions of the evaluation network are the state, relative state, and action set of all agents. The input of the policy network is the state and relative state of the current agent, and the output is its corresponding decision action. All networks can use RELU as the activation function. The training parameters are set as follows: the discount factor is 0.95, the policy learning rate is 0.01, the update factor u is 0.01, and the replay buffer size is 1×10 6 , the batch size is 1024.

[0134] Next, the process of training through the curriculum learning method is introduced.

[0135] In reinforcement learning, the update of the policy network depends on the incentives of positive samples and reward functions, and it is difficult to explore the real expected policy behavior through random exploration in the early stage of training. Therefore, a learning method from simple to complex can be adopted to increase the probability of positive samples and beneficial behaviors in the early stage of training to speed up convergence and obtain better generalization ability.

[0136] In the capture-and-escape scenario, factors that affect the success of the task include the relative mobility between agents, the perception of the capturer, and the complexity of the environment. Therefore, exemplary embodiments of the present disclosure can start training from a relatively simple setting and gradually increase the difficulty of the task.

[0137] For example, the network can be trained through phased curriculum learning until the network converges, and the effectiveness of the obtained strategy can be verified on a physical platform.

[0138] According to an exemplary embodiment of the present disclosure, the training method may include multiple training stages, and each training stage may include multiple training rounds. For example, two training stages may be set. Specifically, in the first training stage in which no obstacles are set in a liquid environment, each capture strategy network is trained based on a first preset situation, wherein the first preset situation is: the maneuverability of the capturer is less than the maneuverability of the escapee, and the difference in maneuverability between the capturer and the escapee is a first preset value, the perception range of the capturer is a second preset value, and the perception range of the capturer is used to determine that the capture mission is successful when the escapee is within the perception range of the capturer; in the second training stage in which obstacles are set in a liquid environment, each capture strategy network is trained based on a second preset situation, wherein the second preset situation is obtained by gradually increasing the first preset value in the first preset situation and gradually decreasing the second preset value.

[0139] Figure 5 A schematic diagram showing the comparison of effects between phased training and continuous training using curriculum learning in an exemplary embodiment of the present disclosure.

[0140] Reference Figure 5 Specifically, in the initial stage of training, i.e., the first training stage (stage I), the mobility difference between the captor and the escapee can be set to be relatively low, and the captor can be set to have a larger perception range. The first training stage can be conducted in a scene without obstacles.

[0141] As an example, the training results show that after about 20,000 rounds of training, the capture team can come up with a cooperative capture strategy. On this basis, in the second training phase (Phase II), the mobility difference between the capturers and the escapees can be gradually increased, the capturers' perception range can be reduced, and complex terrain can be set in the map.

[0142] Figure 5 The training effect was evaluated using the average capture success rate per 1,000 rounds (training generations). The results showed that after 140,000 rounds of training, the average success rate of Phase I of curriculum learning reached 80%. After completing 200,000 rounds of training, the average success rate in Phase II reached 90%. In contrast, without curriculum learning, even after 200,000 rounds of training, the average success rate was less than 65%. Through curriculum learning, the capture team was able to develop an adaptable cooperative capture strategy that maintained a high capture success rate in various obstacle scenarios.

[0143] It is understandable that the specific number of rounds mentioned above may vary depending on the number of agents in the scene, the size and difference of the mobility between the agents, the perception range of the agents, etc.

[0144] Figure 6 , Figure 7 , Figure 8 A screenshot showing a video of an underwater multi-agent capture-escape experiment consisting of two capturers and one escaper in an exemplary embodiment of the present disclosure.

[0145] Reference Figure 6 , Figure 7 , Figure 8 , according to an exemplary embodiment of the present disclosure, wherein the blue track is the capturer P 1 , the green track is the hunter P 2 , the yellow track is the escapee. The perception domain (perception range) of the capturer is represented by a dotted circle. When the perception domain is black, it means that the capture has not been achieved, when it is yellow, it means that there is a high probability of capture, and when it is red, it means that the capture has been successfully achieved. In order to more intuitively show the capture team's capture trend of the escapee, the arrow in the figure represents the capturer's yaw angle, which is used to indicate the capturer's movement at the next moment.

[0146] Figure 6 Demonstrated a cooperative roundup strategy. 1 and P 2 Surround the escapee from two different directions and try to force him into a confined area. This trapping strategy is effective, especially when the escapee is trapped in a corner. The narrow space greatly restricts the escapee's movements, making it almost impossible for him to escape. This process vividly demonstrates the efficient cooperation between the capturers, that is, the capturers cooperate with each other to successfully trap the highly maneuverable escapee in a confined corner.

[0147] Figure 7 The adaptive capture strategy of the capturer in a complex obstacle environment is demonstrated. Figure 7 In (a), when the escapee tries to escape through a narrow gap in the obstacle, the captor P 1 and P 2 Make quick judgments and flexibly adjust their action strategies according to the terrain characteristics. Figure 7 As shown in (b), P 1 and P 2 They chose to bypass obstacles and approach from both sides at the same time to form a siege. In this process, the hunters showed a high degree of adaptability to complex scenarios, ensuring that they could effectively capture the fugitive. Figure 7 In (c), the capturer successfully captured the escapee with precise timing and adaptive strategy adjustment. This process vividly illustrates how the capturer can adaptively adjust the strategy according to the environmental characteristics, thereby improving the capture efficiency and successfully completing the capture mission.

[0148] Figure 8 The role-switching behavior of the capture team in the execution of the roundup mission is demonstrated. Figure 8 In (a), the hunter P 1 and P 2 Initial attempts were made to capture the escapee directly. However, the narrow path of the escapee and the multiple possible escape routes made direct capture almost impossible. Figure 8 In (b), the hunter P 1 According to this situation, the strategy was quickly adjusted, and preventive interception was carried out on the potential escape route of the escapee. Figure 8 As shown in (c), through clear role division and close teamwork, the capturers successfully implemented an effective capture of the escapee. This process vividly demonstrates the role adaptability in the capture strategy. The capturers can autonomously adjust the pursuit strategy according to their respective positions and surrounding environment information, thereby increasing the success rate of the mission.

[0149] According to the exemplary embodiments of the present disclosure, the present disclosure belongs to the field of underwater bionic robot fish swarm intelligence and intelligent decision-making. From the research perspective of multi-underwater intelligent agent game confrontation, a bionic robot fish cluster adaptive collaborative capture strategy based on attraction mechanism is proposed, which utilizes cluster intelligence and team coordination to realize the adaptive cooperative capture of bionic robot fish in complex environments.

[0150] The adaptive collaborative strategy of bionic robot fish swarm based on attraction mechanism proposed in the exemplary embodiment of the present disclosure is specifically divided into underwater collaborative capture task modeling, adaptive capture strategy and multi-agent reinforcement learning method with centralized training and decentralized execution.

[0151] Specifically, an adaptive cooperative capture strategy for multiple underwater bionic robot fishes that combines the artificial potential field method with multi-agent reinforcement learning is designed. The method includes an underwater bionic robot fish capture and escape algorithm based on attraction enhancement, combines the artificial potential field method with the reinforcement learning method, and proposes an adaptive capture strategy that includes target attraction, obstacle repulsion, and additional force, which can realize the collaborative capture and adaptive decision-making among multiple bionic robot fishes. Specifically, the method constructs an underwater multi-agent attraction-enhanced adaptive capture algorithm (MA3UP) by designing a capture method based on the attraction mechanism. By training based on the multi-agent deep deterministic policy gradient network (MADDPG) and using the centralized training decentralized execution (CTDE) training paradigm, an adaptive cooperative capture strategy for bionic robot fish clusters facing underwater scenes can be finally obtained, which can effectively realize the adaptive capture task of bionic robot fishes in different obstacle scenes. Finally, a multi-agent capture strategy that meets actual complex scenes is obtained, which provides theoretical methods and technical support for the intelligent control research of underwater multi-bionic robot fish systems.

[0152] Fig. 9 A flowchart showing a method for rounding up using a roundup strategy network based on multi-agent reinforcement learning in an exemplary embodiment of the present disclosure.

[0153] An exemplary embodiment of the present disclosure also provides a method for rounding up a capture strategy network based on multi-agent reinforcement learning, which is applied to a liquid environment containing at least two capturers and an escapee, wherein the rounding up strategy network is trained by the above-mentioned training method for the rounding up strategy network based on multi-agent reinforcement learning.

[0154] Reference Fig. 9 In step 901, current status information of at least two capturers at the current moment and relative status information of at least two capturers to other capturers and the escaper are respectively obtained.

[0155] In step 902, the state information of at least two capturers and the relative state information of at least two capturers to other capturers and escapers are respectively input into the capture strategy network corresponding to the at least two capturers to obtain the action information of at least two capturers.

[0156] In step 903, the state information and action information of at least two capturers are input into the state transition model to obtain the state information of at least two capturers at the next moment.

[0157] In step 904, at least two capturers perform a task of capturing the escapee based on the state information of each of the at least two capturers at the next moment.

[0158] It can be understood that in the exemplary embodiment of the method for rounding up the roundup strategy network based on multi-agent reinforcement learning, the specific implementation process is roughly the same as the exemplary embodiment of the training method for the roundup strategy network based on multi-agent reinforcement learning, and will not be repeated here.

[0159] Fig.10 A block diagram of a training device for a capture strategy network based on multi-agent reinforcement learning in an exemplary embodiment of the present disclosure is shown.

[0160] Reference Fig.10 An exemplary embodiment of the present disclosure also provides a training device 1000 for a capture strategy network based on multi-agent reinforcement learning, which may include but is not limited to an information acquisition unit 1001, an action decision unit 1002, a state transfer unit 1003, an action evaluation unit 1004, and a network training unit 1005.

[0161] The training device 1000 of the capture strategy network based on multi-agent reinforcement learning is applied to a liquid environment including at least two capturers and escapers.

[0162] Among them, in a training round including multiple capture-escape action moments, for each moment:

[0163] The information acquisition unit 1001 respectively acquires the status information of at least two capturers at the current moment and the relative status information of at least two capturers to other capturers and the escapee.

[0164] The action decision unit 1002 inputs the state information of at least two capturers and the relative state information of at least two capturers to other capturers and the escapee into the capture strategy network corresponding to the at least two capturers, respectively, to obtain the action information of at least two capturers. The action information of at least two capturers refers to the decision information of the capture action of at least two capturers at the current moment.

[0165] The state transfer unit 1003 inputs the state information and action information of at least two capturers into the state transfer model to obtain the state information of at least two capturers at the next moment and the corresponding reward score.

[0166] The action evaluation unit 1004 inputs the global information into the evaluation network to obtain scoring information for the action information of at least two capturers, wherein the global information includes: the state information and action information of at least two capturers at the current moment, and the state information and action information of the escaper at the current moment, wherein the action information of the escaper refers to the decision information of the escape action of the escaper at the current moment.

[0167] In the training round, the network training unit 1005 trains the capture strategy networks corresponding to the at least two capturers respectively based on the reward scores and scoring information obtained by the at least two capturers in the training round.

[0168] It is understandable that in the exemplary embodiment of the training device 1000 for the roundup strategy network based on multi-agent reinforcement learning, the specific implementation process is roughly the same as the exemplary embodiment of the training method for the roundup strategy network based on multi-agent reinforcement learning, and will not be repeated here. The training device 1000 for the roundup strategy network based on multi-agent reinforcement learning can be configured as software, hardware, firmware or any combination of the above items to perform specific functions. For example, these devices may correspond to dedicated integrated circuits, pure software codes, or modules that combine software and hardware. In addition, one or more functions implemented by these devices may also be uniformly executed by components in physical entity devices (e.g., processors, clients, or servers, etc.).

[0169] Fig.11A block diagram of an apparatus for conducting roundups based on a roundup strategy network of multi-agent reinforcement learning in an exemplary embodiment of the present disclosure is shown.

[0170] Reference Fig.11 An exemplary embodiment of the present disclosure also provides a device 1100 for conducting roundups based on a roundup strategy network based on multi-agent reinforcement learning, which may include but is not limited to an information acquisition unit 1101, an action prediction unit 1102, a state transfer unit 1103 and a roundup execution unit 1104.

[0171] The device 1100 for encircling and capturing based on a multi-agent reinforcement learning encirclement strategy network is applied to a liquid environment containing at least two capturers and an escapee, wherein the encirclement strategy network is trained by the training method of the encirclement and capturing strategy network based on multi-agent reinforcement learning.

[0172] The information acquisition unit 1101 may respectively acquire the current state information of at least two capturers at the current moment and the relative state information of at least two capturers to other capturers and the escapee.

[0173] The action prediction unit 1102 can input the state information of at least two capturers and the relative state information of at least two capturers to other capturers and escapers into the capture strategy network corresponding to the at least two capturers to obtain the action information of at least two capturers.

[0174] The state transfer unit 1103 may input the state information and action information of the at least two capturers into the state transfer model to obtain the state information of the at least two capturers at the next moment.

[0175] The capture execution unit 1104 may execute the capture task of the at least two capturers on the escapee based on the state information of the at least two capturers at the next moment.

[0176] It is understandable that in the exemplary embodiment of the device 1100 for rounding up based on the rounding up strategy network of multi-agent reinforcement learning, the specific implementation process is roughly the same as the exemplary embodiment of the method for rounding up based on the rounding up strategy network of multi-agent reinforcement learning, and will not be repeated here. The device 1100 for rounding up based on the rounding up strategy network of multi-agent reinforcement learning can be configured as software, hardware, firmware or any combination of the above items to perform specific functions. For example, these devices may correspond to dedicated integrated circuits, pure software codes, or modules combining software and hardware. In addition, one or more functions implemented by these devices may also be uniformly executed by components in physical entity devices (e.g., processors, clients, or servers, etc.).

[0177] Fig.12 A block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure.

[0178] Reference Fig.12 The electronic device 1200 includes at least one memory 1201 and at least one processor 1202, wherein the at least one memory 1201 stores a set of computer executable instructions. When the computer executable instruction set is executed by the at least one processor 1202, a training method for a hunting strategy network based on multi-agent reinforcement learning and a method for hunting using a hunting strategy network based on multi-agent reinforcement learning according to an exemplary embodiment of the present disclosure are executed.

[0179] As an example, the electronic device 1200 may be a PC, a tablet device, a personal digital assistant, a smart phone, or other device capable of executing the above-mentioned instruction set. Here, the electronic device 1200 is not necessarily a single electronic device, but may also be any device or circuit collection capable of executing the above-mentioned instructions (or instruction sets) individually or in combination. The electronic device 1200 may also be part of an integrated control system or system manager, or may be configured as a portable electronic device interconnected with a local or remote (e.g., via wireless transmission) interface.

[0180] In the electronic device 1200, the processor 1202 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller or a microprocessor. As an example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0181] The processor 1202 may execute instructions or codes stored in the memory 1201, wherein the memory 1201 may also store data. Instructions and data may also be sent and received over a network via a network interface device, wherein the network interface device may employ any known transmission protocol.

[0182] The memory 1201 may be integrated with the processor 1202, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. In addition, the memory 1201 may include a separate device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The memory 1201 and the processor 1202 may be operatively coupled, or may communicate with each other, such as through an I / O port, a network connection, etc., so that the processor 1202 can read files stored in the memory.

[0183] In addition, the electronic device 1200 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 1200 may be connected to each other via a bus and / or a network.

[0184] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein, when the instructions are executed by at least one computing device, the at least one computing device is prompted to execute the above-mentioned training method of the roundup strategy network based on multi-agent reinforcement learning and the method of rounding up the roundup strategy network based on multi-agent reinforcement learning.

[0185] Examples of computer-readable storage media include read-only memory (ROM), random-access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random-access memory (RAM), dynamic random-access memory (DRAM), static random-access memory (SRAM), flash memory, nonvolatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device is configured to store computer programs and any associated data, data files and data structures in a non-transitory manner and provide the computer programs and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system, so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers. It should be noted that the instructions can also be used to execute additional steps in addition to the above steps or to perform more specific processing when executing the above steps. These additional steps and further processing contents have been mentioned in the description of the relevant methods, so they will not be repeated here to avoid repetition.

[0186] Another embodiment of the present disclosure relates to a system comprising at least one computing device and at least one storage device storing instructions, wherein the instructions, when executed by the at least one computing device, prompt the at least one computing device to execute the above-mentioned training method of the roundup strategy network based on multi-agent reinforcement learning and the method of rounding up the roundup strategy network based on multi-agent reinforcement learning.

[0187] It should be noted that the system according to the exemplary embodiments of the present disclosure can completely rely on the execution of computer programs or instructions to realize corresponding functions, that is, each unit corresponds to each step in the functional architecture of the computer program, so that the entire system is called through a special software package (e.g., lib library) to realize the corresponding functions.

[0188] On the other hand, when the above-mentioned system is implemented in software, firmware, middleware or microcode, the program code or code segment for performing the corresponding operation can be stored in a computer-readable medium such as a storage medium, so that at least one processor or at least one computing device can perform the corresponding operation by reading and running the corresponding program code or code segment.

[0189] According to an exemplary embodiment of the present disclosure, the storage device may be integrated with the computing device, for example, RAM or flash memory is arranged within an integrated circuit microprocessor, etc. In addition, the storage device may include an independent device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The storage device and the computing device may be operationally coupled, or may communicate with each other, such as through an I / O port, a network connection, etc., so that the computing device can read instructions stored in the storage device.

[0190] Another embodiment of the present disclosure relates to a computer program product, including a computer program / instruction, which, when executed by a processor, implements the training method of the roundup strategy network based on multi-agent reinforcement learning and the method of rounding up using the roundup strategy network based on multi-agent reinforcement learning as described in any one of the above.

[0191] According to the training method and device of the hunting strategy network based on multi-agent reinforcement learning and the method and device, electronic device, storage medium and computer program product for hunting based on the hunting strategy network provided by the present invention, centralized training is realized by evaluating the processing of global information by the network, and decentralized execution is realized by the hunting strategy networks of the hunters respectively. The paradigm of centralized training and decentralized execution is adopted to ensure the stability and scalability of the algorithm in complex environments, and can further realize the cooperative hunting of bionic robot fish in complex environments. The present invention is not only efficient, but also can meet the task requirements in various complex scenarios.

[0192] In addition, the decision-making mechanism of enhanced attraction is adopted to predict the state of the captor at the next moment, which can ensure the security and adaptability of the strategy while ensuring efficient capture.

[0193] In addition, the curriculum learning method improves the practicality of training and effectively shortens the training cycle by gradually increasing the difficulty of training.

[0194] The above describes various exemplary embodiments of the present disclosure, and it should be understood that the above description is only exemplary and not exhaustive, and the present disclosure is not limited to the disclosed exemplary embodiments. Without departing from the scope and spirit of the present disclosure, many modifications and changes are obvious to those of ordinary skill in the art. Therefore, the scope of protection of the present disclosure should be based on the scope of the claims.

Claims

1. A training method for a hunting strategy network based on multi-agent reinforcement learning, characterized in that: Applied to a liquid environment including at least two capturers and escapers, the training method of the capture strategy network based on multi-agent reinforcement learning includes: In a training episode that includes multiple capture-and-escape moments, for each moment, do the following: Respectively obtain the state information of each of the at least two capturers at the current moment and the relative state information of each of the at least two capturers to other capturers and the escaper, wherein the state information includes the position and yaw angle of the capturer, and the relative state information includes the relative position and relative yaw angle between the capturer and other capturers and the escaper; Inputting the state information of each of the at least two capturers and the relative state information of each of the at least two capturers to other capturers and the escaper into the capture strategy network corresponding to each of the at least two capturers, respectively, to obtain the action information of each of the at least two capturers, wherein the action information of each of the at least two capturers refers to the decision information of the capture action of each of the at least two capturers at the current moment; Inputting the state information and action information of each of the at least two hunters into a state transition model to obtain the state information of each of the at least two hunters at the next moment and the corresponding reward score; and Inputting the global information into the evaluation network to obtain the scoring information for the action information of each of the at least two capturers, wherein the global information includes: the state information and action information of each of the at least two capturers at the current moment, and the state information and action information of the escaper at the current moment, wherein the action information of the escaper refers to the decision information of the escape action of the escaper at the current moment; In the training round, based on the reward scores and the scoring information obtained by each of the at least two hunters in the training round, the hunting strategy networks corresponding to the at least two hunters are trained respectively.

2. The training method of the hunting strategy network based on multi-agent reinforcement learning as claimed in claim 1, characterized in that: The liquid environment also includes obstacles, The step of respectively obtaining the status information of the at least two capturers at the current moment and the relative status information of the at least two capturers to other capturers and the escapee further includes: Respectively obtaining the relative status information between the at least two capturers and the obstacle at the current moment; The state information of each of the at least two capturers and the relative state information of each of the at least two capturers to other capturers and the escaper are respectively input into the corresponding capture strategy network to obtain the action information of each of the at least two capturers, including: Inputting the state information of each of the at least two capturers and the relative state information of each of the at least two capturers and other capturers, the escaper and the obstacle into the corresponding capture strategy network to obtain the action information of each of the at least two capturers; Wherein, the global information also includes: status information of the obstacle.

3. The training method of the hunting strategy network based on multi-agent reinforcement learning as claimed in claim 2, characterized in that: The step of inputting the state information and action information of the at least two capturers into the state transition model to obtain the state information of the at least two capturers at the next moment comprises: For each of the at least two hunters, the following operations are performed: Determining an attractive force corresponding to the capturer based on a distance between the capturer and the escapee; Determining a repulsive force corresponding to the besieger based on a distance between the besieger and the obstacle; Determine an additional force corresponding to the besieger based on the state information of the besieger at the current moment; Adjusting the action information of the captor based on the attractive force, the repulsive force, and the additional force; The state information of the besieger at the current moment and the adjusted action information of the besieger are input into the state transition model to obtain the state information of the besieger at the next moment.

4. The training method of the hunting strategy network based on multi-agent reinforcement learning as claimed in claim 3 is characterized in that: The training method includes multiple training stages, each training stage includes multiple training rounds, and the training of the capture strategy networks corresponding to the at least two capturers respectively includes: In the first training phase in the liquid environment where no obstacles are set, each capture strategy network is trained based on a first preset situation, wherein the first preset situation is: the mobility of the capturer is less than the mobility of the escapee, and the mobility difference between the capturer and the escapee is a first preset value, and the perception range of the capturer is a second preset value, and the perception range of the capturer is used to determine that the capture mission is successful when the escapee is within the perception range of the capturer; In the second training phase in which obstacles are set in the liquid environment, each of the encirclement strategy networks is trained based on a second preset scenario, wherein the second preset scenario is obtained by gradually increasing the first preset value in the first preset scenario and gradually decreasing the second preset value.

5. The training method of the hunting strategy network based on multi-agent reinforcement learning as claimed in claim 3, characterized in that: The bonus score is calculated based on a first bonus score, a second bonus score and a penalty score, wherein the first bonus score is determined by the result of the roundup mission, the second bonus score is determined by the distance between the roundup and the escapee, and the penalty score is determined by the distance between the at least two roundups and the obstacle.

6. The training method of the hunting strategy network based on multi-agent reinforcement learning as claimed in claim 3, characterized in that: The action information of the capturer includes heading change information; The adjusting the action information of the captor based on the attraction, the repulsion, and the additional force includes: Based on the attractive force, the repulsive force, and the additional force, the heading change information of the captor is adjusted, wherein the value of the additional force is determined by the heading change information.

7. The training method of the hunting strategy network based on multi-agent reinforcement learning as claimed in claim 6, characterized in that: The action information of the captor also includes a moving distance, the minimum value of the moving distance is greater than the distance the captor moves in one swing cycle, and the distance moved in one swing cycle is determined by the minimum turning radius of the captor.

8. A method for rounding up based on a rounding up strategy network based on multi-agent reinforcement learning, characterized in that: Applied to a liquid environment including at least two captors and an escapee, the method for encircling using a multi-agent reinforcement learning-based encirclement strategy network includes: Respectively obtain current state information of the at least two capturers at the current moment and relative state information of the at least two capturers to other capturers and the escaper, wherein the state information includes the position and yaw angle of the capturers, and the relative state information includes the relative position and relative yaw angle between the capturers and other capturers and the escaper; Inputting the state information of each of the at least two capturers and the relative state information of each of the at least two capturers to other capturers and the escaper into the capture strategy network corresponding to each of the at least two capturers, respectively, to obtain the action information of each of the at least two capturers; Inputting the state information and action information of the at least two capturers into the state transition model to obtain the state information of the at least two capturers at the next moment; Executing the task of capturing the escapee by the at least two capturers based on the state information of each of the at least two capturers at the next moment; Wherein, the capture strategy network is trained by the training method of the capture strategy network based on multi-agent reinforcement learning as described in any one of claims 1-7.

9. A training device for a hunting strategy network based on multi-agent reinforcement learning, characterized in that: Applied to a liquid environment including at least two capturers and escapers, the training device of the capture strategy network based on multi-agent reinforcement learning comprises: Information acquisition unit; Action decision unit; state transfer unit; Action Assessment Unit; and Network training unit; Among them, in a training round including multiple capture-escape action moments, for each moment: The information acquisition unit respectively acquires the state information of the at least two capturers at the current moment and the relative state information of the at least two capturers to other capturers and the escaper, wherein the state information includes the position and yaw angle of the capturers, and the relative state information includes the relative position and relative yaw angle between the capturers and other capturers and the escaper; The action decision unit inputs the state information of the at least two capturers and the relative state information of the at least two capturers to other capturers and the escaper into the capture strategy network corresponding to the at least two capturers, respectively, to obtain the action information of the at least two capturers, wherein the action information of the at least two capturers refers to the decision information of the capture action of the at least two capturers at the current moment; The state transfer unit inputs the state information and action information of the at least two hunters into the state transfer model to obtain the state information of the at least two hunters at the next moment and the corresponding reward score; and The action evaluation unit inputs the global information into the evaluation network to obtain score information for the action information of each of the at least two capturers, wherein the global information includes: the state information and action information of each of the at least two capturers at the current moment, and the state information and action information of the escaper at the current moment, wherein the action information of the escaper refers to the decision information of the escape action of the escaper at the current moment; In the training round, the network training unit trains the capture strategy networks corresponding to the at least two capturers respectively based on the reward scores and the scoring information obtained by the at least two capturers in the training round.

10. A device for rounding up based on a rounding up strategy network of multi-agent reinforcement learning, characterized in that: Applied to a liquid environment including at least two captors and an escapee, the apparatus for performing encirclement and capture based on a multi-agent reinforcement learning encirclement and capture strategy network comprises: The information acquisition unit is configured to respectively acquire the current state information of the at least two capturers at the current moment and the relative state information of the at least two capturers to other capturers and the escaper, wherein the state information includes the position and yaw angle of the capturers, and the relative state information includes the relative position and relative yaw angle between the capturers and other capturers and the escaper; The action prediction unit is configured to: input the state information of each of the at least two capturers and the relative state information of each of the at least two capturers to other capturers and the escaper into the capture strategy network corresponding to each of the at least two capturers, so as to obtain the action information of each of the at least two capturers; The state transfer unit is configured to: input the state information and action information of the at least two capturers into the state transfer model to obtain the state information of the at least two capturers at the next moment; The capture execution unit is configured to: execute the capture task of the at least two capturers on the escapee based on the state information of each of the at least two capturers at the next moment; Wherein, the capture strategy network is trained by the training method of the capture strategy network based on multi-agent reinforcement learning as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-machine hunting method and device for hierarchical collaborative learning, electronic equipment and medium

    CN117350326A

  • Training method and device for underwater multi-agent surrounding escape game strategy

    CN118070066A