Unmanned ship formation path planning training system based on MAPPO

By improving the MAPPO algorithm framework and combining it with a simulation environment and an orthogonal coded random network distillation algorithm, the path planning capability of unmanned surface vessel formations in complex obstacle environments has been enhanced. This has solved the problems of slow training convergence and low obstacle avoidance success rate, and achieved more efficient learning and obstacle avoidance results.

CN120972937BActive Publication Date: 2026-05-05HUANENG LANCANG RIVER HYDROPOWER CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUANENG LANCANG RIVER HYDROPOWER CO LTD
Filing Date
2025-08-19
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

The existing MAPPO algorithm suffers from slow training convergence, poor path planning performance, and low obstacle avoidance success rate in unmanned surface vessel (USV) formation path planning, making it difficult to effectively cope with complex obstacle environments.

Method used

By introducing a simulation environment, Actor, reward unit, experience pool, critical path unit, and optimization unit, the MAPPO framework is improved. During the training process, critical path information is used to enhance learning efficiency. Combined with the orthogonal coding random network distillation algorithm, intrinsic rewards are provided to optimize the policy network and value network.

Benefits of technology

It improves the path planning capability of unmanned surface vessel formations in complex obstacle environments, enhances learning efficiency and obstacle avoidance success rate, solves the problem that traditional on-policy algorithms cannot utilize previously trained data, and avoids the accumulation of invalid samples in off-policy mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120972937B_ABST
    Figure CN120972937B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of unmanned surface vessel (USV) intelligent control technology, and provides a MAPPO-based USV formation path planning training system, including a simulation environment; multiple Actors, which determine their actions and update their next state based on the action policies and current states stored in their policy networks; a reward unit, which rewards each Actor based on its actions and current state; an experience pool, which stores training samples generated by each Actor executing its current action policy; a value network, which evaluates the value of the training samples; a critical path unit, which determines critical path sample combinations from paths generated by the current action policy and several historical action policies; and an optimization unit, which optimizes the policy network and value network by jointly using the training samples and critical path sample combinations based on the MAPPO optimization algorithm. The technical solution of this invention can significantly improve the path planning training effect in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of unmanned vessel intelligent control technology, specifically relating to an unmanned vessel formation path planning training system based on MAPPO. Background Technology

[0002] When unmanned surface vessels (USVs) navigate on the sea, they face various challenges such as complex marine environments and obstacles. Therefore, path planning is particularly important. Reasonable path planning not only ensures safe navigation but also improves efficiency and reduces resource consumption. Path planning is divided into global path planning and local obstacle avoidance. Traditional algorithms based on electronic nautical charts are suitable for global path planning, while reinforcement learning can flexibly utilize local sensor information to provide new solutions for collaborative navigation and obstacle avoidance of USVs.

[0003] The MAPPO algorithm (Multi-Agent Proximal Policy Optimization) is a reinforcement learning algorithm designed for multi-agent systems. It is based on the single-agent PPO algorithm (Proximal Policy Optimization) and extends it to scenarios of multi-agent cooperation and competition. It supports parallel training of multiple agents and is suitable for large-scale and complex multi-agent environments. Currently, there are technical solutions that introduce the MAPPO algorithm into the path planning of a formation composed of multiple unmanned vessels. Through centralized training and distributed execution, it can effectively improve the learning efficiency and generalization ability of unmanned vessel formation path planning.

[0004] However, when faced with the complex obstacle environment at sea, directly using the simple MAPPO algorithm to train the policy network will result in problems such as slow training convergence, poor path planning performance, and reduced obstacle avoidance success rate. Therefore, it is necessary to improve the existing MAPPO framework to enhance its ability to perform reasonable path planning for unmanned vessel formations in complex obstacle environments. Summary of the Invention

[0005] The purpose of this invention is to provide a training system for unmanned surface vessel (USV) formation path planning based on MAPPO. By improving the existing MAPPO framework, the system enhances the ability of USV formations to plan reasonable paths in complex obstacle environments. The training system includes:

[0006] The simulation environment includes a simulated water area model and the irregular obstacle model contained therein;

[0007] Multiple Actors, corresponding to each unmanned vessel in the unmanned vessel formation, determine the actions to be taken based on the action policies and current states stored in their policy network, and update the next state based on the actions taken.

[0008] The reward unit rewards each Actor based on their actions and current state.

[0009] The experience pool is used to store the training samples generated by each Actor when executing the current action strategy;

[0010] A value network is used to evaluate the value of the training samples;

[0011] Critical path unit, used to determine the critical path sample combination from the paths generated by the current action strategy and several historical action strategies;

[0012] The optimization unit, based on the MAPPO optimization algorithm, jointly optimizes the policy network and value network using the training samples and critical path samples.

[0013] The MAPPO-based unmanned surface vessel (USV) formation path planning training system provided by this invention extracts critical paths by setting critical path units. Since these critical paths largely represent the various continuous actions taken by the USV to avoid obstacles during formation travel, and the resulting continuous changes in state, they contain high-value information that plays a crucial role in improving obstacle avoidance capabilities. Therefore, introducing these high-value critical paths into the training process can effectively improve learning efficiency. In addition, the critical paths are extracted from paths generated by the current strategy and several historical strategies, and include all samples on the critical paths. This not only solves the problem that traditional on-policy algorithms cannot utilize data collected during previous training, but also enables training based on the entire process information of the USV from obstacle detection to decision-making to action to effect. Furthermore, it effectively avoids the problem of invalid / low-value samples accumulating over time due to directly using off-policy mechanisms. Attached Figure Description

[0014] Figure 1 A schematic diagram illustrating a mission scenario for unmanned surface vessel (USV) convoy navigation and obstacle avoidance.

[0015] Figure 2 A diagram illustrating the framework of the MAPPO algorithm for training the path planning capabilities of unmanned surface vessel formations.

[0016] Figure 3 This is a schematic diagram of the framework of an unmanned vessel formation path planning training system based on MAPPO according to an embodiment of the present invention.

[0017] Figure 4 This is a schematic diagram illustrating the movement of an unmanned vessel formation in a specific embodiment.

[0018] Figure 5 A schematic diagram of the architecture of a policy network provided according to an embodiment of the present invention;

[0019] Figure 6 A schematic diagram illustrating the participation of the critical path determination module and critical path storage module in the training process according to an embodiment of the present invention;

[0020] Figure 7 This is a schematic diagram illustrating the implementation process of the orthogonal coded random network distillation algorithm provided in an embodiment of the present invention;

[0021] Figure 8 This is a schematic diagram illustrating how the average reward changes during the training process in a specific embodiment of the present invention;

[0022] Figure 9 This is a schematic diagram illustrating how the TD error changes during the training process in a specific embodiment of the present invention. Detailed Implementation

[0023] The present invention will now be further described based on preferred embodiments and with reference to the accompanying drawings.

[0024] In the description of the embodiments of this invention, it should be noted that if terms such as "upper," "lower," "inner," and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this embodiment is in use, they are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, in order to distinguish different units, terms such as "first" and "second" are used in this specification, but these are not limited by the manufacturing order, nor should they be construed as indicating or implying relative importance. Their names may differ in the detailed description and claims of this invention.

[0025] The terminology used in this specification is for illustrative purposes and is not intended to limit the invention. It should also be noted that, unless otherwise explicitly stated and limited, the terms "set," "connected," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection, a direct connection, or an indirect connection via an intermediate medium; or they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of these terms in this invention.

[0026] Figure 1An exemplary scenario of unmanned surface vessel (USV) formation navigation and obstacle avoidance is shown. The USV formation in the figure includes three USVs, which perform the mission in a cooperative formation. This approach can improve mission efficiency, expand the operational coverage, and enhance the robustness and flexibility of the system. During the cooperative formation process, the USVs achieve autonomous navigation, formation maintenance, and efficient completion of complex tasks through information sharing and division of labor.

[0027] As shown in the figure, when an unmanned surface vessel (USV) formation is sailing toward a target area, it needs to avoid various dynamic and static obstacles on the sea surface in a timely manner and avoid breaking up its formation. Therefore, it is necessary not only to pay attention to the safety of individuals, but also to consider the maintenance of the overall formation. This requires each USV to avoid conflict with its teammates when adjusting its own trajectory, and to restore its formation as soon as possible after avoiding obstacles.

[0028] The motion strategy of unmanned surface vessel (USV) formations also needs to meet real-time requirements, that is, to quickly calculate safe and feasible trajectories in complex environments, while being limited by factors such as computing resources, communication bandwidth, and perception accuracy. At the same time, environmental uncertainties, such as water flow disturbances, sensor measurement errors, and communication delays, will affect the accuracy of obstacle avoidance decisions and increase the complexity of the problem. In addition, as the number of USVs and the complexity of obstacles on the sea surface increase, the difficulty of rationally planning paths will increase significantly.

[0029] As described in the background section, the MAPPO algorithm (Multi-Agent Proximal Policy Optimization) has been widely studied and applied in multi-agent task planning scenarios. Figure 2 The diagram illustrates the framework of the MAPPO algorithm, which is used to train the path planning capability of unmanned vessel formations. This training is conducted on an unmanned vessel formation composed of multiple unmanned vessel models in a simulation environment. The entire training process consists of the upper part, the training sample construction stage, and the lower part, the network training process.

[0030] Specifically, during the training sample construction phase of each round, the Actors corresponding to each unmanned surface vessel store executable action policies through a policy network. In the PPO algorithm, the action policy is the probability (action probability) of the agent taking each action in a certain state, generally expressed as p. t (a t |s t In the form of an action policy, the Actor network determines the current state s. t The next action to be executed is a t The aforementioned actions will interact with the environment and other unmanned surface vessel (USV) models, causing changes in the system state and leading the USV model to enter the next state s.t+1 And receive a reward for the action from the environment (or Actor network).

[0031] The above process is carried out iteratively, allowing each unmanned vessel model to continuously navigate in the simulation environment. At each time step, a set of data of the form (s) will be obtained. t ,a t ,r t ,s t+1 ,d t ,p t The training samples, where d t This is a flag indicating whether a round has ended. Generally, this value can be set to indicate the end of a round when the number of execution steps reaches a preset value or when the unmanned vessel model achieves a preset goal.

[0032] Training samples generated in real time by each unmanned surface vessel are continuously added to the experience pool until the end of the round, at which point the intensive training phase begins. Specifically, the training samples generated in this round are extracted in batches from the experience pool according to certain extraction rules. Subsequently, the value network is based on s t a t Information, by calculating the TD error (in the field of reinforcement learning, the TD error is used to measure the difference between the expected value of the current state and the next state), is applied to the state value V(s). t An evaluation will be conducted, and the actual reward received will be considered. t and the next state s t+1 Value prediction V(s) t+1 ), calculate the dominance function Advantage function This reflects the value of the current action relative to the current state. As a feedback signal to the Actor's action policy optimizer (also known as the learner), the Actor's optimizer updates the policy network based on this feedback information using the policy gradient method. The goal is to maximize the expected reward of the advantage function, thereby making the action selection more reasonable. At the same time, the value network optimizes its value function based on the error between the reward and the value prediction, so as to continuously improve the accuracy of the evaluation of the environment. Through repeated interaction and iterative updates, the Actor network and the value network gradually optimize together, thereby completing one round of training.

[0033] In some specific embodiments, the value network calculates the TD error at each time step based on the different termination conditions of the task in the following way:

[0034]

[0035] After obtaining the TD error, the TD error of a single time step can be used as the dominance function. Alternatively, a more stable dominance function can be obtained by weighted summation of the TD errors at the current time step and several future time steps.

[0036] During the intensive training phase, the core idea for optimizing the policy network is to limit the difference between the old and new policies in each iteration, thereby ensuring the smoothness of the update process. The objective function of PPO is constructed based on the importance sampling technique and can be used to improve the action policy as shown in the following equation:

[0037]

[0038] Among them, E t Let θ represent the expected value, and θ be the policy network parameters. Here, is the estimate of the dominance function at time step t, 'clip' is the cutoff function, 'ò' is a hyperparameter controlling the cutoff range, and ... t (θ) is the action probability ratio, calculated using the following formula:

[0039]

[0040] π θ (a t |s t The numbers ) represent the scenarios where the unmanned vessel is in state s, using the old and updated action strategies, respectively. t At that time, action a is executed. t The probability of.

[0041] The objective function is achieved through a cutoff ratio r. t (θ) limits the magnitude of policy updates, thereby avoiding training instability caused by excessive policy differences. When the advantage... When the probability is positive, the algorithm tends to increase the probability of the action, but the maximum probability will not exceed 1+ò; when the probability is negative, the algorithm will decrease the probability of the action, but the minimum probability will not be lower than 1-ò.

[0042] The objective function for optimizing the value network is shown in the following equation:

[0043]

[0044] Where φ represents the parameters of the valuation network, V φ (o t ) represents the network's value prediction of the current observation state at time t.

[0045] Since MAPPO is a typical on-policy algorithm, meaning it interacts with the environment in real time using the current policy (such as the latest policy network), and the collected trajectories / experiences (states, actions, rewards, next states) are used to directly train and update the same policy network, after completing one round of training sample collection and intensive training, the training samples stored in the experience pool are cleared (i.e., training samples generated by executing the current policy are discarded), and then the above two stages of training are repeated. Furthermore, some existing solutions (such as patent CN115509251A) further subdivide a round into multiple stages (Steps), using the latest training samples to update the policy network and value network in each stage, and clearing the experience pool after the training of that stage, thereby maintaining the freshness of the training samples.

[0046] Clearly, this traditional on-policy algorithm cannot utilize the data collected during previous training, thus ignoring some rare experiences that may be valuable during training. Especially in marine environments with complex static and dynamic obstacles, the various continuous actions taken by unmanned vessels to avoid obstacles during formation navigation, and the resulting continuous changes in state, contain high-value information that plays a crucial role in improving obstacle avoidance capabilities. Therefore, it is necessary to introduce the "high-value" paths generated during the current and historical policy execution into the training process to effectively improve learning efficiency.

[0047] Therefore, this invention provides a novel unmanned surface vessel (USV) formation path planning training system based on MAPPO, such as... Figure 3 As shown, the training system includes a simulation environment, Actors corresponding to each unmanned vessel, a value network, reward units, optimization units, and critical path units.

[0048] The specific implementation methods of the above parts are described in detail below with reference to the accompanying drawings and specific embodiments.

[0049] <Simulation Environment>

[0050] In embodiments of the present invention, such as Figure 1As shown, the simulation environment consists of a simulated water area model and obstacle models included therein. These obstacle models can be static, such as reefs protruding above the water, bridge piers, or various underwater work platforms, or dynamic, such as surface ships or other unmanned surface vessels not belonging to the unmanned surface vessel formation. In some optional embodiments, each obstacle model can move along the same trajectory in different rounds. In other optional embodiments, the obstacle models can navigate in a random manner or under the control of a trained control strategy network, thereby increasing the obstacle avoidance difficulty for the unmanned surface vessel formation and effectively improving the training effect.

[0051] <Motion Process and State Space of Unmanned Surface Vessels>

[0052] In embodiments of the present invention, the unmanned vessel formation may include two, three, or more unmanned vessels. Figure 4 The illustration schematically depicts the movement of an unmanned surface vessel (USV) formation consisting of three USVs. (Reference) Figure 4 , with P i i = 1, 2, 3 represents the i-th unmanned boat. This represents the speed of the i-th unmanned vessel. The velocity heading angle represents the i-th unmanned vessel formation. Let i represent the coordinates of the unmanned vessel i. Then the objective of the unmanned vessel formation can be expressed as the objective function in equation (1):

[0053]

[0054] Where, r G T represents the location of the target point. g The time it takes for the unmanned surface vessel (USV) convoy to reach the target point (the time it takes for the lead USV or any individual USV in the convoy to reach the target point can be used as T) g This formula represents the optimization objective of an unmanned vessel formation to reach the target point with the least amount of effort required while completing obstacle avoidance.

[0055] Accordingly, the kinematic equations of each unmanned vessel can be described by equation (2):

[0056]

[0057] The motion variables of each unmanned vessel include angular velocity ω i and acceleration a i Its constraints are shown in equation (3):

[0058]

[0059] in, That is the minimum speed of the unmanned boat. That is the maximum speed of the unmanned boat; This is the maximum acceleration of the unmanned vessel;

[0060] That is the maximum angular velocity of the unmanned vessel.

[0061] The initial state of each unmanned vessel is shown in equation (4):

[0062]

[0063] In path planning tasks for unmanned surface vessel (USV) swarms, the success of these tasks depends on several key factors, such as the coordinates of any one USV. Achieving equation (5) is a condition for completing the task:

[0064]

[0065] Among them, l g This is the preset distance threshold.

[0066] Furthermore, the motion trajectory of the unmanned vessel must be strictly limited within a predefined motion boundary to avoid exceeding the boundary or losing control. Boundary constraints can be set as shown in equation (6):

[0067]

[0068] Where, x min x max Let x and y be the lower and upper bounds of the unmanned surface vessel's navigation range in the x-direction, respectively. min y max These represent the lower and upper bounds of the unmanned vessel's navigation range in the y-direction, respectively.

[0069] Equations (1) to (6) serve as dynamic models for unmanned vessels, capable of describing the mission objectives, kinematic characteristics, and constraints of unmanned vessel formations. During the training sample construction process described later, each unmanned vessel model will follow the above equations and move in the simulation environment.

[0070] In addition, the state space of the unmanned vessel includes not only its position and velocity, which describe its own motion, but also its relative relationship with the target point and other obstacles. By monitoring this information in real time, the unmanned vessel can dynamically adjust its obstacle avoidance strategy to ensure that it reaches the target point efficiently and safely.

[0071] For any unmanned vessel P i It can obtain its own location information and speed data in real time, that is Meanwhile, to meet the task requirements, the unmanned boat must also have the capabilities of environmental perception and obstacle detection. Therefore, a radar and a forward-looking sonar are respectively installed to perform corresponding functions. However, in the actual environment, the movement speeds of other unmanned boats or obstacles are often difficult to accurately obtain. Therefore, in the present invention, the information that the unmanned boat can obtain includes the position information of the formation teammates The position information S of the obstacle obs =[x obs ,y obs , and the position information S of the target point G G =[x G ,y G are also essential. The state space can finally be expressed as

[0072] <Actor and reward unit>

[0073] The Actor is responsible for selecting the optimal action based on the input state space information using the action policy stored in it, and applying the selected action to the environment to achieve real-time interaction with the environment to update the state of the unmanned boat. Therefore, a corresponding Actor can be matched for each unmanned boat in the unmanned boat formation to achieve decentralized execution under the MAPPO architecture.

[0074] In some specific embodiments, an Actor includes at least one policy network and a preset dynamics model. Among them, the policy network stores the action policy π θ , which is used to determine the next action a t selected when the unmanned boat is in the state s t ; the controller uses the dynamics model described by, for example, equations (1) to (6) set in advance to realize the interaction between the unmanned boat and the environment, and update to obtain the state s t+1 of the unmanned boat.

[0075] The reward unit gives the corresponding reward r t based on the preset reward function according to each state s t in the operation process of the unmanned boat, so as to provide numerical feedback for the unmanned boat and guide its optimal decision-making.

[0076] Same as the embodiment shown in Figure 1 , in the training sample generation stage (or called the execution stage) of each round, the Actor corresponding to each unmanned boat continuously performs the process of "arriving at the current state -> obtaining the reward and determining the next action -> arriving at the new state" by executing its current policy, so as to continuously obtain (s t ,a t ,r t ,s t+1,d t ,p t As the voyage progresses, each Actor continuously stores the training samples it obtains into the experience pool until the training process ends.

[0077] Figure 5 The architecture of a specific policy network is illustrated. This network includes stacked layers, convolutional layers, pooling layers, flattening layers, multi-head self-attention layers, normalization layers, fully connected layers, and activation function layers. In the stacked layers, the state information of the unmanned surface vessels (USVs) is processed by stacking multiple consecutive frames of observation data. Stacking multiple frames provides more temporal information, helping the network capture the interactions and dynamic changes between USVs and improving the robustness of the policy. The processed data then passes through the multi-head self-attention layer, which dynamically adjusts the attention allocation among the multiple USV states, focusing on the region or USV with the greatest impact on the decision at the current moment, further improving the accuracy and flexibility of the decision-making process. The output of the self-attention mechanism is then processed again through fully connected layers to extract higher-level features.

[0078] MAPPO, as an efficient policy gradient method, typically selects an appropriate probability distribution to output the action policy in a continuous action space. Traditionally, the Gaussian distribution is often used to model continuous action policies, but in unmanned surface vessel (USV) path planning tasks, due to the strict limitations on the action space, the Beta distribution shows a more significant advantage.

[0079] The Beta distribution is a continuous probability distribution defined on the interval [0,1], and its probability density function (PDF) takes the following form:

[0080]

[0081] Where α>0 and β>0 are shape parameters, and B(α,β) is the Beta function. The shape of its distribution is determined by its parameters α and β, and it can flexibly represent unimodal, bimodal, and symmetrical or asymmetrical distribution forms.

[0082] In an embodiment of the invention, the output of the policy network is a two-dimensional vector, representing the shape parameters α and β in the Beta distribution, which determine the probability distribution of actions. During training, based on the generated Beta distribution, the algorithm obtains specific action decisions through sampling, and then updates the parameters of the policy network. During this process, the shape of the Beta distribution can be flexibly adjusted to better constrain the action selection of the unmanned surface vessel, ensuring that the output action is within a reasonable range. During testing, to simplify the decision-making process, the mean of the Beta distribution can be directly output as the action, ensuring that the output action is both stable in actual operation and can effectively cope with different environmental changes.

[0083] <Critical Path Unit>

[0084] After obtaining a preset number of training samples, or after the Actor completes a preset objective, the intensive training phase can begin, followed by comparison. Figure 1 and Figure 3 As can be seen, the training system provided by this invention, in addition to setting up an experience pool to store the training samples generated by the Actor executing the current action strategy, also adds a critical path unit. The critical path unit is used to determine the critical path combination from the paths generated by the Actor executing the current action strategy and several historical action strategies. During the intensive training phase, the optimization unit will jointly use the training samples in the training pool and the critical path sample combination determined by the critical path unit to optimize the policy network and the value network.

[0085] In embodiments of the present invention, a critical path refers to a complete path with significant training value obtained by an actor through executing the current action strategy or by executing several historical action strategies. The critical path sample combination is a combination of training samples corresponding to each time step on the complete path. That is, each critical path sample combination includes all the training samples experienced by an unmanned vessel during its navigation along the critical path.

[0086] For example, once the complete path τ traversed by an actor in a certain round is determined as the critical path, the corresponding critical path sample set is the collection of training samples obtained by that actor at each time step of that path:

[0087]

[0088] Where T is the total number of time steps traversed by path τ, and the state of the Actor at the end of the round is s. τ,T .

[0089] In some embodiments of the present invention, the critical path unit further includes a critical path determination module and a critical path storage module. The critical path determination module is used to evaluate the importance of each path generated by the current behavior strategy and several historical behavior strategies based on the overall path trajectory and determine the critical path. The critical path storage module is used to store the collection of training samples obtained at each time step of the critical path, i.e., the critical path sample combination.

[0090] Figure 6 The diagram illustrates the flowchart of the training process in some embodiments, showing the participation of the critical path determination module and the critical path storage module. Figure 6 As shown, during the strategy execution phase of each round, each Actor stores the collected training samples of the unmanned ships into the training pool, and also summarizes the training samples contained in each complete driving trajectory of each unmanned ship into path sample combinations. The critical path is determined by calculating the priority of each path sample combination, and the critical path (the sample combinations contained in these critical paths are called critical path sample combinations) is then stored in the critical path storage module according to the priority.

[0091] During the training phase, the combination of critical path samples stored in the critical path storage module is sampled by the optimization unit according to a certain probability, and then combined with the training samples sampled from the experience pool to optimize the policy network and value network based on the MAPPO algorithm.

[0092] In some preferred embodiments, the critical path determination module can select the critical path from the paths aggregated by each Actor through the following steps:

[0093] First, for any path τ, the dominance function at each time step along the path is calculated using the following formula.

[0094]

[0095] Where, δ τ,t to δ τ,T-1 Let λ be the TD error corresponding to time step t to T-1. τ As a weight parameter, preferably, to maintain consistency with the MAPPO algorithm framework, λ can be set to... τ =γλ, where γ is the discount factor and λ is the trade-off parameter.

[0096] Determine the dominance function at each time step (0 to T-1). Then, the trajectory priority P of path τ is determined by the following formula. τ P τ The importance of the overall path trajectory used to determine the path τ:

[0097] P τ =λ1·std(A τ,t )+λ2·TD τ +λ3·Entropy τ (11),

[0098] (11) In the formula, std() is the standard deviation function, TD τ Let be the mean TD error of path τ at each time step, i.e., mean(δ) τ,t Entropy τ Let mean(p) be the average probability of action at each time step along path τ. τ,t ), where λ1, λ2, and λ3 are weighting coefficients.

[0099] As can be seen from equation (11), in the embodiments of the present invention, for any path, the "scarcity of the overall trajectory" is evaluated from three aspects. The first term is used to measure the degree of advantage fluctuation of the unmanned ship in different states-actions during navigation. The larger the fluctuation value, the more special the area where the advantage function changes drastically during navigation. This drastic change is very likely caused by the unmanned ship encountering complex obstacles during navigation and the strategy network not yet forming a stable and effective response strategy. The second term is the average difference between the predicted value and the actual return, which represents the degree of "deviation" in the value network's "understanding" of the overall trajectory of the path. The larger the value, the more the value network has not yet learned to effectively evaluate the experience. The strategy entropy is the average information entropy of the action probability distribution in the path. High entropy means that the action strategy adopted by the strategy network is highly uncertain (action selection is scattered), indicating that the strategy to deal with some or all states in the path is still in the exploratory stage.

[0100] Therefore, the aforementioned priority evaluation criteria can assess the probability of each complete path encountering complex obstacles from three dimensions: action strategy, value assessment, and action probability. It can extract paths with a high probability of encountering complex obstacles. Furthermore, compared to existing algorithms that prioritize and sample single-step training samples, by extracting all training sample combinations experienced by the critical path, it can obtain a series of decision-action-state information of the unmanned vessel under the drive of the action strategy. Therefore, it can be trained based on the entire process information of the unmanned vessel discovering obstacles, making decisions, taking actions, and producing effects, thereby effectively improving the training efficiency of strategies adopted in complex obstacle sea conditions.

[0101] At the same time, it should also be pointed out that, such as Figure 6As shown, in the embodiments of the present invention, the training samples stored in the experience pool are generated by the Actor by executing the current action policy, while the critical path sample combination stored in the critical path storage module includes the current action policy and the sample combination corresponding to the critical path generated by several historical action policies. Obviously, the training sample storage mechanism of the experience pool is the conventional mechanism used by the on-policy algorithm, while the sample storage mechanism of the critical path storage module combines the samples generated by the current action policy and the historical action policies.

[0102] The reason for adopting the above mechanism is that not all Actors encounter complex obstacles during the sample generation process in each round, which may lead to an insufficient number of high-value critical paths, making it difficult to achieve efficient training on complex obstacle sea conditions. Therefore, in this invention, the on-policy and off-policy mechanisms are combined: the experience pool only stores each single-step sample generated by the current policy, while the critical path storage module stores critical path sample combinations in a first-in-first-out manner with a preset capacity threshold (for example, in some embodiments, the capacity threshold is 100, that is, the number of critical paths does not exceed 100). When the capacity reaches the upper limit, that is, while storing the critical path sample combination generated by the current action policy, the critical path sample combination generated by the earliest historical action policy is discarded. By "storing a limited number of complete training sample combinations of high-value paths and updating them in real time", the problem that the simple on-policy mechanism cannot obtain enough high-value samples can be solved, and the problem that invalid / low-value samples accumulate over time due to the direct use of the off-policy mechanism can be avoided.

[0103] Back Figure 6 During the intensive training phase, the optimization unit will simultaneously extract several training samples from the training pool and several key path sample combinations from the key path storage unit. Then, by combining the above samples, the MAPPO algorithm is used to optimize the value network and the policy network in each Actor. The MAPPO optimization strategy and the objective function used have been introduced above and will not be repeated here.

[0104] Training can be conducted in multiple batches, with training samples drawn from the training pool each time in a sequential or random manner. In some preferred embodiments, for any critical path τ... n The samples will be extracted by the optimization unit according to the following sampling probabilities for use in optimizing the policy network and the value network:

[0105]

[0106] Among them, S τ Let P be the set of critical paths.τ,n For the critical path τ n Trajectory priority, P τ,m For S τ The trajectory priority of the m-th critical path in the middle.

[0107] <Reward Mechanism>

[0108] As mentioned earlier, the reward unit rewards each Actor based on the actions taken and their current state. In the reinforcement learning framework, rewards are used to provide numerical feedback to the Actor, guiding them to optimize their decision-making. Reasonable reward settings can effectively improve learning efficiency, strategy quality, and task completion.

[0109] In some specific embodiments, the reward unit provides a reward r to each Actor. t This includes extrinsic rewards, which are the rewards obtained based on task performance during the execution of the action strategy. These rewards can be constructed by comprehensively considering factors such as target achievement, formation maintenance, collision avoidance, and path optimization. For example, they can consist of one or more of the following: formation rewards, collision penalties, task rewards, boundary violations, and step penalties. The expression for this is shown below:

[0110] r t =r ext,t =r t f +r t c +r t g +r t v +r t s (13),

[0111] Where, r ext,t Indicates extrinsic reward, r t f r t c r t g r t v and r t s These are formation rewards, collision penalties, mission rewards, boundary crossing penalties, and step penalties.

[0112] The settings for the above items can be customized based on the number of Actors, the task, the environment, and other specific factors. For example, for formation rewards, the maximum reward can be set when the distance between any two Actors is exactly equal to the preset formation distance. When the distance between Actors deviates from the formation distance, whether too large or too small, the reward value gradually decreases. Alternatively, to encourage Actors to choose the shortest path, a fixed negative reward r can be applied for each step. t s If the Actor chooses the option with more steps, r t s The accumulation of negative rewards will obviously be greater.

[0113] Observing equation (13), we can see that the external reward consists of two types of rewards. Continuous rewards provide feedback at each time step, guiding the Actor to gradually optimize the strategy, such as giving rewards based on changes in target distance or maintaining formation. Sparse rewards are only provided at a specific time step, such as successfully reaching the target point or touching an obstacle. If only sparse reward items such as collision penalties are set, it will lead to low exploration efficiency and make it difficult for the Actor to learn effective strategies efficiently. Therefore, it is necessary to increase the motivation of each Actor to explore the new environment based on equation (13).

[0114] Therefore, in some preferred embodiments of the present invention, the reward unit provides a reward r to each Actor. t It also includes intrinsic reward r int,t The source of intrinsic reward does not depend on the performance of the task, but rather comes from the Actor's own "curiosity" or "exploration motivation" to encourage the Actor to explore new areas.

[0115] Random Network Distillation (RND) is an exploration-reward mechanism applied in the field of reinforcement learning. Its core idea is to provide an agent with an "intrinsic reward" signal in an unsupervised manner, thereby encouraging the agent to more actively explore unseen or novel states in the environment, which helps to solve the exploration problem in sparse reward tasks.

[0116] The execution framework of RND consists of two main components:

[0117] 1) Fixed random target network f(s): A neural network whose parameters are randomly initialized and remain unchanged throughout the training process. Its input is the current environmental state s of the agent, and its output is the feature vector of s determined by the fixed network parameters of f(s).

[0118] 2) Prediction Network A trainable neural network with network parameters is used to fit the output of f(s) to the same input state s as closely as possible.

[0119] At each time step during task execution, the current state of the Actor is input into these two networks, and the error between their outputs is calculated:

[0120]

[0121] For prediction networks The training aims to minimize L(θ) as the objective function, such that... To achieve the best possible fit to f(s), L(θ) is also used as an intrinsic reward, i.e.:

[0122] r int,t (s)=Ll(θ).

[0123] The reason for using L(θ) to reward the Actor is that the larger the error, the less frequently the state is accessed during training. Therefore, this error is fed back to the Actor as an intrinsic reward signal to encourage it to explore these rare states. For states that have been accessed frequently, the prediction network can fit the output of the target network well, the error decreases, and the intrinsic reward also decreases.

[0124] The advantages of RNDs lie in their simplicity and effectiveness; however, when applied to the MAPPO framework, RNDs present the following problems:

[0125] First, using only a single target network and a prediction network can easily lead to a rapid decrease in error. Once a state is visited multiple times, the prediction network can quickly fit the target network, causing the intrinsic reward to disappear rapidly and the exploration behavior to stagnate. Second, if the prediction network learns a "speculative" optimal solution (such as predicting only a certain pattern) to suppress the overall error, the RND mechanism is at risk of failure. In addition, RND is sensitive to the structure of the environment. If the state space structure is complex, a single error scale cannot accurately characterize novelty.

[0126] Therefore, in a preferred embodiment of the present invention, an orthogonal coding random network distillation algorithm is provided based on improvements to existing RNDs. This algorithm enhances the novelty representation capability by introducing orthogonality between multiple target networks and making the prediction network approximate the combined expression of multiple target outputs during training, while also combating the problem of rapid overfitting of the prediction network.

[0127] Figure 7 The diagram illustrates the implementation flow of the orthogonal coded random network distillation algorithm in some specific embodiments. (Refer to...) Figure 7 The orthogonal coded random network distillation algorithm includes the following operations:

[0128] Operation 1: Initialize several target networks f1, f2... and their corresponding prediction networks. The parameters of each target network are randomly set and remain unchanged throughout the training process. The number of target networks and prediction networks can be reasonably set by comprehensively considering factors such as the size of the navigation environment, the size of the unmanned vessel formation, the complexity of obstacles, and the processing capacity of the system. For example, two, three, or more pairs of target networks and prediction networks can be set, but it must be ensured that the number of target networks and prediction networks does not exceed a certain upper limit (such as 10 pairs).

[0129] Operation 2: At each time step of the training sample generation phase, repeatedly execute the following steps:

[0130] The first step is to input the states generated by each Actor into all target networks and prediction networks, and determine the intrinsic reward for each Actor based on the deviation between the outputs of all target networks and their corresponding prediction networks.

[0131] Specifically, the expression for intrinsic reward is as follows:

[0132]

[0133] Where K represents the number of target networks and prediction networks.

[0134] The second step is to optimize each prediction network based on the deviation between the outputs of all target networks and their corresponding prediction networks, as well as the orthogonality penalty term. The orthogonality penalty term is obtained by summing the squares of the pairwise inner products of the outputs of each target network.

[0135] Specifically, the objective function is set as shown in the following formula:

[0136]

[0137] objective function L total It consists of two terms, where the first term represents the deviation between the output of the current prediction network and the output of the target network, and the second term L... orth An orthogonal penalty term is constructed by summing the squares of the pairwise inner products of the outputs of each target network to enhance the diversity among the state feature vectors. This term is then multiplied by a weighting coefficient λ. orth Then add it to the first term to get L total and with L total Optimizing the prediction network with the minimum objective enables multi-dimensional representation of state features, effectively overcoming the problems of excessively fast convergence and dynamic failure in exploring new environments that exist in single-objective networks and prediction networks.

[0138] The intrinsic reward r generated by the above orthogonal coded random network distillation algorithm is... int,tExternal rewards r ext,t Perform weighted summation (e.g., for r) int,t Assign weights λ int By doing so, you can obtain the total reward r for each Actor. t =r ext,t +λ int r int,t Using this to reward each Actor can make each Actor's action strategy both consider the task objective and generate sufficient motivation to explore new environments.

[0139] After training the policy network using the MAPPO-based unmanned vessel formation path planning system provided by this invention, it can be ported to the unmanned vessel formation control system for automatic path planning of each unmanned vessel in the formation.

[0140] In some specific embodiments, the unmanned vessel formation control system includes a sampling unit, a communication unit, and a path planning unit.

[0141] The sampling unit is used to collect the status of each unmanned vessel in the unmanned vessel formation and its environment. In some specific embodiments, it can be composed of various information collection devices equipped on each unmanned vessel. For example, a GPS positioning system can obtain high-precision location information in real time; attitude sensors, gyroscopes and inertial navigation systems (INS) can work together to accurately measure heading angle and speed; radar and forward-looking sonar can be used to detect the position and speed of obstacles.

[0142] The communication unit is used to communicate between the various unmanned vessels to ensure the coordination and information sharing of the formation. Specifically, depending on the target mission and navigation environment, each unmanned vessel can be equipped with wireless communication modules such as WIFI, 4G / 5G, satellite or radio to realize remote data transmission and reception.

[0143] The control unit includes a strategy network corresponding to each unmanned surface vessel (USV), used to control the USVs to navigate in formation. The strategy network is trained using the aforementioned MAPPO-based USV formation path planning training system. Specifically, the trained strategy network can be installed in the host computer of each USV. Each strategy network receives the state and local environment information of its own USV through a sampling unit, and simultaneously receives the state and local environment information of other USVs through a communication unit. Based on this input data, it executes the trained action strategy to output control variables for the USVs, which are then executed by the USVs' propulsion systems, thereby achieving effective control of the USV formation navigation.

[0144] <Specific Implementation>

[0145] This embodiment verifies the training effect of the MAPPO-based unmanned vessel formation path planning training system through simulation experiments. The experimental environment is as follows: Figure 1 As shown, in this simulation experiment, three unmanned surface vessels (USVs) were selected to form a formation. Each USV had to pass through a restricted channel in sequence and finally safely reach the designated destination. During the experiment, the USVs had to avoid all obstacles while ensuring their own navigation trajectory remained stable, avoiding collisions, and maintaining an equilateral triangle structure at all times to ensure formation stability and collaborative operation capabilities in complex environments.

[0146] In this embodiment, the training parameters are set with a discount factor γ = 0.98, a maximum step size per round of tasks of 1000, a batch sampling size of 96, and a training process of 2000 rounds. Each training round has 1000 maximum time steps. At the beginning of each round, the initial positions of the unmanned vessel and obstacles are randomly reset. Figure 6 The training process is shown, and the reward function is constructed using a weighted sum of extrinsic and intrinsic rewards.

[0147] Table 1 lists the specific simulation parameters for this embodiment.

[0148] Table 1 Simulation parameters for specific embodiments

[0149]

[0150] Tables 2 and 3 list the structures of the policy network and the value network, respectively.

[0151] Table 2 Strategy Network Structure

[0152] Neural network layers enter Output effect Stacked layers (3,60,60) (9,60,60) 3-frame dynamic stacking Convolutional layer 1 (9,60,60) (32,60,60) Local feature extraction Pooling layer 1 (32,60,60) (32,30,30) Dimensional reduction Convolutional layer 2 (32,30,30) (64,30,30) Deep feature extraction Pooling layer 2 (64,30,30) (64,15,15) Dimensional reduction Flattening layer (64,15,15) (64,225) Transform into a sequence Attention layer (64,225) (64,225) Extract key information Normalization layer (64,225) (64,225) Normalized data distribution Fully connected layer (64,225) 256 Advanced Feature Extraction Alpha layer 256 Action dimension Predict the Beta distribution parameter α Beta layer 256 Action dimension Predict the Beta distribution parameter β Softplus layer Action dimension Action dimension Ensure α,β > 0

[0153] Table 3 Value Network Structure

[0154]

[0155]

[0156] Table 4 lists the structures of the target network and the prediction network used to generate intrinsic rewards.

[0157] Table 4 Target Network and Predicted Network Structure

[0158] Neural network layers enter Output effect Stacked layers (3,60,60) (9,60,60) 3-frame dynamic stacking Convolutional layer 1 (9,60,60) (32,60,60) Local feature extraction Pooling layer 1 (32,60,60) (32,30,30) Dimensional reduction Convolutional layer 2 (32,30,30) (64,30,30) Deep feature extraction Pooling layer 2 (64,30,30) (64,15,15) Dimensional reduction Flattening layer (64,15,15) (64,225) Transform into sequence input Fully connected layer (64,225) 256 Advanced Feature Extraction Fully connected layer 256 128 Calculate intrinsic reward

[0159] Figure 8 , Figure 9 The average reward curve and average TD error curve during the training process using the training system of the present invention are shown respectively. For comparison, the corresponding results using the conventional MAPPO algorithm are also shown in the two figures.

[0160] pass Figure 8 , Figure 9 As can be seen, the training system of the present invention trains the unmanned vessel formation path planning capability, which not only significantly improves the final reward value, but also has a faster and more stable curve convergence speed.

[0161] The specific embodiments of the present invention have been described in detail above. For those skilled in the art, several improvements and modifications can be made to the present invention without departing from the principle of the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A training system for unmanned surface vessel (USV) formation path planning based on MAPPO, characterized in that, include: The simulation environment includes a simulated water area model and the irregular obstacle model contained therein; Multiple Actors, corresponding to each unmanned vessel in the unmanned vessel formation, determine the actions to be taken based on the action policies and current states stored in their policy network, and update the next state based on the actions taken. The reward unit rewards each Actor based on their actions and current state. The experience pool is used to store the training samples generated by each Actor when executing the current action strategy; A value network is used to evaluate the value of the training samples; Critical path unit, used to determine the critical path sample combination from the paths generated by the current action strategy and several historical action strategies; The optimization unit, based on the MAPPO optimization algorithm, jointly uses the training samples and critical path samples to optimize the policy network and value network. The critical path unit includes: The critical path determination module is used to assess the importance of each path generated by the current behavior strategy and several historical behavior strategies based on the overall path trajectory and determine the critical path. A critical path storage module is used to store the critical path sample combination; For any path The importance of its overall path trajectory is determined by the following formula: , in, For path At each time step The advantage function, For path The total number of time steps elapsed. The trajectory priority for this path. to For time step to The corresponding TD error, For weight parameters, It is a function of standard deviation. For path The mean of TD error at each time step, For path The mean of the action probability at each time step. , , These are the weighting coefficients.

2. The MAPPO-based unmanned vessel formation path planning training system according to claim 1, characterized in that, The training samples include the state, actions, immediate rewards, end flags, action probabilities, and updated state after each action is performed for each unmanned vessel. The critical path sample set includes all training samples experienced by the unmanned vessel during its navigation along the critical path.

3. The unmanned vessel formation path planning training system based on MAPPO according to claim 1, characterized in that, Any critical path The network is selected by the optimization unit according to the sampling probability shown in the following formula for optimizing the policy network and the value network: , in, The set of critical paths. Critical path Trajectory priority, for The Middle Trajectory priority of critical paths.

4. The unmanned vessel formation path planning training system based on MAPPO according to claim 1, characterized in that, At any given time, the number of paths identified as critical paths does not exceed a preset capacity limit, and the critical path sample combination stored in the critical path storage module is updated in a first-in-first-out manner.

5. The unmanned vessel formation path planning training system based on MAPPO according to claim 1, characterized in that, The rewards provided by the reward unit to each Actor include external rewards, which consist of one or more of the following: formation rewards, collision penalties, task rewards, boundary crossing penalties, and step penalties.

6. The MAPPO-based unmanned vessel formation path planning training system according to claim 5, characterized in that, The reward provided by the reward unit to each Actor also includes an intrinsic reward, which is generated based on the orthogonal coding random network distillation algorithm.

7. The MAPPO-based unmanned surface vessel formation path planning training system according to claim 6, characterized in that, The orthogonal coded random network distillation algorithm includes: Initialize several pairs of target networks and prediction networks, wherein the parameters of the target networks are randomly set and remain unchanged during training; And, at each time step of the training sample generation phase, the following steps are executed iteratively: The states generated by each Actor are input into all target networks and prediction networks, and the intrinsic reward for each Actor is determined based on the deviation between the outputs of all target networks and their corresponding prediction networks. Each prediction network is optimized based on the deviation between the outputs of all target networks and their corresponding prediction networks, as well as an orthogonal penalty term. The orthogonal penalty term is obtained by summing the squares of the pairwise inner products of the outputs of each target network.

Citation Information

Patent Citations

  • Multi-AGV path planning model training method based on MAPPO algorithm and path planning method

    CN117707063A