A path planning method for cooperative cruising of unmanned ships
By introducing a weighted objective function and a dual-delay network into the deep deterministic policy gradient algorithm, the overestimation problem of the value function in the cooperative cruise path planning of unmanned vessels is solved, achieving more efficient and higher-quality path planning and improving the autonomy and self-cooperation capabilities of unmanned vessels.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINESE PEOPLES LIBERATION ARMY ARMY BORDER & COASTAL DEFENSE ACAD
- Filing Date
- 2022-10-11
- Publication Date
- 2026-07-21
AI Technical Summary
Existing unmanned surface vessel (USV) cooperative cruise path planning suffers from overestimation of value functions, which makes it impossible to perform path planning efficiently and reasonably in complex mission environments, thus affecting the autonomy and self-cooperation capabilities of USVs.
By employing the Deep Deterministic Policy Gradient (DDPG) algorithm combined with a weighted objective function and a dual-delay network, the overestimation problem of the value function is mitigated and the cooperative cruise strategy is optimized by training an unmanned vessel cooperative cruise path planning system model.
It improves the efficiency and quality of collaborative cruising of unmanned vessels in complex environments, and enhances the accuracy of path planning and the fault tolerance of the system.
Smart Images

Figure CN115729236B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of military information technology, and in particular to a path planning method for cooperative cruising of unmanned vessels. Background Technology
[0002] Unmanned surface vessels (USVs) are a crucial component of future intelligent and unmanned warfare, playing a vital role in future naval battles. As more USVs participate in increasingly complex military missions, the demands for their autonomy are also rising. Incorporating artificial intelligence (AI) technology into the autonomous systems of USVs enables them to perform more complex tasks, increasing their autonomy and self-coordination capabilities.
[0003] While existing deep reinforcement learning methods used in unmanned surface vessels (USVs) cooperative cruise path planning have effectively addressed the challenges of cooperative cruise operations in high-latitude, highly dynamic battlefield environments, and are well-suited to the characteristics and patterns of unmanned combat platform path planning tasks, they suffer from overestimation of the value function. This overestimation directly hinders efficient and rational path planning in complex mission environments. Therefore, how to rationally utilize limited computational resources, mitigate the overestimation problem, improve the algorithm's final performance, and enable efficient path planning to enhance the efficiency and quality of cooperative cruise operations for USVs in complex environments has become an urgent problem to be solved.
[0004] Therefore, it is necessary to improve one or more of the problems existing in the above-mentioned related technical solutions in order to improve the global estimation accuracy and fault tolerance performance of the navigation system.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this disclosure is to provide a path planning method for cooperative cruising of unmanned vessels, which can alleviate the problem of overestimation of value functions and improve the efficiency and quality of cooperative cruising of unmanned vessels.
[0007] This disclosure provides a path planning method for cooperative cruising of unmanned vessels, the method comprising the following steps: A path planning system model for cooperative cruising of unmanned vessels is constructed. The path planning system model for cooperative cruising of unmanned vessels includes a set of states of multiple unmanned vessels at any time, a set of actions of multiple unmanned vessels at any time, and environmental feedback obtained by multiple unmanned vessels. The path planning system model for the cooperative cruise of the unmanned vessels is trained using a deep deterministic policy gradient algorithm. During the training process, a weighted objective function and a dual-delay network are introduced, and the training results are updated to obtain multiple cooperative cruise strategies for the unmanned vessels. The obtained cooperative cruise strategy of multiple unmanned vessels was deployed in a real mission environment.
[0008] In an exemplary embodiment of this disclosure, the step of constructing a path planning system model for cooperative cruising of unmanned vessels includes: multiple unmanned vessels at any given time... Set of environmental states Multiple unmanned vessels are grouped according to the environmental conditions. The set of actions consisting of the corresponding actions performed in each of the states. ; and environmental feedback obtained from multiple unmanned vessels. .
[0009] In an exemplary embodiment of this disclosure, the set of environmental states The formulas include: in, and The x and y coordinates represent the state of each unmanned surface vessel, respectively. , ; n This indicates the number of the unmanned vessel.
[0010] In an exemplary embodiment of this disclosure, the action set The formulas include: in, This indicates that the unmanned vessel is maneuvering in a due east direction; This indicates that the unmanned vessel is maneuvering in a southeast direction; This indicates that the unmanned vessel is maneuvering in a due south direction; This indicates that the unmanned vessel is maneuvering in a southwest direction. This indicates that the unmanned vessel is maneuvering in a due west direction; This indicates that the unmanned vessel is maneuvering in a northwest direction; This indicates that the unmanned vessel is maneuvering in a due north direction; The unmanned vessel maneuvers in a northeast direction; the range of values for the maneuver includes... 0 indicates no action.
[0011] In an exemplary embodiment of this disclosure, the environmental feedback This includes rewards for unmanned vessels reaching any location, penalties for unmanned vessels encountering obstacles, penalties for unmanned vessels touching boundaries, and penalties for unmanned vessels colliding with each other.
[0012] In an exemplary embodiment of this disclosure, the unmanned vessel cooperative cruise path planning system model is trained using a deep deterministic policy gradient algorithm; and a weighted objective function and a dual-delay network are introduced to update the training results to obtain multiple cooperative cruise strategies for the unmanned vessels. In this step, the training process of the unmanned vessel cooperative cruise path planning system model includes multiple parameters, which include: Multiple evaluation networks, the multiple evaluation networks including , , ; An action network ; Multiple target evaluation networks, the multiple target evaluation networks including , , ; A target action network ; Memory storage unit The memory storage unit The capacity is The discount factor is... Batch processing capacity is The cutoff coefficient is The weighted stability coefficient is ; The parameter copy interval is The parameter replication factor is The noise is The number of loops is .
[0013] In an exemplary embodiment of this disclosure, the path planning system model for cooperative cruising of the unmanned vessels is trained using a deep deterministic policy gradient algorithm; during the training process, a weighted objective function and a dual-delay network are introduced, and the training results are updated to obtain the cooperative cruising strategies of multiple unmanned vessels. Initialize the multiple parameters; For action Select to initialize the noise The noise obey Distribution, where ~ indicates obedience; clip indicates clipping. Represents a normal distribution; Initialization time Accept the initial state of the unmanned vessel ; According to the action network and the noise The action selected by the unmanned vessel for ; Perform the action Observe the new environmental conditions The environmental feedback is obtained. ; To the memory storage unit The environmental feedback is stored in the middle. Rewards Group ; From the memory storage unit Random sampling is performed to obtain sample group data. , ; Select the action described at the next moment. ; make ; By minimizing the loss method The evaluation networks are updated accordingly; among them, Update the above Evaluation of the network: ; Update the above Evaluation of the network: ; Update the above Evaluation of the network: ; in, This indicates an assignment operation; At the time t The parameter copy interval cannot be divided evenly. When, if the stated t The time value is equal to the cycle time. T According to the action network The cooperative cruise strategy for multiple unmanned vessels is obtained, and the formula of the cooperative cruise strategy includes: ,in, Indicates the environmental state. Indicates mobile network Network parameters, Indicates the final output strategy; A Indicates action The set, Represents an action network, abbreviated as .
[0014] In an exemplary embodiment of this disclosure, when the time... t Able to divide the parameter copy interval At this time, the following steps also need to be performed: The algorithmic formula for the gradient of the depth-deterministic policy is as follows: Update the action network ; Indicates the batch processing capacity. To find the derivative, Indicates an action, Indicates the evaluation network, Represents the set of environmental states The action network; Represents the gradient of a deep deterministic policy; Update the target evaluation network and the target action network respectively. The parameters, where, ; ; This indicates an assignment operation; If the above t The time value is equal to the cycle time. T According to the action network The cooperative cruise strategy of multiple unmanned vessels is obtained. ; In an exemplary embodiment of this disclosure, when the t The time interval can be divided by the parameter copy interval. At this time, the following steps also need to be performed: The algorithmic formula for the gradient of the depth-deterministic policy is as follows: Update the action network ; Update the target evaluation network and the target action network respectively. The parameters, where, ; ; If the above t The time value is not equal to the loop time. T Then it is necessary to From that moment on, return to the steps according to the action network. and the noise The action selected by the unmanned vessel The Continue the training process.
[0015] The technical solution provided in this disclosure may include the following beneficial effects: In this embodiment, a path planning method for cooperative cruising of unmanned vessels is proposed. This method introduces a weighted objective function and a dual-delay network during the training process of the deep deterministic policy gradient algorithm, thereby alleviating the overestimation problem of the value function during the training process of the deep deterministic policy gradient algorithm and avoiding the underestimation problem, thus improving the final performance of the algorithm and enhancing the efficiency and quality of the cooperative cruising strategy of unmanned vessels. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0017] Figure 1 A schematic diagram illustrating the steps of a path planning method for cooperative cruising of unmanned vessels in an exemplary embodiment of this disclosure; Figure 2 A flowchart illustrating a path planning method for cooperative cruising of unmanned vessels in an exemplary embodiment of this disclosure; Figure 3 This diagram illustrates an exemplary embodiment of the present disclosure that incorporates a weighted objective function and a dual-delay network during the training process of a deep deterministic policy gradient algorithm. Detailed Implementation
[0018] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0019] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0020] This example implementation provides a path planning method for cooperative cruising of unmanned vessels, referring to... Figure 1 As shown, the method may include the following steps: Step S101: Construct a path planning system model for cooperative cruising of unmanned vessels; wherein, the path planning system model for cooperative cruising of unmanned vessels includes the state set of multiple unmanned vessels at any time, the action set of multiple unmanned vessels at any time, and the environmental feedback obtained by multiple unmanned vessels. Step S102: Train the path planning system model for cooperative cruise of unmanned vessels using the deep deterministic policy gradient algorithm; introduce a weighted objective function and a dual-delay network during the training process, and update the training results to obtain cooperative cruise strategies for multiple unmanned vessels. Step S103: Deploy the obtained cooperative cruise strategy of multiple unmanned vessels in a real mission environment.
[0021] This embodiment proposes a path planning method for cooperative cruising of unmanned vessels. This method introduces a weighted objective function and a dual-delay network during the training process of the deep deterministic policy gradient algorithm, thereby alleviating the overestimation problem of the value function during the training process of the deep deterministic policy gradient algorithm and avoiding the underestimation problem, thus improving the final performance of the algorithm and enhancing the efficiency and quality of the cooperative cruising strategy of unmanned vessels.
[0022] The following is a more detailed explanation of each step of the calculation method in this example implementation.
[0023] The Deep Deterministic Policy Gradient (DDPG) algorithm, when solving the cooperative cruise path planning problem for unmanned surface vessels (USVs), leverages the continuous interaction between the USV and its environment, adjusting its behavioral strategy based on reward and penalty information from environmental feedback, ultimately learning the optimal strategy. However, this algorithm suffers from overestimation during training, reducing its learning efficiency and final performance. This disclosure addresses this issue by introducing a weighted objective function and a dual-delay network during the algorithm's training process. This helps mitigate the overestimation problem while avoiding underestimation, thereby optimizing the cooperative cruise strategy generated by the algorithm. Specifically: like Figure 2 As shown, in step S101, a path planning system model for unmanned vessel cooperative cruising is constructed.
[0024] In this step, the path planning method for cooperative cruising of unmanned surface vessels proposed in this disclosure is described using the example of how to enable as many unmanned surface vessels as possible to reach the target location in a complex battlefield environment.
[0025] Unmanned surface vessels are responsible for patrol missions within the controlled area. According to the patrol and reconnaissance plan, the unmanned surface vessels gather and depart from the pre-set sea area. In the complex sea environment, there are several unknown islands and reefs and obstacles. Multiple unmanned surface vessels need to patrol and coordinate autonomously in the complex environment to reach as many target locations as possible within a limited time and complete the patrol and reconnaissance mission within the controlled sea area.
[0026] In this example implementation, the unmanned surface vessels are divided into three groups, each representing an operational unit. The three groups are pre-deployed in the northwest corner of a designated area; the total patrol duration is preset to 5000 time steps; and 100 target locations are preset to simulate a real maritime battlefield environment. The collaborative patrol mission terminates when the total patrol duration is reached, or when the unmanned surface vessels reach and reconnoiter all preset target locations.
[0027] Thus, at any time Set of environmental states for: ,in, and These represent the x and y coordinates of each unmanned surface vessel's position in the environment. Although they are only used to represent the position of the unmanned surface vessels, they also represent the environmental state set. This can include various environmental factors, such as the current direction and speed of the water flow at the current location of the unmanned surface vessel (USV), visibility, and wind speed. Subscripts 1, 2, and 3 represent the numbers of the three groups of unmanned surface vessels. The value ranges for each dimension of the state variables are as follows: , .
[0028] Multiple unmanned vessels assemble based on environmental conditions The set of actions consisting of the corresponding actions performed in each of the states. for, ,in, This indicates that the unmanned vessel is maneuvering in a due east direction; This indicates that the unmanned vessel is maneuvering in a southeast direction; This indicates that the unmanned vessel is maneuvering in a due south direction; This indicates that the unmanned vessel is maneuvering in a southwest direction. This indicates that the unmanned vessel is maneuvering in a due west direction; This indicates that the unmanned vessel is maneuvering in a northwest direction; This indicates that the unmanned vessel is maneuvering in a due north direction; The unmanned vessel maneuvers in a northeast direction; the range of values for the maneuver includes... The entire specified area is treated as a coordinate system, where 0 represents no action and 50 represents a total of 50 action zones.
[0029] Environmental feedback received include: Rewards for unmanned surface vessels reaching any target location: = +100; Penalties for unmanned surface vessels encountering obstacles: = -50; Penalties for unmanned surface vessels crossing the boundary: = -1000; Penalties for collisions between unmanned surface vessels: = -100.
[0030] like Figure 3 As shown, in step S102, the path planning system model for cooperative cruising of unmanned vessels is trained using the deep deterministic policy gradient algorithm; a weighted objective function and a dual-delay network are introduced during the training process, and the training results are updated to obtain cooperative cruising strategies for multiple unmanned vessels.
[0031] In this step, the training process of the unmanned vessel cooperative cruise path planning system model includes multiple parameters: Three evaluation networks, including , , ; can be abbreviated as , , ; An action network ; can be abbreviated as ; Three target evaluation networks, and multiple target evaluation networks including , , ; A target action network ; These networks all employ fully connected neural networks with two hidden layers (64 neurons).
[0032] Memory storage unit The memory storage unit The capacity is = 10000; Discount factor is =0.99, batch processing capacity is =32, cutoff coefficient is =0.5, the weight stability coefficient is =1; parameter copy interval is =5, parameter replication factor is =0.5, noise is The cycle time is =20000.
[0033] The process of a coordinated cruise strategy involving multiple unmanned vessels includes: Initialize multiple parameters; For action Select initialization noise ,noise obey Distribution, where ~ indicates obedience; clip indicates clipping. Represents a normal distribution; Initialization time Accept the initial state of the unmanned vessel ; According to the Action Network and noise The action selected by the unmanned vessel , ; Execute action Observe the new environmental conditions Receive environmental feedback ; to memory storage unit Data stored in From memory storage unit Random sampling is performed to obtain sample group data. , ; Choose the next action ; make ; This section introduces a weighted objective function and a dual-delay network, employing a loss minimization method. Update multiple evaluation networks separately , and ;in, Update the above Evaluation of the network: ; Update the above Evaluation of the network: ; Update the above Evaluation of the network: ; in, This indicates an assignment operation, which means to assign a value to another value. The size of the right-hand part of the value is assigned The left side.
[0034] like t The size of the time interval is equal to the loop time of 20000, according to the action network. The cooperative cruise strategy for multiple unmanned vessels is obtained, and the formula for the cooperative cruise strategy is: ;in, Indicates the environmental state. Indicates mobile network Network parameters, Indicates the final output strategy; A Indicates action The set, Represents the set of environmental states. Action network, abbreviated as .
[0035] when t If the time interval is divisible by the parameter 5, then the following steps are also required: The algorithm formula using the gradient of a deep deterministic policy is as follows: Update mobile network ; Indicates the batch processing capacity. To find the derivative, Indicates an action, Indicates the evaluation network, Represents the set of environmental states The action network; Represents the gradient of a deep deterministic policy; Update multiple target evaluation networks and the target action network respectively. The parameters, where, ; ; This indicates assignment; the assignment operation here means to assign a value to another value. The size of the right-hand part of the value is assigned The left side.
[0036] like t If the time interval is not equal to the loop time of 20000, then it is necessary to... From that moment on, return to the steps according to the action network. and the noise The action selected by the unmanned vessel The The training process continues until the value at a certain moment equals the cycle time of 20,000, thus obtaining the cooperative cruise strategy for multiple unmanned vessels. .
[0037] It should be noted that although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that these steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps. Furthermore, it is readily understood that these steps may be executed synchronously or asynchronously, for example, in multiple modules / processes / threads.
[0038] It should be noted that although several units of the system for executing actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units. Some or all of the units can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without any inventive effort.
[0039] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
Claims
1. A path planning method for cooperative cruising of unmanned vessels, characterized in that, Includes the following steps: A path planning system model for cooperative cruising of unmanned vessels is constructed. The path planning system model for cooperative cruising of unmanned vessels includes a set of states of multiple unmanned vessels at any time, a set of actions of multiple unmanned vessels at any time, and environmental feedback obtained by multiple unmanned vessels. The path planning system model for the cooperative cruise of the unmanned vessel is trained using a deep deterministic policy gradient algorithm; A weighted objective function and a dual-delay network are introduced during the training process, and the training results are updated to obtain a cooperative cruise strategy for multiple unmanned vessels. The training process of the unmanned vessel cooperative cruise path planning system model includes multiple evaluation networks, which include... , , ; make ; By minimizing the loss method The evaluation networks are updated accordingly; among them, Update the above Evaluation of the network: ; Update the above Evaluation of the network: ; Update the above Evaluation of the network: ; in, This indicates that multiple unmanned vessels are at any given time. The set of environmental states, This represents the environmental feedback received by multiple unmanned vessels. This represents a set of multiple unmanned vessels based on the environmental state. A set of actions consisting of the corresponding actions performed in each of the states. Indicates the discount factor. Represents the weighted stability coefficient. Represents the target evaluation network, ← indicates the assignment operation, i = 1, 2, ..., C; Indicates the batch processing capacity; The obtained cooperative cruise strategy of multiple unmanned vessels was deployed in a real mission environment.
2. The path planning method for cooperative cruising of unmanned vessels according to claim 1, characterized in that, The set of environmental states The formulas include: in, and The x and y coordinates represent the state of each unmanned surface vessel, respectively. , ; n This indicates the number of the unmanned vessel.
3. The path planning method for cooperative cruising of unmanned vessels according to claim 2, characterized in that, The set of actions The formulas include: in, This indicates that the unmanned vessel is maneuvering in a due east direction; This indicates that the unmanned vessel is maneuvering in a southeast direction; This indicates that the unmanned vessel is maneuvering in a due south direction; This indicates that the unmanned vessel is maneuvering in a southwest direction. This indicates that the unmanned vessel is maneuvering in a due west direction; This indicates that the unmanned vessel is maneuvering in a northwest direction; This indicates that the unmanned vessel is maneuvering in a due north direction; The unmanned vessel maneuvers in a northeast direction; the range of values for the maneuver includes... 0 indicates no action.
4. The path planning method for cooperative cruising of unmanned vessels according to claim 3, characterized in that, The environmental feedback This includes rewards for unmanned vessels reaching any location, penalties for unmanned vessels encountering obstacles, penalties for unmanned vessels touching boundaries, and penalties for unmanned vessels colliding with each other.
5. The path planning method for cooperative cruising of unmanned vessels according to claim 1, characterized in that, In the step of training the path planning system model for the cooperative cruise of unmanned vessels using a deep deterministic policy gradient algorithm, and updating the training results by introducing a weighted objective function and a dual-delay network to obtain multiple cooperative cruise strategies for the unmanned vessels, the training process of the path planning system model for the cooperative cruise of unmanned vessels includes multiple parameters, including: Multiple evaluation networks, the multiple evaluation networks including , , ; An action network ; Multiple target evaluation networks, the multiple target evaluation networks including , , ; A target action network ; Memory storage unit The memory storage unit The capacity is The cutoff coefficient is ; The parameter copy interval is The parameter replication factor is The noise is The cycle time is T.
6. The path planning method for cooperative cruising of unmanned vessels according to claim 5, characterized in that, The path planning system model for the cooperative cruise of the unmanned vessels is trained using a deep deterministic policy gradient algorithm. During training, a weighted objective function and a dual-delay network are introduced, and the training results are updated to obtain the cooperative cruise strategies for multiple unmanned vessels. The process includes: Initialize the multiple parameters; For action Select to initialize the noise The noise obey Distribution, where ~ indicates obedience; clip indicates clipping. Represents a normal distribution; Initialization time Accept the initial state of the unmanned vessel ; According to the action network and the noise The action selected by the unmanned vessel for ; Perform the action Observe the new environmental conditions The environmental feedback is obtained. ; To the memory storage unit The environmental feedback is stored in the middle. Rewards Group ; From the memory storage unit Random sampling is performed to obtain sample group data. , ; Select the action described at the next moment. ; At the time t The parameter copy interval cannot be divided evenly. When, if the stated t The time value is equal to the cycle time. T According to the action network The cooperative cruise strategy for multiple unmanned vessels is obtained, and the formula of the cooperative cruise strategy includes: ,in, Indicates the environmental state. Indicates mobile network Network parameters, Indicates the final output strategy; A Indicates action The set, Represents an action network, abbreviated as .
7. The path planning method for cooperative cruising of unmanned vessels according to claim 6, characterized in that, At the time t Able to divide the parameter copy interval At this time, the following steps also need to be performed: The algorithmic formula for the gradient of the depth-deterministic policy is as follows: Update the action network ; Indicates the batch processing capacity. To find the derivative, Indicates an action, Indicates the evaluation network, Represents the set of environmental states The action network; Represents the gradient of a deep deterministic policy; Update the target evaluation network and the target action network respectively. The parameters, where, ; ; This indicates an assignment operation; If the above t The time value is equal to the cycle time. T According to the action network The cooperative cruise strategy of multiple unmanned vessels is obtained. .
8. The path planning method for cooperative cruising of unmanned vessels according to claim 7, characterized in that, When the t The time interval can be divided by the parameter copy interval. At this time, the following steps also need to be performed: The algorithmic formula for the gradient of the depth-deterministic policy is as follows: Update the action network ; Update the target evaluation network and the target action network respectively. The parameters, where, ; ; If the above t The time value is not equal to the loop time. T Then it is necessary to From that moment on, return to the steps according to the action network. and the noise The action selected by the unmanned vessel The Continue the training process.
9. The path planning method for cooperative cruising of unmanned vessels according to claim 6, characterized in that, The weight stability coefficient Calculation formula include: 。