A multi-ship collision avoidance training method and system based on centralized critics

CN122816253APending Publication Date: 2026-09-25汉江国家实验室
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610889744.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

现有机器人协同方法假设智能体动力学相似,无法直接处理这种异构性约束

Benefits of technology

1、自动化生产:直接生成安全、协同的避碰动作序列,决策过程自动化,并能处理人类驾驶员难以短时间计算的复杂多船局面;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816253A_ABST
    Figure CN122816253A_ABST
Patent Text Reader

Abstract

The application provides a multi-ship collision avoidance training method and system based on a centralized critic, comprising: through a quantifiable reward function deeply fused with COLREGs rules, responsibility division, avoidance direction, straight sailing ship obligation and the like are converted into calculable reward and punishment signals, so that the centralized critic not only evaluates "safety", but also evaluates "compliance" and "coordination"; through a centralized-distributed architecture adapted to heterogeneous ships, ship type codes and maneuverability parameters are explicitly input into the centralized critic network, so that the centralized critic network can output personalized strategy gradients for different ship types; during execution, each ship can generate compliant and complementary collision avoidance actions only based on its own local observation (including its own maneuvering characteristics); and through a generalized actor network: based on massive AIS data training, a shared actor network architecture can dynamically adapt to different types of ships and new ships not seen during training, realizing zero-shot transfer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent ship scheduling technology, and in particular to a multi-ship collision avoidance training method and system based on centralized critics. Background Technology

[0002] In distributed multi-agent reinforcement learning, each agent (ship) learns only from its own perspective, and the environment is non-stationary for each agent (because other agents are also learning), resulting in extremely unstable training and difficulty in learning cooperative strategies.

[0003] Existing reinforcement learning-based methods are mostly distributed architectures, where each ship only performs local optimization from its own perspective. Their decisions may conflict because they cannot predict the reactions of other ships, and they are prone to deadlock or other risks in complex encounter situations.

[0004] In recent years, the centralized training, distributed execution (CTDE) framework has been proposed and applied in the field of robot cooperative control. By having a centralized critic evaluate the global joint actions during the training phase, it alleviates the problem of environmental non-stationarity to some extent. However, none of these methods consider two special constraints in the field of ship collision avoidance: (1) Mandatory regulatory constraints: The International Regulations for Preventing Collisions at Sea (COLREGs) have clear legal provisions on the behavior of ships in different encounter situations (head-on, cross, overtaking) (who gives way, who goes straight, avoidance direction, etc.). The reward function of the general CTDE framework usually only includes collision risk penalties, which cannot guide the agent to learn compliant collision avoidance behavior and is prone to producing dangerous actions that are "physically safe but violate navigation regulations", which is unacceptable in real shipping.

[0005] (2) Ship dynamics and heterogeneity: Ships have characteristics such as large inertia, underactuation, and turning delay, and the maneuverability of different ship types (large merchant ships, small fishing boats, and high-speed boats) varies greatly. Their collision avoidance strategies must adapt to these heterogeneous characteristics. Existing robot cooperative methods assume that the dynamics of intelligent agents are similar, which cannot directly handle this heterogeneity constraint.

[0006] Therefore, how to deeply integrate regulatory logic with heterogeneous ship dynamics into the CTDE training framework is a core technical challenge that urgently needs to be solved. Summary of the Invention

[0007] This invention provides a multi-ship collision avoidance training method and system based on centralized critics to address the deficiencies in existing technologies.

[0008] In a first aspect, the present invention provides a multi-ship collision avoidance training method based on centralized critics, comprising: For each ship, an actor network and a centralized critic network are constructed. The actor network is used to output collision avoidance actions based on the ship's local observation information, and the centralized critic network is used to evaluate the expected cumulative reward based on the global state information and joint action information of all ships during the training phase. A multi-objective reward function is determined to embed the responsibility quantification of collision avoidance rules. The multi-objective reward function includes a safety reward component and a rule responsibility reward component. The rule responsibility reward component is based on the judgment of the encounter situation of ships and the role allocation of the give-way vessel and the straight-going vessel. Differentiated rewards and penalties are given for whether the collision avoidance behavior of different roles complies with the collision avoidance rules. In the simulation environment, all ships generate and execute actions according to their respective actor networks, acquire global experience and reward signals for each ship, and store the global experience in a shared experience replay pool. Samples are taken from the experience replay pool, the reward signal calculated using the multi-objective reward function is used to update the centralized critic network, and the actor network is updated using policy gradients based on the evaluation value output by the centralized critic network. After training is completed, during the deployment phase, only the actor network of each ship is retained, and each ship independently generates collision avoidance maneuvers based on its own local observation information.

[0009] According to the present invention, a multi-ship collision avoidance training method based on a centralized critic is provided, wherein the rule-based responsibility reward component includes: Based on the relative bearings and course intersection angles between the vessels, the type of encounter situation the vessel is currently in is determined. The types of encounter situations include face-to-face situations, cross encounter situations, and overtaking situations. The roles of the give-way vessel and the straight-ahead vessel are determined based on the type of encounter situation described. If the yielding vessel takes an avoidance maneuver that complies with the collision avoidance rules, it will be given a positive reward; if it takes an avoidance maneuver that violates the collision avoidance rules, it will be given a strong negative reward. If the straight-going vessel maintains its course and speed under safe conditions, a positive reward will be given; if it makes a turn that exceeds the preset range when it is not necessary, a slight negative reward will be given. If the giving vessel fails to take timely action and the straight vessel is forced to give way in violation of regulations, an additional negative reward will be imposed on the giving vessel.

[0010] According to the present invention, a multi-ship collision avoidance training method based on centralized critics is provided, wherein the avoidance actions conforming to the collision avoidance rules include: In situations where the two sides meet or cross each other, turn right to avoid the situation, and maintain a safe lateral distance when overtaking. Correspondingly, the avoidance actions that violate the collision avoidance rules include: In a situation where two or more people meet, turn left to avoid them.

[0011] According to the multi-ship collision avoidance training method based on a centralized critic provided by the present invention, the multi-objective reward function further includes a collaborative reward component, which includes: When multiple vessels engage in alternating and orderly avoidance maneuvers, each participating vessel will receive an additional team bonus. When a collision avoidance stalemate is detected, a negative reward is given to each of the relevant vessels; wherein, the collision avoidance stalemate includes situations in which both vessels turn left at the same time, both vessels turn right at the same time, or both vessels wait for each other without taking any action.

[0012] According to the present invention, a multi-ship collision avoidance training method based on a centralized critic is provided, wherein the multi-objective reward function further includes an efficiency reward component, which is used to give a slight negative reward for unnecessary detour or deceleration behavior; The total reward value of the multi-objective reward function is the weighted sum of the safety reward component, the rule responsibility reward component, the collaboration reward component, and the efficiency reward component. The weight of each component is adjustable, and the weight of the rule responsibility reward component is set to a higher value in the early stage of training.

[0013] According to the present invention, a multi-ship collision avoidance training method based on a centralized critic is provided, wherein the local observation information of the actor network includes the ship's own state information and other ship information; The vessel status information includes position, speed, heading, length, beam, maximum speed, maximum turning rate, vessel type code, and maneuverability index. The information about other vessels includes relative bearing, distance, speed, heading, vessel type, nearest encounter distance, and nearest encounter time, obtained through radar or automatic identification systems. The maneuverability index includes a turning radius coefficient and a stopping distance coefficient, which are used to characterize the differences in maneuverability of different vessels; The actor network, by inputting the maneuverability index, enables the same network architecture to generate tailored, differentiated collision avoidance strategies for vessels with different maneuverability.

[0014] According to the present invention, a multi-ship collision avoidance training method based on a centralized critic network is provided, wherein the input of the centralized critic network includes the global state splicing of all ships and the joint action splicing of all ships. The Q-value output by the centralized critic network represents the expected cumulative return that the ship can obtain in the future under the global state and the joint action conditions.

[0015] According to the present invention, a multi-ship collision avoidance training method based on a centralized critic is provided, wherein the training phase further includes: The actor network is pre-trained using historical Automatic Identification System (AIS) data to learn common maritime behavior patterns and cooperation strategies. After training, compliance fine-tuning is performed using the multi-objective reward function.

[0016] Secondly, the present invention also provides a multi-ship collision avoidance training system based on centralized critics, comprising: The module is used to build an actor network and a centralized critic network for each ship. The actor network is used to output collision avoidance actions based on the local observation information of the ship. The centralized critic network is used to evaluate the expected cumulative reward based on the global state information and joint action information of all ships during the training phase. The determination module is used to determine the multi-objective reward function for embedding collision avoidance rule responsibility quantification. The multi-objective reward function includes a safety reward component and a rule responsibility reward component. The rule responsibility reward component is based on the judgment of the ship encounter situation and the role allocation of the give-way vessel and the straight-going vessel, and gives differentiated rewards and penalties for whether the collision avoidance behavior of different roles complies with the collision avoidance rules. The reward module is used in the simulation environment to generate and execute actions according to their respective actor networks, obtain global experience and reward signals of each ship, and store the global experience in a shared experience replay pool. The training module is used to sample from the experience replay pool, update the centralized critic network using the reward signal calculated by the multi-objective reward function, and update the actor network through policy gradient based on the evaluation value output by the centralized critic network. The deployment module is used to retain only the actor network of each ship during the deployment phase after training is completed, and each ship independently generates collision avoidance actions based on its own local observation information.

[0017] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multi-ship collision avoidance training method based on a centralized critic as described above.

[0018] The multi-ship collision avoidance training method and system based on centralized critics provided by this invention have the following beneficial effects: 1. Automated production: Directly generates safe and coordinated collision avoidance action sequences, automates the decision-making process, and can handle complex multi-ship situations that are difficult for human drivers to calculate in a short time; 2. Overall coordination: When executing in a multi-ship situation, ensure that the actions of each ship are complementary and optimal as a whole, fundamentally avoiding risks caused by decision-making conflicts; 3. Operational efficiency: If the reward function incorporates a fuel consumption model, the agent can spontaneously learn the most economical strategy; in addition, the route process is smoother and there are fewer conflicts, making the arrival time of ships more predictable. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating the multi-ship collision avoidance training method based on centralized critics provided by the present invention. Figure 2 This is the overall flowchart provided by the present invention; Figure 3 This is a flowchart of the training execution process provided by the present invention; Figure 4 This is the rule-based reward quantification logic diagram provided by the present invention; Figure 5 This is a schematic diagram of the structure of the multi-ship collision avoidance training system based on a centralized critic provided by the present invention; Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0022] Figure 1 This is a flowchart illustrating the multi-ship collision avoidance training method based on centralized critics provided in an embodiment of the present invention, as shown below. Figure 1 As shown, it includes: Step 100: Construct an actor network and a centralized critic network for each ship. The actor network is used to output collision avoidance actions based on the ship's local observation information, and the centralized critic network is used to evaluate the expected cumulative reward during the training phase based on the global state information and joint action information of all ships. Step 200: Determine the multi-objective reward function for embedding collision avoidance rule responsibility quantification. The multi-objective reward function includes a safety reward component and a rule responsibility reward component. The rule responsibility reward component is based on the judgment of the ship encounter situation and the role allocation of the give-way vessel and the straight-going vessel. Differentiated rewards and penalties are given for whether the collision avoidance behavior of different roles complies with the collision avoidance rules. Step 300: In the simulation environment, all ships generate and execute actions according to their respective actor networks, acquire global experience and reward signals for each ship, and store the global experience in a shared experience replay pool. Step 400: Sample from the experience replay pool, update the centralized critic network using the reward signal calculated by the multi-objective reward function, and update the actor network through policy gradient based on the evaluation value output by the centralized critic network; Step 500: After training is completed, during the deployment phase, only the actor network of each ship is retained, and each ship independently generates collision avoidance actions based on its own local observation information.

[0023] Specifically, such as Figure 2 As shown, the steps of this embodiment of the invention are as follows: I. Agent Design Each agent i (each ship) contains two core neural networks for the training phase; only the actor network is retained for the execution phase.

[0024] 1) Actor Network: Input: Local observations o_i of agent i, specifically including: Vessel status: position, speed, course, length, beam, maximum speed, maximum turning rate, vessel type code (merchant / fishing / military, etc.), maneuverability indices (such as turning radius coefficient, stopping distance coefficient). Information on other vessels (obtained via radar / AIS, up to the nearest N vessels): relative bearing, distance, speed, heading, vessel type, CPA (closest encounter distance), TCPA (closest encounter time). Environmental information: water depth, distance to channel boundary (optional) Output: The action a_i of agent i (e.g., turning angle, speed, etc.) Policy function: π_i Function: Local observation determines the action to be taken. 2) Critics Network: Input: The concatenation of the states of all agents s = {s1, s2, …, sN} + the concatenation of the actions of all agents a = {a1, a2, …, aN} Output: Q-value Q_i(s, a1, a2, …, aN), representing the expected cumulative reward for agent i in the future when all ships take joint actions (a1, …, aN) in global state s. Function: To assess the overall situation and the effectiveness of coordinated actions. Note: The critic network is only used during the training phase, and its input information is not available during the execution phase.

[0025] II. Experience Replay Pool A shared memory is used to store the experience tuples of all agents: (s, a1, a2, ..., aN, r1, r2, ..., rN, s', o1', o2', ..., oN') in: s: Current global state a_i: Actions of each agent r_i: The reward received by each agent S': Global state at the next moment o_i': Local observations of each agent at the next time step III. Training Phase, such as Figure 3 As shown: 1) Multi-objective reward function incorporating COLREGs rules The reward function r_i is calculated by superimposing four components: safety, rules, cooperation, and efficiency, as follows: Figure 4 As shown, the design of each component is as follows: Security rewards Based on CPA / TCPA: When CPA < security threshold and TCPA < time threshold, a negative reward is given, which increases exponentially as CPA and TCPA decrease.

[0026] Collision terminates the event: A huge negative reward is given for a collision.

[0027] Rules, Responsibilities, and Rewards Real-time assessment of the encounter situation and roles of each pair of ships, and corresponding rewards and penalties: Situation assessment: Based on relative bearing and heading intersection angle, the encounter / intersection / overtaking scenarios are quantified.

[0028] Role assignment: Determine the give-way vessel and the stand-on vessel.

[0029] Reward and punishment logic: If the yielding vessel takes an evasive maneuver that conforms to COLREGs (such as turning right in an encounter / crossing situation, or maintaining a safe lateral distance in an overtaking situation), a positive reward is given; if it turns left (in violation), a strong negative reward is given.

[0030] If a ship sailing straight maintains its course and speed safely, it will receive a positive reward; if it makes a large, unnecessary turn (which may disrupt coordination), it will receive a slight negative reward.

[0031] If the giving vessel fails to take action and the direct vessel is forced to give way illegally, the giving vessel will receive an additional negative reward.

[0032] This reward function allows centralized critics to assess whether each vessel in a joint operation complies with maritime regulations, something that has never been publicly stated or implied in general robotic collaborative methods.

[0033] Collaborative rewards Additional team rewards are given when multiple vessels take turns giving way in an orderly manner (e.g., the giving-way vessel turns right in turn while the straight-going vessel maintains its course).

[0034] When a collision avoidance stalemate is detected (such as both ships turning left, both turning right, or waiting for each other), each ship is given a negative reward.

[0035] Efficiency Rewards A slight negative incentive is given for unnecessary detours or slowdowns to encourage economy.

[0036] Total Rewards: r_i = w_1 r_i_security+ w_2 r_i_rule+ w_3 r_i_cooperation+ w_4 r_i efficiency The w_2 (rule weight) can be set higher in the early stages of training to ensure that compliant behavior is learned first.

[0037] 2) Training Process Initialization: Initialize the actor network π_i and the critic network Q_i, as well as the corresponding target network, for each agent.

[0038] Environmental interaction: a. For each agent i, its actor network selects action a_i based on the current local observation o_i (adding noise if necessary). b. All actions are performed in the simulation environment. The environment transitions to the next state s', and each agent receives a reward r_i and a new local observation o_i'. c. Store the global experience (s, a, r, s', o') into the experience replay pool. Model update: a. Randomly sample a small batch of experience from the playback pool.

[0039] b. Update critics: a) For each agent i, calculate the target Q value. b) Minimize the loss function of the critic network: L(θ_{Q_i}) c) When calculating Q_i and Q_i^{target}, the input is the actions of all ships. c. Update actors: a) For each agent i, the optimization objective of its actor network is to maximize the Q-value evaluated by its critic network. b) Gradient calculation of the actor network: _{θ_{π_i}} J c) The gradient of the critic network Q_i on the action a_i of agent i is backpropagated through the actor network π_i, thereby updating the parameters of the actor network. d. Update the target network Using a large amount of AIS data covering various ship types and scenarios, a powerful and universal actor network is trained, enabling it to learn common maritime knowledge and cooperation strategies.

[0040] IV. Implementation Phase 1) During deployment, only the actor network for each agent needs to be loaded. 2) Each ship i independently calculates its action a_i based on its local observations o_i obtained from its own sensors, through its own actor network π_i(o_i), and executes the action. 3) The observation o_i should contain a "ship characteristic description vector", which may include: [ship length, ship width, length, maximum speed, maximum turning rate, ship type...] 4) The critic network has been completely abandoned; the system is distributed and requires no central communication.

[0041] Understandably, this invention, by deeply integrating the quantifiable reward function of the COLREGs rules, transforms the division of responsibilities, avoidance directions, and straight-line vessel obligations in the International Regulations for Preventing Collisions at Sea (ICP-1) into calculable reward and penalty signals. This allows centralized critics to assess not only "safety" but also "compliance" and "cooperation." Furthermore, through a centralized-distributed architecture adapted to heterogeneous vessels, the invention explicitly inputs vessel type codes and maneuvering parameters into the centralized critic network, enabling it to output personalized strategy gradients for different vessel types. During execution, each vessel can generate compliant and complementary collision avoidance maneuvers based solely on its own local observations (including its own maneuvering characteristics). Additionally, through a generalized actor network trained on massive amounts of AIS data, a shared actor network architecture can dynamically adapt to different types of vessels and new vessels not seen during training, achieving zero-sample transfer.

[0042] The following describes the multi-ship collision avoidance training system based on a centralized critic provided by the present invention. The multi-ship collision avoidance training system based on a centralized critic described below can be referred to in correspondence with the multi-ship collision avoidance training method based on a centralized critic described above.

[0043] Figure 5 This is a schematic diagram of the structure of a multi-ship collision avoidance training system based on a centralized critic, as provided in an embodiment of the present invention. Figure 5 As shown, it includes: a construction module 51, a determination module 52, a reward module 53, a training module 54, and a deployment module 55, wherein: The construction module 51 is used to construct an actor network and a centralized critic network for each vessel. The actor network outputs collision avoidance actions based on the vessel's local observation information, and the centralized critic network evaluates the expected cumulative reward during the training phase based on the global state information and joint action information of all vessels. The determination module 52 is used to determine a multi-objective reward function that embeds collision avoidance rule responsibility quantification. The multi-objective reward function includes a safety reward component and a rule responsibility reward component. The rule responsibility reward component is based on the judgment of the vessel encounter situation and the role allocation between the yielding vessel and the straight-ahead vessel, and assigns a difference based on whether the collision avoidance behavior of different roles conforms to the collision avoidance rules. Alienation of rewards and penalties; Reward module 53 is used to generate and execute actions according to their respective actor networks in the simulation environment, obtain global experience and reward signals of each ship, and store the global experience in a shared experience replay pool; Training module 54 is used to sample from the experience replay pool, update the centralized critic network using the reward signals calculated by the multi-objective reward function, and update the actor network through policy gradient based on the evaluation value output by the centralized critic network; Deployment module 55 is used to retain only the actor networks of each ship in the deployment phase after training is completed, and each ship independently generates collision avoidance actions based on its own local observation information.

[0044] Figure 6An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a multi-ship collision avoidance training method based on a centralized critic. This method includes: constructing an actor network and a centralized critic network for each ship; the actor network outputs collision avoidance actions based on the ship's local observation information; and the centralized critic network evaluates the expected cumulative reward during the training phase based on the global state information and joint action information of all ships. It also determines a multi-objective reward function embedding collision avoidance rule responsibility quantification, the multi-objective reward function including a safety reward component and a rule responsibility reward component, the rule responsibility reward component being based on the judgment of the ship encounter situation and the give-way and straight-ahead navigation. Ships are assigned roles, and differentiated rewards and penalties are given based on whether their collision avoidance behavior conforms to the collision avoidance rules. In the simulation environment, all ships generate and execute actions according to their respective actor networks, acquire global experience and reward signals for each ship, and store the global experience in a shared experience replay pool. Sampling is performed from the experience replay pool, and the reward signals calculated using the multi-objective reward function are used to update the centralized critic network. The actor network is then updated based on the evaluation value output by the centralized critic network through policy gradient. After training, only the actor networks of each ship are retained during the deployment phase, and each ship independently generates collision avoidance actions based on its own local observation information.

[0045] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0046] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0047] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-ship collision avoidance training method based on centralized critics, characterized in that, include: For each ship, an actor network and a centralized critic network are constructed. The actor network is used to output collision avoidance actions based on the ship's local observation information, and the centralized critic network is used to evaluate the expected cumulative reward based on the global state information and joint action information of all ships during the training phase. A multi-objective reward function is determined to embed the responsibility quantification of collision avoidance rules. The multi-objective reward function includes a safety reward component and a rule responsibility reward component. The rule responsibility reward component is based on the judgment of the encounter situation of ships and the role allocation of the give-way vessel and the straight-going vessel. Differentiated rewards and penalties are given for whether the collision avoidance behavior of different roles complies with the collision avoidance rules. In the simulation environment, all ships generate and execute actions according to their respective actor networks, acquire global experience and reward signals for each ship, and store the global experience in a shared experience replay pool. Samples are taken from the experience replay pool, the reward signal calculated using the multi-objective reward function is used to update the centralized critic network, and the actor network is updated using policy gradients based on the evaluation value output by the centralized critic network. After training is completed, during the deployment phase, only the actor network of each ship is retained, and each ship independently generates collision avoidance maneuvers based on its own local observation information.

2. The multi-ship collision avoidance training method based on centralized critics according to claim 1, characterized in that, The rule-based responsibility reward components include: Based on the relative bearings and course intersection angles between the vessels, the type of encounter situation the vessel is currently in is determined. The types of encounter situations include face-to-face situations, cross encounter situations, and overtaking situations. The roles of the give-way vessel and the straight-ahead vessel are determined based on the type of encounter situation described. If the yielding vessel takes an avoidance maneuver that complies with the collision avoidance rules, it will be given a positive reward; if it takes an avoidance maneuver that violates the collision avoidance rules, it will be given a strong negative reward. If the straight-going vessel maintains its course and speed under safe conditions, a positive reward will be given; if it makes a turn that exceeds the preset range when it is not necessary, a slight negative reward will be given. If the giving vessel fails to take timely action and the straight vessel is forced to give way in violation of regulations, an additional negative reward will be imposed on the giving vessel.

3. The multi-ship collision avoidance training method based on centralized critics according to claim 2, characterized in that, The avoidance actions that comply with the collision avoidance rules include: In situations where the two sides meet or cross each other, turn right to avoid the situation, and maintain a safe lateral distance when overtaking. Correspondingly, the avoidance actions that violate the collision avoidance rules include: In a situation where two or more people meet, turn left to avoid them.

4. The multi-ship collision avoidance training method based on centralized critics according to claim 1, characterized in that, The multi-objective reward function further includes a collaborative reward component, which includes: When multiple vessels engage in alternating and orderly avoidance maneuvers, each participating vessel will receive an additional team bonus. When a collision avoidance stalemate is detected, a negative reward is given to each of the relevant vessels; wherein, the collision avoidance stalemate includes situations in which both vessels turn left at the same time, both vessels turn right at the same time, or both vessels wait for each other without taking any action.

5. The multi-ship collision avoidance training method based on centralized critics according to claim 1, characterized in that, The multi-objective reward function also includes an efficiency reward component, which is used to give a slight negative reward for unnecessary detours or decelerations. The total reward value of the multi-objective reward function is the weighted sum of the safety reward component, the rule responsibility reward component, the collaboration reward component, and the efficiency reward component. The weight of each component is adjustable, and the weight of the rule responsibility reward component is set to a higher value in the early stage of training.

6. The multi-ship collision avoidance training method based on centralized critics according to claim 1, characterized in that, The local observation information of the actor network includes the ship's own status information and information about other ships; The vessel status information includes position, speed, heading, length, beam, maximum speed, maximum turning rate, vessel type code, and maneuverability index. The information about other vessels includes relative bearing, distance, speed, heading, vessel type, nearest encounter distance, and nearest encounter time, obtained through radar or automatic identification systems. The maneuverability index includes a turning radius coefficient and a stopping distance coefficient, which are used to characterize the differences in maneuverability of different vessels; The actor network, by inputting the maneuverability index, enables the same network architecture to generate tailored, differentiated collision avoidance strategies for vessels with different maneuverability.

7. The multi-ship collision avoidance training method based on centralized critics according to claim 1, characterized in that, The input to the centralized critic network includes a global state patch of all ships and a joint action patch of all ships. The Q-value output by the centralized critic network represents the expected cumulative return that the ship can obtain in the future under the global state and the joint action conditions.

8. The multi-ship collision avoidance training method based on centralized critics according to claim 1, characterized in that, The training phase also includes: The actor network is pre-trained using historical Automatic Identification System (AIS) data to learn common maritime behavior patterns and cooperation strategies. After training, compliance fine-tuning is performed using the multi-objective reward function.

9. A multi-ship collision avoidance training system based on centralized critics, characterized in that, include: The module is used to build an actor network and a centralized critic network for each ship. The actor network is used to output collision avoidance actions based on the local observation information of the ship. The centralized critic network is used to evaluate the expected cumulative reward based on the global state information and joint action information of all ships during the training phase. The determination module is used to determine the multi-objective reward function for embedding collision avoidance rule responsibility quantification. The multi-objective reward function includes a safety reward component and a rule responsibility reward component. The rule responsibility reward component is based on the judgment of the ship encounter situation and the role allocation of the give-way vessel and the straight-going vessel, and gives differentiated rewards and penalties for whether the collision avoidance behavior of different roles complies with the collision avoidance rules. The reward module is used in the simulation environment to generate and execute actions according to their respective actor networks, obtain global experience and reward signals of each ship, and store the global experience in a shared experience replay pool. The training module is used to sample from the experience replay pool, update the centralized critic network using the reward signal calculated by the multi-objective reward function, and update the actor network through policy gradient based on the evaluation value output by the centralized critic network. The deployment module is used to retain only the actor network of each ship during the deployment phase after training is completed, and each ship independently generates collision avoidance actions based on its own local observation information.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the multi-ship collision avoidance training method based on a centralized critic as described in any one of claims 1 to 8.