Multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios

CN122569043APending Publication Date: 2026-08-14BOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]现有避碰决策技术仍存在明显不足,难以适配复杂海上强交互场景

Benefits of technology

本发明提高了自主船舶在复杂海上交通环境中的避碰决策能力和环境适应能力,为船舶的海上智能化航行提供了新的技术启示。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122569043A_ABST
    Figure CN122569043A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios, relating to the fields of intelligent shipping, autonomous vessel decision-making, and maritime traffic safety. The method includes the following steps: acquiring maritime traffic scenario information to construct a semi-cooperative phase mechanism for the target vessel, and implementing phase transitions based on the local interaction conditions between the target vessel and the vessel itself; updating the target vessel's motion state based on its current semi-cooperative phase to construct a highly interactive maritime traffic scenario with dynamic behavioral evolution characteristics; and constructing a multi-agent semi-cooperative evolutionary decision-making model and its reward function based on the highly interactive maritime traffic scenario, and obtaining an autonomous decision-making strategy through model training to control the vessel's autonomous collision avoidance in the presence of the target vessel and / or static obstacles. This invention improves the collision avoidance decision-making capability and environmental adaptability of autonomous vessels in complex maritime traffic environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent shipping, autonomous ship decision-making, and maritime traffic safety, and more specifically, to a multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios. Background Technology

[0002] With the rapid development of intelligent shipping and autonomous vessel technology, autonomous collision avoidance decision-making has become a key research direction in the field of intelligent navigation. Autonomous vessels need to adjust their course and speed in real time according to the maritime environment, balancing navigation safety and efficiency. The actual maritime traffic environment is complex, with frequent encounters, crossings, and overtaking between vessels. Under high traffic density conditions, frequent interactions between multiple vessels and dynamically changing risks create highly uncertain multi-vessel encounter scenarios. Compared to simple single-vessel or two-vessel encounters, multi-vessel scenarios with strong interaction are highly coupled and have many variables, placing extremely high demands on the risk assessment and collision avoidance decision-making capabilities of the vessel's autonomous decision-making system.

[0003] Current autonomous collision avoidance decision-making methods for ships can be mainly divided into four categories: rule-based decision-making methods based on maritime collision avoidance rules, risk assessment methods based on geometric relationships, path planning methods based on trajectory optimization, and intelligent decision-making methods based on reinforcement learning. Among them, rule-based decision-making methods are only applicable to simple encounter scenarios; geometric risk assessment methods rely on fixed distance and time indicators to judge collision risk; path planning methods achieve passive collision avoidance by optimizing the trajectory; and reinforcement learning methods learn decision-making strategies through interaction with the environment, which is the mainstream research direction for intelligent collision avoidance.

[0004] Existing collision avoidance decision-making technologies still have significant shortcomings and are difficult to adapt to complex maritime scenarios with strong interactions. On the one hand, most existing methods assume that the target vessel's course and speed are fixed, using simple models to describe vessel motion, ignoring the dynamic differences in vessel avoidance behavior in real-world scenarios. This makes them unable to accurately characterize the real-time risks of multi-vehicle interactions, easily leading to decision-making errors. On the other hand, existing technologies tend to focus on collision avoidance results, ignoring the evolving characteristics of the traffic environment caused by vessel dynamic behavior. This results in a large deviation between the algorithm training environment and the actual scenario, leading to poor generalization ability and scenario adaptability of the decision-making model, and an inability to effectively cope with complex multi-vehicle encounters. Therefore, there is an urgent need for an autonomous decision-making scheme that can accurately characterize the dynamic interaction behavior of vessels and adapt to the dynamic evolution of the maritime environment to meet the safety collision avoidance requirements of autonomous vessels in complex scenarios. Summary of the Invention

[0005] To address the aforementioned issues, the present invention aims to provide a multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios, thereby enhancing the collision avoidance decision-making capabilities and environmental adaptability of autonomous vessels in complex maritime traffic environments.

[0006] To achieve the above technical objectives, this application provides a multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios, including the following steps: Acquire information about maritime traffic scenarios to construct a semi-cooperative phase mechanism for the target vessel, and realize phase transition based on the local interaction conditions between the target vessel and the vessel itself; Based on the target vessel's current semi-cooperative phase, update the target vessel's motion state to construct a highly interactive maritime traffic scenario with dynamic behavioral evolution characteristics. Based on highly interactive maritime traffic scenarios, a multi-agent semi-cooperative evolutionary decision-making model and its reward function are constructed. By training the model, an autonomous decision-making strategy is obtained to control the ship to autonomously avoid collisions with target ships and / or static obstacles.

[0007] Preferably, when acquiring maritime traffic scene information, the information acquired includes the ship's own information, the target ship's information, and static obstacle information, which together form the maritime traffic scene information. The ship's own information includes the ship's current position, speed, course, and target point position; the target ship's information includes the target ship's position, speed, course, and target point position; and the static obstacle information includes the obstacle's position and shape.

[0008] Preferably, when constructing the target vessel semi-cooperative phase mechanism, the target vessel semi-cooperative phase mechanism includes a normal navigation phase, an implicit cooperation phase, a weak cooperation adjustment phase, a cooperation degradation phase, and a recovery phase.

[0009] Preferably, when acquiring local interaction conditions, the local interaction conditions include one or more of the following: relative distance, relative bearing, relative speed, collision risk, ship domain constraints, and collision avoidance liability relationship.

[0010] Preferably, during the execution phase transition, the target vessel's speed change is generated based on the current semi-cooperative phase: in, Indicates the current speed; Indicates the initial velocity; This represents the velocity perturbation corresponding to the current semi-cooperative phase; Indicates the velocity recovery coefficient; These represent the minimum and maximum speeds allowed for the target vessel, respectively.

[0011] The target vessel's course change is generated based on the current semi-cooperative phase: in, For the actual course, The heading deviation generated during the semi-cooperative phase.

[0012] Preferably, when obtaining the heading deviation, the heading deviation is updated in the following manner: in, This represents the heading disturbance corresponding to the current semi-cooperative phase. This is the heading recovery coefficient.

[0013] Preferably, when constructing a multi-agent semi-cooperative evolutionary model, the multi-agent semi-cooperative evolutionary model is established based on a partially observable Markov decision process.

[0014] Preferably, when constructing the reward function, the reward function is constructed based on progress reward, yaw reward, collision penalty reward, safety margin reward, and emergency avoidance reward.

[0015] Preferably, reinforcement learning algorithms are used for model training.

[0016] Preferably, when acquiring an autonomous decision-making strategy, for the ship, through continuous interaction with a highly interactive maritime traffic scenario that includes different semi-cooperative evolutionary stages, the strategy parameters are updated based on navigation progress, course control, collision events, safety margins, and rewards for emergency avoidance, thereby obtaining an autonomous decision-making strategy that generates collision avoidance actions for the ship based on the current traffic scenario state.

[0017] The present invention discloses the following technical effects: This invention improves the collision avoidance decision-making ability and environmental adaptability of autonomous ships in complex maritime traffic environments, and provides new technological inspiration for intelligent navigation of ships at sea. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of the method described in this invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0021] Example 1: As Figure 1 As shown, the present invention provides a method applicable to autonomous navigation decision-making under conditions of multiple ship encounters, uncertain target ship behavior, dynamic changes in the degree of cooperation, and continuous evolution of local risks.

[0022] S100 acquires maritime traffic scene information, which includes information about the ship itself, information about the target ship, and information about static obstacles.

[0023] S200 establishes a semi-cooperative phase mechanism for target vessels, dividing the behavior of target vessels into a normal navigation phase, an implicit cooperation phase, a weak cooperation adjustment phase, a cooperation degradation phase, and a recovery phase, and realizing phase transitions based on the local interaction conditions between the target vessel and the vessel itself.

[0024] In one embodiment, the semi-cooperative phase mechanism includes the normal navigation phase, implicit cooperation phase, weak cooperation adjustment phase, cooperation degradation phase, and recovery phase in the target vessel behavior mechanism.

[0025] For example, a semi-cooperative phase set is established for each target vessel: in This indicates that during the normal navigation phase, the target vessel maintains its original course and speed. This indicates an implicit cooperation phase, where the target vessel exhibits a slight tendency to cooperate, such as slight deceleration or small-angle turning; This indicates a weak cooperation adjustment phase, during which the target vessel makes limited short-term maneuvers but does not completely give way; This indicates a phase of declining cooperation, during which the target vessel may suddenly accelerate, suddenly veer off course, or fail to respond effectively to the risk of collision. This indicates the recovery phase, where, as local risks decrease, the target vessel gradually returns to a stable navigation state.

[0026] In one embodiment, the phase transfer of the target vessel is determined by local interaction conditions, which include relative distance, relative bearing, relative speed, collision risk, vessel domain constraints, and collision avoidance liability relationships.

[0027] For example, for the first The target ship, the phased transfer is represented as: in, For the first Local encounter conditions between the target vessel and the vessel itself. When the relative distance is large and the risk is low, the target vessel tends to maintain normal navigation; when the relative distance decreases or the risk increases, the target vessel may enter an implicit or weak cooperation adjustment phase; when the traffic environment becomes congested, the responsibility for the encounter is unclear, or the target vessel's response is unstable, the target vessel may enter a cooperation degradation phase; when the risk decreases, the target vessel enters a recovery phase and gradually returns to a stable navigation state.

[0028] Based on the target vessel's current semi-cooperative phase, S300 generates corresponding speed and heading changes, updates the target vessel's motion state, and constructs a highly interactive maritime traffic scenario with dynamic behavioral evolution characteristics.

[0029] In one implementation, for each target vessel, speed and heading changes are generated based on its current semi-cooperative phase.

[0030] For example, suppose the first The initial speed of the target ship is The current speed is Then the velocity evolves as follows: in, This represents the velocity perturbation corresponding to the current semi-cooperative phase. and These are the minimum and maximum speeds allowed for the target ship, respectively. The velocity recovery coefficient, This is the amplitude limiting function.

[0031] For example, the target vessel's course change is generated based on the current semi-cooperative phase and updated as follows: in, For the actual course, This is the heading deviation caused by the semi-cooperative phase.

[0032] For example, the heading deviation evolves as follows: in, This represents the heading disturbance corresponding to the current semi-cooperative phase. This is the heading recovery coefficient.

[0033] For example, the target ship's position is updated as follows: .

[0034] For example, through the above mechanism, different target ships can be in different semi-cooperative stages at the same time, forming an asynchronous, diverse and continuously evolving multi-ship strong interactive traffic scenario.

[0035] Based on the highly interactive maritime traffic scenario, S400 establishes a multi-agent semi-cooperative evolutionary decision-making model and constructs a state space, action space, and partially observable space.

[0036] In one implementation, the multi-agent semi-cooperative evolution model is based on a partially observable Markov decision process.

[0037] For example, the state space is: in, Indicates the position of the own ship; Indicates the ship's speed; Indicates the ship's heading angle; This indicates the distance between the ship's current position and the target point; Location of the target ship. For the target ship speed, For the target ship's course, The distance between the target ship and its target point. The target vessel is currently in a semi-cooperative phase. This refers to the heading deviation of the target vessel during the semi-cooperative phase. For the first The center coordinates of each obstacle Let be the radius of the obstacle.

[0038] For example, the action space is a continuous heading adjustment action: ,in, Indicates the maximum left turn. Indicates the maximum right turn. This indicates that the current course will be maintained.

[0039] For example, action value This will be mapped to actual changes in heading. : in This represents the change in the ship's heading angle within a single time step. For the maximum permissible steering angle, This refers to the normalized actions output by this ship.

[0040] For example, this ship is The course change at any given time is as follows: ; The ship's position will then be updated as follows: ; Part of the observable space is: in, For the first The relative positions of the target vessels; The normalized relative heading; It is the relative velocity; To maintain the current speed and course, in the future By predicting the relative coordinates of the location, this invention encodes each target vessel into a token vector, which simultaneously contains current information and short-term predictions in time.

[0041] For example, in order to highlight hazardous targets, a hazard candidate set is defined. Choose the one closest to your ship Boat: .

[0042] For example, the danger mask vector is designed as follows: If the first The boat belongs to ,but ,otherwise At the same time, from Select the most dangerous ship (e.g., the target with the shortest distance) and use its token as a separate danger feature vector. The data is being integrated into the observation process.

[0043] In one implementation, for static obstacles, the shortest vector from the ship's center point to the nearest static obstacle boundary is extracted and used to set the safety margin between the ship and the static obstacle: in, For the first The position of the center of a static obstacle; The radius of the static obstacle; This is the equivalent radius of the ship (calculated using the ship's length and width). for The position vector of the ship at that moment.

[0044] S500 builds a multi-dimensional reward function on this basis.

[0045] In one embodiment, the multi-agent semi-cooperative evolutionary model is established based on a partially observable Markov decision process, and the reward function is: .

[0046] For example, As a progress reward, to encourage the ship to gradually approach the target point: in, This is the Euclidean distance from the ship to the target point; This is the progress reward coefficient; This is the progress trimming threshold.

[0047] For example, Bonus for yaw: in, This represents the yaw error between the ship and the target direction. It is a gate function that decays with distance; For heading reward weighting.

[0048] For example, As a collision penalty reward, if the ship successfully reaches the target point, a reward is given; if a collision occurs, a penalty is applied, as shown in the formula: For example, As a safety margin bonus, it quantifies the collision risk between the vessel and the target vessel: in, This is the safety margin bonus coefficient.

[0049] For example, The maximum risk index is the time required to reach the nearest encounter point. for: in, For this ship and the first The distance to the target vessel; For safe distances based on the ship's domain; This refers to the time margin between this vessel and the target vessel. The equivalent radius of the target vessel is defined here, where both the vessel and the target vessel are considered as equivalent circles, with their radii determined by the diagonals of the vessel's length and breadth; a safety factor is also introduced. This allows for greater clearance for ships to maneuver, enabling the calculation of safe distances based on the ship's domain. ; This represents the time margin threshold. Define this ship and the first. The relative positions of the target vessels are as follows: The relative speed is ,but The calculation can be solved using the formula: .

[0050] For example, Rewards for emergency avoidance: in The danger level of the most dangerous target at present; For the actions of this ship; The weights are respectively for moving away from and towards danger.

[0051] The S600 employs reinforcement learning algorithms to train a multi-agent semi-cooperative evolutionary decision-making model. Utilizing the trained autonomous decision-making strategy, it generates ship control actions based on the current traffic scenario state, enabling autonomous collision avoidance decisions in target ship and static obstacle environments.

[0052] In one embodiment, during model training, the present invention employs the Truncated Quantile Critics (TQC) algorithm to train the autonomous decision-making model. The model takes the ship's motion state, relative motion information of the target ship, future trajectory prediction information, hazard target markers, and static obstacle boundary information as inputs, and the ship's heading control actions as outputs.

[0053] For example, during the training process, the ship continuously interacts with the highly interactive maritime traffic scenario that evolves from the semi-cooperative phase of the target ship, and updates the network parameters based on navigation progress rewards, course control rewards, collision event rewards, safety margin rewards, and emergency avoidance rewards.

[0054] For example, after training, the mapping relationship between the traffic scene state and the ship's control actions is obtained, enabling the ship to generate collision avoidance decision actions in real time according to the current traffic environment.

[0055] Example 2: This invention provides a multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios, comprising the following steps: S1. Obtain maritime traffic scene information, which includes information about the vessel itself, target vessels, and static obstacles. The vessel information includes its current position, speed, course, and target point location. The target vessel information includes its position, speed, course, and target point location. The static obstacle information includes the location and shape of the obstacles. Further, this embodiment uses a two-dimensional maritime traffic scene to construct a simulation environment, which includes the vessel itself, several target vessels, and several static obstacles. The vessel needs to reach the target point while ensuring navigational safety, and simultaneously avoid collisions with target vessels and static obstacles.

[0056] S2. In order to describe the characteristics of the target ship’s behavior changes, the present invention establishes a semi-cooperative phase mechanism for the target ship.

[0057] Specifically, a semi-cooperative phase set is established for each target vessel: in This indicates that during the normal navigation phase, the target vessel maintains its original course and speed. This indicates an implicit cooperation phase, where the target vessel exhibits a slight tendency to cooperate, such as slight deceleration or small-angle turning; This indicates a weak cooperation adjustment phase, during which the target vessel makes limited short-term maneuvers but does not completely give way; This indicates a phase of declining cooperation, during which the target vessel may suddenly accelerate, suddenly veer off course, or fail to respond effectively to the risk of collision. This indicates the recovery phase, where, as local risks decrease, the target vessel gradually returns to a stable navigation state.

[0058] Furthermore, the phase transfer of the target vessel is determined by local interaction conditions, including relative distance, relative bearing, relative speed, collision risk, vessel domain constraints, and collision avoidance liability relationships. For the first... The target ship, the phased transfer is represented as: in, For the first Local encounter conditions between the target vessel and the vessel itself. When the relative distance is large and the risk is low, the target vessel tends to maintain normal navigation; when the relative distance decreases or the risk increases, the target vessel may enter an implicit or weak cooperation adjustment phase; when the traffic environment becomes congested, the responsibility for the encounter is unclear, or the target vessel's response is unstable, the target vessel may enter a cooperation degradation phase; when the risk decreases, the target vessel enters a recovery phase and gradually returns to a stable navigation state.

[0059] S3. Generate speed and heading changes based on the target vessel's current semi-cooperative phase. For the first... The initial speed of the target ship is The current speed is Then the velocity evolves as follows: in, This represents the velocity perturbation corresponding to the current semi-cooperative phase. and These are the minimum and maximum speeds allowed for the target ship, respectively. The velocity recovery coefficient, This is the amplitude limiting function.

[0060] The target vessel's course change is generated based on the current semi-cooperative phase and updated as follows: in, For the actual course, This is the heading deviation caused by the semi-cooperative phase. The heading deviation evolves as follows: in, This represents the heading disturbance corresponding to the current semi-cooperative phase. This is the heading recovery coefficient. The target ship's position is then updated as follows: .

[0061] Through the above mechanism, the present invention enables different target ships to be in different semi-cooperative stages at the same time, forming an asynchronous, diverse and continuously evolving multi-ship strong interactive traffic scenario.

[0062] S4. This invention models the multi-agent semi-cooperative evolutionary model as a partially observable Markov decision process. Specifically, it constructs a state space, action space, observation space, and reward function.

[0063] Furthermore, the state space is defined as: in, Indicates the position of the own ship; Indicates the ship's speed; Indicates the ship's heading angle; This indicates the distance between the ship's current position and the target point. Location of the target ship. For the target ship speed, For the target ship's course, The distance between the target ship and its target point. The target vessel is currently in a semi-cooperative phase. This refers to the heading deviation of the target vessel during the semi-cooperative phase. For the first The center coordinates of each obstacle Let be the radius of the obstacle.

[0064] Furthermore, the action space is defined as continuous heading adjustment actions: .in, Indicates the maximum left turn. Indicates the maximum right turn. Indicates maintaining the current heading. Action value This will be mapped to actual changes in heading. : in, This represents the change in the ship's heading angle within a single time step. For the maximum permissible steering angle, This refers to the normalized actions output by this ship. Therefore, this ship is... The course change at any given time is as follows: ; The ship's position will then be updated as follows: .

[0065] Furthermore, the partially observable space is defined as: in, For the first The relative positions of the target vessels; The normalized relative heading; It is the relative velocity; To maintain the current speed and course, in the future By predicting the relative coordinates of the location, this invention encodes each target vessel into a token vector, which simultaneously contains current information and short-term predictions in time.

[0066] To highlight hazardous targets, this invention defines a hazardous candidate set. Choose the one closest to your ship Boat: Furthermore, the present invention designs the danger mask vector as follows: If the first The boat belongs to ,but ,otherwise At the same time, from Select the most dangerous ship (e.g., the target with the shortest distance) and use its token as a separate danger feature vector. The data is being integrated into the observation process.

[0067] For static obstacles, this invention extracts the shortest vector from the ship's center point to the nearest static obstacle boundary, which is used to set the safety margin between the ship and the static obstacle: in, For the first The position of the center of a static obstacle; The radius of the static obstacle; This is the equivalent radius of the ship (calculated using the ship's length and width). for The position vector of the ship at that moment.

[0068] S5. In order to guide the ship to learn safe and effective collision avoidance strategies, this invention constructs a reward function consisting of progress reward, yaw reward, collision event reward, safety margin reward and emergency avoidance reward.

[0069] The total reward function is defined as: .

[0070] in, As a progress reward, to encourage the ship to gradually approach the target point: In the formula, This is the Euclidean distance from the ship to the target point; This is the progress reward coefficient; The progress trimming threshold; Bonus for yaw: In the formula, This represents the yaw error between the ship and the target direction. It is a gate function that decays with distance; Weighting of course reward; As a collision penalty reward, if the ship successfully reaches the target point, a reward is given; if a collision occurs, a penalty is applied, as shown in the formula: ; As a safety margin bonus, it quantifies the collision risk between the vessel and the target vessel: In the formula, This is the safety margin bonus coefficient; The maximum risk index is the time required to reach the nearest encounter point. for: In the formula, For this ship and the first The distance to the target vessel; For safe distances based on the ship's domain; This refers to the time margin between this vessel and the target vessel. The equivalent radius of the target vessel is defined here, where both the vessel and the target vessel are considered as equivalent circles, with their radii determined by the diagonals of the vessel's length and breadth; a safety factor is also introduced. This allows for greater clearance for ships to maneuver, enabling the calculation of safe distances based on the ship's domain. ; This represents the time margin threshold. Define this ship and the first. The relative positions of the target vessels are as follows: The relative speed is ,but The calculation can be solved using the formula: ; Rewards for emergency avoidance: In the formula, The danger level of the most dangerous target at present; For the actions of this ship; The weights are respectively for moving away from and towards danger.

[0071] S6. During model training, this invention employs the Truncated Quantile Critics (TQC) algorithm to train the autonomous decision-making model. The model takes the ship's motion state, relative motion information of the target ship, future trajectory prediction information, hazard target markers, and static obstacle boundary information as input, and the ship's heading control actions as output. During training, the ship continuously interacts with the highly interactive maritime traffic scenario evolving from the semi-cooperative phase of the target ship, and updates the network parameters based on navigation progress rewards, heading control rewards, collision event rewards, safety margin rewards, and emergency avoidance rewards. After training, the mapping relationship between the traffic scenario state and the ship's control actions is obtained, enabling the ship to generate collision avoidance decisions in real time based on the current traffic environment.

[0072] After training is completed, once the vessel acquires the current traffic environment status, the decision-making strategy outputs course control actions in real time based on the input status. The vessel adjusts its course and trajectory according to the control actions, thereby completing autonomous collision avoidance of target vessels and static obstacles, and safely navigating to the target point.

[0073] In summary, this invention discloses a multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios. First, it acquires maritime traffic scenario information including the vessel itself, the target vessel, and static obstacles. Then, it constructs a semi-cooperative phase mechanism for the target vessel, dividing its behavior into normal navigation, implicit cooperation, weak cooperation adjustment, cooperative degradation, and recovery phases, and implementing phase transitions based on local interaction conditions. Based on the target vessel's current phase, it generates speed and heading changes, constructing a highly interactive maritime traffic scenario with dynamic behavioral evolution characteristics. On this basis, it establishes an autonomous decision-making model based on a partially observable Markov decision process, constructing a multi-agent semi-cooperative evolutionary state space, action space, observation space, and a multi-dimensional reward function, and trains the model using a truncated quantile critic algorithm. After training, it establishes a mapping relationship between the traffic scenario state and the vessel's control actions, enabling the vessel to generate collision avoidance decisions in real time based on the current traffic environment state, achieving autonomous navigation and collision avoidance decision-making in complex maritime traffic environments. This invention can improve the autonomous vessel's adaptability to changes in target vessel behavior and its collision avoidance decision-making capabilities in complex scenarios.

[0074] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0075] In the description of this invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0076] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios, characterized in that: Includes the following steps: Acquire information about maritime traffic scenarios to construct a semi-cooperative phase mechanism for the target vessel, and realize phase transition based on the local interaction conditions between the target vessel and the vessel itself; Based on the target vessel's current semi-cooperative phase, update the target vessel's motion state to construct a highly interactive maritime traffic scenario with dynamic behavioral evolution characteristics. Based on the highly interactive maritime traffic scenario, a multi-agent semi-cooperative evolutionary decision-making model and its reward function are constructed. By training the model, an autonomous decision-making strategy is obtained to control the ship to autonomously avoid collisions with the target ship and / or static obstacles.

2. The multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios according to claim 1, characterized in that: When acquiring maritime traffic scene information, the information of the ship itself, the target ship, and static obstacles are acquired as the maritime traffic scene information. The information of the ship itself includes the ship's current position, speed, course, and target point position. The information of the target ship includes the target ship's position, speed, course, and target point position. The information of static obstacles includes the position and shape information of the obstacles.

3. The multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios according to claim 1, characterized in that: When constructing the target vessel semi-cooperative phase mechanism, the target vessel semi-cooperative phase mechanism includes a normal navigation phase, an implicit cooperation phase, a weak cooperation adjustment phase, a cooperation degradation phase, and a recovery phase.

4. The multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios according to claim 1, characterized in that: When obtaining local interaction conditions, the local interaction conditions include one or more of the following: relative distance, relative bearing, relative speed, collision risk, ship domain constraints, and collision avoidance liability relationship.

5. The multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios according to claim 1, characterized in that: During the execution phase transition, the speed change of the target vessel is generated based on the current semi-cooperative phase: in, Indicates the current speed; Indicates the initial velocity; This represents the velocity perturbation corresponding to the current semi-cooperative phase; Indicates the velocity recovery coefficient; These represent the minimum and maximum speeds allowed for the target vessel, respectively. The target vessel's course change is generated based on the current semi-cooperative phase: in, For the actual course, The heading deviation generated during the semi-cooperative phase.

6. The multi-agent semi-cooperative evolutionary decision-making method for strongly interactive maritime traffic scenarios according to claim 5, characterized in that: When obtaining the heading deviation, the heading deviation is updated as follows: in, This represents the heading disturbance corresponding to the current semi-cooperative phase. This is the heading recovery coefficient.

7. The multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios according to claim 1, characterized in that: When constructing a multi-agent semi-cooperative evolutionary model, the model is established based on a partially observable Markov decision process.

8. The multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios according to claim 1, characterized in that: When constructing the reward function, the reward function is constructed based on the progress reward, yaw reward, collision penalty reward, safety margin reward, and emergency avoidance reward.

9. The multi-agent semi-cooperative evolutionary decision-making method for highly interactive maritime traffic scenarios according to claim 1, characterized in that: Reinforcement learning algorithms are used for model training.

10. The multi-agent semi-cooperative evolutionary decision-making method for strongly interactive maritime traffic scenarios according to claim 1, characterized in that: When acquiring an autonomous decision-making strategy, for the vessel, the autonomous decision-making strategy is obtained by continuously interacting with a highly interactive maritime traffic scenario that includes different semi-cooperative evolutionary stages, updating the strategy parameters based on navigation progress, course control, collision events, safety margins, and emergency avoidance rewards, and generating collision avoidance decision actions for the vessel based on the current traffic scenario state.