An unmanned ship collision avoidance control method and device, a terminal device and a medium

CN117406716BActive Publication Date: 2026-08-28CENT SOUTH UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311263298.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-27
Publication Date
2026-08-28
Estimated Expiration
2043-09-27

AI Technical Summary

Technical Problem

这种算法通常适用于相对简单的水域环境,但在复杂多变的情况下可能效果有限

Benefits of technology

本申请提供的无人艇避碰控制方法,根据状态信息构建目标无人艇的动力学模型,再分别构建目标无人艇的速度障碍区域和速度可行区域,然后根据状态信息、速度障碍区域以及速度可行区域,构建控制奖励函数,通过该控制奖励函数能够选择出更优秀的控制动作,避免无人艇发生碰撞,提高避碰控制效果;此外,该方法不依赖水域环境参数,能够适用于复杂的水域环境,具备较强的泛化能力,有利于提高避碰控制效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117406716B_ABST
    Figure CN117406716B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of water unmanned surface vehicle control, and provides an unmanned surface vehicle collision avoidance control method and device, terminal equipment and medium, state information is collected, a dynamic model of a target unmanned surface vehicle is constructed, a speed obstacle area of the target unmanned surface vehicle is constructed, a speed feasible area of the target unmanned surface vehicle is constructed, a control reward function is constructed according to the state information, the speed obstacle area and the speed feasible area, and the target unmanned surface vehicle is controlled by using a control action corresponding to a maximum reward value. The application can improve the unmanned surface vehicle collision avoidance control effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of unmanned surface vessel control technology, and particularly relates to an unmanned surface vessel collision avoidance control method, device, terminal equipment and medium. Background Technology

[0002] With the rapid advancement of technology, unmanned surface vessels (USVs) have become revolutionary tools in the marine field, sparking widespread interest in their broad application prospects. USV swarms possess remarkable versatility. In research, scientists can use USVs to explore the deep ocean, monitor water quality, meteorological conditions, and ecosystems, providing valuable data for climate change research. Furthermore, USV swarms have demonstrated significant potential in environmental monitoring, enabling real-time monitoring of pollution, marine life migration, and natural disasters, supporting environmental protection efforts and resource management. Military applications also benefit greatly; USV swarms can be used for maritime patrols, surface target tracking, and underwater combat support missions. They provide intelligence, surveillance, and reconnaissance capabilities, contributing to maintaining maritime security. In addition, USV swarms play a crucial role in maritime search and rescue. USVs can respond rapidly to disaster events, searching for missing vessels or personnel and providing life-saving support.

[0003] However, realizing these applications faces many complex challenges, one of which is collision avoidance control for unmanned surface vessel (USV) swarms. Multiple USVs performing missions simultaneously in waterways need to maintain a safe distance between them to avoid collisions. This necessitates efficient collaborative operations and collision avoidance strategies.

[0004] With the development of unmanned surface vessel (USV) swarm technology, many methods have emerged to address the collision avoidance control problem of USV swarms, including: Method 1: Artificial Potential Field Method The artificial potential field method uses a virtual potential field to describe the dynamic environment of unmanned surface vessels (USVs). Each USV is treated as a point with both repulsive and attractive forces, and the goal is to minimize the potential energy function. This method can achieve path planning while avoiding collisions and is suitable for relatively simple aquatic environments.

[0005] Method 2: Model Predictive Control (MPC) Model predictive control (MPC) is a control method based on mathematical models that can be used for collision avoidance in unmanned surface vessel (USV) swarms. By predicting the trajectory of USVs over a future period, the MPC algorithm can select the optimal control strategy to avoid collisions. This method requires sophisticated dynamic modeling of the system but performs well in complex environments.

[0006] Method 3: Rule-based algorithm Rule-based algorithms use a set of predefined rules and strategies to guide the behavior of unmanned surface vessels (USVs) to avoid collisions. For example, avoidance rules can be defined, requiring USVs to change course or speed when encountering other vessels to avoid a collision. Such algorithms are generally suitable for relatively simple aquatic environments, but may have limited effectiveness in complex and variable conditions.

[0007] However, due to the underactuated nature of unmanned surface vessels (USVs), the uncertainty of dynamic waters, and environmental perception errors, the above-mentioned USV collision avoidance control methods have poor control effects and the results in actual application are not ideal. Summary of the Invention

[0008] This application provides a collision avoidance control method, device, terminal equipment, and medium for unmanned surface vessels (USVs), which can improve the collision avoidance control effect of USVs.

[0009] Firstly, this application provides a collision avoidance control method for unmanned surface vessels, including: Collect the state information of the target unmanned surface vessel and construct a dynamic model of the target unmanned surface vessel based on the state information; Based on the pre-set collision radius and dynamic model of the target unmanned surface vessel (USV), a speed obstacle zone for the USV is constructed; when the speed of the USV is within the speed range corresponding to the speed obstacle zone, the USV collides with other USVs. Based on the speed obstacle zone, construct the speed feasible zone for the target unmanned surface vessel; when the speed of the target unmanned surface vessel is within the speed range corresponding to the speed feasible zone, the target unmanned surface vessel will not collide with other unmanned surface vessels; Based on the state information, the speed obstacle area, and the speed feasible area, a control reward function is constructed. The control reward function is used to calculate the reward value corresponding to different control actions, which include the speed change of the target unmanned surface vessel at each moment. Control the target unmanned surface vessel using the control action corresponding to the maximum reward value.

[0010] Optionally, the expression for the dynamic model is as follows:

[0011]

[0012] in, This represents the state vector of the target unmanned surface vessel. , This indicates the position of the target unmanned surface vessel in the world coordinate system. Indicates the yaw angle of the target unmanned surface vessel. This represents the transformation matrix from the ship's coordinate system to the world coordinate system. , This represents the velocity vector of the target unmanned surface vessel in the ship's coordinate system. , This represents the linear velocity of the target unmanned surface vessel in the ship's coordinate system. This represents the angular velocity of the target unmanned surface vessel in the ship's coordinate system. express The derivative of express The derivative of This represents the rotational inertia matrix of the target unmanned surface vessel. This represents the three-axis torque of the target unmanned surface vessel. , This represents the thrust along the U-axis of the ship's hull. This represents the thrust along the ship's v-axis. This represents the torque along the w direction. This represents the unmodeled portion of the target unmanned surface vessel's control, which includes random disturbances and surface resistance.

[0013] Optionally, the expression for the speed barrier region is as follows:

[0014] in, Indicates the target unmanned surface vessel Other unmanned surface vessels Speed ​​barrier zone between Indicates the target unmanned surface vessel The velocity vector, Indicates the starting point is , direction is rays, express and Minkowski and, unmanned surface vessel The collision set, Indicates the target unmanned surface vessel The collision set, , This represents an element in A.

[0015] Optionally, based on the speed obstacle zone, a speed-feasible region for the target unmanned surface vessel can be constructed, including: A time window is introduced into the speed obstacle region to obtain a speed obstacle region with a time window; the expression for the speed obstacle region with a time window is: , This indicates a time window used to control the cycle. Indicates the center of the circle is , radius is A circle; Draw a velocity plane diagram corresponding to the velocity obstacle area with a time window, for use in escape speed calculation. The direction is perpendicular and passes through The perpendicular lines divide the velocity plane diagram, and... The region of direction is considered as the feasible region for velocity. The escape speed represents the minimum speed change of the target unmanned surface vessel as it escapes the obstacle zone.

[0016] Optionally, the expression controlling the reward function is as follows:

[0017]

[0018] in, Indicates the target unmanned surface vessel at time Collision avoidance bonus value Represents the regional reward function. This indicates the collision reward; when a target unmanned surface vessel collides with the target, the collision reward is negative.

[0019] The arrival reward indicates that when the target unmanned surface vessel (USV) reaches the desired location, the arrival reward is a positive reward. The desired location is the pre-set coordinates of the position where the USV is expected to reach.

[0020]

[0021] The vibration reward indicates a negative reward when the target unmanned surface vessel experiences vibration. Indicates an adjustable parameter. , Indicates the estimated collision time.

[0022] Secondly, this application provides a collision avoidance control device for unmanned surface vessels, comprising: The information acquisition module is used to collect the state information of the target unmanned surface vessel and construct a dynamic model of the target unmanned surface vessel based on the state information; The first region construction module is used to construct the speed obstacle region of the target unmanned surface vessel (USV) based on the pre-set collision radius and dynamic model of the target USV; when the speed of the target USV is within the speed range corresponding to the speed obstacle region, the target USV collides with other USVs. The second region construction module is used to construct the speed-feasible region of the target unmanned surface vessel (USV) based on the speed obstacle region; when the speed of the target USV is within the speed range corresponding to the speed-feasible region, the target USV will not collide with other USVs. The reward calculation module is used to construct a control reward function based on state information, speed obstacle area, and speed feasible area. The control reward function is used to calculate the reward value corresponding to different control actions, including the speed change of the target unmanned surface vessel at each moment. The collision avoidance control module is used to control the target unmanned surface vessel by utilizing the control action corresponding to the maximum reward value.

[0023] Thirdly, this application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described unmanned surface vessel collision avoidance control method.

[0024] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned unmanned surface vessel collision avoidance control method.

[0025] The above-mentioned solution in this application has the following beneficial effects: The collision avoidance control method for unmanned surface vessels (USVs) provided in this application constructs a dynamic model of the target USV based on state information, and then constructs the speed obstacle region and speed feasible region of the target USV. Based on the state information, the speed obstacle region, and the speed feasible region, a control reward function is constructed. This control reward function can select better control actions to avoid collisions and improve the collision avoidance control effect. Furthermore, this method does not depend on aquatic environmental parameters, is applicable to complex aquatic environments, and has strong generalization ability, which is beneficial for improving the collision avoidance control effect.

[0026] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 A flowchart of an unmanned surface vessel collision avoidance control method provided in an embodiment of this application; Figure 2This is a schematic diagram of the coordinate system of an unmanned surface vessel in one embodiment of this application; Figure 3a This is a schematic diagram of the speed obstacle area of ​​an unmanned surface vessel in one embodiment of this application; Figure 3b This is a schematic diagram of the feasible speed region for an unmanned surface vessel in one embodiment of this application; Figure 4 This is a comparison chart of network training results in one embodiment of this application; Figure 5 This is a schematic diagram of the collision avoidance control simulation results provided in one embodiment of this application; Figure 6 This is a schematic diagram of the structure of an unmanned surface vessel collision avoidance control device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0029] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0030] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0031] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0032] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0033] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0034] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0035] To address the poor control performance of current unmanned surface vessel (USV) collision avoidance control methods, this application provides a USV collision avoidance control method. It constructs a dynamic model of the target USV based on state information, then establishes a velocity obstacle region and a velocity feasible region for the target USV. Finally, based on the state information, the velocity obstacle region, and the velocity feasible region, a control reward function is constructed. This control reward function can select superior control actions to avoid collisions and improve collision avoidance control performance. Furthermore, this method is independent of aquatic environmental parameters, is applicable to complex aquatic environments, and possesses strong generalization ability, which further enhances collision avoidance control performance.

[0036] The collision avoidance control method for unmanned surface vessels provided in this application is illustrated below.

[0037] like Figure 1 As shown, the unmanned surface vessel collision avoidance control method provided in this application includes the following steps: Step 11: Collect the state information of the target unmanned surface vessel and construct a dynamic model of the target unmanned surface vessel based on the state information.

[0038] It should be noted that the status information of the target unmanned surface vessel (USV) can be obtained through sensors installed on the USV, or through other common methods such as the Global Positioning System (GPS) and Inertial Measurement Unit (IMU). Furthermore, the target USV can use its onboard LiDAR and monocular camera to perceive the status data of other USVs within a fixed range.

[0039] Specifically, the expression for the target unmanned surface vessel dynamics model is as follows:

[0040]

[0041] in, This represents the state vector of the target unmanned surface vessel. , This indicates the position of the target unmanned surface vessel in the world coordinate system. Indicates the yaw angle of the target unmanned surface vessel. This represents the transformation matrix from the ship's coordinate system to the world coordinate system. , This represents the velocity vector of the target unmanned surface vessel in the ship's coordinate system. , This represents the linear velocity of the target unmanned surface vessel in the ship's coordinate system. This represents the angular velocity of the target unmanned surface vessel in the ship's coordinate system. express The derivative of express The derivative of This represents the rotational inertia matrix of the target unmanned surface vessel. This represents the three-axis torque of the target unmanned surface vessel. , This represents the thrust along the U-axis of the ship's hull. This represents the thrust along the ship's v-axis. This represents the torque along the w direction. This represents the unmodeled portion of the target unmanned surface vessel's control, which includes random disturbances and surface resistance.

[0042] It should be noted that in the fields of navigation and unmanned surface vessels (USVs), a point on the Earth's surface is typically used as the origin of the world coordinate system, such as a point on the equator or a certain international standard meridian. The latitude and longitude information of the USV represents its coordinates. The origin of the hull coordinate system is usually located at the center of mass of the USV or a fixed point on the hull, and moves relative to the hull as it moves. The axes of the hull coordinate system are usually aligned with the hull's direction; see references for details. Figure 2 .

[0043] Step 12: Construct the speed obstacle zone of the target unmanned surface vessel based on the pre-set collision radius and dynamic model of the target unmanned surface vessel.

[0044] It should be noted that when the target unmanned surface vessel's speed falls within the speed limit zone, it will collide with other unmanned surface vessels. Therefore, in actual control, efforts should be made to keep the target unmanned surface vessel away from the speed limit zone as much as possible.

[0045] In the embodiments of this application, the expression for the speed barrier region is as follows:

[0046] in, Indicates the target unmanned surface vessel Other unmanned surface vessels Speed ​​barrier zone between Indicates the target unmanned surface vessel The velocity vector, Indicates the starting point is , direction is rays, express and Minkowski and, unmanned surface vessel The collision set, Indicates the target unmanned surface vessel The collision set, , This represents an element in A.

[0047] Figure 3a This is a schematic diagram of the speed obstacle area of ​​the target unmanned surface vessel in one embodiment of this application.

[0048] Step 13: Construct the speed-feasible area of ​​the target unmanned surface vessel based on the speed obstacle area.

[0049] It should be noted that when the target unmanned surface vessel's speed is within the speed range corresponding to the feasible speed zone, the target unmanned surface vessel will not collide with other unmanned surface vessels. Accordingly, in actual control, the target unmanned surface vessel should be kept within the feasible speed zone as much as possible.

[0050] Figure 3b This is a schematic diagram of the feasible speed region of the target unmanned surface vessel in one embodiment of this application.

[0051] Step 14: Construct a control reward function based on the state information, the speed obstacle area, and the speed feasible area.

[0052] It should be noted that the control reward function is used to calculate the reward value corresponding to different control actions, which include the speed change of the target unmanned surface vessel at each moment.

[0053] In the embodiments of this application, the expression for the control reward function is as follows:

[0054]

[0055] in, Indicates the target unmanned surface vessel at time Collision avoidance bonus value Represents the regional reward function. This indicates the collision reward. When a collision occurs with the target unmanned surface vessel, the collision reward is negative; otherwise, the collision reward is positive.

[0056] This indicates the arrival reward. A positive reward is given when the target unmanned surface vessel reaches the desired location; otherwise, a negative reward is given.

[0057] The vibration reward indicates that when the target unmanned surface vessel vibrates, the vibration reward is negative; otherwise, the vibration reward is positive. Indicates an adjustable parameter. , This indicates the estimated time of collision.

[0058]

[0059] The regional reward function constructed in this application will give a large positive reward when the speed of the unmanned surface vessel is in the speed feasible region; when the speed of the unmanned surface vessel is in the speed obstacle region but the collision risk is low, a small positive reward will be given, the more other unmanned surface vessels around the target unmanned surface vessel, the higher the collision risk, and vice versa; when the speed of the unmanned surface vessel is in the speed obstacle region and the collision risk is high, a negative reward will be given.

[0060] It is worth mentioning that the control reward function constructed in this application fully considers the region where the target unmanned surface vessel's speed is located (speed feasible region or speed obstacle region), the gap between the actual position and the expected position of the target unmanned surface vessel, and the safety performance of the target unmanned surface vessel itself (whether oscillation occurs), in order to obtain better control actions and improve the control effect.

[0061] It should be noted that the above-mentioned control reward function is applied to the process of deep reinforcement learning models searching for the optimal control action.

[0062] The following is an exemplary description of constructing the speed-feasible region of the target unmanned surface vessel based on the speed obstacle region in step 14.

[0063] Specifically, a time window is introduced into the speed obstacle zone to obtain a speed obstacle zone with a time window.

[0064] The expression for the velocity obstacle region with a time window is: , This indicates a time window used to control the cycle. Indicates the center of the circle is , radius is A circle.

[0065] Draw a velocity plane diagram corresponding to the velocity obstacle area with a time window, for use in escape speed calculation. The direction is perpendicular and passes through The perpendicular lines divide the velocity plane diagram, and... The region of direction is considered as the feasible region for velocity. .

[0066] Escape speed refers to the minimum speed change of the target unmanned surface vessel when escaping the speed obstacle zone.

[0067] The following is an exemplary description of the process of obtaining the optimal control action using a deep reinforcement learning model in the embodiments of this application.

[0068] First, define the control actions for controlling the target unmanned surface vessel as follows: ,in, Indicates the target unmanned surface vessel exist Controlling actions at all times Indicates the target unmanned surface vessel exist Moment velocity components The change Indicates the target unmanned surface vessel exist Moment velocity components The change in quantity.

[0069] Then, an initial policy network and an initial value network are constructed and trained.

[0070] The policy network is used to generate control actions, and the value network is used to calculate the reward value corresponding to the control actions.

[0071] Finally, the control action corresponding to the maximum reward value is taken as the optimal control action.

[0072] The following is an exemplary description of the training process for the initial policy network and the initial value network in the embodiments of this application.

[0073] Step i: Generate control actions using the initial policy network, and store the observation data (state data) of the target unmanned surface vessel after executing the control actions into the experience playback pool.

[0074] Step ii: When the data in the experience replay pool exceeds the length required for one training session, a set of observation data is randomly selected from the experience replay pool for training.

[0075] Step iii: Input the set of observation data into a bidirectional gating recurrent unit (BIGRU) to obtain the feature vector of the set of observation data. Step iv: Input the feature vector of the set of observation data into the initial policy network and the initial value network respectively, and update the network parameters of the initial policy network and the initial value network respectively until the network parameters of the initial policy network and the initial value network reach the optimal.

[0076] In the embodiments of this application, the policy network and value network have the same architecture, both containing two hidden layers, each containing 256 neurons. The learning rates of the policy network and value network are respectively... and .

[0077] Furthermore, KL divergence is introduced during network updates to constrain the difference between the new and current policies. Updates are stopped when the difference between the new and current policies becomes too large. This prevents excessively large policy updates from causing instability in the training results.

[0078] The criterion for determining whether the network parameters of the initial policy network and the initial value network have reached their optimal values ​​is that the number of parameter updates reaches the preset maximum number of updates.

[0079] In the embodiments of this application, the losses of the reward, policy network, and value network obtained through the above training steps are also compared with those of Deep Reinforcement Learning Recurrent Neural Network (DRLRNN) and Deep Reinforcement Learning Long Short-Term Memory Network (DRLLSTM), and the results are as follows. Figure 4 As shown.

[0080] Depend on Figure 4 It is evident that the training method provided in this application can achieve higher rewards and the network loss can decay more quickly.

[0081] Step 15: Use the control action corresponding to the maximum reward value to control the target unmanned surface vessel.

[0082] For example, the control action is converted into a control signal and input into the control system of the target unmanned surface vessel to achieve control of the target unmanned surface vessel. As can be seen from the previous text, using the control action corresponding to the maximum reward value to control the target unmanned surface vessel minimizes the risk of collision.

[0083] To verify the effectiveness of the unmanned surface vessel collision avoidance control method provided in this application, numerical simulation experiments and virtual reality scene experiments were also conducted.

[0084] The unmanned surface vessel (USV) had a radius of 2 meters; its perception range was 5 meters; the map size was a 10m x 10m square; and the control cycle was 0.01 seconds. The upper bound of the KL divergence was set to 0.01. The training process lasted 500 rounds, with the target USV interacting with the environment 1000 steps in each round. Simulation results are as follows: Figure 5 As shown.

[0085] Depend on Figure 5 As can be seen, the unmanned surface vessel (USV) collision avoidance control method provided in this application can enable each USV to reach the desired position without collision, thus verifying the effectiveness of the USV collision avoidance control method provided in this application.

[0086] In summary, the collision avoidance control method for unmanned surface vessels (USVs) provided in this application constructs a dynamic model of the target USV based on state information, then constructs the speed obstacle region and speed feasible region of the target USV, and finally constructs a control reward function based on the state information, speed obstacle region, and speed feasible region. This method can select better control actions, avoid collisions with the USV, and improve the collision avoidance control effect. In addition, this method does not depend on aquatic environmental parameters, can be applied to complex aquatic environments, and has strong generalization ability, which is conducive to improving the collision avoidance control effect.

[0087] The following is an exemplary description of the unmanned surface vessel collision avoidance control device provided in this application.

[0088] like Figure 6 As shown, the unmanned surface vessel collision avoidance control device 600 includes: The information acquisition module 601 is used to acquire the state information of the target unmanned surface vessel and construct a dynamic model of the target unmanned surface vessel based on the state information; The first region construction module 602 is used to construct a speed obstacle region for the target unmanned surface vessel (USV) based on the pre-set collision radius and dynamic model of the target USV; when the speed of the target USV is within the speed range corresponding to the speed obstacle region, the target USV collides with other USVs. The second region construction module 603 is used to construct a speed-feasible region for the target unmanned surface vessel based on the speed obstacle region; when the speed of the target unmanned surface vessel is within the speed range corresponding to the speed-feasible region, the target unmanned surface vessel will not collide with other unmanned surface vessels. The reward calculation module 604 is used to construct a control reward function based on the state information, the speed obstacle area, and the speed feasible area. The control reward function is used to calculate the reward value corresponding to different control actions, including the speed change of the target unmanned surface vessel at each moment. The collision avoidance control module 605 is used to control the target unmanned surface vessel by utilizing the control action corresponding to the maximum reward value.

[0089] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0090] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0091] like Figure 7 As shown, embodiments of this application provide a terminal device, such as... Figure 7 As shown, the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 7 The diagram shows only one processor, a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 executes the computer program D102 to implement the steps in any of the above method embodiments.

[0092] Specifically, when the processor D100 executes the computer program D102, it collects the state information of the target unmanned surface vessel (USV) and constructs a dynamic model of the USV based on the state information. Then, based on the pre-set collision radius of the USV and the dynamic model, it constructs a speed obstacle region for the USV, and then constructs a speed feasible region for the USV based on the speed obstacle region. Subsequently, it constructs a control reward function based on the state information, the speed obstacle region, and the speed feasible region. Finally, it uses the control action corresponding to the maximum reward value to control the USV. The method of constructing a dynamic model of the USV based on the state information, then constructing the speed obstacle region and the speed feasible region, and finally constructing a control reward function based on the state information, the speed obstacle region, and the speed feasible region, can select better control actions, avoid collisions, and improve collision avoidance control performance. Furthermore, this method does not depend on aquatic environmental parameters, is applicable to complex aquatic environments, and has strong generalization ability, which is beneficial for improving collision avoidance control performance.

[0093] The processor D100 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0094] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device D10. Furthermore, the memory D101 may include both internal and external storage units of the terminal device D10. The memory D101 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory D101 can also be used to temporarily store data that has been output or will be output.

[0095] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0096] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0097] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to the unmanned surface vessel collision avoidance control device / terminal device, recording media, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0098] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0099] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0100] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0101] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0102] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A collision avoidance control method for unmanned surface vessels, characterized in that, include: Collect the state information of the target unmanned surface vessel and construct a dynamic model of the target unmanned surface vessel based on the state information; Based on the pre-set collision radius of the target unmanned surface vessel (USV) and the dynamic model, a speed obstacle zone for the USV is constructed; when the speed of the USV is within the speed range corresponding to the speed obstacle zone, the USV collides with other USVs. Based on the speed obstacle area, a speed feasible region for the target unmanned surface vessel is constructed; when the speed of the target unmanned surface vessel is within the speed range corresponding to the speed feasible region, the target unmanned surface vessel will not collide with other unmanned surface vessels. Based on the state information, the speed obstacle region, and the speed feasible region, a control reward function is constructed; the control reward function is used to calculate the reward value corresponding to different control actions, and the control actions include the speed change of the target unmanned surface vessel at each moment. The target unmanned surface vessel is controlled using the control action corresponding to the maximum reward value; The expression for the control reward function is as follows: in, Indicates that the target unmanned surface vessel is at time Collision avoidance bonus value Represents the regional reward function. This indicates a collision reward; when the target unmanned surface vessel collides with another object, the collision reward is negative. The arrival reward indicates that a positive reward is given when the target unmanned surface vessel (USV) reaches the desired location. The desired location is a pre-set coordinate system indicating the coordinates at which the target USV is expected to reach. The vibration reward is negative when the target unmanned surface vessel vibrates. Indicates an adjustable parameter. , Indicates the estimated collision time. 。 2. The unmanned surface vessel collision avoidance control method according to claim 1, characterized in that, The expression for the dynamic model is as follows: in, This represents the state vector of the target unmanned surface vessel. , This indicates the position of the target unmanned surface vessel in the world coordinate system. This indicates the yaw angle of the target unmanned surface vessel. This represents the transformation matrix from the ship's coordinate system to the world coordinate system. , This represents the velocity vector of the target unmanned surface vessel in the ship's coordinate system. , This represents the linear velocity of the target unmanned surface vessel in the ship's coordinate system. This represents the angular velocity of the target unmanned surface vessel in the ship's coordinate system. express The derivative, express The derivative of This represents the rotational inertia matrix of the target unmanned surface vessel. This represents the three-axis torque of the target unmanned surface vessel. , This represents the thrust along the U-axis of the hull. This represents the thrust along the ship's v-axis. This represents the torque along the w direction. This refers to the unmodeled portion of the target unmanned surface vessel's control, which includes random disturbances and surface resistance.

3. The unmanned surface vessel collision avoidance control method according to claim 2, characterized in that, The expression for the speed barrier region is as follows: in, Indicates the target unmanned surface vessel and unmanned surface vessels Speed ​​barrier zone between Indicates the target unmanned surface vessel The velocity vector, Indicates the starting point is , direction is rays, express and Minkowski and, unmanned surface vessel The collision set, Indicates the target unmanned surface vessel The collision set, , This represents an element in A.

4. The unmanned surface vessel collision avoidance control method according to claim 3, characterized in that, Constructing the speed-feasible region of the target unmanned surface vessel based on the speed obstacle region includes: A time window is introduced into the speed obstacle region to obtain a speed obstacle region with a time window; wherein, the expression for the speed obstacle region with a time window is: , This indicates a time window used to control the cycle. Indicates the center of the circle is , radius is A circle; Draw a velocity plane diagram corresponding to the velocity obstacle area with the time window, for use in determining escape speed. The direction is perpendicular and passes through The perpendicular line divides the velocity plane diagram, and... The region of direction is used as the feasible region of velocity. Wherein, the escape speed represents the minimum speed change of the target unmanned surface vessel as it escapes the obstacle zone.

5. A collision avoidance control device for unmanned surface vessels, characterized in that, include: The information acquisition module is used to acquire the state information of the target unmanned surface vessel and construct a dynamic model of the target unmanned surface vessel based on the state information. The first region construction module is used to construct a speed obstacle region for the target unmanned surface vessel (USV) based on the pre-set collision radius and the dynamic model; when the speed of the target USV is within the speed range corresponding to the speed obstacle region, the target USV collides with other USVs. The second region construction module is used to construct a speed-feasible region for the target unmanned surface vessel based on the speed obstacle region; when the speed of the target unmanned surface vessel is within the speed range corresponding to the speed-feasible region, the target unmanned surface vessel will not collide with other unmanned surface vessels. The reward calculation module is used to construct a control reward function based on the state information, the speed obstacle area, and the speed feasible area; the control reward function is used to calculate the reward value corresponding to different control actions, and the control actions include the speed change of the target unmanned surface vessel at each moment; The collision avoidance control module is used to control the target unmanned surface vessel using the control action corresponding to the maximum reward value; The expression for the control reward function is as follows: in, Indicates that the target unmanned surface vessel is at time Collision avoidance bonus value Represents the regional reward function. This indicates a collision reward; when the target unmanned surface vessel collides with another object, the collision reward is negative. The arrival reward indicates that when the target unmanned surface vessel (USV) reaches the desired position, the arrival reward is a positive reward. The desired position is a pre-set coordinate system indicating the coordinates at which the target USV is expected to reach. The vibration reward is negative when the target unmanned surface vessel vibrates. Indicates an adjustable parameter. , Indicates the estimated collision time. 。 6. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the unmanned surface vessel collision avoidance control method as described in any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the unmanned surface vessel collision avoidance control method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Collision avoidance decision-making method for multiple unmanned ships

    CN115718497A