Ship automatic berthing method and system based on deep reinforcement learning
Through the method based on deep reinforcement learning, an automatic control model is built, which solves the complexity of ship berthing operations in complex seaport environments, realizes automated control, and improves the safety and accuracy of ship berthing operations.
Patent Information
- Application Number
- CN202510144415.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively deal with the complex and changeable seaport environment, resulting in complex operation of ship berths and difficult to automatically control.
Using a deep reinforcement learning method, an automatic control model is constructed by training a neural network, using port maps and obstacle data, combining two Critic networks and an Actor network to realize automated control of ship berthing processes.
It has achieved the ability to dynamically adapt to complex environments and has automated decision-making capabilities, which has significantly improved the safety and accuracy of ship mooring operations.
Smart Images

Figure CN119987375A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ship automation control, and in particular to a ship automatic mooring method and system based on deep reinforcement learning. Background Art
[0002] With the advancement of globalization, the number of ships in large seaports has increased year by year, making ship berthing operations more and more difficult. Berthing is one of the most complex operations in ship control, requiring rapid decision-making and processing of multiple dynamic parameters. Existing methods mostly use traditional control algorithms, which are difficult to cope with the complex and changing seaport environment. Summary of the invention
[0003] In view of the above technical problems, the present invention provides a method and system for automatic ship berthing based on deep reinforcement learning, which realizes automatic control of the ship berthing process by training neural networks.
[0004] Other features and advantages of the present invention will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.
[0005] The present invention discloses a method for automatic ship mooring based on deep reinforcement learning, the method comprising:
[0006] An automatic control model is constructed based on deep reinforcement learning training. The environment of the automatic control model is based on the port map and combines the location data of static and dynamic obstacles. Its state includes the current position, speed, and heading of the ship. Its actions include adjusting the speed and heading of the ship. The automatic control model also includes two Critic networks and one Actor network.
[0007] Acquire the current coordinates, target coordinates and speed of the ship in the environment in real time, and calculate the safe distance at which the ship can approach the target coordinates, wherein the safe distance is the maximum approach distance at which the ship has enough space to reduce the speed to the required speed and maneuver during the mooring period;
[0008] When the distance between the ship and the target coordinates is less than or equal to the safety distance, the current state information of the ship is read, and the state information includes the current coordinates, speed, obstacle data, and heading. The speed and heading to be taken at the current moment are generated through the two Critic networks and the Actor network, and input into the ship controller to execute the set speed and heading.
[0009] Furthermore, the automatic control model, when being trained, includes:
[0010] For a given environment, the initial coordinates and target coordinates are randomly assigned, and the optimal route from the initial coordinates to the target coordinates is drawn based on the D-star algorithm or expert experience, and the image is saved as a reference sample for selecting the route and speed;
[0011] The Critic network is learned based on the reference sample. During learning, operations are performed along the optimal route of the reference sample, and the result of each execution is recorded. If deviations or errors occur from the optimal route during execution, back propagation is performed to adjust the weights of the Critic network.
[0012] Furthermore, when calculating the safety distance, it includes:
[0013] Based on the dynamic equation, the relationship between the ship's speed and time is obtained, and the variables are separated and integrated to obtain the relationship between the speed and time;
[0014] Obtaining a braking time based on the change relationship, and integrating the speed based on the braking time to obtain a braking distance;
[0015] The total distance required for the ship to stop from the current speed is calculated, and the braking distance is subtracted from the total distance to obtain the safety distance.
[0016] Furthermore, the method further comprises:
[0017] After the two Critic networks generate the speed and heading respectively, the motion of the ship is predicted and described based on the motion mathematical model, which includes the position and heading of the ship, the rotation matrix, the mass matrix, the control input matrix, the control input vector, the resistance vector, and the external interference force vector.
[0018] Furthermore, the structure of the Critic network specifically includes:
[0019] A Q network, the Q network is used to approximate a Q value function, and after inputting the state information, outputs the Q value of each action under the state information;
[0020] An experience replay mechanism, which is used to store interaction data with the environment;
[0021] A target network, which is a copy of the Q network, is used to stabilize the training process.
[0022] Furthermore, the Critic network and the Actor network, when being trained, include:
[0023] Initialize the Critic network and the Actor network, initialize the parameters of the target network, observe the current state of the environment, select an action, execute an action, store state transfer, check whether to update the network, calculate the target action and target value, update the Critic network, delay updating the Actor network, update the parameters of the target network, and determine whether the training is completed.
[0024] According to another aspect of the present invention, a ship automatic mooring system based on deep reinforcement learning is provided, the system comprising:
[0025] A model generation module is used to build an automatic control model based on deep reinforcement learning training. The environment of the automatic control model is based on the port map and combines the location data of static and dynamic obstacles. Its state includes the current position, speed, and heading of the ship. Its actions include adjusting the speed and heading of the ship. The automatic control model also includes two Critic networks and one Actor network.
[0026] A state monitoring module is used to obtain the current coordinates, target coordinates and speed of the ship in the environment in real time, and calculate the safe distance at which the ship can approach the target coordinates, where the safe distance is the maximum approach distance at which the ship has enough space to reduce the speed to the required speed and maneuver during the mooring period;
[0027] The automatic control module is used to read the current state information of the ship when the distance between the ship and the target coordinates is less than or equal to the safety distance. The state information includes the current coordinates, speed, obstacle data, and heading. The speed and heading to be taken at the current moment are generated through the two Critic networks and the Actor network, and input into the ship controller to execute the set speed and heading.
[0028] The technical solution disclosed in this disclosure has the following beneficial effects:
[0029] Dynamically adapt to complex environments: can quickly respond to dynamically changing environments and adapt to complex port berthing needs;
[0030] Strong automated decision-making capabilities: Based on the Actor network and dual-Critic network, it has the ability to make automatic decisions in unknown environments, reducing human operational errors;
[0031] Efficient path planning: It is efficient and robust in path planning, which can significantly improve the safety and accuracy of berthing operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 This is a flow chart of a method for automatic ship mooring based on deep reinforcement learning in an embodiment of this specification;
[0033] Figure 2 This is a structural block diagram of a ship automatic mooring system based on deep reinforcement learning in an embodiment of this specification;
[0034] Figure 3 A device for executing a method for automatic ship mooring based on deep reinforcement learning in an embodiment of this specification;
[0035] Figure 4 A computer-readable storage medium storing a method for automatic ship mooring based on deep reinforcement learning in an embodiment of this specification. DETAILED DESCRIPTION
[0036] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein; on the contrary, these embodiments are provided so that the present disclosure will be more comprehensive and complete, and the concepts of the example embodiments are fully conveyed to those skilled in the art. The described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, systems, steps, etc. may be adopted. In other cases, known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0037] In addition, the accompanying drawings are only schematic illustrations of the present disclosure. The same reference numerals in the figures represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor systems and / or microcontroller systems.
[0038] like Figure 1 As shown, the embodiment of this specification provides a method for automatic ship mooring based on deep reinforcement learning, which is applied to a ship communication system, wherein the ship communication system includes a plurality of ships with integrated wireless modules and a server for coordinating ship communication spectrum, and the server may specifically be a base station. The method may specifically include the following steps S101 to S103:
[0039] In step S101, an automatic control model is constructed based on deep reinforcement learning training. The environment of the automatic control model is based on a port map and combined with location data of static and dynamic obstacles. Its state includes the current position, speed, and heading of the ship. Its actions include adjusting the speed and heading of the ship. The automatic control model also includes two Critic networks and one Actor network. The two Critic networks are used to estimate the value function of the current state-action pair to reduce the bias of overestimation, and the Actor network is used to generate the best action for the current state.
[0040] Among them, the port map can provide real spatial restrictions and boundary adjustments. In the automatic control model, static obstacles can be fixed facilities, such as fixed positions in the port, and dynamic obstacles can be moving ships, which are provided by the positioning systems of other moving ships. The speed of the ship can be longitudinal speed, lateral speed and angular velocity. The specific actions are the propeller speed control and rudder angle control of the ship. The Actor network is responsible for generating actions, and the two Critic networks can calculate the dual value function and select the minimum value to improve the stability of training, that is, they are responsible for evaluating the value of these actions. Through the above training, the construction of an automatic control model with the port environment as the core, combined with obstacle data and reinforcement learning, was completed.
[0041] The structure of the Critic network specifically includes: a Q network, which is used to approximate the Q value function and output the Q value of each action under the state information after inputting the state information; an experience replay mechanism, which is used to store interaction data with the environment; and a target network, which is a copy of the Q network and is used to stabilize the training process.
[0042] Among them, the automatic control model is trained on two critic networks. and an Actor network μ θ During training, the Critic network estimates the state-action value Q(s, a) to judge the quality of an action, and the Actor network generates action a, which is the optimal strategy to execute under the current state. 2 , θ is the random initial weight of the network, and the target network is a delayed copy of the Critic and Actor networks, which is used to stabilize the training process. During the training process, it mainly includes: initializing the Critic and Actor networks, initializing the target network parameters, observing the current state S of the environment, selecting action a, executing action a, storing state transfer, checking whether to update the network, calculating the target action and target value, updating the Critic network, delaying the update of the Actor network, updating the target network parameters, and judging whether the training is finished, among which:
[0043] In the current state S of the observed environment, obtain the state information of the ship, such as the current position, speed, heading, etc.
[0044] In the selection action a, a=clip(μ θ (s)+∈,a low , a high ), Actor network μ θ Generate the optimal action in the current state, add noise ∈~N to enhance exploration, and limit the action range to [a low , a high ].
[0045] In executing action a, perform action a in the environment and obtain: new state s′, reward r, and whether to end signal d.
[0046] In the storage state transfer, based on the experience replay mechanism, (s, a, r, s′, d) is stored in the experience replay buffer D for subsequent training.
[0047] When checking whether to update the network, a test is performed. If the update conditions are met, batch training is started and a batch of data B is randomly sampled from the buffer D.
[0048] In calculating the target action and target value, the target action a′(s′) is generated based on the target Actor network and regularized noise is added. The target value y is:
[0049] Among them, γ is a discount factor, which is used to measure the importance of future rewards, and then the minimum value in the dual critic network is used to reduce the over-estimation bias.
[0050] In updating the Critic network, the target value y and the estimated value are minimized by the gradient descent method. The error between:
[0051]
[0052] In the delayed update Actor network, the Actor network is updated every certain number of time steps. During the update, the policy gradient is calculated to obtain:
[0053]
[0054] In updating the target network parameters, the target network approaches the main network parameters in a soft update manner:
[0055] φ targ,i ←ρφ targ,i +(1-ρ)φ i , i=1,2;
[0056] θtarg ←ρθ targ +(1-ρ)θ;
[0057] Where ρ is the soft update coefficient, which is usually set to a small value (such as 0.005).
[0058] In determining whether training is finished, if the maximum number of training time steps is reached or the model converges, the training is finished.
[0059] In the above, the model is strengthened by training in the continuous action space and the over-estimation bias is reduced by the dual critic network to enhance stability. Specifically, the Actor network is used to generate actions, the Critic network is used to evaluate the action quality, and the replay buffer and target network are used to stabilize the training process.
[0060] In step S102, the current coordinates, target coordinates and speed of the ship in the environment are obtained in real time, and the safe distance at which the ship can approach the target coordinates is calculated, and the safe distance is the maximum approach distance at which the ship has enough space to reduce the speed to the required speed and maneuver during mooring. The current coordinates, target coordinates and speed of the ship in the environment are obtained in real time, and the safe distance at which the ship can approach the target coordinates is calculated, and the safe distance is the maximum approach distance at which the ship has enough space to reduce the speed to the required speed and maneuver during mooring. The current coordinates, target coordinates and speed of the ship in the environment are obtained in real time, and the safe distance at which the ship can approach the target coordinates is calculated, and the safe distance is the maximum approach distance at which the ship has enough space to reduce the speed to the required speed and maneuver during mooring.
[0061] After the model is trained, it can be combined with the ship's navigation system or sensor to perform mooring navigation. The first step is to obtain the ship's real-time status and target information. The current coordinates are obtained through sensors or navigation systems. The coordinates of the ship in the environment are expressed as (x current ,y current ); The current speed is obtained through the sensor, including the longitudinal speed, lateral speed, and angular speed of the ship, and then the combined speed of the ship is calculated; the target information includes the target coordinates and target speed. The target coordinates are expressed as (x target ,y target ), provided by the map in the environment, the target speed is the low speed required for berthing, usually close to zero. The second step is to calculate the distance from the ship to the target, which can be calculated by calculating the Euclidean distance:
[0062]
[0063] The Euclidean distance defines the straight-line distance between the ship and the target point and is the basis for deciding whether to enter the speed reduction zone.
[0064] The safety distance is the minimum distance required for the ship to decelerate from the current speed to the target speed, plus the maneuvering space required for berthing adjustment, to ensure that the ship can safely decelerate and adjust its position and direction during berthing to avoid collision. Assuming that the ship follows a uniform deceleration motion during deceleration, the basic deceleration formula is:
[0065]
[0066] Among them, v current is the current speed of the ship, v target is the target speed, a safe is the safety deceleration, which depends on the ship dynamic parameters, d safe is the basic distance required for deceleration. By transforming the basic formula, we get:
[0067]
[0068] After obtaining the basic distance required for deceleration, the course, speed, etc. can be adjusted by monitoring the status of the ship.
[0069] Additionally, an adjustment distance needs to be added to the basic distance to make it safer, wherein the adjustment distance is determined based on the complexity of parking (such as the number and dynamics of obstacles).
[0070] In step S103, when the distance between the ship and the target coordinates is less than or equal to the safety distance, the current state information of the ship is read, and the state information includes the current coordinates, speed, obstacle data, and heading. The speed and heading to be taken at the current moment are generated through the two Critic networks and the Actor network, and input into the ship controller to execute the set speed and heading.
[0071] Among them, the Actor network is a strategy network, which generates the control action of the ship according to the current state, including: speed control, which is used to adjust the propeller speed n, and heading control, which is used to adjust the rudder angle δ. The formula is a=(n, δ)=μ θ (s). The two Critic networks are value function networks used to evaluate the quality of the action in the current state (i.e., Q value). The dual network structure of the Critic network is designed to reduce estimation bias and improve the stability of the strategy. The goal is to maximize the Q value and guide the Actor to optimize:
[0072] Q(s, a)=r+γ·Q(s′, a′);
[0073] r is the current reward, γ is the discount factor, Q(s′, a′) is the Q value of the next state, s is the state, and a is the action. The Q value is calculated by two Critic networks, and the minimum value is taken to guide the Actor network optimization.
[0074] The Actor network generates speed and heading actions according to the state, which are input into the ship controller as control signals to adjust the propeller speed and rudder angle to achieve speed and heading control of the ship.
[0075] In one embodiment, when training, the automatic control model includes: randomly assigning initial coordinates and target coordinates to a given environment, drawing the optimal route from the initial coordinates to the target coordinates based on the D-star algorithm or expert experience, and saving the image as a reference sample for selecting the route and speed; learning the Critic network based on the reference sample, and during learning, performing operations along the optimal route of the reference sample, and recording the results of each execution. If a deviation or error occurs from the optimal route during execution, back propagation is performed to adjust the weights of the Critic network.
[0076] As explained, the training of the automatic control model can be achieved by randomly generating initial coordinates and target coordinates to simulate different operation scenarios, which can enhance the model's adaptability to diverse tasks. In each simulated environment, the model uses the D-star algorithm or relies on expert experience to generate the optimal path from the initial position to the target position. Among them, the D-star algorithm is a dynamic path planning method that can efficiently find feasible paths in complex environments while responding to the dynamic changes of obstacles. The generated path is saved in the form of an image together with the environmental data as an important reference for subsequent training of the Critic network. This reference sample not only provides a visual representation of the optimal path, but also provides guidance for the path and speed selection of the ship in a complex environment. Based on these reference samples, the Critic network begins to learn how to evaluate the value of actions under different states. During the learning process, the ship is guided to operate along the optimal route. This guidance ensures that the initial learning process of the ship is efficient and stable. After each operation, the ship records the current state, action and its feedback results, including whether it deviates from the optimal path. In this way, the ship can attribute feedback to the success and failure of path selection, thereby continuously adjusting its strategy. When the ship's operation deviates from the optimal path in the reference sample or produces erroneous behavior, the training process triggers an adjustment mechanism. Through the back-propagation algorithm, the deviation is passed as a feedback signal to the Critic network to update its internal weights. This adjustment mechanism ensures that the Critic network can better evaluate the value of actions in similar scenarios in the future and avoid similar deviations from happening again. Through this iterative optimization, the Critic network gradually improves its ability to judge dynamic paths in complex environments.
[0077] In one embodiment, when calculating the safety distance, it includes: based on the dynamic equation, obtaining the relationship between the speed of the ship and the time, separating the variables and integrating them to obtain the relationship between the speed and time; obtaining the braking time based on the changing relationship, integrating the speed based on the braking time to obtain the braking distance; calculating the total distance required for the ship to stop from the current speed, subtracting the braking distance from the total distance to obtain the safety distance.
[0078] As an explanation, in this embodiment, the safe distance of the ship is calculated based on its dynamic characteristics. First, the relationship between the ship's speed and time is described by the ship's dynamic equation. The dynamic equation takes into account factors such as the resistance, propulsion and inertia force of the ship in the water, thereby accurately describing how the ship's speed changes over time. In order to simplify the calculation, the speed and time relationship is separated by variable operation, and then the separated variable equation is integrated to derive an analytical expression for the change of the ship's speed over time. This process lays a mathematical foundation for the subsequent calculation of the braking time. Using the obtained relationship between the speed and time, the time required for the ship to decelerate from the current speed to stop, that is, the braking time, is determined. After the braking time is determined, the speed function is integrated again to calculate the displacement experienced by the ship during the braking process, that is, the braking distance. The core is to accumulate the distance information through the relationship between the speed and time, which reflects the distance requirement of the ship from kinetic energy to static state. Finally, the total distance required for the ship to stop completely from the current speed is taken as the starting point, and the braking distance calculated above is deducted to obtain the safe distance. Safety distance refers to the additional space that a ship needs to leave outside the braking distance to ensure safe speed adjustment and attitude control during berthing or obstacle avoidance.
[0079] In one embodiment, the method further includes: after the two Critic networks generate the speed and heading respectively, predicting and describing the motion of the ship based on a motion mathematical model, wherein the motion mathematical model includes the position and heading of the ship, a rotation matrix, a mass matrix, a control input matrix, a control input vector, a resistance vector, and an external interference force vector.
[0080] In this embodiment, not only the speed and heading generated by the two Critic networks are relied upon, but also the motion of the ship is predicted and described by introducing a motion mathematical model to further improve the accuracy and reliability of the model. The motion mathematical model comprehensively describes the dynamic behavior of the ship, and specifically includes the following parts: First, the position and heading of the ship are used as state variables to characterize the spatial state of the ship in real time. Specifically, the position and heading states characterize the position (x, y) and heading angle δ of the ship in the ground coordinate system. These states, combined with the generated speed and heading signals, provide initial conditions for subsequent motion prediction. By introducing a rotation matrix, the coordinate transformation from the hull coordinate system to the ground coordinate system can be achieved, so that the model can accurately predict the path of the ship in the actual space. The rotation matrix is expressed as:
[0081]
[0082] The mass matrix takes into account the inertial characteristics and added mass effect of the ship, which is crucial in ship dynamics because the acceleration and deceleration of the ship are not only controlled by thrust, but also significantly affected by inertial drag. With the control input matrix and control input vector, the propeller thrust and rudder angle signals can be converted into the actual power output of the ship, which directly affects the speed and heading adjustment of the ship. The mass matrix is expressed as:
[0083]
[0084] Among them, m ij It represents the mass of the ship and the additional mass, reflecting the influence of inertia force on acceleration.
[0085] Control input matrix and control input vector. The control input is realized by rudder angle and propeller thrust. The formula is:
[0086]
[0087] Where X n , Y δ 、N δ Represents the contribution of the propeller and rudder to the longitudinal force, lateral force and moment.
[0088] The drag vector describes the effect of drag on the motion of the ship:
[0089]
[0090] Among them, X d , Y d 、N d Represents the resistance terms in the longitudinal, lateral and rotational directions.
[0091] The external interference force vector comprehensively considers the resistance (such as water flow and air resistance) and environmental forces (such as wind and wave action) encountered by the ship when it moves in the water. The existence of these external factors makes the motion prediction of the ship closer to the actual operation scenario. The external interference force vector is expressed as:
[0092]
[0093] Among them, τ x , τ y , τ ψ It indicates the external force acting on the ship in the longitudinal, transverse and rotational directions.
[0094] Taking the above factors into consideration, the dynamic model of the ship is expressed as:
[0095]
[0096] in, is the acceleration vector, v is the ship velocity vector, and C(v) is the Coriolis force matrix, which takes into account the effect of ship rotation on inertia force.
[0097] The ship's dynamic model is used to describe the dynamic behavior of the ship under control. By inputting the speed and heading generated by the Critic network into the dynamic model, the future state of the ship can be predicted to ensure the rationality and stability of the control signal in a dynamic and complex environment. This combines the decision-making ability of reinforcement learning with the accuracy of the dynamic model, effectively improving the reliability of ship path planning and control tasks.
[0098] By combining the above elements, the motion mathematical model can predict the future position and heading of the ship under a specific control signal and provide an accurate description of the ship's motion. This prediction result is not only used to verify the rationality of the currently generated speed and heading, but also can be used as a basis for subsequent action optimization and path adjustment. In this way, the method further enhances the accuracy and reliability of ship path planning and control in a dynamic and complex environment, providing technical support for the efficient completion of safe berthing tasks.
[0099] Based on the same idea, Figure 2 As shown, the embodiment of this specification also provides a ship automatic mooring system based on deep reinforcement learning, the system comprising:
[0100] The model generation module 201 is used to build an automatic control model based on deep reinforcement learning training. The environment of the automatic control model is based on the port map and combines the location data of static and dynamic obstacles. Its state includes the current position, speed, and heading of the ship. Its action includes adjusting the speed and heading of the ship. The automatic control model also includes two Critic networks and one Actor network.
[0101] The state monitoring module 202 is used to obtain the current coordinates, target coordinates and speed of the ship in the environment in real time, and calculate the safe distance at which the ship can approach the target coordinates, where the safe distance is the maximum approach distance at which the ship has enough space to reduce the speed to the required speed and maneuver during the mooring period;
[0102] The automatic control module 203 is used to read the current state information of the ship when the distance between the ship and the target coordinates is less than or equal to the safety distance. The state information includes the current coordinates, speed, obstacle data, and heading. The speed and heading to be taken at the current moment are generated through the two Critic networks and the Actor network, and input into the ship controller to execute the set speed and heading.
[0103] The advantages of the system are as follows:
[0104] Dynamically adapt to complex environments: can quickly respond to dynamically changing environments and adapt to complex port berthing needs;
[0105] Strong automated decision-making capabilities: Using deep Q-learning algorithms, it has the ability to make automatic decisions in unknown environments, reducing human operational errors;
[0106] Efficient path planning: It is efficient and robust in path planning, which can significantly improve the safety and accuracy of berthing operations.
[0107] Based on the same idea, the embodiment of this specification also provides a ship automatic mooring device based on deep reinforcement learning, such as Figure 3 shown.
[0108] The automatic ship mooring equipment based on deep reinforcement learning may be the terminal device or server provided in the above-mentioned embodiment.
[0109] The automatic mooring equipment for ships based on deep reinforcement learning may have relatively large differences due to different configurations or performances, and may include one or more processors 301 and memory 302, and one or more storage applications or data may be stored in the memory 302. Among them, the memory 302 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) and / or a cache memory unit, and may further include a read-only storage unit. The application stored in the memory 302 may include one or more program modules (not shown in the figure), such program modules include but are not limited to: an operating system, one or more applications, other program modules and program data, each of these examples or some combination may include the implementation of a network environment. Furthermore, the processor 301 may be configured to communicate with the memory 302 to execute a series of computer executable instructions in the memory 302 on the automatic mooring equipment for ships based on deep reinforcement learning. The automatic mooring device for ships based on deep reinforcement learning may also include one or more power supplies 303, one or more wired or wireless network interfaces 304, one or more I / O interfaces (input and output interfaces) 305, one or more external devices 306 (such as keyboards, pointing devices, Bluetooth devices, etc.), and may also communicate with one or more devices that enable a user to interact with the device, and / or communicate with any device that enables the device to communicate with one or more other computing devices (such as routers, modems, etc.). Such communication may be performed through the I / O interface 305. In addition, the device may also communicate with one or more networks (such as a local area network (LAN)) through the wired or wireless interface 304.
[0110] Specifically, in this embodiment, the automatic mooring device for ships based on deep reinforcement learning includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer executable instructions for the automatic mooring device for ships based on deep reinforcement learning, and the one or more programs are configured to be executed by one or more processors, including the following computer executable instructions:
[0111] An automatic control model is constructed based on deep reinforcement learning training. The environment of the automatic control model is based on the port map and combines the location data of static and dynamic obstacles. Its state includes the current position, speed, and heading of the ship. Its actions include adjusting the speed and heading of the ship. The automatic control model also includes two Critic networks and one Actor network.
[0112] Acquire the current coordinates, target coordinates and speed of the ship in the environment in real time, and calculate the safe distance at which the ship can approach the target coordinates, wherein the safe distance is the maximum approach distance at which the ship has enough space to reduce the speed to the required speed and maneuver during the mooring period;
[0113] When the distance between the ship and the target coordinates is less than or equal to the safety distance, the current state information of the ship is read, and the state information includes the current coordinates, speed, obstacle data, and heading. The speed and heading to be taken at the current moment are generated through the two Critic networks and the Actor network, and input into the ship controller to execute the set speed and heading.
[0114] Based on the same idea, the exemplary embodiment of the present disclosure also provides a computer-readable storage medium on which a program product capable of implementing the above method of the present specification is stored. In some possible implementations, various aspects of the present disclosure can also be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Automatic mooring method for ships based on deep reinforcement learning" section of the present specification.
[0115] refer to Figure 4 As shown, a program product 400 for implementing the above method according to an exemplary embodiment of the present disclosure is described, which can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, system or device.
[0116] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, system or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0117] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, wherein readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, system, or device.
[0118] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0119] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0120] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal system, or a network device, etc.) to execute the method according to the exemplary implementation of the present disclosure.
[0121] In addition, the above-mentioned figures are only schematic illustrations of the processes included in the method according to the exemplary embodiment of the present disclosure, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.
[0122] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the exemplary embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0123] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.
Claims
1. A method for automatic ship mooring based on deep reinforcement learning, characterized in that: The method comprises: An automatic control model is constructed based on deep reinforcement learning training. The environment of the automatic control model is based on the port map and combines the location data of static and dynamic obstacles. Its state includes the current position, speed, and heading of the ship. Its actions include adjusting the speed and heading of the ship. The automatic control model also includes two Critic networks and one Actor network. Acquire the current coordinates, target coordinates and speed of the ship in the environment in real time, and calculate the safe distance at which the ship can approach the target coordinates, wherein the safe distance is the maximum approach distance at which the ship has enough space to reduce the speed to the required speed and maneuver during the mooring period; When the distance between the ship and the target coordinates is less than or equal to the safety distance, the current state information of the ship is read, and the state information includes the current coordinates, speed, obstacle data, and heading. The speed and heading to be taken at the current moment are generated through the two Critic networks and the Actor network, and input into the ship controller to execute the set speed and heading.
2. According to claim 1, a method for automatic ship mooring based on deep reinforcement learning is characterized in that: The automatic control model, when being trained, includes: For a given environment, the initial coordinates and target coordinates are randomly assigned, and the optimal route from the initial coordinates to the target coordinates is drawn based on the D-star algorithm or expert experience, and the image is saved as a reference sample for selecting the route and speed; The Critic network is learned based on the reference sample. During learning, operations are performed along the optimal route of the reference sample, and the result of each execution is recorded. If deviations or errors occur from the optimal route during execution, back propagation is performed to adjust the weights of the Critic network.
3. The method for automatic ship mooring based on deep reinforcement learning according to claim 1, characterized in that: When calculating the safety distance, include: Based on the dynamic equation, the relationship between the ship's speed and time is obtained, and the variables are separated and integrated to obtain the relationship between the speed and time; Obtaining a braking time based on the change relationship, and integrating the speed based on the braking time to obtain a braking distance; The total distance required for the ship to stop from the current speed is calculated, and the braking distance is subtracted from the total distance to obtain the safety distance.
4. The method for automatic ship mooring based on deep reinforcement learning according to claim 1, characterized in that: The method further comprises: After the two Critic networks generate the speed and heading respectively, the motion of the ship is predicted and described based on the motion mathematical model, which includes the position and heading of the ship, the rotation matrix, the mass matrix, the control input matrix, the control input vector, the resistance vector, and the external interference force vector.
5. The method for automatic ship mooring based on deep reinforcement learning according to claim 1, characterized in that: The structure of the Critic network specifically includes: A Q network, the Q network is used to approximate a Q value function, and after inputting the state information, outputs the Q value of each action under the state information; An experience replay mechanism, which is used to store interaction data with the environment; A target network, which is a copy of the Q network, is used to stabilize the training process.
6. The method for automatic ship mooring based on deep reinforcement learning according to claim 5, characterized in that: The Critic network and the Actor network, when being trained, include: Initialize the Critic network and the Actor network, initialize the parameters of the target network, observe the current state of the environment, select an action, execute an action, store state transfer, check whether to update the network, calculate the target action and target value, update the Critic network, delay updating the Actor network, update the parameters of the target network, and determine whether the training is completed.
7. A ship automatic mooring system based on deep reinforcement learning, characterized in that: The system comprises: A model generation module is used to build an automatic control model based on deep reinforcement learning training. The environment of the automatic control model is based on the port map and combines the location data of static and dynamic obstacles. Its state includes the current position, speed, and heading of the ship. Its actions include adjusting the speed and heading of the ship. The automatic control model also includes two Critic networks and one Actor network. A state monitoring module is used to obtain the current coordinates, target coordinates and speed of the ship in the environment in real time, and calculate the safe distance at which the ship can approach the target coordinates, where the safe distance is the maximum approach distance at which the ship has enough space to reduce the speed to the required speed and maneuver during the mooring period; The automatic control module is used to read the current state information of the ship when the distance between the ship and the target coordinates is less than or equal to the safety distance. The state information includes the current coordinates, speed, obstacle data, and heading. The speed and heading to be taken at the current moment are generated through the two Critic networks and the Actor network, and input into the ship controller to execute the set speed and heading.