Intelligent ship navigation control method and system based on maximum entropy inverse reinforcement learning
By constructing a ship motion simulation model and inversely deriving the reward function through maximum entropy inverse reinforcement learning, the problem of existing intelligent ship control methods being unable to adaptively integrate complex environments and expert experience is solved, thus achieving safe and efficient navigation control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN UNIV OF TECH
- Filing Date
- 2026-03-11
- Publication Date
- 2026-07-03
Smart Images

Figure CN122331371A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of ship control technology, and in particular to an intelligent ship navigation control method and system based on maximum entropy inverse reinforcement learning. Background Technology
[0002] Autonomous navigation control of intelligent ships is a key technology for the shipping industry to move towards intelligence. Existing control methods, such as traditional control algorithms and model-based optimization methods, heavily rely on manually defined rules and precise ship motion simulation models, making it difficult to adaptively integrate complex navigation rules, dynamic environments, and the expert experience of human captains. While reinforcement learning-based methods have potential, their performance bottleneck lies in the need for manually designed reward functions, a subjective and difficult process that often results in the learned strategies failing to achieve the optimal balance between safety, efficiency, and rule compliance. Summary of the Invention
[0003] The main objective of this application is to propose an intelligent ship navigation control method and system based on maximum entropy inverse reinforcement learning, so as to generate navigation control commands that are safe, efficient and deeply aligned with the wisdom of maritime practice.
[0004] To achieve the above objectives, one aspect of this application proposes an intelligent ship navigation control method based on maximum entropy inverse reinforcement learning, comprising the following steps: Collect expert demonstration trajectory data, preprocess the expert demonstration trajectory data to obtain expert trajectory dataset, and then construct a ship motion simulation model based on the expert trajectory dataset; Based on the maximum entropy inverse reinforcement learning algorithm, the optimal reward function is obtained according to the expert trajectory dataset and the ship motion simulation model; Based on the reinforcement learning algorithm, the optimal policy network is obtained according to the optimal reward function and the ship motion simulation model; The ship status data is acquired, and ship motion control commands are generated based on the optimal strategy network and the ship status data.
[0005] In some embodiments, the process of collecting expert demonstration trajectory data and preprocessing the expert demonstration trajectory data to obtain an expert trajectory dataset specifically includes: Define the ship's state as a state variable, and define the ship's actions as action variables; Define the ship's state-action basis functions based on the state variables and the action variables; Collect the expert demonstration trajectory data; Based on the state-action basis function, the expert demonstration trajectory data is cleaned and smoothed to obtain the expert trajectory dataset; The expert trajectory dataset includes several driving trajectory data.
[0006] In some embodiments, the expert trajectory dataset includes several driving trajectory data, and the step of constructing a ship motion simulation model based on the expert trajectory dataset specifically includes: Calculate the state transition amount corresponding to the ship's action at each moment in each of the aforementioned driving trajectory data; The ship's actions and state transitions at each moment are paired and summarized to obtain an offline database; Based on the offline database, a ship motion simulation model is constructed based on database query and similarity measurement. The ship motion simulation model is used to predict the ship state at the next moment based on the ship state at the current moment and the ship's actions.
[0007] In some embodiments, the method of obtaining the optimal reward function based on the maximum entropy inverse reinforcement learning algorithm, according to the expert trajectory dataset and the ship motion simulation model, specifically includes: Construct a reward function neural network and initialize the reward function weight parameters in the reward function neural network; Calculate the expected value of the experts based on the expert trajectory dataset and the reward function neural network; Based on the reinforcement learning algorithm, the current optimal navigation strategy is calculated according to the reward function neural network; Based on the ship motion simulation model, a first trajectory dataset corresponding to the current optimal navigation strategy is generated; Calculate the expected value of the strategy based on the first trajectory dataset and the reward function neural network; Calculate the gradient of the objective function based on the expert's expected value and the strategy's expected value; The weight parameters of the reward function are updated based on the gradient of the objective function; When the updated reward function weight parameters reach the preset first convergence condition, the reward function neural network corresponding to the reward function weight parameters is taken as the optimal reward function.
[0008] In some embodiments, obtaining the optimal policy network based on the reinforcement learning algorithm, according to the optimal reward function and the ship motion simulation model, specifically includes: Initialize the value network and policy network; Obtain the current navigation strategy and generate a second trajectory dataset corresponding to the current navigation strategy based on the ship motion simulation model; Based on the optimal reward function, calculate the instantaneous reward at each time step in the second trajectory dataset to obtain an empirical sample; Based on the empirical samples, the value network and the policy network are optimized, and the parameters of the value network and the policy network are updated. When the parameters of the updated policy network reach the preset second convergence condition, the optimal policy network is obtained.
[0009] In some embodiments, optimizing the value network and the policy network based on the empirical samples, and updating the parameters of the value network and the policy network, specifically includes: Calculate the cumulative discounted return corresponding to the instantaneous reward at each time step in the experience sample; Calculate the advantage function based on the cumulative return from the discount and the value network; Calculate the policy loss based on the advantage function; Calculate the value loss based on the cumulative return from the discount and the valuation of the value network output; Based on the strategy loss and the value loss, the total loss function is obtained; Based on the total loss function, the value network and the policy network are optimized, and the parameters of the value network and the policy network are updated.
[0010] In some embodiments, acquiring ship state data and generating ship motion control commands based on the optimal strategy network and the ship state data specifically includes: The ship status data is acquired, including the ship's position, heading, speed, and angular velocity. The ship's position, heading, speed, and angular velocity are input into the optimal strategy network, and the ship's motion control commands are output, including the port and starboard rudder angles and the port and starboard main engine speeds.
[0011] To achieve the above objectives, another aspect of this application proposes an intelligent ship navigation control system based on maximum entropy inverse reinforcement learning, comprising: The ship motion simulation module is used to collect expert demonstration trajectory data, preprocess the expert demonstration trajectory data to obtain an expert trajectory dataset, and then construct a ship motion simulation model based on the expert trajectory dataset. The inverse reinforcement learning module is used to obtain the optimal reward function based on the expert trajectory dataset and the ship motion simulation model using the maximum entropy inverse reinforcement learning algorithm. The reinforcement learning module is used to obtain the optimal policy network based on the reinforcement learning algorithm, the optimal reward function, and the ship motion simulation model. The ship motion control module is used to acquire ship status data and generate ship motion control commands based on the optimal strategy network and the ship status data.
[0012] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0013] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.
[0014] The embodiments of this application include at least the following beneficial effects: The intelligent ship navigation control method and system based on maximum entropy inverse reinforcement learning of this application first collects expert demonstration trajectory data, preprocesses the expert demonstration trajectory data to obtain an expert trajectory dataset, and then constructs a ship motion simulation model based on the expert trajectory dataset; next, based on the maximum entropy inverse reinforcement learning algorithm, the optimal reward function is obtained according to the expert trajectory dataset and the ship motion simulation model; then, based on the reinforcement learning algorithm, the optimal policy network is obtained according to the optimal reward function and the ship motion simulation model; finally, ship state data is acquired, and ship motion control commands are generated according to the optimal policy network and the ship state data. This application does not rely on manually preset rewards. By introducing maximum entropy inverse reinforcement learning to model the driving behavior of maritime experts, it can inversely derive the implicit, comprehensive multi-objective reward function from limited expert demonstration trajectory data, thereby generating navigation control commands that are both safe and efficient, and deeply aligned with the wisdom of maritime practice, providing a solid technical foundation for truly realizing autonomous navigation of intelligent ships with human expert-level decision-making capabilities. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments of this application are described below. It should be understood that the drawings described below are only for the purpose of clearly illustrating some embodiments of the technical solutions in this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0016] Figure 1 A flowchart illustrating the steps of an intelligent ship navigation control method based on maximum entropy inverse reinforcement learning, provided in one embodiment of this application. Figure 2 A schematic diagram of the maximum entropy inverse reinforcement learning process provided in one embodiment of this application; Figure 3A schematic diagram of the structure of an intelligent ship navigation control system based on maximum entropy inverse reinforcement learning provided in one embodiment of this application; Figure 4 This is a schematic diagram illustrating an embodiment of an intelligent ship navigation control system based on maximum entropy inverse reinforcement learning provided in this application. Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in one embodiment of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0019] Autonomous navigation control of intelligent ships is a key technology for the shipping industry to move towards intelligence. Existing control methods, such as traditional control algorithms and model-based optimization methods, heavily rely on manually defined rules and precise ship motion simulation models, making it difficult to adaptively integrate complex navigation rules, dynamic environments, and the expert experience of human captains. While reinforcement learning-based methods have potential, their performance bottleneck lies in the need for manually designed reward functions, a subjective and difficult process that often results in the learned strategies failing to achieve the optimal balance between safety, efficiency, and rule compliance.
[0020] Among the currently popular reinforcement learning methods, the performance ceiling is strictly constrained by the subjective challenge of "reward function design." For an agent to learn a superior navigation control strategy through interactive trial and error, algorithm designers must pre-construct a reward function that can accurately quantify the "quality" of navigation. However, a high-quality trajectory is a multi-dimensional objective, requiring a complex trade-off between absolute safety, strict rule compliance, fuel economy, navigation timeliness, and equipment wear and tear costs. Designers must rely on limited experience and subjective judgment to assign appropriate weights to these diverse and often conflicting objectives. Any tiny weight deviation can be amplified dramatically in the "exploration-exploitation" reinforcement learning loop, leading to a final learned strategy that deviates significantly from actual needs.
[0021] In view of this, this application proposes an intelligent ship navigation control method based on maximum entropy inverse reinforcement learning. First, expert demonstration trajectory data is collected and preprocessed to obtain an expert trajectory dataset. Then, a ship motion simulation model is constructed based on the expert trajectory dataset. Next, based on the maximum entropy inverse reinforcement learning algorithm, the optimal reward function is obtained according to the expert trajectory dataset and the ship motion simulation model. Then, based on the reinforcement learning algorithm, the optimal policy network is obtained according to the optimal reward function and the ship motion simulation model. Finally, ship state data is acquired, and ship motion control commands are generated according to the optimal policy network and the ship state data. This application does not rely on manually preset rewards. By introducing maximum entropy inverse reinforcement learning to model the driving behavior of maritime experts, it can inversely derive an implicit, comprehensive multi-objective reward function from limited expert demonstration trajectory data. This generates navigation control commands that are safe, efficient, and deeply aligned with the wisdom of maritime practice, providing a solid technical foundation for truly realizing autonomous navigation of intelligent ships with human expert-level decision-making capabilities.
[0022] Reference Figure 1 , Figure 1 This is a flowchart illustrating the steps of an intelligent ship navigation control method based on maximum entropy inverse reinforcement learning according to an embodiment of this application. This application proposes an intelligent ship navigation control method based on maximum entropy inverse reinforcement learning, which may include, but is not limited to, the following steps S101 to S104: Step S101: Collect expert demonstration trajectory data, preprocess the expert demonstration trajectory data to obtain expert trajectory dataset, and then construct a ship motion simulation model based on the expert trajectory dataset; Specifically, high-quality expert demonstration trajectory data is collected, and the collected raw data is preprocessed by cleaning, filtering, and smoothing to construct a standardized expert trajectory dataset containing state-action sequences, providing reliable training samples for subsequent reward function learning; then the state space and action space are defined to construct a ship motion simulation model.
[0023] As an optional implementation, the step of collecting expert demonstration trajectory data and preprocessing the expert demonstration trajectory data to obtain an expert trajectory dataset can be further divided into the following steps S1011 to S1014: Step S1011: Define the ship's state as a state variable and define the ship's actions as action variables; Step S1012: Define the ship's state-action basis functions based on the state variables and action variables; Specifically, define the ship's state-action basis functions. Define state variables The ship's status, including its current position. and True course of the ship Ship speed Ship angular velocity Action variables To perform actions for a ship, which may include the port and starboard rudder angles. and Left and right main unit speeds and .
[0024] Step S1013: Collect expert demonstration trajectory data; Step S1014: Based on the state-action basis function, clean and smooth the expert demonstration trajectory data to obtain the expert trajectory dataset; The expert trajectory dataset includes several driving trajectory data.
[0025] Specifically, collecting In berthing and unberthing operations, the driving data of experienced crew members is used as expert demonstration trajectory data for reward function inference and policy training in the inverse reinforcement learning model. Next, [the text abruptly ends here]. The crew's driving data was cleaned and smoothed using Gaussian filtering to remove abnormal noise and improve signal quality. Then, effective data segments were extracted based on key nodes for subsequent analysis, resulting in an expert trajectory dataset. ,in , indicating the first The driving trajectory data includes the complete status. and actions The sequence, where t represents time t. This indicates the length of the trajectory.
[0026] As an optional implementation, the expert trajectory dataset includes several driving trajectory data. The step of constructing a ship motion simulation model based on the expert trajectory dataset can be further divided into the following steps S1015 to S1017: Step S1015: Calculate the state transition amount corresponding to the ship's action at each moment in each driving trajectory data; Step S1016: Pair and summarize the ship's actions and state transitions at each moment to obtain an offline database; In some optional embodiments, the ship's actions and state transitions at each moment are extracted from all expert driving trajectories. , This refers to the difference between the ship's state at the next moment and its state at the current moment. Specifically, the state transition quantity is calculated for each moment. and the action at that moment Pairing them together creates an action-state transition data pair. Finally, the data pairs at every moment of all the expert driving trajectories are aggregated to obtain a comprehensive dataset. An offline database containing data pairs. ,in , indicating the first A set of data, including the actions performed by the ship at a specific moment. and the ship's state transition at that moment .
[0027] Step S1017: Based on the offline database, construct a ship motion simulation model based on database query and similarity measurement. The ship motion simulation model is used to predict the ship state at the next moment based on the ship state at the current moment and the ship's actions.
[0028] Specifically, given the current ship status and ship operations Under the premise of querying the offline database using database queries or similarity measurement methods. Obtain the current status of the ship Take action below This will cause a change in the state. Thus through Output This allows us to obtain the ship's status at the next moment.
[0029] (1) Database query: Enter the current ship action. If there are data pairs in the database , If the action performed by the ship is exactly the same as the current ship action, then the data is read. As a change of state of a ship, through Output the ship's state at the next moment; if no data in the database meets this requirement, perform approximate matching and estimation.
[0030] (2) Approximate matching and estimation: To measure the similarity between actions, the weighted Manhattan distance is used as a metric. The actions in the database are... With the current ship's actions The distance calculation formula is: ; Where 4 represents the dimension of the action vector. and These are the currently executed actions of the ship, respectively. Actions in the database The 1 eigenvalue, For example, a change in rudder angle of the same magnitude usually has a more significant impact on the ship's state transition than a change in main engine speed, so it can be assigned a higher weight.
[0031] Selected by Manhattan distance One of the current ship actions The closest action, as the candidate set ( ), denoted as And obtain the corresponding ship state transition quantities. Calculate the initial estimate based on the distance between candidate sets. : ; in, for The weighting coefficients, Temperature is a parameter used to control the concentration of the weight distribution. Depend on and Decision. To ensure the estimated ship state transition. In accordance with the basic laws of ship motion, it is also necessary to... Perform physical constraint corrections. The physical constraint corrections performed include those made at that moment. , The constraints include kinematic continuity constraints on the rate of change of information, energy constraints on the upper limit of energy supply, and dynamic smoothing constraints for abrupt changes. The final result is a modified... Final calculation This serves as the output of the ship's status at the next moment.
[0032] It should be noted that, through steps S1011 to S1017 described above, this embodiment of the application prepares high-quality expert driving data for berthing and unberthing tasks, ensuring that subsequent analysis is based on reliable and clean expert driving data. It also constructs an offline ship motion simulation model for algorithm training, laying a solid foundation for subsequent inverse derivation of the reward function and strategy optimization. The maximum entropy inverse reinforcement learning algorithm can robustly learn the implicit multi-objective reward function from high-quality expert demonstrations, providing a reliable guarantee for the subsequent generation of intelligent navigation strategies that combine safety, economy, and comfort.
[0033] Step S102: Based on the maximum entropy inverse reinforcement learning algorithm, the optimal reward function is obtained according to the expert trajectory dataset and the ship motion simulation model; Specifically, a comprehensive reward standard is derived from expert demonstration trajectory data through maximum entropy probabilistic modeling. Based on the maximum entropy principle, a trajectory probability distribution model is constructed. Under the constraint of matching the expected value of expert behavior, a probabilistic framework with reward weights as parameters is established. By maximizing the likelihood probability of the expert trajectory through an optimization algorithm, the reward function parameters that best explain the expert decision-making pattern are solved in reverse, thus obtaining an objective reward function that balances multiple objectives and does not require manual pre-setting of weights.
[0034] As an optional implementation, step S102 can be further divided into the following steps S1021 to S1028: Step S1021: Construct a reward function neural network and initialize the reward function weight parameters in the reward function neural network; Specifically, such as Figure 2 The diagram shows the flowchart of maximum entropy inverse reinforcement learning. First, a reward function neural network is constructed. Its input is the ship's status. Actions performed by the ship The output is a scalar reward value, where The weight parameters of the reward function are also the target we hope to obtain through inverse reinforcement learning. Next, we define the likelihood distribution of the trajectory, which follows the maximum entropy principle. Assuming that the driver prefers to choose strategies with higher cumulative rewards—that is, considering a quantitative comprehensive evaluation under conditions such as safety and economy—to select the decision rule for the ship's actions, the trajectory likelihood distribution is: ; in, It is a discount factor. This indicates the weight parameters of the entire trajectory in the reward function. The cumulative rewards below. It is a partition function. Its function is to calculate the unnormalized probability of all possible trajectories. Summation is performed to ensure that the sum of all trajectory probabilities is 1, thus forming an effective distribution. The partition function is typically not directly computable, therefore this embodiment cleverly avoids this in subsequent steps. The core idea of trajectory likelihood distribution is that, in berthing and unberthing tasks, the reward is quantified by comprehensively considering factors such as navigation safety, economy, and navigation stability. The higher the cumulative reward of a navigation trajectory, the greater its probability of occurrence. Expert driving data has the highest likelihood in this distribution, and maximizing the trajectory likelihood distribution is the process of solving for the reward function that best suits the expert driving strategy.
[0035] Step S1022: Calculate the expected value of the experts based on the expert trajectory dataset and the reward function neural network; Specifically, expected value is a core indicator for evaluating a ship's long-term performance under a given navigation strategy. It is an estimated value of all cumulative rewards that can be obtained throughout the entire voyage, starting from the current state and following a specific navigation strategy. The fundamental purpose of calculating expert expected value is to provide a clear and quantifiable optimization objective for the learning algorithm. Specifically, the reward function obtained from inverse reinforcement learning must ensure that the expected value of the intelligent ship's control strategy aligns with the expected value of the experienced pilot's behavior. for: ; in, This represents the number of trajectories in the expert data.
[0036] Step S1023: Based on the reinforcement learning algorithm, calculate the current optimal navigation strategy according to the reward function neural network; Step S1024: Based on the ship motion simulation model, generate the first trajectory dataset corresponding to the current optimal navigation strategy; Specifically, based on the current reward function weight parameters The optimal navigation strategy is determined using reinforcement learning algorithms. Next, in the existing offline ship motion simulation model, starting from the preset initial state, the current optimal navigation strategy is applied. Generate ship execution actions Then, the state transition is advanced through the constructed offline ship motion simulation model, and finally the first trajectory dataset containing complete state-action data is output. This first trajectory dataset can be used as the current optimal navigation strategy. It is an important basis for the evaluation and expected value calculation.
[0037] Step S1025: Calculate the expected value of the policy based on the first trajectory dataset and the reward function neural network; Specifically, the current optimal navigation strategy Expected value of strategy for: ; The expected navigation score that can be obtained by the current navigation control strategy is calculated and compared with the expected value of the expert navigation trajectory. The purpose is to guide the intelligent agent to continuously adjust and optimize its driving strategy, so that it can eventually match the decision-making level of human experts in a statistical sense, thereby learning the implicit and efficient navigation trade-off wisdom of experts.
[0038] Step S1026: Calculate the gradient of the objective function based on the expert's expected value and the strategy's expected value; Specifically, the log-likelihood of the expert trajectory is defined as:
[0039]
[0040] ; This form expresses the core idea of maximum entropy inverse reinforcement learning: maximizing the likelihood function is equivalent to maximizing the uncertainty of the trajectory distribution under the constraint that the expected value of the current navigation control strategy matches the expected value of the trajectory data of the expert driver, so that the finally learned strategy has adaptability and robustness.
[0041] Next, the gradient of the objective function is calculated as follows: ; in, This indicates the reward gradient for the expert driving trajectory. Let represent the logarithmic gradient of the partition function, which indicates that increasing expert rewards while decreasing average rewards is beneficial. However, since the logarithmic gradient of the partition function cannot be calculated precisely, other methods are needed to approximate its value. (The last sentence appears to be incomplete and possibly refers to a problem or inconsistency.) ,because: = ; Therefore, we get:
[0042] ; Therefore, the objective gradient function can be transformed into: ; in, Let represent the reward gradient for m expert trajectories. This represents the trajectory probability distribution derived under the maximum entropy principle. Indicates the parameters of the current reward function. The expectation of all possible trajectories under the induced measurement slight distribution. Indicates the current strategy The expected reward gradient below.
[0043] get: ; Because the expected value of the current strategy can be expressed as: ; In summary, the gradient of the objective function can be written as: ; Where, gradient The gradient refers to the direction of the reward value of the state-action pair experienced by the driver. This points in the direction of the reward value of the state-action pair experienced by the current smart ship's strategy. Therefore... This indicates that the gradient vector accurately characterizes the local rate of change and the steepest ascent direction of the objective function at the current parameter value. Within the framework of maximum entropy inverse reinforcement learning, this operation aims to quantify the deviation between the current reward function parameters and the optimal solution. By comparing the gradients of the intelligent ship's and the pilot's heading data, a set of reward criteria that can explain the pilot's decision-making logic is derived, enabling the intelligent ship to learn to make safe and efficient decisions similar to those of the pilot.
[0044] Step S1027: Update the weight parameters of the reward function according to the gradient of the objective function; Step S1028: When the updated reward function weight parameters reach the preset first convergence condition, the reward function neural network corresponding to the reward function weight parameters is taken as the optimal reward function.
[0045] Specifically, the weight parameters of the reward function are updated using the calculated gradient of the objective function. : ; in, The learning rate is the step size for parameter updates relative to the gradient direction.
[0046] Finally, determine whether the first convergence condition has been met (such as the change in the reward function approaching zero or reaching the maximum number of iterations). If convergence is achieved, the loop ends, and the reward function neural network corresponding to the weight parameters of the reward function is taken as the optimal reward function. If convergence is not achieved, the policy and expectation under the new reward function are recalculated, and the next loop begins.
[0047] It should be noted that, through steps S1021 to S1028, this embodiment systematically completes the derivation process of the inverse reinforcement learning reward function based on the maximum entropy principle. The unique advantage of the maximum entropy method lies in its ability to maintain maximum uncertainty regarding all compatible behavioral patterns when analyzing driver behavior, even without fully determining the driver's intentions, thereby avoiding arbitrary assumptions. Within this probabilistic framework, this embodiment, considering safety, economy, and other factors, constructs a parameterized reward function. By defining the log-likelihood distribution of expert trajectories, the probability of occurrence of a continuous navigation trajectory is correlated with its corresponding reward value. During iterative optimization, through precise calculation of the difference between the expected performance of expert driving data and the expected performance of the control strategy generated by the current reward function, a gradient update mechanism based on the log-likelihood function is constructed. This gradually adjusts the reward function parameters, enabling the intelligent ship's behavior to continuously approach the expert's navigation level. The maximum entropy principle ensures that the derivation process can fully fit the expert navigation data and that the learned reward function possesses stronger generalization ability and robustness, meaning that it can still guide the intelligent ship to make reasonable, safe, and efficient navigation decisions under different sea areas, weather conditions, and ship loads. This rigorous mathematical derivation provides a theoretical guarantee for robustly learning complex reward functions from a limited sample of experts, laying a solid foundation for generating intelligent decision-making strategies that are safe, economical, and comfortable.
[0048] Step S103: Based on the reinforcement learning algorithm, the optimal policy network is obtained according to the optimal reward function and the ship motion simulation model; Specifically, the learned reward function is transformed into an executable navigation strategy. Using the aforementioned derived reward function as the optimization objective, the agent can learn the optimal strategy through continuous interaction with the environment in a virtual simulation environment or through interaction with an offline dataset. Advanced reinforcement learning algorithms are employed to iteratively optimize the strategy through value function estimation, strategy evaluation, and strategy improvement, thereby gradually improving the performance of the strategy and ultimately obtaining the optimal strategy that can adapt to complex navigation conditions and balance multi-dimensional performance indicators.
[0049] As an optional implementation, step S103 can be further divided into the following steps S1031 to S1035: Step S1031: Initialize the value network and policy network; Step S1032: Obtain the current navigation strategy and generate the second trajectory dataset corresponding to the current navigation strategy based on the ship motion simulation model; Step S1033: Calculate the instantaneous reward at each time step in the second trajectory dataset according to the optimal reward function to obtain the experience sample; Specifically, using the current navigation strategy Multiple complete navigation trajectory data were obtained by running the simulation in an offline ship motion model, and a second trajectory dataset was collected. ,in Next, using the current reward function... For each time step Provide instant rewards The calculation provides training samples, including reward signals, for subsequent strategy evaluation and optimization.
[0050] Step S1034: Optimize the value network and policy network based on empirical samples, and update the parameters of the value network and policy network. As an optional implementation, step S1034 can be further divided into the following steps S10341 to S10346: Step S10341: Calculate the cumulative discounted return corresponding to the instantaneous reward at each time step in the experience sample; Specifically, for each time step The instant rewards are calculated using a discount-based cumulative return calculation method, as follows: ; in, This represents the offset at a future time step. Indicates time step Instant rewards received This represents the discount factor for this step. To start from time step Initial discounts and rewards.
[0051] Step S10342: Calculate the advantage function based on the cumulative return of the discount and the value network; Specifically, the generalized advantage (GAE) method is used to calculate the advantage function. : ; ; in, Indicates time step Instant rewards received Representing state The value function under the current ship state, i.e. Initially, the expected total return achievable by following the current strategy reflects the long-term expected performance of following the strategy under the current ship condition. The parameters of the value function are represented in the formula. This represents a value estimate of the state at the next moment. This represents the value estimate at the current moment. Indicates a single action performed by the ship. Compared with expected returns difference. Indicates time step The advantage function estimation, i.e., the current ship state The following actions were taken by the ship. Compared with the average performance of the current strategy (i.e., the value function) In comparison, if it is a positive number, it indicates that at this time... It is better than the average movement. This represents the GAE smoothing discount factor, which is used to control for the additional discount introduced by variance. A smaller value indicates lower variance and higher bias. The product represents the decay rate. Indicates the time step from the current time. Forward steps Indicates the first Step into the future TD error The composite discount factor represents the weight of the exponential decay of future TD errors. The essence of GAE is: using an exponentially decaying weight... The degree of "unexpectedness" for all moments, both now and in the future. By performing weighted summation, flexibly setting the evaluation perspective, and selecting an equilibrium point suitable for the nautical scenario, a dominance estimate that achieves a good balance between bias and variance can be obtained. .
[0052] Step S10343: Calculate the policy loss based on the advantage function; Step S10344: Calculate the value loss based on the cumulative return from the discount and the valuation of the value network output; Step S10345: Obtain the total loss function based on the strategy loss and value loss; Specifically, from the second trajectory dataset A small batch of samples is extracted for iterative optimization. First, the probability ratio is calculated: ; in, Indicates the state of the new strategy Select action The probability, Indicates the old policy state Select action The probability, This represents the change in behavior between the old and new strategies. A value greater than 1 indicates that the probability of the new strategy choosing this action has increased, suggesting that the new strategy prefers this action. Conversely, a value less than 1 indicates that the new strategy dislikes this action. This allows us to use data collected from the old strategy to evaluate and optimize the new data. Final strategy loss. for: ; in, This represents a hyperparameter, which strictly limits the magnitude of policy changes in a single update. This indicates the unpruned objective, i.e., the standard policy gradient objective. This represents the target after pruning, limiting the update magnitude. It employs pessimistic pruning, selecting the more conservative and smaller result from the pre- and post-pruning outcomes. Overall, this pruning mechanism relies on... It provides direction, but doesn't fully trust the extent of its updates, forcing each update to be small.
[0053] Then, calculate the value loss. : ; in, This represents the predicted value of the value network. Indicates from state The initial actual discount return. This formula represents the value network. The predicted value should be as close as possible to the actual discount report. A more accurate value network can calculate a more reliable advantage estimate. .
[0054] Finally, the total loss function is calculated based on the strategy loss and value loss. : ; in, This refers to the primary driving force, responsible for optimizing intelligent ship navigation strategies to obtain greater returns; Responsible for improving the current ship condition, following existing strategies, and the expected total return that can be obtained from future voyages; Entropy represents the distribution of strategies. The higher the entropy, the stronger the exploratory nature, which prompts intelligent ships to explore a wider range of navigation and movement spaces and discover better navigation strategies. and This represents the hyperparameters used to balance the relative importance of these three terms in the total loss.
[0055] Step S10346: Optimize the value network and policy network according to the total loss function, and update the parameters of the value network and policy network.
[0056] Step S1035: When the parameters of the updated policy network reach the preset second convergence condition, the optimal policy network is obtained.
[0057] Specifically, by minimizing the total loss function, the parameters of the policy network are updated simultaneously and collaboratively using gradient descent. and parameters of the value network This causes the navigation strategy, guided by the current reward function, to evolve towards obtaining higher returns. After completing all mini-batch samples, Updated to This prepares for the next round of updates until convergence, yielding the optimal policy network. .
[0058] It should be noted that, through the aforementioned steps S1031 to S1035 and S10341 to S10346, the embodiments of this application systematically complete policy optimization based on the maximum entropy inverse reinforcement learning reward function. Specifically, the reward function learned by the IRL, which contains expert navigation preferences and multidimensional constraints, is transformed into a directly executable high-level control policy. This strategy works in any state. Below, the output action Instead of abstract path points, these are concrete, continuous ship motion control commands. This design allows the execution of actions to directly impact the ship's dynamics model. The agent continuously outputs optimal control command sequences, performing closed-loop rolling optimization in the state space. Essentially, this executes an adaptive model predictive control: each action aims to instantaneously optimize a long-term goal defined by a reward function, and the resulting continuous state transitions naturally synthesize and update in real-time an achievable optimal trajectory that conforms to ship dynamics. The advantage of this framework lies in the fact that, expressed through a reward function, the control strategy directly embeds an understanding and trade-off of the expert's complex intentions, ensuring that the generated navigation control commands are not only dynamically feasible but also fundamentally mimic and generalize the expert's comprehensive trade-offs in collision avoidance, navigation efficiency, and comfort. This rigorous transformation from reward to control commands provides a reliable core decision-making and control foundation for intelligent ships to achieve expert-like, safe, and efficient autonomous navigation.
[0059] Step S104: Obtain ship status data, and generate ship motion control commands based on the optimal strategy network and ship status data.
[0060] Specifically, the trained optimal policy network is deployed into the actual ship control system to generate reference trajectories. Through three stages—state perception, decision reasoning, and trajectory generation—autonomous navigation control of the ship is achieved, ensuring that the learned policies can be effectively executed in real navigation environments.
[0061] As an optional implementation, step S104 can be further divided into the following steps S1041 to S1042: Step S1041: Obtain ship status data, which includes ship position, heading, speed and angular velocity; Step S1042: Input the ship's position, heading, speed, and angular velocity into the optimal strategy network, and output the ship motion control commands, which include the port and starboard rudder angles and the port and starboard main engine speeds.
[0062] Specifically, ship status data, including ship position, is obtained by fusing data from shipborne sensors. and True course of the ship Ship speed and ship angular velocity These raw data are preprocessed to a state consistent with that of the training phase, and used as the optimal policy network. The system generates ship motion control commands from the input, including target values for the port and starboard main engine speeds and rudder angles, which serve as the input reference for the tracking controller. The tracking controller, through calculation and closed-loop adjustment, outputs low-level control signals that minimize tracking errors and directly drive the main engine and steering gear. The ship receives and executes these low-level commands, enabling it to stably and accurately complete complex berthing and unberthing maneuvers.
[0063] The above describes the intelligent ship navigation control method based on maximum entropy inverse reinforcement learning according to embodiments of this application. It can be understood that embodiments of this application model the driving behavior of maritime experts using a maximum entropy probabilistic framework, enabling robust learning of implicit multi-objective reward functions from limited demonstration data. This framework does not impose a specific form on the reward function, but rather maximizes policy randomness, accommodating reasonable ambiguities that experts may have in their decisions under complex hydrological and traffic environments, thus more completely capturing the comprehensive decision-making logic of human drivers in scenarios such as berthing / unberthing and navigation in narrow waterways. Furthermore, the resulting strategy not only reproduces the typical behavioral patterns demonstrated by experts but also exhibits good generalization ability to new, unseen scenarios. This probabilistic modeling approach allows the strategy to maintain reasonable decision diversity in the face of uncertainty, avoiding the overfitting problem that is prone to occur in traditional methods. Compared with traditional path planning methods that rely on manually designed reward weights, embodiments of this application automatically extract optimal navigation decision criteria in a data-driven manner, significantly reducing the heavy reliance on domain knowledge and the burden of engineering parameter tuning, and improving the intelligence level of trajectory generation. Ultimately, the navigation strategy generated by this method achieves a good balance between safety and reliability, energy economy and human-like driving. Its behavior not only conforms to international maritime collision avoidance rules and port navigation practices, but also reflects the driving skills of experienced drivers, providing key technical support for ships to achieve advanced autonomous navigation.
[0064] Reference Figure 3 This application also provides an intelligent ship navigation control system based on maximum entropy inverse reinforcement learning, including: The ship motion simulation module is used to collect expert demonstration trajectory data, preprocess the expert demonstration trajectory data to obtain expert trajectory dataset, and then construct a ship motion simulation model based on the expert trajectory dataset. The inverse reinforcement learning module is used to obtain the optimal reward function based on the maximum entropy inverse reinforcement learning algorithm, according to the expert trajectory dataset and the ship motion simulation model. The reinforcement learning module is used to obtain the optimal policy network based on the reinforcement learning algorithm, the optimal reward function, and the ship motion simulation model. The ship motion control module is used to acquire ship status data and generate ship motion control commands based on the optimal strategy network and ship status data.
[0065] Specifically, such as Figure 4 The diagram illustrates the implementation of an intelligent ship navigation control system based on maximum entropy inverse reinforcement learning. Based on the aforementioned intelligent ship navigation control system, the flow of the intelligent ship navigation control method based on maximum entropy inverse reinforcement learning in this embodiment is as follows: Step 1: Data preparation and ship motion simulation modeling.
[0066] Acquire expert demonstration trajectory data to prepare a high-quality expert trajectory dataset for inverse reinforcement learning. including The established trajectory ensures that subsequent analysis is based on reliable and clean expert driving data, and an offline ship motion simulation model is constructed to simulate any given state-action pair. As input, the output is the next state the ship will enter after this action is performed. This lays a solid foundation for subsequent inverse derivation of the reward function and policy optimization. An expert trajectory dataset was constructed, and a ship motion simulation based on database query and similarity measurement was built for subsequent inverse derivation of the reward function and policy optimization.
[0067] Step 2: Obtain the optimal strategy .
[0068] (1) Constructing a reward function neural network Initialize neural network parameters .
[0069] Constructed reward function neural network The goal is to provide a parameterized reward representation for downstream reinforcement learning tasks, judging the merits of taking a certain action in the current state to guide learning. The state includes the ship's lateral position, longitudinal position, heading angle (bow), speed, and angular velocity relative to the berth. Actions correspond to control commands for port / starboard rudder angles and main engine speed. The reward function is constructed by penalizing errors in the aforementioned states while simultaneously encouraging smooth maneuvering, low energy consumption, and compliance with safety regulations. The network parameters... This directly constitutes the optimization objective of the inverse reinforcement learning process. Maximum entropy inverse reinforcement learning obtains a multi-objective reward function that has both high generalization ability and can reasonably explain expert behavior under the constraint that the expected features of the generated policy match the expected value of the expert demonstration. Thus, in highly complex scenarios such as ship berthing and unberthing with extremely high safety requirements, it guides the agent to learn robust manipulation strategies that adapt to different conditions, while ensuring that the decision-making process conforms to navigation conventions and operational experience.
[0070] (2) Calculate the expected value of experts.
[0071] Experts' expected value is: ; From limited expert driving data, the average cumulative performance is statistically analyzed to serve as a benchmark for the intelligent ship strategy in subsequent optimization processes. This step aggregates the cumulative performance of driving data on specific features into a low-dimensional and statistically representative vector, namely the expected value of experience. This vector encapsulates the multi-objective trade-offs implicit in excellent driving patterns at the data level. Therefore, this expected value of experience features provides a clear data-driven optimization objective for the inverse derivation of the reward function. That is, the learned reward function should enable the agent to generate a cumulative expected feature value on the same features that approximates the expert benchmark as closely as possible, thereby ensuring that the learned navigation strategy statistically reproduces the reliability and effectiveness of expert driving.
[0072] (3) Solving for the optimal strategy .
[0073] Based on reinforcement learning algorithms, using current reward function neural networks As the reward function, the current optimal navigation strategy is solved. .
[0074] (4) Generation Strategy The trajectory below.
[0075] Based on the current optimal navigation strategy Generate with the existing offline ship motion simulation model Strategy The trajectories are used to systematically reflect the evolution of the ship's motion state under the action of the strategy. These trajectories fully record the navigation sequence from start to finish, providing a crucial data foundation for subsequent strategy performance evaluation and calculation of expected feature accumulation, and supporting the quantitative analysis of the effectiveness, safety, and efficiency of the strategy navigation.
[0076] (5) Calculate the current optimal navigation strategy Expected value.
[0077] Current Strategy The expected value is: ; The behavior of the intelligent ship under this strategy is aligned and compared with the behavior in expert demonstration driving data.
[0078] (6) Calculate the log-likelihood of the expert trajectory.
[0079] The definition of the log-likelihood of an expert trajectory is:
[0080] ; Maximizing the likelihood function is equivalent to maximizing the trajectory distribution entropy corresponding to the policy, under the constraint that the expected value of the model's prediction matches the expected value of the expert trajectory data. The learning process requires the agent not only to statistically reproduce features of expert driving, such as heading stability, energy efficiency, and collision avoidance tendency, but also to avoid making unfounded preference assumptions about behavioral patterns not reflected in the expert data, thereby maintaining the maximum randomness and generalization ability of the policy within the constraints.
[0081] (7) Calculate the gradient of the objective function .
[0082] The gradient of the objective function is: ; It can be simplified to: ; The gradient of the objective function corresponds to the direction that increases the likelihood of the expert trajectory and makes the policy-induced feature expectation more consistent with the expert feature expectation; its magnitude characterizes the local sensitivity of the objective function in that direction, thus providing a quantitative basis for determining the update step size for each step in gradient descent algorithms. This provides a theoretical basis for determining the parameter update direction and magnitude in optimization algorithms.
[0083] (8) Execute parameter update.
[0084] Update reward function parameters : ; By using the gradient ascent method, the parameters of the reward function are iteratively updated to maximize the expected return of the intelligent ship strategy, gradually approaching the characteristic expectation of expert driving data, and ensuring that the reward function can continuously guide the navigation strategy to evolve in a direction that is closer to the expert driving mode.
[0085] (9) Obtain the optimal reward function .
[0086] Check the convergence condition; if the change in the reward function approaches zero or the maximum number of iterations is reached, then the final result will be... As parameters of the optimal reward function The corresponding reward function neural network The optimal reward function is the output of maximum entropy inverse reinforcement learning. This reward function incorporates the multi-objective trade-off preferences and behavioral patterns implicit in expert driving data, thus providing accurate and generalizable reward guidance for subsequent intelligent ship navigation strategy learning and optimization.
[0087] (10) Solve for the optimal policy network .
[0088] Based on reinforcement learning algorithms, according to the optimal reward function Solve for the optimal policy network By inputting real-time, multi-source environmental information into a trained optimal policy network, the model will directly and instantly generate the safest and most economical control commands for the current state, namely the left and right main engine speeds and left and right rudder angles, which directly correspond to the ship's operating mode, so as to achieve safe and economical autonomous navigation in berthing and unberthing scenarios.
[0089] Step 3: Real-time perception and construction of environmental status.
[0090] Ship status, including ship position, is obtained by fusing data from shipborne sensors. True course of the ship speed of the boat ship angular velocity These raw data are preprocessed to form a state consistent with that of the training phase, and then used as input to the optimal policy network.
[0091] Step 4: Ship control.
[0092] Based on ship status data, through the optimal strategy network The system generates ship motion control commands, including target values for the port and starboard main engine speeds and rudder angles, which serve as the input reference for the tracking controller. Through calculation and closed-loop adjustment, the tracking controller outputs low-level control signals that minimize tracking errors and directly drive the main engine and steering gear. The ship receives and executes these low-level commands, enabling it to stably and accurately complete complex berthing and unberthing maneuvers.
[0093] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0094] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0095] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0096] Please see Figure 5 , Figure 5 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1002 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1002 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001 using the methods described in the embodiments of this application. Input / output interface 1003 is used to implement information input and output; The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004); The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0097] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0098] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0099] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0100] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0101] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0102] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0103] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0104] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0105] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0106] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0107] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0108] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0109] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0110] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0111] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0112] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. An intelligent ship navigation control method based on maximum entropy inverse reinforcement learning, characterized in that, Includes the following steps: Collect expert demonstration trajectory data, preprocess the expert demonstration trajectory data to obtain expert trajectory dataset, and then construct a ship motion simulation model based on the expert trajectory dataset; Based on the maximum entropy inverse reinforcement learning algorithm, the optimal reward function is obtained according to the expert trajectory dataset and the ship motion simulation model; Based on the reinforcement learning algorithm, the optimal policy network is obtained according to the optimal reward function and the ship motion simulation model; The ship status data is acquired, and ship motion control commands are generated based on the optimal strategy network and the ship status data.
2. The method according to claim 1, characterized in that, The process of collecting expert demonstration trajectory data and preprocessing the expert demonstration trajectory data to obtain an expert trajectory dataset specifically includes: Define the ship's state as a state variable, and define the ship's actions as action variables; Define the ship's state-action basis functions based on the state variables and the action variables; Collect the expert demonstration trajectory data; Based on the state-action basis function, the expert demonstration trajectory data is cleaned and smoothed to obtain the expert trajectory dataset; The expert trajectory dataset includes several driving trajectory data.
3. The method according to claim 1, characterized in that, The expert trajectory dataset includes several driving trajectory data. The step of constructing a ship motion simulation model based on the expert trajectory dataset specifically includes: Calculate the state transition amount corresponding to the ship's action at each moment in each of the aforementioned driving trajectory data; The ship's actions and state transitions at each moment are paired and summarized to obtain an offline database; Based on the offline database, a ship motion simulation model is constructed based on database query and similarity measurement. The ship motion simulation model is used to predict the ship state at the next moment based on the ship state at the current moment and the ship's actions.
4. The method according to claim 1, characterized in that, The maximum entropy inverse reinforcement learning algorithm, based on the expert trajectory dataset and the ship motion simulation model, obtains the optimal reward function, specifically including: Construct a reward function neural network and initialize the reward function weight parameters in the reward function neural network; Calculate the expected value of the experts based on the expert trajectory dataset and the reward function neural network; Based on the reinforcement learning algorithm, the current optimal navigation strategy is calculated according to the reward function neural network. Based on the ship motion simulation model, a first trajectory dataset corresponding to the current optimal navigation strategy is generated; Calculate the expected value of the policy based on the first trajectory dataset and the reward function neural network; Calculate the gradient of the objective function based on the expert's expected value and the strategy's expected value; The weight parameters of the reward function are updated based on the gradient of the objective function; When the updated reward function weight parameters reach the preset first convergence condition, the reward function neural network corresponding to the reward function weight parameters is taken as the optimal reward function.
5. The method according to claim 1, characterized in that, The method for obtaining the optimal policy network based on the reinforcement learning algorithm, according to the optimal reward function and the ship motion simulation model, specifically includes: Initialize the value network and policy network; Obtain the current navigation strategy and generate a second trajectory dataset corresponding to the current navigation strategy based on the ship motion simulation model; Based on the optimal reward function, calculate the instantaneous reward at each time step in the second trajectory dataset to obtain an empirical sample; Based on the empirical samples, the value network and the policy network are optimized, and the parameters of the value network and the policy network are updated. When the parameters of the updated policy network reach the preset second convergence condition, the optimal policy network is obtained.
6. The method according to claim 5, characterized in that, The step of optimizing the value network and the policy network based on the empirical samples, and updating the parameters of the value network and the policy network, specifically includes: Calculate the cumulative discounted return corresponding to the instantaneous reward at each time step in the experience sample; Calculate the advantage function based on the cumulative return from the discount and the value network; Calculate the policy loss based on the advantage function; Calculate the value loss based on the cumulative return from the discount and the valuation of the value network output; Based on the strategy loss and the value loss, the total loss function is obtained; Based on the total loss function, the value network and the policy network are optimized, and the parameters of the value network and the policy network are updated.
7. The method according to claim 1, characterized in that, The process of acquiring ship status data and generating ship motion control commands based on the optimal strategy network and the ship status data specifically includes: The ship status data is acquired, including the ship's position, heading, speed, and angular velocity. The ship's position, heading, speed, and angular velocity are input into the optimal strategy network, and the ship's motion control commands are output, which include the port and starboard rudder angles and the port and starboard main engine speeds.
8. An intelligent ship navigation control system based on maximum entropy inverse reinforcement learning, characterized in that, include: The ship motion simulation module is used to collect expert demonstration trajectory data, preprocess the expert demonstration trajectory data to obtain an expert trajectory dataset, and then construct a ship motion simulation model based on the expert trajectory dataset. The inverse reinforcement learning module is used to obtain the optimal reward function based on the expert trajectory dataset and the ship motion simulation model using the maximum entropy inverse reinforcement learning algorithm. The reinforcement learning module is used to obtain the optimal policy network based on the reinforcement learning algorithm, the optimal reward function, and the ship motion simulation model. The ship motion control module is used to acquire ship status data and generate ship motion control commands based on the optimal strategy network and the ship status data.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method of any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.