Predictable navigation obstacle avoidance control method and system for underactuated unmanned ship under dynamic environment interference

By performing dynamic model simulation and strategic model optimization on under-driven surface unmanned boats, the challenge of obstacle avoidance navigation in dynamic surface environments is solved, and predictable navigation and stable obstacle avoidance of unmanned boats are achieved.

CN120029293APending Publication Date: 2025-05-23SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510172677.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

There are several challenges in existing under-driven surface unmanned boats’ obstacle avoidance navigation methods in dynamic surface environments, including relying on precise mathematical models and dynamic parameters, modular systems lead to error accumulation, and lack of global chart information lead to navigation planning difficulties.

Method used

By simulating the dynamic model of under-driven water surface unmanned boats, dynamic characteristics and interference characteristics are obtained, action distribution is generated using the strategy model, state prediction model is constructed, and actions are screened and optimized to achieve predictable navigation obstacle avoidance control.

Benefits of technology

It improves the navigation stability and robustness of unmanned boats in dynamic environments, reduces the cost and time of strategy training, enhances the diversity and dynamic characteristics capture capabilities, and ensures stable obstacle avoidance navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029293A_ABST
    Figure CN120029293A_ABST
Patent Text Reader

Abstract

The invention provides a predictable navigation obstacle avoidance control method and system for an underactuated unmanned surface vehicle under dynamic environment interference. The predictable navigation obstacle avoidance control method comprises the following steps: S1, acquiring dynamic characteristics of the underactuated unmanned surface vehicle and interference characteristics in a specific environment; s2, using the strategy model to generate action distribution, and performing action sampling based on the action distribution; s3, constructing a state prediction model of the underactuated unmanned surface vehicle, and predicting the state of the underactuated unmanned surface vehicle under the action of the sampled action; s4, based on the predicted state of the underactuated water surface unmanned ship, screening the state of the underactuated water surface unmanned ship which does not meet the preset requirement and the corresponding action; and S5, evaluating the state of the underactuated unmanned surface vehicle which does not meet the preset requirement and the corresponding action of the underactuated unmanned surface vehicle by using the strategy evaluation model to obtain an evaluation value, and when the evaluation value does not meet the preset requirement, optimizing the strategy model, and repeatedly triggering the steps S2 to S5 until the evaluation value meets the preset requirement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of obstacle avoidance navigation technology, and in particular to a predictable navigation obstacle avoidance control method and system for an underactuated unmanned boat under dynamic environmental interference. Background Art

[0002] In recent years, unmanned surface vehicles have attracted widespread attention for their ability to perform various maritime missions, such as hydrographic surveys, marine resource exploration, and marine transportation, especially those with long endurance. Unmanned surface vehicles have been widely used in scientific research, commercial fields, and civil navigation lights. In order to meet the mission requirements in different marine environments, obstacle avoidance is a vital capability of underactuated surface vehicles. In complex reef areas, underactuated surface vehicles must navigate through environmental disturbances such as wind, waves, and currents, and rely on limited sensor data to avoid obstacles and reach their destination smoothly.

[0003] However, the existing methods for avoiding obstacles in dynamic water environments still have some difficulties, including:

[0004] 1) It relies on precise mathematical models and various assumptions, while the real-world ocean environment is highly uncertain, and actual conditions such as dynamic environmental disturbances do not always meet these assumptions;

[0005] 2) It relies on the accurate dynamic model and hydrodynamic parameters of the target unmanned boat. However, in fact, the hydrodynamic parameters of the unmanned boat are different in different navigation states, which are difficult to obtain accurately or even completely unknown.

[0006] 3) In order to ensure the stability of the system, a modular mechanism is usually used to divide the navigation system into sub-modules such as perception, planning and control, and each module performs its task independently. However, this modular approach may lead to error accumulation and coupling problems between modules, because the planned path may be inconsistent with the dynamic characteristics of the underactuated surface vehicle or the corresponding control strategy;

[0007] 4) The existing unmanned boat obstacle avoidance navigation method requires global nautical chart information to plan a complete navigation route in advance. However, in some unknown waters, complete and accurate nautical chart information is difficult to obtain.

[0008] Considering the above challenges, the obstacle avoidance and navigation technology of underactuated surface unmanned vehicles in dynamic water environments still needs further exploration and research. Summary of the invention

[0009] In view of the defects in the prior art, the purpose of the present invention is to provide a predictable navigation and obstacle avoidance control method and system for an under-actuated unmanned boat under dynamic environmental interference.

[0010] According to the present invention, a predictable navigation and obstacle avoidance control method for an underactuated unmanned boat under dynamic environmental interference is provided, comprising:

[0011] Step S1: simulating the dynamic model of the under-actuated unmanned surface vehicle, and obtaining the dynamic characteristics of the under-actuated unmanned surface vehicle and the interference characteristics under a specific environment based on the simulation of the dynamic model of the under-actuated unmanned surface vehicle;

[0012] Step S2: Based on the current state of the underactuated unmanned surface vehicle and the dynamic characteristics of the underactuated unmanned surface vehicle, the strategy model is used to generate an action distribution, and action sampling is performed based on the action distribution;

[0013] Step S3: constructing a state prediction model of the underactuated unmanned surface vehicle, and predicting the state of the underactuated unmanned surface vehicle under the action of the sampled actions based on the acquired dynamic characteristics of the underactuated unmanned surface vehicle and the interference characteristics in a specific environment;

[0014] Step S4: based on the predicted states of the under-actuated unmanned surface vehicles, screening the states of the under-actuated unmanned surface vehicles that do not meet the preset requirements and their corresponding actions;

[0015] Step S5: Use the strategy evaluation model to evaluate the state of the under-actuated unmanned surface vehicle that does not meet the preset requirements and its corresponding actions to obtain an evaluation value. When the evaluation value does not meet the preset requirements, optimize the strategy model and repeatedly trigger steps S2 to S5 until the evaluation value meets the preset requirements.

[0016] Preferably, the step S1 comprises:

[0017] Step S1.1: Based on the simulation of the underactuated unmanned surface vehicle dynamics model, the dynamic characteristics of the underactuated unmanned surface vehicle are extracted through a long short-term memory deep learning network; wherein the training objectives of the long short-term memory deep learning network include:

[0018]

[0019] Among them, MSE represents mean square error loss; Decoder representing the long short-term memory deep learning network; Represents the encoder of the long short-term memory deep learning network; s t ' represents the dynamic characteristics of underactuated unmanned surface vehicle;

[0020] Step S1.2: Based on the simulation of the underactuated unmanned surface vehicle dynamics model, multi-layer perception is used to capture the interference characteristics in the current environment;

[0021] The training objectives of the multi-layer perceptron include:

[0022]

[0023] Among them, MSE represents mean square error loss; represents a decoder; Indicates encoder; s″ t Indicates interference characteristics.

[0024] Preferably, step S2 comprises:

[0025] Step S2.1: construct a strategy model based on a multi-layer feedforward network;

[0026] Step S2.2: Based on the current state of the underactuated unmanned surface vehicle and the dynamic characteristics of the underactuated unmanned surface vehicle, the strategy model is used to generate the action distribution N(μ t ,σ t 2 );

[0027] Step S2.3: Perform action sampling based on the probability density of the action distribution.

[0028] Preferably, step S3 comprises:

[0029] Step S3.1: constructing a state prediction model of an underactuated unmanned surface vehicle with a multi-layer perception mechanism based on a multi-layer feedforward network; wherein the training objectives of the state prediction model of the underactuated unmanned surface vehicle include:

[0030]

[0031] Among them, MSE represents the mean square error loss function; F p represents the prediction network; F m represents the motion learning network, f di,t represents the dynamic characteristics of the underactuated unmanned surface vehicle at time t, f dy,t represents the interference feature at time t, a t Indicates the sampling action; s t+1 Indicates the predicted state at time t+1;

[0032] Step S3.2: generating motion features based on the acquired dynamic characteristics of the underactuated unmanned surface vehicle and the interference characteristics in a specific environment;

[0033] f m,t =F v (F q (f dy,t )·f k (f di,t ))

[0034] Among them, f m,t Indicates motion characteristics; F v Represents the value network; f dy,trepresents the interference feature at time t; F q represents the query network; F k represents a key-value network; f di,t represents the dynamic characteristics of the underactuated unmanned surface vehicle at time t;

[0035] Step S3.3: splicing the generated motion features and the sampled actions, inputting the state prediction model of the under-actuated unmanned surface vehicle, and predicting the state of the under-actuated unmanned surface vehicle;

[0036] s t+1 =F p (a t ,f m,t )

[0037] Among them, s t+1 Indicates the predicted state at time t+1; F p represents the state prediction model of underactuated unmanned surface vehicle; a t Indicates the sampling action; f m,t Indicates motion characteristics.

[0038] Preferably, step S5 comprises:

[0039] Step S5.1: construct a multi-layer perceptron based on a multi-layer feedforward network to build a strategy evaluation model; wherein the training objectives of the strategy evaluation model include:

[0040]

[0041] in, Express expectations, represents the sample library, Q φ represents the strategy evaluation model network, r represents the reward function, and γ represents the decay coefficient; represents the expected value of the policy evaluation network;

[0042] Step S5.2: Use the strategy evaluation model to evaluate the state of the underactuated unmanned surface vehicle that does not meet the preset requirements and its corresponding actions to obtain an evaluation value. When the evaluation value does not meet the preset requirements, optimize the strategy model and repeatedly trigger steps S2 to S5 until the evaluation value meets the preset requirements; wherein the optimization objectives of the strategy model include:

[0043]

[0044] Among them, α represents the hyperparameter coefficient and θ represents the parameters of the strategy module.

[0045] According to the present invention, a predictable navigation and obstacle avoidance control system for an underactuated unmanned boat under dynamic environmental interference is provided, comprising:

[0046] Module M1: Simulate the dynamic model of the under-actuated unmanned surface vehicle, and obtain the dynamic characteristics of the under-actuated unmanned surface vehicle and the interference characteristics in a specific environment based on the simulation of the dynamic model of the under-actuated unmanned surface vehicle;

[0047] Module M2: Based on the current state of the underactuated unmanned surface vehicle and the dynamic characteristics of the underactuated unmanned surface vehicle, the strategy model is used to generate action distribution, and action sampling is performed based on the action distribution;

[0048] Module M3: Construct a state prediction model for an underactuated unmanned surface vehicle. Based on the acquired dynamic characteristics of the underactuated unmanned surface vehicle and the interference characteristics in a specific environment, the state of the underactuated unmanned surface vehicle is predicted under the action of the sampled actions.

[0049] Module M4: based on the predicted state of the under-actuated unmanned surface vehicle, screen the state of the under-actuated unmanned surface vehicle that does not meet the preset requirements and its corresponding actions;

[0050] Module M5: Use the strategy evaluation model to evaluate the state of the under-actuated surface unmanned vehicle that does not meet the preset requirements and its corresponding actions to obtain an evaluation value. When the evaluation value does not meet the preset requirements, the strategy model is optimized and modules M2 to M5 are repeatedly triggered until the evaluation value meets the preset requirements.

[0051] Preferably, the module M1 comprises:

[0052] Module M1.1: Based on the simulation of the underactuated unmanned surface vehicle dynamics model, the dynamic characteristics of the underactuated unmanned surface vehicle are extracted through the long short-term memory deep learning network; wherein the training objectives of the long short-term memory deep learning network include:

[0053]

[0054] Among them, MSE represents mean square error loss; Decoder representing the long short-term memory deep learning network; Represents the encoder of the long short-term memory deep learning network; s t ' represents the dynamic characteristics of underactuated unmanned surface vehicle;

[0055] Module M1.2: Based on the simulation of the dynamic model of the underactuated unmanned surface vehicle, multi-layer perception is used to capture the interference characteristics in the current environment;

[0056] The training objectives of the multi-layer perceptron include:

[0057]

[0058] Among them, MSE represents mean square error loss; represents a decoder; represents the encoder; s' t ' indicates interference characteristics.

[0059] Preferably, the module M2 comprises:

[0060] Module M2.1: Building a strategy model based on a multi-layer feedforward network;

[0061] Module M2.2: Based on the current state of the underactuated unmanned surface vehicle and the dynamic characteristics of the underactuated unmanned surface vehicle, the strategy model is used to generate the action distribution N(μ t ,σ t 2 );

[0062] Module M2.3: Action sampling based on the probability density of action distribution.

[0063] Preferably, the module M3 comprises:

[0064] Module M3.1: Constructing a state prediction model of an underactuated unmanned surface vehicle with a multi-layer perception mechanism based on a multi-layer feedforward network; wherein the training objectives of the state prediction model of the underactuated unmanned surface vehicle include:

[0065]

[0066] Among them, MSE represents the mean square error loss function; F p represents the prediction network; F m represents the motion learning network, f di,t represents the dynamic characteristics of the underactuated unmanned surface vehicle at time t, f dy,t represents the interference feature at time t, a t Indicates the sampling action; s t+1 Indicates the predicted state at time t+1;

[0067] Module M3.2: Generate motion features based on the acquired dynamic characteristics of the underactuated surface unmanned vehicle and the interference characteristics in a specific environment;

[0068] f m,t =F v (F q (f dy,t )·f k (f di,t ))

[0069] Among them, f m,t Indicates motion characteristics; F v Represents the value network; f dy,t represents the interference feature at time t; F q represents the query network; F k represents a key-value network; fdi,t represents the dynamic characteristics of the underactuated unmanned surface vehicle at time t;

[0070] Module M3.3: Splice the generated motion features and sampled actions, input them into the state prediction model of the under-actuated unmanned surface vehicle, and predict the state of the under-actuated unmanned surface vehicle;

[0071] s t+1 =F p (a t ,f m,t )

[0072] Among them, s t+1 Indicates the predicted state at time t+1; F p represents the state prediction model of underactuated unmanned surface vehicle; a t Indicates the sampling action; f m,t Indicates motion characteristics.

[0073] Preferably, the module M5 comprises:

[0074] Module M5.1: Construct a strategy evaluation model based on a multi-layer feedforward network and a multi-layer perceptron; wherein the training objectives of the strategy evaluation model include:

[0075]

[0076] in, Express expectations, represents the sample library, Qφ represents the strategy evaluation model network, r represents the reward function, and γ represents the attenuation coefficient; represents the expected value of the policy evaluation network;

[0077] Module M5.2: Use the strategy evaluation model to evaluate the state of the underactuated unmanned surface vehicle that does not meet the preset requirements and its corresponding actions to obtain an evaluation value. When the evaluation value does not meet the preset requirements, optimize the strategy model and repeatedly trigger modules M2 to M5 until the evaluation value meets the preset requirements; wherein the optimization objectives of the strategy model include:

[0078]

[0079] Among them, α represents the hyperparameter coefficient and θ represents the parameters of the strategy module.

[0080] Compared with the prior art, the present invention has the following beneficial effects:

[0081] 1. The present invention uses a pre-training mechanism of the prediction model. When facing a new unmanned boat type, there is no need to start training from scratch. Instead, training can be performed based on the pre-training model, thereby effectively reducing the cost and time of strategy training;

[0082] 2. The present invention can effectively improve the deployment capability of the strategy in an unmanned boat platform with unknown dynamics through the unmanned boat state prediction model; the dynamic information missing due to unknown dynamics is actually supplemented by the prediction model, so even if the dynamic characteristics of the target unmanned boat are unknown, the proposed strategy can be well deployed;

[0083] 3. The present invention can effectively improve the diversity of strategies, that is, in each mission environment, the unmanned boat can obtain a variety of feasible strategies instead of only being able to generate one specific strategy;

[0084] 4. The present invention can effectively capture the unmanned boat's own dynamic characteristics and external interference, thereby improving the unmanned boat's robustness to environmental interference in a dynamic environment;

[0085] 5. The present invention can generate a variety of effective control behaviors under the influence of dynamic environmental interference to ensure stable obstacle avoidance navigation of the unmanned boat. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0087] Figure 1 This is a flow chart of the predictable navigation and obstacle avoidance control method for an underactuated unmanned boat under dynamic environmental interference.

[0088] Figure 2 During the strategy exploration process, the under-actuated surface unmanned vehicle will perform obstacle avoidance and will give you a coordinate system diagram of the task.

[0089] Figure 3 This is a flow chart of the predictable navigation and obstacle avoidance control method for an underactuated unmanned boat under dynamic environmental interference.

[0090] Figures 4a to 4e This is a schematic diagram of the state prediction results of the state prediction module for the underactuated surface unmanned vehicle under dynamic environmental interference.

[0091] Figure 5a to Figure 5c It is the thrust and torque control command curve of the under-actuated surface unmanned vehicle under dynamic environmental interference in the simulation experiment, as well as the schematic diagram of the obstacle avoidance navigation of the under-actuated surface unmanned vehicle under dynamic environmental interference. DETAILED DESCRIPTION

[0092] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0093] Example 1

[0094] According to the present invention, a predictable navigation and obstacle avoidance control method for an underactuated unmanned boat under dynamic environmental interference is provided. Figures 1 to 3 As shown, including:

[0095] Step S1: simulating the dynamic model of the under-actuated unmanned surface vehicle, and obtaining the dynamic characteristics of the under-actuated unmanned surface vehicle and the interference characteristics under a specific environment based on the simulation of the dynamic model of the under-actuated unmanned surface vehicle;

[0096] Step S2: Based on the current state of the underactuated unmanned surface vehicle and the dynamic characteristics of the underactuated unmanned surface vehicle, the strategy model is used to generate an action distribution, and action sampling is performed based on the action distribution;

[0097] Step S3: constructing a state prediction model of the underactuated unmanned surface vehicle, and predicting the state of the underactuated unmanned surface vehicle under the action of the sampled actions based on the acquired dynamic characteristics of the underactuated unmanned surface vehicle and the interference characteristics in a specific environment;

[0098] Step S4: based on the predicted states of the under-actuated unmanned surface vehicles, screening the states of the under-actuated unmanned surface vehicles that do not meet the preset requirements and their corresponding actions;

[0099] Step S5: Use the strategy evaluation model to evaluate the state of the under-actuated unmanned surface vehicle that does not meet the preset requirements and its corresponding actions to obtain an evaluation value. When the evaluation value does not meet the preset requirements, optimize the strategy model and repeatedly trigger steps S2 to S5 until the evaluation value meets the preset requirements.

[0100] Specifically, the step S1 includes:

[0101] The dynamic model of the underactuated unmanned surface vehicle is simulated. Based on the simulation of the dynamic model of the underactuated unmanned surface vehicle, the learning module is used to obtain the dynamic characteristics of the underactuated unmanned surface vehicle and the interference characteristics in a specific environment; the navigation state change of the underactuated unmanned surface vehicle after performing a given action is affected by its dynamic characteristics and the interference characteristics of the surrounding environment. Therefore, these two characteristics are extracted by analyzing the historical time behavior and state changes of the target underactuated unmanned surface vehicle. The dynamic characteristics are related to the structure of the underactuated unmanned surface vehicle. When the ship's attitude does not change significantly, the dynamic characteristics are usually considered to be stable. However, due to strong environmental interference, there is a deviation between the actual motion response of the real underactuated unmanned surface vehicle and the dynamic characteristics. Although the environmental interference shows strong dynamics in the short term, it can usually be approximated by a normal distribution over a longer time range. Therefore, the dynamic characteristics of the underactuated unmanned surface vehicle are implicitly described in the long-term sequence state.

[0102] Wherein, the learning module includes: dynamic learning module F dy and interference extraction module F di The dynamic learning module extracts the dynamic features of the underactuated unmanned surface vehicle from the long-term state. dy =F dy (st′), where s t ′=[s t-m ,s t-m+1 ,…,s t ] represents the long-term state, and m represents the sequence length of the long-term state. Environmental disturbances, including wind, waves, and currents, introduce changing external factors into the navigation process, showing a high degree of randomness and uncertainty. However, some external disturbances, such as wind and current, maintain a certain regularity in the short term. Therefore, a disturbance extraction module F is developed. di , which is used to learn the impact of current dynamic disturbance on the behavioral response of the unmanned boat and generate disturbance features from the short-term state.

[0103] More specifically, the dynamic learning module and the interference extraction module are both pre-trained through the autoencoder mechanism. Specifically, each module consists of an encoder and a decoder. The encoder takes the sequence state as input to generate implicit features, while the decoder reconstructs the input state using the implicit features. The LSTM method is used as the backbone network of the two modules. The pre-training objectives of the two modules are defined as:

[0104]

[0105] in, and denote the encoder and decoder of the dynamic learning module, respectively. and denote the encoder and decoder of the interference learning module respectively, and MSE denotes the mean square error loss. The generated features and As dynamic features f dy,t and interference feature f di,t .

[0106] Specifically, the step S2 includes:

[0107] Step S2.1: Based on the multi-layer feedforward network, a multi-layer perceptron is constructed to establish a strategy model. The strategy model includes: an input layer, a hidden layer, and an output layer. The input layer is responsible for receiving the original input data, the hidden layer further extracts the features of the input data, and the output layer generates the final output. Each layer of the network contains a large number of neurons, and its operation process can be expressed as follows:

[0108]

[0109] in, represents the weighted output of the ith neuron in the lth layer, represents the weight from the jth neuron to the ith neuron in the l-1th layer, represents the activation value of the jth neuron in the l-1th layer, represents the bias term of the i-th neuron in the l-th layer, n l-1 Represents the number of neurons in the l-1th layer of the network. The network outputs of the input layer and the hidden layer will then be input into the Relu activation function to obtain the output x of the network in this layer. In the policy model, there are one layer each of the input layer, hidden layer, and output layer. The number of neurons in the input layer is 45, the number of neurons in the hidden layer is 16, and the number of neurons in the output layer is 2.

[0110] In this embodiment, every 10 time steps, the policy module is trained according to the following formula:

[0111]

[0112] Step S2.2: At each time step, the current state s t Obtained from the environment and input into the policy module to generate the action distribution N(μ t ,σ t 2 );

[0113] Step S2.3: Establish a behavior sampling mechanism to extract the behavior from the distribution a t ~N(μ t ,σ t 2 ) to sample multiple actions.

[0114] In this embodiment, in order to enhance the diversity of actions, a novel action sampling mechanism is introduced. During the training process, not only one action a is sampled at each time step. t , but sampling multiple potential actions a t =[a 1,t ,…,a i,t ,…,a i,l ], where i represents the index of the sampled action and l represents the number of all sampled actions.

[0115] Specifically, step S3 includes:

[0116] A state prediction model of an underactuated surface unmanned vehicle is constructed by building a multi-layer perceptron based on a multi-layer feedforward network; the number of neurons in the input layer is 34, the number of neurons in the hidden layer is 32, and the number of neurons in the output layer is 13; the training objective can be expressed as:

[0117]

[0118] Using attention mechanism to fuse dynamic features f dy and interference feature f di Specifically, a query network F is constructed q , a key-value network F k and a value network F v Given the dynamic features, the query network and the value network generate query features f q,t =F q (f dy,t ) and value characteristics f v,t =F v (f dy,t ), while the key-value network uses interference features to generate f k,t =F k (f di,t ). Next, generate the attention feature f att,t =softmax(f q,t f k,t ), which is used to describe the influence of environmental interference on the dynamic model, where · represents the dot product process. In the generated attention feature f att,t Based on this, a motion feature f is generated to describe the current motion characteristics of the target underactuated surface unmanned vehicle. m,t , the formula is:

[0119] f m,t =F v (F q (f dy,t )·F k (f di,t ))

[0120] The motion feature f m With the generated action a t Concatenate and input into the prediction module to predict the next state s t+1 =F p (a t ,f m,t ), where F p Represents the prediction network and then all sampled actions and the current state s t will be input to the prediction module to predict the corresponding next time step state s i,t+1 =F p (a i,t ,f m,t ).

[0121] Specifically, step S4 includes:

[0122] Through the reward function r(s t ) Calculate the reward for each predicted state; choose the action with the smallest reward as the action to be executed:

[0123]

[0124] In most cases, all actions obtained through sampling are associated with higher rewards, indicating that they can bring higher rewards. When the action with the lowest reward can still obtain a considerable evaluation value, it means that the potential evaluation value of other sampled actions is even higher. This shows that the sampled diverse actions can each effectively complete the task.

[0125] Specifically, step S5 includes:

[0126] Step S5.1: Construct a multi-layer perceptron based on a multi-layer feedforward network to build a strategy evaluation model; the number of neurons in the input layer is 47, the number of neurons in the hidden layer is 16, and the number of neurons in the output layer is 1; the training objective can be expressed as:

[0127]

[0128] In order to verify the implementation effect of the present invention, a simulation experiment was conducted. The target ship type of the simulation is Cybership II, which has a length of 1.255m and is a typical under-actuated unmanned surface vessel.

[0129] Figures 4a to 4e The state prediction result of the unmanned boat state prediction module is shown in Figure 1. The black solid line represents the actual navigation state of the unmanned boat, and the red dotted line represents the state prediction result of the unmanned boat state prediction module. The figure shows the predicted state of the one-step forward prediction of 10 consecutive time steps of multiple states such as propulsion speed, lateral speed and trajectory. It can be proved from the prediction result figure that the state prediction module can accurately predict the future navigation state of the unmanned boat.

[0130] Figure 5a to Figure 5c The control command curve of the unmanned boat in the process of navigation and obstacle avoidance, as well as the navigation comparison results with other methods. It can be seen from the figure that the generated control command can effectively realize the robust navigation task under dynamic environmental interference. Under diverse environmental interference, the three classic reinforcement learning methods of DDPG, SAC, and DQN are difficult to resist the influence of environmental interference and collide, but the method proposed in this invention can avoid obstacles well and reach the destination.

[0131] The present invention also provides a predictable navigation and obstacle avoidance control system for an under-actuated unmanned boat under dynamic environmental interference. The predictable navigation and obstacle avoidance control system for an under-actuated unmanned boat under dynamic environmental interference can be realized by executing the process steps of the predictable navigation and obstacle avoidance control method for an under-actuated unmanned boat under dynamic environmental interference, that is, those skilled in the art can understand the predictable navigation and obstacle avoidance control method for an under-actuated unmanned boat under dynamic environmental interference as a preferred implementation mode of the predictable navigation and obstacle avoidance control system for an under-actuated unmanned boat under dynamic environmental interference.

[0132] Example 2

[0133] Embodiment 2 is a preferred embodiment of Embodiment 1

[0134] According to the present invention, a predictable navigation and obstacle avoidance control method for an underactuated unmanned boat under dynamic environmental interference is provided, comprising:

[0135] The present invention proposes a predictive deep reinforcement learning method for the control of surface unmanned vehicles, aiming to cope with the problems of navigation under dynamic environmental disturbances and obstacle avoidance under local observations. The present invention introduces a motion prediction module for predicting the future state of the surface unmanned vehicle. The motion of the underactuated surface unmanned vehicle is affected by both its own dynamics and environmental disturbances. Although the dynamics of the target surface unmanned vehicle are relatively stable, the environmental disturbances show a high degree of randomness in the long term. Therefore, the prediction module learns the dynamic characteristics of the surface unmanned vehicle from its long-term historical behavior, while extracting the influence of environmental disturbances from its short-term temporal state.

[0136] On the one hand, a state prediction module for underactuated unmanned surface vehicles under dynamic disturbance is established, including:

[0137] Establish an underactuated unmanned surface vehicle dynamics learning module to learn the inherent motion response characteristics of the unmanned surface vehicle;

[0138] Establish a dynamic environment interference feature extraction module to determine the impact of the current dynamic environment interference on the under-actuated unmanned surface vehicle;

[0139] Construct a motion learning module, integrate the obtained unmanned boat dynamic characteristics and environmental interference effects, and fit the current unmanned real-time dynamic response mode;

[0140] Build a state prediction model to predict the future navigation state of the unmanned boat after a given action is executed based on the real-time dynamic response mode and control actions.

[0141] Construct an entropy maximization strategy optimization mechanism, including:

[0142] Build a policy learning module to generate behavior distribution based on state input;

[0143] Construct a strategy evaluation module to evaluate the quality of control instructions obtained by the strategy to guide the optimization of the strategy module;

[0144] Combined with the entropy maximization reinforcement learning method, the policy module outputs the mean and variance of the action distribution. During the training process, multiple actions are sampled from the distribution and input into the prediction module to predict the corresponding future state. In the policy optimization process, we should not only pursue better task performance, but also obtain greater action sampling entropy to achieve more diverse navigation behavior generation.

[0145] A behavior sampling mechanism is proposed, and a reward function is designed for the obstacle avoidance navigation task to evaluate the performance of the predicted future state, thereby selecting the behavior with the worst performance. Through continuous optimization, the worst behavior sampled from the behavior distribution can also achieve good performance, thereby improving the task stability of the strategy and its robustness to dynamic environmental interference.

[0146] Those skilled in the art know that, in addition to implementing the system, device and its various modules provided by the present invention in a purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Therefore, the system, device and its various modules provided by the present invention can be considered as a hardware component, and the modules included therein for implementing various programs can also be considered as structures within the hardware component; the modules for implementing various functions can also be considered as both software programs for implementing the method and structures within the hardware component.

[0147] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A predictable navigation and obstacle avoidance control method for an underactuated unmanned boat under dynamic environmental interference, characterized in that: include: Step S1: simulating the dynamic model of the under-actuated unmanned surface vehicle, and obtaining the dynamic characteristics of the under-actuated unmanned surface vehicle and the interference characteristics under a specific environment based on the simulation of the dynamic model of the under-actuated unmanned surface vehicle; Step S2: Based on the current state of the underactuated unmanned surface vehicle and the dynamic characteristics of the underactuated unmanned surface vehicle, the strategy model is used to generate an action distribution, and action sampling is performed based on the action distribution; Step S3: constructing a state prediction model of the underactuated unmanned surface vehicle, and predicting the state of the underactuated unmanned surface vehicle under the action of the sampled actions based on the acquired dynamic characteristics of the underactuated unmanned surface vehicle and the interference characteristics in a specific environment; Step S4: based on the predicted states of the under-actuated unmanned surface vehicles, screening the states of the under-actuated unmanned surface vehicles that do not meet the preset requirements and their corresponding actions; Step S5: Use the strategy evaluation model to evaluate the state of the under-actuated unmanned surface vehicle that does not meet the preset requirements and its corresponding actions to obtain an evaluation value. When the evaluation value does not meet the preset requirements, optimize the strategy model and repeatedly trigger steps S2 to S5 until the evaluation value meets the preset requirements.

2. The predictable navigation and obstacle avoidance control method for an underactuated unmanned vehicle under dynamic environmental interference according to claim 1 is characterized in that: The step S1 comprises: Step S1.1: Based on the simulation of the underactuated unmanned surface vehicle dynamics model, the dynamic characteristics of the underactuated unmanned surface vehicle are extracted through a long short-term memory deep learning network; wherein the training objectives of the long short-term memory deep learning network include: Among them, MSE represents mean square error loss; Decoder representing the long short-term memory deep learning network; Represents the encoder of the long short-term memory deep learning network; s t ' represents the dynamic characteristics of underactuated unmanned surface vehicle; Step S1.2: Based on the simulation of the underactuated unmanned surface vehicle dynamics model, multi-layer perception is used to capture the interference characteristics in the current environment; The training objectives of the multi-layer perceptron include: Among them, MSE represents mean square error loss; represents a decoder; represents the encoder; s' t ' indicates interference characteristics.

3. The predictable navigation and obstacle avoidance control method for an underactuated unmanned vehicle under dynamic environmental interference according to claim 1 is characterized in that: The step S2 comprises: Step S2.1: construct a strategy model based on a multi-layer feedforward network; Step S2.2: Based on the current state of the underactuated unmanned surface vehicle and the dynamic characteristics of the underactuated unmanned surface vehicle, the strategy model is used to generate the action distribution N(μ t ,σ t 2 ); Step S2.3: Perform action sampling based on the probability density of the action distribution.

4. The predictable navigation and obstacle avoidance control method for an underactuated unmanned vehicle under dynamic environmental interference according to claim 1 is characterized in that: The step S3 comprises: Step S3.1: constructing a state prediction model of an underactuated unmanned surface vehicle with a multi-layer perception mechanism based on a multi-layer feedforward network; wherein the training objectives of the state prediction model of the underactuated unmanned surface vehicle include: Among them, MSE represents the mean square error loss function; F p represents the prediction network; F m represents the motion learning network, f di,t represents the dynamic characteristics of the underactuated unmanned surface vehicle at time t, f dy,t represents the interference feature at time t, a t Indicates the sampling action; s t+1 Indicates the predicted state at time t+1; Step S3.2: generating motion features based on the acquired dynamic characteristics of the underactuated unmanned surface vehicle and the interference characteristics in a specific environment; f m,t =F v (F q (f dy,t )·F k (f di,t )) Among them, f m,t Indicates motion characteristics; F v Represents the value network; f dy,t represents the interference feature at time t; F q represents the query network; F k represents a key-value network; f di,t represents the dynamic characteristics of the underactuated unmanned surface vehicle at time t; Step S3.3: splicing the generated motion features and the sampled actions, inputting the state prediction model of the under-actuated unmanned surface vehicle, and predicting the state of the under-actuated unmanned surface vehicle; s t+1 =F p (a t ,f m,t ) Among them, s t+1 Indicates the predicted state at time t+1; F p represents the state prediction model of underactuated unmanned surface vehicle; a t Indicates the sampling action; f m,t Indicates motion characteristics.

5. The predictable navigation and obstacle avoidance control method for an underactuated unmanned vehicle under dynamic environmental interference according to claim 1, characterized in that: The step S5 comprises: Step S5.1: construct a multi-layer perceptron based on a multi-layer feedforward network to build a strategy evaluation model; wherein the training objectives of the strategy evaluation model include: in, Express expectations, represents the sample library, Q φ represents the strategy evaluation model network, r represents the reward function, and γ represents the decay coefficient; represents the expected value of the policy evaluation network; Step S5.2: Use the strategy evaluation model to evaluate the state of the underactuated unmanned surface vehicle that does not meet the preset requirements and its corresponding actions to obtain an evaluation value. When the evaluation value does not meet the preset requirements, optimize the strategy model and repeatedly trigger steps S2 to S5 until the evaluation value meets the preset requirements; wherein the optimization objectives of the strategy model include: Among them, α represents the hyperparameter coefficient and θ represents the parameters of the strategy module.

6. A predictable navigation and obstacle avoidance control system for an underactuated unmanned boat under dynamic environmental interference, characterized in that: include: Module M1: Simulate the dynamic model of the under-actuated unmanned surface vehicle, and obtain the dynamic characteristics of the under-actuated unmanned surface vehicle and the interference characteristics in a specific environment based on the simulation of the dynamic model of the under-actuated unmanned surface vehicle; Module M2: Based on the current state of the underactuated unmanned surface vehicle and the dynamic characteristics of the underactuated unmanned surface vehicle, the strategy model is used to generate action distribution, and action sampling is performed based on the action distribution; Module M3: Construct a state prediction model for an underactuated unmanned surface vehicle. Based on the acquired dynamic characteristics of the underactuated unmanned surface vehicle and the interference characteristics in a specific environment, the state of the underactuated unmanned surface vehicle is predicted under the action of the sampled actions. Module M4: based on the predicted state of the under-actuated unmanned surface vehicle, screen the state of the under-actuated unmanned surface vehicle that does not meet the preset requirements and its corresponding actions; Module M5: Use the strategy evaluation model to evaluate the state of the under-actuated surface unmanned vehicle that does not meet the preset requirements and its corresponding actions to obtain an evaluation value. When the evaluation value does not meet the preset requirements, the strategy model is optimized and modules M2 to M5 are repeatedly triggered until the evaluation value meets the preset requirements.

7. The predictable navigation and obstacle avoidance control system for an underactuated unmanned boat under dynamic environmental interference according to claim 6, characterized in that: The module M1 comprises: Module M1.1: Based on the simulation of the underactuated unmanned surface vehicle dynamics model, the dynamic characteristics of the underactuated unmanned surface vehicle are extracted through the long short-term memory deep learning network; wherein the training objectives of the long short-term memory deep learning network include: Among them, MSE represents mean square error loss; Decoder representing the long short-term memory deep learning network; Represents the encoder of the long short-term memory deep learning network; s t ' represents the dynamic characteristics of underactuated unmanned surface vehicle; Module M1.2: Based on the simulation of the dynamic model of the underactuated unmanned surface vehicle, multi-layer perception is used to capture the interference characteristics in the current environment; The training objectives of the multi-layer perceptron include: Among them, MSE represents mean square error loss; represents a decoder; represents the encoder; s' t ' indicates interference characteristics.

8. The predictable navigation and obstacle avoidance control system for an underactuated unmanned boat under dynamic environmental interference according to claim 6, characterized in that: The module M2 comprises: Module M2.1: Building a strategy model based on a multi-layer feedforward network; Module M2.2: Based on the current state of the underactuated unmanned surface vehicle and the dynamic characteristics of the underactuated unmanned surface vehicle, the strategy model is used to generate the action distribution N(μ t ,σ t 2 ); Module M2.3: Action sampling based on the probability density of action distribution.

9. The predictable navigation and obstacle avoidance control system for an underactuated unmanned boat under dynamic environmental interference according to claim 6, characterized in that: The module M3 comprises: Module M3.1: Constructing a state prediction model of an underactuated unmanned surface vehicle with a multi-layer perception mechanism based on a multi-layer feedforward network; wherein the training objectives of the state prediction model of the underactuated unmanned surface vehicle include: Among them, MSE represents the mean square error loss function; F p represents the prediction network; F m represents the motion learning network, f di,t represents the dynamic characteristics of the underactuated unmanned surface vehicle at time t, f dy,t represents the interference feature at time t, a t Indicates the sampling action; s t+1 Indicates the predicted state at time t+1; Module M3.2: Generate motion features based on the acquired dynamic characteristics of the underactuated surface unmanned vehicle and the interference characteristics in a specific environment; f m,t =F v (F q (f dy,t )·f k (f di,t )) Among them, f m,t Indicates motion characteristics; F v Represents the value network; f dy,t represents the interference feature at time t; F q represents the query network; F k represents a key-value network; f di,t represents the dynamic characteristics of the underactuated unmanned surface vehicle at time t; Module M3.3: Splice the generated motion features and sampled actions, input them into the state prediction model of the under-actuated unmanned surface vehicle, and predict the state of the under-actuated unmanned surface vehicle; s t+1 =F p (a t ,f m,t ) Among them, s t+1 Indicates the predicted state at time t+1; F p represents the state prediction model of underactuated unmanned surface vehicle; a t Indicates the sampling action; f m,t Indicates motion characteristics.

10. The predictable navigation and obstacle avoidance control system for an underactuated unmanned boat under dynamic environmental interference according to claim 6, characterized in that: The module M5 comprises: Module M5.1: Construct a strategy evaluation model based on a multi-layer feedforward network and a multi-layer perceptron; wherein the training objectives of the strategy evaluation model include: in, Express expectations, represents the sample library, Q φ represents the strategy evaluation model network, r represents the reward function, and γ represents the decay coefficient; represents the expected value of the policy evaluation network; Module M5.2: Use the strategy evaluation model to evaluate the state of the underactuated unmanned surface vehicle that does not meet the preset requirements and its corresponding actions to obtain an evaluation value. When the evaluation value does not meet the preset requirements, optimize the strategy model and repeatedly trigger modules M2 to M5 until the evaluation value meets the preset requirements; wherein the optimization objectives of the strategy model include: Among them, α represents the hyperparameter coefficient and θ represents the parameters of the strategy module.

Citation Information

Cited By

  • Unmanned aerial vehicle detection and countering method and system based on deep learning

    CN121000332A