An intelligent control method for unmanned boats based on improved low-entropy strategy and network structure
By introducing improved low-entropy strategies and network structure into the intelligent control method of unmanned boats, combining the low-entropy action selection mechanism and improved LSTM network, the problem that unmanned boats are difficult to achieve multi-task goals in complex marine environments is solved, and more efficient autonomous control and path planning are achieved.
Patent Information
- Application Number
- CN202510360666.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The existing intelligent control methods of unmanned boats are difficult to achieve safe, fast and energy-saving integrated autonomous navigation in complex marine environments, and cannot take into account the needs of multi-task targets.
The intelligent control method of unmanned boats based on improved low entropy strategy and network structure is adopted. By setting the action space and state space, the reward function for multi-task targets is designed, and the unmanned boat controller combining the low entropy action selection mechanism and the improved LSTM network are used for intelligent control.
It realizes intelligent control of unmanned boats to meet the needs of multi-task targets in complex waters, improves the autonomous control capabilities and path planning efficiency of unmanned boats, and enhances the adaptability to complex environments.
Smart Images

Figure CN119902533B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent control and ship and ocean engineering, and in particular to an intelligent control method for an unmanned boat based on an improved low-entropy strategy and network structure. Background Technique
[0002] In recent years, with the rapid development of unmanned intelligent control technology, surface unmanned boats have received increasing attention. Compared with drones and unmanned vehicles, the research on unmanned boats is still in the rapid development stage. As a small surface operation task platform with capabilities such as autonomous planning and autonomous navigation, in a complex and changeable water area environment, its movement is non-linear and the interference it receives is uncertain. Therefore, the autonomous navigation and intelligent control technology of unmanned boats is particularly important.
[0003] Patent CN114077258B discloses an unmanned boat pose control method based on the reinforcement learning PPO2 algorithm. This method trains the pose controller of the unmanned boat using the PPO2 algorithm by setting the action space and state space, setting the reward function, and designing a deep neural network architecture, and can adaptively adjust its pose. Patent CN115454092B discloses an unmanned boat path planning method and system based on an improved RRT algorithm. This method uses the Collision function to detect obstacle collisions for the new node Qnew; uses the collision function to detect obstacle collisions for the new node Qnew, and selects the termination detection method of the algorithm according to the end connection probability q to backtrack the path from the end point Qgoal to the starting point Qstart in the random tree T, which can reduce the unmanned boat path planning time. Patent application CN117991780A discloses an unmanned boat path optimization method based on the GWO algorithm. This method includes: selecting a mixture of a great circle route and a rhumb line as the initial; constructing the objective function; optimizing with the GWO algorithm, and can achieve global optimal path planning, etc.
[0004] However, the machine learning methods adopted in the above methods are not only difficult to control unmanned boats to meet the comprehensive autonomous navigation requirements of safety, speed and energy saving in a complex marine environment, but also unable to take into account the multi-task objective requirements of unmanned boats at the same time.
[0005] Therefore, there is an urgent need for an improved intelligent control method for unmanned boats to meet the multi-task objective requirements of unmanned boats in a complex marine environment and achieve the optimal intelligent control effect at the same time. Summary of the Invention
[0006] The purpose of the present invention is to provide an intelligent control method for an unmanned boat based on an improved low-entropy strategy and network structure that can achieve multi-task objectives in complex waters.
[0007] The object of the present invention can be achieved by the following technical solutions:
[0008] An intelligent control method for an unmanned boat based on an improved low-entropy strategy and network structure, comprising the following steps:
[0009] Set the action space and state space of the unmanned boat, set the reward function according to the multi-task target requirements of the unmanned boat, and perform intelligent control using an unmanned boat controller that combines a low-entropy action selection mechanism and an improved LSTM network until reaching the target point to complete the intelligent control process of the unmanned boat;
[0010] Among them, the unmanned boat controller includes an Actor network, a Critic network, and an improved LSTM network. The process of the unmanned boat controller performing intelligent control includes:
[0011] Collect a batch of data from the experience replay pool, input it into the improved LSTM network for processing, then output the control action through the Actor network, and use the low-entropy action selection mechanism to select the low-entropy action to obtain the control output u(t) to control the operation of the unmanned boat, where the experience replay pool is formed by interacting with the environment and combining the action space, state space, and reward function;
[0012] Collect a batch of data from the experience replay pool, input it into the Critic network for value estimation, and guide the Actor network to output the control action;
[0013] Repeat the above control process until the unmanned boat reaches the target point.
[0014] Further, the action space of the unmanned boat includes the left motor control rate u l (t), the right motor control rate u r (t), the distance d t from the target point, and the distance d o from the obstacle. The state space includes the linear velocity v of the unmanned boat, the angular velocity φ, the distance d t from the target point, the distance d o from the obstacle, the left motor control rate u l (t), and the right motor control rate u r (t).
[0015] Further, the expression of the reward function is:
[0016] R = R 1 + R 2 + R 3 + R 4
[0017]
[0018] Wherein:
[0019] R 1 is the reward function for the unmanned boat to reach the target point, and the expression is: R 1 = w 1 r reach ,
[0020] R 2 is the reward function for the distance of the unmanned boat away from the obstacle, and the expression is:
[0021] R 3 is the reward function for the linear velocity and angular velocity control of the unmanned boat, and the expression is:
[0022] R 3 = -W 3 ||v current - v previous || - w 5 ||φ current - φ previous ||,
[0023] R 4 is the reward function for the sailing time control of the unmanned boat, and the expression is: R 4 = - w 4 v current - w 6 r d ,
[0024] In the formula, R is the total reward function, w 1 is the weight of the reward for the unmanned boat to reach the target point in the total reward, r reach is the reward for the unmanned boat to reach the target point, w 2 is the weight of the reward for the distance of the unmanned boat away from the obstacle in the total reward, d is the distance when the unmanned boat is within the set distance from the obstacle, r o is the reward for the unmanned boat to collide with the obstacle, w 3 is the weight of the reward for the linear velocity stability of the unmanned boat in the total reward, v current is the current linear velocity of the unmanned boat, v previous is the linear velocity of the unmanned boat at the previous moment, w 4 is the weight of the reward for the current linear velocity of the unmanned boat in the total reward, w 5 is the weight of the reward for the angular velocity stability of the unmanned boat in the total reward, φ current is the current angular velocity of the unmanned boat, φ previous is the angular velocity of the unmanned boat at the previous moment, w 6 is the weight of the reward for the distance between the unmanned boat and the target point in the total reward; r d is the reward for the distance between the unmanned boat and the target point.
[0025] Furthermore, the Actor network is a dual-policy network, including an Actor_1 sub-network and an Actor_2 sub-network. The Actor_1 sub-network and the Actor_2 sub-network interact with the environment in parallel and output control actions.
[0026] Furthermore, the Critic network includes a Critic_1 sub-network and a Critc_2 sub-network. The Critic_1 sub-network and the Critc_2 sub-network perform the value estimation in parallel, where the value estimation results include the state value function V(s t ) and the state-action value function Q(s t , a t ), and the expressions are respectively:
[0027]
[0028] In the formula, is the expectation of the action a t under the policy π, α is the temperature coefficient, and log(π(a|s)) is the logarithmic probability of the policy π selecting the action a in the state s. is the expectation of the next moment state s t+1 under the transfer function D, r(s t , a t ) is the immediate reward for executing the action a t in the state s t , γ is the discount factor, and V(s t+1 ) is the state value estimation of the next moment state.
[0029] Furthermore, the operation expression for selecting a low-entropy action using the low-entropy action selection mechanism is:
[0030]
[0031] In the formula, H(π(.|s t )) is the entropy of the policy π in the state s t , is to select the minimum entropy value in the policy π in the state s t , i is the expectation of the action a under the policy π i , log(π t (a i |s t )) is the logarithmic probability of the policy π t selecting the action a i in the state s t , t and is the policy πi At state s t select action a t with a weight, where a is the action with the lowest entropy value, is to select the policy π with the lowest entropy value i and select action a.
[0032] Furthermore, the improved LSTM network improves the update method of the cell state during the input gate calculation process. The updated expression of the improved cell state is:
[0033] c t = f t · c t-1 + i t · tanh(W c [h t-1 , x t + b c ) + M t ,
[0034] M t = sigmoid(M t-1 · c t-1 ),
[0035] In the formula, c t is the cell state of the improved LSTM network structure at the current time step, f t is the input of the forget gate, c t-1 is the cell state at the previous moment, i t is the output of the input gate, tanh is the hyperbolic tangent function, W c is the weight matrix of the hidden state, which is updated by the backpropagation algorithm, h t-1 is the hidden state at the previous moment, x t is the input at the current time step, b c is the bias term, which is updated by the backpropagation algorithm, M t is the memory matrix at the previous moment, sigmoid is the sigmoid function operation, M t-1 is the memory matrix at the previous moment.
[0036] Furthermore, the guidance for the Actor network to output control actions means updating the network parameters based on the value evaluation results to select the optimal control actions. The specific steps include:
[0037] For each batch of data, after being processed by the improved LSTM network and the Actor network, the network parameters of the unmanned boat controller are updated in combination with the value evaluation results. The expression of the loss function for network parameter update is:
[0038] where, L Q (ω) is the loss function of the neural network parameter ω, and ω is the neural network parameter; (s t , a t , r t , s t+1 ) is the sample data sampled from the experience replay pool, i.e., R, s t is the state at the current time step, a t is the action at the current time step, r t is the immediate reward for executing the action a t , s t+1 is the state at the next moment; E is the expectation of sampling (s t , a t , r t , s t+1 ) from the experience replay pool, i.e., R and taking the expectation under the policy π θ , R is all the data stored in the experience replay pool, a t+1 is the action at the next moment, π θ (.|s t+1 ) is the policy π with parameter θ at the state s t+1 , Q ω (s t , a t ) is the Q value at the current time step, γ is the discount factor, is the expected return of the Q function with a smaller Q value at the next moment, α is the temperature coefficient, log(π(a t+1 |s t+1 )) is the logarithmic probability of the policy π selecting the action a t+1 at the state s t+1 .
[0039] Further, the objective function for updating the unmanned boat controller is:
[0040]
[0041] where, J(π) is the objective function for policy optimization, E is the expectation of the target value of the state-action group (s t , a t ) that follows the joint distribution of states and actions under the policy π, ρ π is the joint distribution of states and actions under the policy π; γ is the discount factor, s t is the state at the current time step, a t is the action at the current time step, r(s t , a t ) is the immediate reward for executing the action a t , α is the temperature coefficient, H(.|s t ) is the policy π at the state s tEntropy under
[0042] Furthermore, the training process of the unmanned boat controller includes:
[0043] 1) Add wind and current disturbances to the environment, where the expression of wind disturbance in the environment is:
[0044]
[0045] The expression of current disturbance in the environment is:
[0046]
[0047] In the formula, F wind is the wind disturbance, F flow is the current disturbance, C wind is the wind force coefficient, θ is the heading angle of the unmanned boat, V wind is the wind speed, C flow is the water current coefficient, β is the angle between the water current direction and the heading of the unmanned boat, V flow is the water current velocity;
[0048] 2) Set the training period N;
[0049] 3) Interact with the environment, and put the interaction information into the experience replay pool in chronological order;
[0050] 4) Judge whether the experience replay pool is full. If it is, extract a batch of data from the experience replay pool to iteratively update the network parameters. If not, execute step 2);
[0051] 5) Judge whether the set training period N is reached,
[0052] If it is, further judge whether the multi-objective requirements are met. If so, output the control result and obtain the control effect. If not, execute step 3);
[0053] If not, execute step 3) until the training period N is reached and the multi-objective requirements are met, and the training process is completed.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] (1) Based on the self - learning ability of the reinforcement learning framework, the present invention designs a reward function for the unmanned boat to meet multi - objective requirements, so as to achieve the cooperative control of multi - task objectives of the unmanned boat. An improved LSTM network structure is introduced into the reinforcement learning framework to capture the data dependency relationship, thereby improving the sample efficiency. And a low - entropy action selection mechanism is used in the reinforcement learning framework to select low - entropy actions, solving problems such as algorithm instability and difficulty in convergence caused by over - estimation of entropy value and blind exploration of more actions, improving the information processing ability of the unmanned boat in complex waters. Therefore, the present invention can achieve the intelligent control of the unmanned boat to meet the multi - task objective requirements in complex waters.
[0056] (2) The present invention uses an improved LSTM network structure to integrate spatio - temporal data, and improves the way of updating the cell state in the input gate through a memory matrix, so that the improved LSTM network structure can better utilize historical data information, further improve the sample efficiency, and improve the autonomous control ability of the unmanned boat.
[0057] (3) In the Actor - Critic network structure of the present invention, the parallel execution of Actor_1 and Actor_2, Critic_1 and Critic_2 is adopted for reinforcement learning, which can interact and learn with the environment simultaneously, accelerate the learning process, and can reduce the policy fluctuations and enhance the robustness caused by the randomness or instability of a single Actor. The Actor - Critic network structure of the present invention is more suitable for complex waters. By combining the low - entropy action selection mechanism, the unmanned boat controller further improves its performance of autonomous control and decision - making, as well as the ability of faster and safer path planning and obstacle avoidance.
[0058] (4) The present invention also considers the wind and current interference in the actual environment, and can automatically adjust the control strategy by interacting with the wind and current interference in the environment, enabling the unmanned boat to quickly adapt to the environmental changes in complex waters and improve the task efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is a schematic flow chart of the method of the present invention;
[0060] Figure 2 is a schematic diagram of the principle of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manner and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0062] This example provides an intelligent control method for an unmanned boat based on an improved low - entropy strategy and network structure, asFigure 1 As shown in the figure, it includes the following steps:
[0063] Step S1: Set the action space and state space of the unmanned boat.
[0064] Specifically, it includes the following: S101: Set the action space of the unmanned boat, including: the control rate u l (t) of the left motor of the unmanned boat, the control rate u r (t) of the right motor, the distance d t from the target point, and the distance d o from the obstacle; the action spaces of the control rate u l (t) of the left motor and the control rate u r (t) of the right motor of the unmanned boat are respectively [-20, 20], the action space of the distance d t from the target point is [0, 1000], and the action space of the distance d o from the obstacle is [0, 200].
[0065] S102: Set the state space of the unmanned boat, including: the linear velocity v of the unmanned boat, the angular velocity φ, the distance d t from the target point, the distance d o from the obstacle, the control rate u l (t) of the left motor and the control rate u r (t) of the right motor; the state space of the linear velocity v of the unmanned boat is [-1, 1], the state space of the angular velocity φ is [-1, 1], the state spaces of the control rate u l (t) of the left motor and the control rate u r (t) of the right motor are respectively [-1, 1], the state space of the distance d t from the target point is [0, 1000], and the state space of the distance d o from the obstacle is [0, 200].
[0066] Step S2: Set the reward function according to the multi-task target requirements of the unmanned boat.
[0067] Specifically, it includes the following:
[0068] S201: Set the reward function for the unmanned boat to reach the target point. To make the unmanned boat reach the target point, set the reward function R 1 of task objective one as:
[0069] R 1 = w 1 r reach ,
[0070] where w 1 is the weight of the reward for the unmanned boat to reach the target point in the total reward; r reach is the reward for the unmanned boat to reach the target point.
[0071] S202. Set the reward function for the unmanned boat to keep a distance from obstacles. To make the unmanned boat stay more than 50 m away from obstacles, set the reward function R of task objective 2 as follows: 2 Set as:
[0072]
[0073] where w 2 is the weight of the reward of task objective 2 in the total reward; d is the distance when the unmanned boat is within 50 m of the obstacle; r o is the reward for the unmanned boat colliding with the obstacle.
[0074] S203. Set the reward function for controlling the linear velocity and angular velocity of the unmanned boat. To make the control of the unmanned boat as stable as possible, set the reward function R of task objective 3 as follows: 3 Set as:
[0075] R 3 =-w 3 ||v current -v previous ||-w 5 ||φ current -φ previous ||,
[0076] where W 3 is the weight of the reward for the linear velocity stability of the unmanned boat in the total reward; v current is the current linear velocity of the unmanned boat, v previous is the linear velocity of the unmanned boat at the previous moment; w 5 is the weight of the reward for the angular velocity stability of the unmanned boat in the total reward; φ current is the current angular velocity of the unmanned boat, φ previous is the angular velocity of the unmanned boat at the previous moment.
[0077] S204. Set the reward function for controlling the sailing time of the unmanned boat. To make the unmanned boat reach the target point quickly, set the reward function R of task objective 4 as follows: 4 Set as:
[0078] R 4 =-w 4 v current -w 6 r d ,
[0079] where w 4 is the weight of the reward for the current linear velocity of the unmanned boat in the total reward; v current is the current linear velocity of the unmanned boat; w 6 is the weight of the reward for the distance between the unmanned boat and the target point in the total reward; r dIt is the reward for the distance between the unmanned boat and the target point.
[0080] S205. Set the multi-task objective demand reward function for the unmanned boat to reach the target point, the distance from obstacles, the linear velocity, the angular velocity, and the sailing time, and set the reward function as follows:
[0081] R = R 1 + R 2 + R 3 + R 4
[0082]
[0083] Among them, R is the total reward for multi-task objective demands; w 1 is the weight of the reward for the unmanned boat to reach the target point in the total reward, set to 0.3; r reach is the reward for the unmanned boat to reach the target point; w 2 is the weight of the reward for the distance between the unmanned boat and the obstacle in the total reward, set to 0.2; d is the distance when the unmanned boat is within 50m of the obstacle; r o is the reward for the collision between the unmanned boat and the obstacle; w 3 is the weight of the reward for the linear velocity stability of the unmanned boat in the total reward, set to 0.1; v current is the current linear velocity of the unmanned boat; v previous is the linear velocity of the unmanned boat at the previous moment; w 4 is the weight of the reward for the current linear velocity of the unmanned boat in the total reward, set to 0.1; w 5 is the weight of the reward for the angular velocity stability of the unmanned boat in the total reward, set to 0.1; φ current is the current angular velocity of the unmanned boat; φ previous is the angular velocity of the unmanned boat at the previous moment; w 6 is the weight of the reward for the distance between the unmanned boat and the target point in the total reward, set to 0.2; r d is the reward for the distance between the unmanned boat and the target point.
[0084] Step S3. Combine the low-entropy action selection mechanism and the improved LSTM network structure to improve the intelligent control method of the unmanned boat.
[0085] Specifically, it includes the following:
[0086] S301. Combine the low-entropy action selection mechanism to improve the intelligent control method of the unmanned boat. The low-entropy action selection mechanism is set as:
[0087]
[0088] Among them, H(π(.|s t )) is in the state s tThe entropy of the lower policy π; is the one that selects the minimum entropy value among those in the state s t under the policy π i ; is the expectation of the action a i under the policy π t ; log(π i (a t |s t )) is the logarithmic probability of the policy π i selecting the action a t in the state s t ; is the weight of the policy π i selecting the action a t in the state s t ; a is the action with the lowest entropy value; is the action a selected by choosing the policy π with the lowest entropy value i .
[0089] S302. Combine with the improved LSTM network structure to improve the intelligent control method of the unmanned boat. Improve the input gate calculation process through the improved LSTM network structure, and the specific settings are as follows:
[0090] c t = f t · c t-1 + i t · tanh(W c [h t-1 , x t + b c )+ M t ,
[0091] M t = sigmoid(M t-1 · c t-1 ),
[0092] where c t is the cell state of the improved LSTM network structure at the current time step; f t is the input of the forget gate; c t-1 is the cell state at the previous moment; i t is the output of the input gate; tanh is the hyperbolic tangent function; W c is the weight matrix of the hidden state, which is updated by the backpropagation algorithm; h t-1 is the hidden state at the previous moment; x t is the input at the current time step; b c is the bias term, which is updated by the backpropagation algorithm; M t is the memory matrix at the current time step; sigmoid is the sigmoid function operation; Mt-1 is the memory matrix at the previous moment.
[0093] S303, as Figure 2 shown, the unmanned boat intelligent controller includes an Actor network structure, a Critic network structure, and an improved LSTM network structure. The Actor network structure includes Actor_1 and Actor_2 sub-networks, each sub-network executes in parallel and has a three-layer structure; according to the requirements of the unmanned boat intelligent controller, the input layer is designed with 4 nodes, the hidden layer is 256 nodes, and the output layer is 2 nodes; the left motor control rate u l (t), the right motor control rate u r (t), the distance d t from the target point, and the distance d o from the obstacle are collected from the experience replay pool. After being processed by the improved LSTM network structure, control actions are output through Actor_1 and Actor_2, and the entropy values of each control action are compared to select the low-entropy action. After obtaining the control output u(t), anti-normalization is required to obtain the actual motor speed. The left motor control rate u l (t) and the right motor control rate u r (t) of the unmanned boat are controlled. The Critic network structure includes Critic_1 and Critc_2 sub-networks, each sub-network executes in parallel and has a three-layer structure. The input layer has 6 nodes, the hidden layer has 256 nodes, and the output layer has 1 node. The linear velocity v, angular velocity φ of the unmanned boat, the distance d t from the target point, the distance d o from the obstacle, the left motor control rate u l (t) and the right motor control rate u r (t) are collected from the experience replay pool, and the state value function V(s t ) and the state-action value function Q(s t , a t ) are output, which are used for parameter update of the Actor network so that the Actor network can learn a better strategy.
[0094] After the unmanned boat intelligent controller collects a batch of data, its Critic network structure outputs the state value function V(s t ) and the state-action value function Q(s t , a t ), and the formula is as follows:
[0095]
[0096] Among them, is the action a t under the policy πCalculate the expectation; α is the temperature coefficient, usually 0.2; log(π(a|s)) is the logarithmic probability of the policy π selecting action a in state s; is the expectation for the next state s under the transition function D t+1 Calculate the expectation; r(s t , a t ) is the immediate reward for executing action a in state s t ; γ is the discount factor, usually 0.99; V(s t ) is the state value estimate of the next state. t+1 ) is the state value estimate of the next state.
[0097] After the intelligent controller of the unmanned ship collects a batch of data, it is input to the improved LSTM network structure for processing, and then input to the Actor network structure to update the parameters of the Actor_Critic network structure. Its loss function is as follows:
[0098]
[0099] Among them, L Q (ω) is the loss function of the neural network parameter ω; ω is the neural network parameter; (s t , a t , r t , s t+1 ) is the sample data sampled from the experience replay pool, i.e., R. s t is the state at the current time step, a t is the action at the current time step, r t is the immediate reward for executing action a t , s t+1 is the state at the next time step; E is the expectation for sampling (s t , a t , r t , s t+1 ) from the experience replay pool, i.e., R and calculating the expectation under the policy π θ ; R is all the data stored in the experience replay pool; a t+1 is the action at the next time step; π θ (.|s t+1 ) is the policy π with parameter θ in state s t+1 ; Q ω (s t , a t ) is the Q value at the current time step; γ is the discount factor, usually 0.99; is the expected return of the Q function with a smaller Q value at the next time step; α is the temperature coefficient, usually 0.2; log(π(a t+1 |s t+1 )) is the logarithmic probability of the policy π selecting action a in state s t+1Lower selection action a t+1 Log probability of
[0100] The objective function for the final update of the unmanned boat intelligent controller is:
[0101]
[0102] Among them, J(π) is the objective function for policy optimization; E is the expectation of the objective value for the state-action pair (s t , a t ) that follows the joint distribution of states and actions under policy π; ρ π is the joint distribution of states and actions under policy π; γ is the discount factor, usually 0.99; s t is the state at the current time step, a t is the action at the current time step; r(s t , a t ) is the immediate reward for executing action a t ; α is the temperature coefficient, generally 0.2; H(.|s t ) is the entropy of policy π at state s t .
[0103] Step S4: Train the unmanned boat intelligent controller based on the improved unmanned boat intelligent control method.
[0104] Specifically, it includes the following:
[0105] S401: Add wind and current disturbances to the environment. The formula for the wind disturbance F wina in the environment is:
[0106]
[0107] Among them, C wind is the wind force coefficient, usually in the range of 1.0 to 1.5 according to different boats and wind forces; θ is the heading angle of the unmanned boat; V wind is the wind speed.
[0108] The formula for the current disturbance F flow in the environment is:
[0109]
[0110] Among them, C flow is the water current coefficient, usually in the range of 0.01 to 0.05 according to different boats and water currents; β is the angle between the water current direction and the heading of the unmanned boat; V flow is the water current velocity;
[0111] S402: Set the number of training cycles N;
[0112] S403. The unmanned boat intelligent controller based on the improved unmanned boat intelligent control method interacts with the environment and stores the interaction information in the experience replay pool in chronological order.
[0113] S404. Determine whether the experience replay pool is full; if not, continue to execute step S402; if so, take out the data to iterate the neural network parameters.
[0114] S405. Determine whether the set number of training cycles N is reached; if not, continue to execute step S403; if so, observe the control effect of the unmanned boat intelligent controller.
[0115] Step S5. Determine whether the trained unmanned boat intelligent controller meets the multi-task objective requirements; if not, continue to execute step S4; if so, end and save the model and neural network parameters for the improved unmanned boat intelligent controller to call.
[0116] As Figure 2 shown, it is the schematic diagram of the principle of the present invention. In the specific implementation, the unmanned boat intelligent controller selects low-entropy actions and outputs the control rate u l (t) of the left motor of the unmanned boat and the control rate u r (t) of the right motor to control the unmanned boat in the actual sea conditions to obtain the state information v, φ, d o , d t , u l (t), u r (t) of the unmanned boat. Finally, the state information v, φ, d o , d t , u l (t), u r (t) of the unmanned boat are input to the unmanned boat intelligent controller; through the steps S1 - S5 of the above specific implementation, continue to select low-entropy actions and feedback the new state information v, φ, d o , d t , u l (t), u r (t) of the unmanned boat to the unmanned boat intelligent controller to continuously optimize the intelligent control effect.
[0117] When the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0118] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program codes. The solutions in the embodiments of the present invention can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0119] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0120] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0121] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 or steps for implementing the functions specified in one block or multiple blocks.
[0122] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0123] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. An intelligent control method for an unmanned boat based on an improved low entropy strategy and a network structure, characterized in that: The following steps are involved: Set the action space and state space of the unmanned boat, set the reward function according to the multi-task target requirements of the unmanned boat, and use the unmanned boat controller that combines the low entropy action selection mechanism and the improved LSTM network for intelligent control until the unmanned boat reaches the target point, completing the intelligent control process of the unmanned boat; The unmanned boat controller includes an Actor network, a Critic network and an improved LSTM network. The process of intelligent control of the unmanned boat controller includes: A batch of data is collected from the experience replay pool, input into the improved LSTM network for processing, and then the control action is output through the Actor network, and the low entropy action selection mechanism is used to select the low entropy action to obtain the control output u(t) to control the operation of the unmanned boat, wherein the experience replay pool is formed by interacting with the environment and combining the action space, state space and reward function; Collect a batch of data from the experience replay pool, input it into the Critic network for value estimation, and guide the Actor network to output control actions; Repeat the above control process until the unmanned boat reaches the target point.
2. According to claim 1, the unmanned boat intelligent control method based on improved low entropy strategy and network structure is characterized in that: The action space of the unmanned boat includes the control rate u of the left motor of the unmanned boat l (t), right motor control rate u r (t), distance d from the target point t and the distance d from the obstacle o The state space includes the unmanned boat linear velocity v, angular velocity φ, and distance d from the target point t , distance to obstacle d o , left motor control rate u l (t) and right motor control rate u r (t).
3. According to claim 1, the unmanned boat intelligent control method based on improved low entropy strategy and network structure is characterized in that: The expression of the reward function is: R=R1+R2+R3+R4 R=w1r reach -w2(1-d / 50)r o -w3||v current -v previous ||-w4v current -w5||φ current -f previous ||-w6r d , in: R1 is the reward function for the unmanned boat to reach the target point, expressed as: R1 = w1r reach , R2 is the reward function for the distance of the unmanned boat away from obstacles, expressed as: R2 = -w2(1-d / 50)r o , R3 is the reward function for the control of the linear velocity and angular velocity of the unmanned boat, and its expression is: R3=-w3||v current -v previous ||-w5||φ current -f previous ||, R4 is the reward function for the unmanned boat’s navigation time control, and its expression is: R4=-w4v current -w6r d , In the formula, R is the total reward function, w1 is the weight of the reward of the unmanned boat reaching the target point in the total reward, and r reach is the reward for the unmanned boat to reach the target point, w2 is the weight of the reward for the distance of the unmanned boat away from the obstacle in the total reward, d is the distance when the unmanned boat is within the set distance from the obstacle, r o is the reward for the collision between the unmanned boat and the obstacle, w3 is the weight of the reward for the stability of the unmanned boat's linear speed in the total reward, and v current is the current linear velocity of the unmanned boat, v previous is the linear velocity of the unmanned boat at the previous moment, w4 is the weight of the reward of the current linear velocity of the unmanned boat in the total reward, w5 is the weight of the reward of the angular velocity stability of the unmanned boat in the total reward, φ current is the current angular velocity of the unmanned boat, φ previous is the angular velocity of the unmanned boat at the previous moment, w6 is the weight of the reward of the distance between the unmanned boat and the target point in the total reward; r d It is the reward for the distance between the unmanned boat and the target point.
4. The unmanned boat intelligent control method based on improved low entropy strategy and network structure according to claim 1 is characterized in that: The Actor network is a dual-strategy network, including an Actor_1 subnetwork and an Actor_2 subnetwork. The Actor_1 subnetwork and the Actor_2 subnetwork interact with the environment in parallel and output control actions.
5. The unmanned boat intelligent control method based on improved low entropy strategy and network structure according to claim 1 is characterized in that: The Critic network includes a Critic_1 subnetwork and a Critic_2 subnetwork, and the Critic_1 subnetwork and the Critic_2 subnetwork perform the value estimation in parallel, wherein the value estimation result includes a state value function V(s t ) and the state-action value function Q(s t , a t ), the expressions are: In the formula, is the action a under the strategy π t Find the expectation, α is the temperature coefficient, log(π(a|s)) is the logarithmic probability of strategy π choosing action a in state s, It is the state s at the next moment under the transfer function D t+1 Find the expectation, r(s t , a t ) is in state s t Next, perform action a t The instant reward, γ is the discount factor, V(s t+1 ) is the state value estimate of the state at the next moment.
6. The unmanned boat intelligent control method based on improved low entropy strategy and network structure according to claim 1 is characterized in that: The operation expression for selecting a low entropy action using the low entropy action selection mechanism is: In the formula, H(π(.|s t )) is in state s t The entropy of the next strategy π, To select in state s t Next strategy π i The minimum entropy value in In strategy π i Next action a t The expectation of log(π i (a t |s t )) is the strategy π i In status t Next select action a t The logarithmic probability of is strategy π i In status t Next select action a t The weight of a is the action with the lowest entropy value. is to choose the strategy with the lowest entropy value π i Selected action a.
7. The unmanned boat intelligent control method based on improved low entropy strategy and network structure according to claim 1 is characterized in that: The improved LSTM network improves the updating method of the cell state in the input gate calculation process. The updated expression of the improved cell state is: c t =f t ·c t-1 +i t ·tanh(W c [h t-1 ,x t ]+b c )+M t , M t =sigmoid(M t-1 ·c t-1 ), In the formula, c t is the cell state of the current time step of the improved LSTM network structure, f t is the input of the forget gate, c t-1 is the cell state at the previous moment, i t is the output of the input gate, tanh is the hyperbolic tangent function, W c is the weight matrix of the hidden state, which is updated by the back-propagation algorithm. t-1 is the hidden state at the previous moment, x t is the input of the current time step, b c is the bias term, which is updated by the back-propagation algorithm. t is the memory matrix of the previous moment, sigmoid is the sigmoid function operation, M t-1 is the memory matrix of the previous moment.
8. The unmanned boat intelligent control method based on improved low entropy strategy and network structure according to claim 1 is characterized in that: The guiding of the Actor network output control action refers to updating the network parameters based on the value evaluation result to select the optimal control action. The specific steps include: For each batch of data, after being processed by the improved LSTM network and Actor network, the network parameters of the unmanned boat controller are updated in combination with the value evaluation results, wherein the expression of the loss function of the network parameter update is: Where, L Q (ω) is the loss function of the neural network parameter ω, ω is the neural network parameter; (s t , a t , r t ,s t+1 ) is the sample data sampled from the experience replay pool, i.e., R, s t is the state of the current time step, a t is the action at the current time step, r t Is to perform action a t Instant rewards, s t+1 is the state at the next moment; E is the sample from the experience replay pool, i.e. R (s t , a t , r t ,s t+1 ) and in strategy π θ Find the expectation below, R is the total data stored in the experience replay pool, a t+1 is the action at the next moment, π θ (·|s t+1 ) is in state s t+1 The strategy π, Q under the parameter θ ω (s t , a t ) is the Q value of the current time step, γ is the discount factor, is the expected return of the Q function with a smaller Q value at the next moment, α is the temperature coefficient, log(π(a t+1 |s t+1 )) is the strategy π in state s t+1 Next select action a t+1 The logarithmic probability of .
9. The unmanned boat intelligent control method based on improved low entropy strategy and network structure according to claim 1 is characterized in that: The objective function updated by the unmanned boat controller is: Where J(π) is the objective function of policy optimization, and E is the state-action set (s) that obeys the joint distribution of state and action under policy π. t , a t ) to find the expected target value, ρ π is the joint distribution of states and actions under policy π; γ is the discount factor, s t is the state of the current time step, a t is the action at the current time step, r(s t , a t ) is to perform action a t The instant reward, α is the temperature coefficient, H(·|s t ) is the policy π in state s t The entropy under .
10. The unmanned boat intelligent control method based on improved low entropy strategy and network structure according to claim 1 is characterized in that: The training process of the unmanned boat controller includes: 1) Add wind and flow interference to the environment, where the expression of wind interference in the environment is: The expression of flow interference in the environment is: In the formula, F wind is wind disturbance, F flow is the flow interference, C wind is the wind force coefficient, θ is the heading angle of the unmanned boat, V wind is the wind speed, C flow is the water flow coefficient, β is the angle between the water flow direction and the heading of the unmanned boat, V flow is the water flow velocity; 2) Set the training cycle N; 3) Interact with the environment and put the interaction information into the experience replay pool in chronological order; 4) Determine whether the experience replay pool is full. If so, extract a batch of data from the experience replay pool to iteratively update the network parameters. If not, execute step 2); 5) Determine whether the set training cycle N has been reached, If yes, then further determine whether the multi-objective requirements are met. If yes, output the control result to obtain the control effect. If no, execute step 3); If not, execute step 3) until the training cycle N is reached and the multi-objective requirements are met, and the training process is completed.
Citation Information
Patent Citations
A path planning method and system for unmanned boat based on improved RRT algorithm
CN115454092B
Unmanned ship path optimization method based on GWO algorithm
CN117991780A
Low-altitude logistics multi-unmanned aerial vehicle cooperative distribution method based on deep reinforcement learning
CN119671425A
A method for determining a trajectory for a marine vessel under environmental disturbance
EP4432263A1