Vector propulsion UUV smart motion control method based on improved maximum entropy reinforcement learning

By improving the maximum entropy reinforcement learning method, combining deep neural networks and traditional controllers, the problem of three-dimensional space sensitive motion control in UUV in complex environments is solved, efficient vector propulsion UUV motion control is achieved, and the safety and adaptability of control are improved.

CN120276236APending Publication Date: 2025-07-08HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510393905.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing traditional nonlinear motion control method is difficult to cope with parameter perturbation and environmental interference in UUV three-dimensional motion control, especially in large-scale motion state space, and model-free reinforcement learning has low sample utilization rate in high-dimensional state space and action space, and poor environmental adaptability.

Method used

The improved maximum entropy reinforcement learning method is adopted to establish a vector-promoting UUV six-degree of freedom sensitive motion model, transform the control problem into a Markov decision-making process, combine the deep neural network and traditional nonlinear controller, and improve the SAC algorithm to train the strategy network and evaluate the network, and use the n-step TD algorithm to balance the variance and deviation of the equilibrium state-action value estimates, and select the appropriate controller for actual control output.

Benefits of technology

It improves the three-dimensional space sensitive motion control effect of UUV under parameter perturbation and environmental interference conditions, reduces sampling cost, improves sample utilization and convergence speed, and enhances the safety and adaptability of motion control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276236A_ABST
    Figure CN120276236A_ABST
Patent Text Reader

Abstract

The invention discloses a vector propulsion UUV smart motion control method based on improved maximum entropy reinforcement learning, and belongs to the technical field of underwater vehicle motion control. The method comprises the following steps: establishing a vector propulsion UUV six-degree-of-freedom smart motion model; the deep neural network model is combined with an SAC algorithm framework, the sampling step number n is selected according to the difference between the deep neural network model and an actual environment model, and an n-step TD algorithm is adopted to train an evaluation network through imaginary empirical data; the environment change degree is judged according to the difference between the deep neural network model and the actual environment model, and then a strategy network or a backup controller is selected for actual control output; by utilizing the powerful learning ability of deep reinforcement learning and the interpretability of the traditional nonlinear control method, the three-dimensional space smart motion control of the vector propulsion UUV under the conditions of parameter perturbation or completely unknown, unknown environment interference and state constraint can be realized; and engineering implementation and deployment are carried out on the premise of ensuring that the control method is safe and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of underwater vehicle motion control, and particularly relates to a vector propulsion UUV intelligent motion control method based on improved maximum entropy reinforcement learning. Background Art

[0002] With the development of social economy, the value of the ocean has been increasingly emphasized by humans, and people have developed and utilized it from different perspectives. UUVs have the capabilities of perception, decision-making, action, and even cooperation, and have gradually become the main direction of the development of marine equipment. As a marine operation tool, UUVs are widely used in many fields such as deep-sea exploration, biological sampling, and pipeline inspection. With the increasing complexity of marine operation tasks, we require UUVs to be able to navigate at high speed and maneuverability and also be able to change direction quickly at low speed. In short, UUVs need to have both mobility and agility.

[0003] According to different drive configuration forms, UUVs can be roughly divided into the following five categories: fin-rudder type, fin-rudder + auxiliary propulsion type, vector propulsion type, multi-propeller type, and bionic type. Among them, the vector propulsion type UUV not only has the same large-thrust propeller at the tail and streamlined low-drag shape as the fin-rudder type UUV, meeting the mobility requirements; but also can provide sufficient torque to quickly perform attitude transformation by deflecting the propeller direction at low speed, meeting the agility requirements.

[0004] The three-dimensional motion control of UUVs is one of the basic prerequisites for implementing marine operation tasks. Subject to many constraints from the UUV itself and the external environment, the three-dimensional motion control of UUVs still poses a considerable challenge:

[0005] (1) The three-dimensional motion model of UUVs has high nonlinearity and coupling;

[0006] (2) The hydrodynamic parameters of UUVs are difficult to accurately measure and are affected by unknown ocean currents and waves from the complex marine environment;

[0007] (3) The energy carried by UUVs themselves is very limited and it is difficult to replenish due to the energy form and environment;

[0008] (4) The input of UUV actuators is limited, and the risk of damage will increase due to frequent saturation actions;

[0009] (5) In addition, complex underwater operation tasks also require controlling the UUV to operate in a large range of motion state spaces such as large trim motion and large curvature turning motion.

[0010] Current research on traditional non - linear motion control methods focuses on solving the problems of parameter perturbation and environmental interference in the non - linear coupling model of UUVs. However, there is very little research on control within a large - scale motion state space. Most traditional non - linear motion control methods are based on accurate system motion models. When the actual environment differs greatly from the system model, the control effect is poor, and usually a complex and time - consuming parameter tuning process is required.

[0011] With the development of optimal control theory, computer control technology, and artificial intelligence theory, reinforcement learning methods have gradually matured and been widely applied in the fields of robot control and autonomous driving. Whether in the aspect of strategy games or in the classical control field, reinforcement learning (RL) methods have demonstrated their powerful decision - making ability. Especially when combined with neural networks to form deep reinforcement learning (DRL) methods, they can handle problems in higher - dimensional state spaces and action spaces, making the control strategy more adaptable. This feature exactly meets the requirements of agile motion for a large number of dimensions and a large range in the motion state space. However, although model - free deep reinforcement learning can train the optimal strategy in a completely unknown environment, the problems of sparse rewards and slow convergence speed brought about by high - dimensional state spaces and action spaces limit its application in the design of motion controllers. Especially when the influence of system parameter perturbation and environmental time - varying interference is large, the sample utilization rate of model - free reinforcement learning is low, and the environmental adaptability is poor. Summary of the Invention

[0012] The present invention provides a vector - propulsion UUV agile motion control method based on improved maximum - entropy reinforcement learning, which can achieve the goal of vector - propulsion UUV agile motion control and can be implemented and deployed in engineering on the premise of ensuring the safety and reliability of the control method.

[0013] The present invention provides a vector - propulsion UUV agile motion control method based on improved maximum - entropy reinforcement learning, including:

[0014] Step 1: Establish a six - degree - of - freedom agile motion model of the vector - propulsion UUV;

[0015] Step 2: Transform the vector - propulsion UUV agile motion control problem into a Markov decision process;

[0016] Step 3: Use the improved SAC algorithm to train the intelligent motion control task of the vector propulsion UUV; the network model based on the improved SAC algorithm includes a deep neural network for fitting the six-degree-of-freedom intelligent motion model of the vector propulsion UUV, an evaluation network Critic, and a policy network Actor; select the sampling step number n according to the error between the deep neural network and the actual environment model, and train the evaluation network Critic through the n-step TD algorithm; train the policy network Actor according to the state-action value output by the evaluation network Critic and the entropy of the policy network Actor;

[0017] Step 4: After training, the state output by the actual environment model is input into the controller, and according to the difference between the deep neural network and the actual environment model, select the controller to output the control action in real time; when the environment changes little, the difference between the deep neural network and the actual environment model is small, that is, the loss value Loss of the deep neural network is less than the selection threshold T switch , at this time the controller selects the policy network Actor for actual control output; when the environment changes greatly, the difference between the deep neural network and the actual environment model is large, that is, the loss value Loss of the deep neural network is greater than the selection threshold T switch , at this time the controller selects the backup controller for actual control output.

[0018] Furthermore, in the said step 2, the six-degree-of-freedom intelligent motion model of the vector propulsion UUV includes a kinematic model and a dynamic model;

[0019] The kinematic model:

[0020]

[0021] Wherein, is the generalized pose vector; is the generalized velocity vector; is the linear velocity coordinate transformation matrix; is the angular velocity coordinate transformation matrix;

[0022] The dynamic model:

[0023]

[0024] Wherein, M is the body inertia matrix; C and D are the Coriolis force matrices; g is the gravity and buoyancy vector; is the thrust vector; is the propeller deflection angle vector; is the environmental disturbance force vector.

[0025] Furthermore, in the step 2, the Markov decision process consists of an action space, a state space, and a reward function; the state space includes the deviation between the current actual pose and the desired pose of the UUV, the current actual speed of the UUV, and the actual deflection angle of the vector thruster; the action space includes the thrust of the vector thruster and the desired deflection angle; the reward function R is designed as follows:

[0026] R(S t ,A t ,S t+1 )=c1r error +c2r action +c3r smooth +c4r alive

[0027] where r error is the error-related term; r action is the action-related term; r smooth is the smoothing term; r alive is the survival term; c i is the weight of the four rewards.

[0028] Furthermore, in the step 3, the specific structures of the deep neural network, the policy network Actor, and the evaluation network Critic are as follows:

[0029] The deep neural network consists of 3 fully connected layers, the activation function of the hidden layer is Relu, and the activation function of the output layer is Tanh;

[0030] The evaluation network Critic consists of 3 fully connected layers, the activation function of the hidden layer is Relu, and the activation function of the output layer is the identity function;

[0031] The policy network Actor consists of 2 RNN layers and 2 fully connected layers, the activation functions of the hidden layers are all Relu, and the activation function of the output layer is Tanh.

[0032] Furthermore, the step 3 specifically includes the following steps:

[0033] Step 3.1: Initialize the policy network Actor and the deep neural network; interact the backup controller with the six-degree-of-freedom model of the vector UUV to obtain sub-optimal experience data, and save it in the sub-optimal experience buffer; collect batchSize sub-optimal experience tuples from the sub-optimal experience buffer, and train the deep neural network model to approximate the six-degree-of-freedom model of the vector-propelled UUV and train the policy network Actor to approximate the backup controller by minimizing the mean square error; the backup controller is a traditional non-linear controller;

[0034] Step 3.2: Interact with the actual environment through the policy network Actor to obtain real experience data and store it in the real experience buffer; the actual environment model parameters consist of the six-degree-of-freedom agile motion model parameters of the vector propulsion UUV and random quantities; design the interaction termination condition done. If the termination condition done is true, stop the interaction process between the policy network Actor and the actual environment, and store the obtained real experience tuple (S t , A t , S t+1 , R t+1 ) in the real experience buffer; if the termination condition done is false, add the real experience tuple (S t , A t , S t+1 , R t+1 ) to the real experience buffer and update the time step to the next moment; when the number of tuples in the real experience buffer exceeds the batch size batchSize, then update the parameters of the deep neural network, the evaluation network Critic, and the policy network Actor at each time step in each subsequent round of interaction.

[0035] Step 3.3: Update the deep neural network: Collect batchSize real experience tuples from the real experience buffer, train the deep neural network model by calculating the loss value Loss to approximate the actual environment, and select the sampling step number n according to the error between the deep neural network and the actual environment.

[0036] Step 3.4: Update the evaluation network Critic: If the sampling step number n = 0, use the real experience tuple (S t , A t , S t+1 , R t+1 ) in the real experience buffer to calculate the 1-step time difference target and update the evaluation network Critic using the 1-step TD algorithm; if n > 0, after inputting the sampled batchSize real experience tuples into the deep neural network model, start the n-step interaction between the policy network Actor and the deep neural network model, calculate the cumulative reward value at each time step, and store the imaginary experience tuple (S t , A t , S’ t+n , R t+1,...,t+n ) in the imaginary experience buffer at the nth time step; input the imgBatchSize imaginary experience tuples in the imaginary experience buffer into the evaluation network Critic and then output the predicted state-action value Q w (S t , A t ), calculate the n-step time difference target and update the evaluation network Critic using the n-step TD algorithm.

[0037] Step 3.5: Update the policy network Actor: After inputting the sampled imgBatchSize imaginary experience tuples into the policy network Actor, output and the entropy value Update the policy network Actor according to the SAC algorithm.

[0038] Furthermore, step 3.2 is specifically as follows: First, randomly initialize the state of the actual environment model as S t , input S t into the policy network Actor and then output the action A t , input A t into the actual environment model and then output the state value S at the next moment t+1 and the reward value R t+1 , and judge the termination condition done; the termination condition includes that the interaction reaches the maximum number of time steps or the pose deviation reaches the maximum threshold.

[0039] Furthermore, in step 3.3, randomly sample batchSize experience tuples from the real buffer and input them into the deep neural network, and calculate the loss value Loss by minimizing the mean square error:

[0040]

[0041] where bs is batchSize; δS t is the output value of the deep neural network model;

[0042] The sampling step n:

[0043]

[0044] where k loss and c loss are scaling coefficients; n max is the maximum sampling step.

[0045] Furthermore, in step 3.4, the 1-step temporal difference target

[0046]

[0047] where Q w is the state-action value; γ is the return discount rate; is calculated through the policy network Actor; H is the entropy of the policy network Actor; β is a regularization coefficient used to control the importance of the entropy;

[0048] The loss function L Q1 (w):

[0049]

[0050] The n-step temporal difference target

[0051]

[0052] The loss function L of the n-step TD algorithm Qn (w):

[0053]

[0054] where ibs is imgBatchSize.

[0055] Furthermore, the loss function L of the policy π (θ):

[0056]

[0057] Furthermore, in step 4, when selecting a backup controller for actual control output, continue to optimize the deep neural network, the policy network Actor, and the evaluation network Critic for the next control movement.

[0058] The present invention also provides a computer device / system, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, it implements the steps of the vector propulsion UUV agile motion control method based on improved maximum entropy reinforcement learning described in any one of the above.

[0059] The present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, it implements the steps of the vector propulsion UUV agile motion control method based on improved maximum entropy reinforcement learning described in any one of the above.

[0060] The present invention also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, it implements the steps of the vector propulsion UUV agile motion control method based on improved maximum entropy reinforcement learning described in any one of the above.

[0061] The beneficial effects of the present invention are as follows:

[0062] The present invention proposes a vector propulsion UUV agile motion control method based on improved maximum entropy reinforcement learning, aiming to achieve three-dimensional space agile motion control of vector propulsion UUV under parameter perturbation or complete unknown, unknown environmental interference, and state constraint conditions. Based on SAC, this method combines the powerful fitting ability of the deep neural network model to solve the problem of completely unknown parameters of the vector propulsion UUV agile motion model, thereby reducing the actual sampling cost. Embedding the deep neural network model into the reinforcement learning training process not only improves the sample utilization rate but also uses the n-step TD algorithm to balance the variance and bias of the state-action value estimation, further improving the convergence speed. The present invention introduces the difference between the deep neural network model and the actual environmental model, which is not only used to select the sampling step number n in the n-step TD algorithm but also used to select whether the policy network Actor or the backup controller is used for actual control output, improving the safety of motion control and even enabling the ability of adaptive task switching. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 is a flowchart of the vector propulsion UUV agile motion control method based on improved maximum entropy reinforcement learning according to the present invention;

[0064] Figure 2 is a schematic diagram of the coordinate system of the vector propulsion UUV according to the present invention;

[0065] Figure 3 is a schematic diagram of the deep neural network structure according to the present invention;

[0066] Figure 4 is a schematic diagram of the evaluation network structure according to the present invention;

[0067] Figure 5 is a schematic diagram of the policy network structure according to the present invention;

[0068] Figure 6 is a schematic diagram of the policy initialization process according to the present invention;

[0069] Figure 7 is a schematic diagram of the policy optimization process according to the present invention;

[0070] Figure 8 is a schematic diagram of the policy deployment process according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0071] The following further describes the present invention with reference to the accompanying drawings.

[0072] The present invention discloses a vector propulsion UUV agile motion control method based on improved maximum entropy reinforcement learning, as Figure 1 shown, including:

[0073] Step 1: Establish a six-degree-of-freedom agile motion model of the vector propulsion UUV: Assume that the gravity and buoyancy forces acting on the vector propulsion UUV cancel each other out or are relatively small; Assume that the deflection angle response of the vector thruster has a certain delay and inertia and is regarded as a first-order inertial system; Assume that the vector propulsion UUV only has left-right and up-down symmetries; Considering system parameter perturbations and unknown environmental disturbances, establish a six-degree-of-freedom agile kinematic model and a six-degree-of-freedom agile dynamic model of the vector propulsion UUV based on quaternion attitude representation.

[0074] Step 2: Transform the agile motion control problem of the vector propulsion UUV into a Markov decision process:

[0075] (2.1) Considering the target-oriented principle, design the state space to include the deviation between the current actual pose and the desired pose of the UUV, the current actual velocity of the UUV, and the actual deflection angle of the vector thruster.

[0076] (2.2) Considering the inertia of the vector thruster, the action space includes the thrust of the vector thruster and the desired deflection angle.

[0077] (2.3) Considering the target-oriented principle, energy-saving requirements, and the chattering phenomenon of the vector thruster, design the reward function to include an error-related term, an action-related term, a smoothing term, and a survival term.

[0078] Step 3: Establish one deep neural network, two identical structure policy networks Actor, and four identical structure evaluation networks Critic:

[0079] (3.1) Both the deep neural network and the evaluation network Critic are composed of three fully connected layers, and the activation function of the hidden layer is Relu. The difference is that the activation function of the output layer of the deep neural network is Tanh, while the activation function of the output layer of the evaluation network Critic is the identity function.

[0080] (3.2) The policy network Actor is composed of two RNN layers and two fully connected layers, the activation function of the hidden layer is Relu, and the activation function of the output layer is Tanh.

[0081] Step 4: Initialize the policy network Actor and the deep neural network, that is, the policy initialization stage:

[0082] (4.1) Design a traditional non-linear controller (backup controller), interact with the six-degree-of-freedom model of the vector propulsion UUV to obtain sub-optimal experience data, and save it in the sub-optimal experience buffer.

[0083] (4.2) Collect data from the sub-optimal experience buffer, and train a deep neural network model to approximate the six-degree-of-freedom model of the vector propulsion UUV by minimizing the mean square error (MSE), and train the policy network Actor to approximate the traditional non-linear controller (backup controller).

[0084] Step 5: Train the deep neural network, the policy network Actor, and the evaluation network Critic, that is, the policy optimization stage:

[0085] (5.1) Interact with the actual environment through the policy network Actor to obtain real experience data, and save it in the real experience buffer;

[0086] (5.2) Collect data from the real experience buffer, and train a deep neural network model to approximate the actual environment by minimizing the mean square error (MSE), and select the sampling step number n according to the error between the deep neural network model and the actual environment;

[0087] (5.3) Collect the current state S in the real experience buffer t As the initial state, interact with the deep neural network model through the policy network Actor to obtain imaginary experience data, and save it in the imaginary experience buffer;

[0088] (5.4) Collect data from the imaginary experience buffer, and train the evaluation network Critic through the n-step TD algorithm;

[0089] (5.5) According to the state-action value output by the evaluation network Critic, and the entropy of the policy network Actor, train the policy network Actor through the SAC algorithm;

[0090] (5.6) Repeat (5.1) to (5.5) until a satisfactory training effect is obtained or the maximum number of training times is reached.

[0091] Step 6: Establish a smart motion controller for the vector propulsion UUV, that is, the policy deployment stage: Judge the degree of environmental change according to the difference between the deep neural network model and the actual environment model. When the environmental change is small, select the policy network Actor for actual control output; when the environmental change is large, select the backup controller for actual control output, and return to Step 5 to continue training.

[0092] Example 1

[0093] Such as Figure 2As shown in the figure, the vector propulsion UUV only has left - right and up - down symmetry, but not front - back symmetry. The vector thruster can deflect in two directions: pitch and yaw. Assuming that the deflection angle response of the vector thruster has a certain delay and inertia, it is regarded as a first - order inertial system. The origin of the body - fixed coordinate system {b} is located at the center of buoyancy of the body. The X - axis points to the bow of the vehicle, the Y - axis points to the starboard side of the body, and the Z - axis points to the belly of the body. The inertial coordinate system {e} is the NED coordinate system, whose origin is a fixed point on the earth. The X - axis points due north, the Y - axis points due east, and the Z - axis is perpendicular to the sea level and points downward.

[0094] To meet the requirements of motion agility, the UUV needs to have no static stability or weak static stability. Therefore, there are two ways to set the center of gravity: one is that the center of gravity coincides exactly with the center of buoyancy; the other is that the center of gravity does not coincide with the center of buoyancy, but is located on the Z - axis and is close to the center of buoyancy. The six - degree - of - freedom agile kinematic model of the vector propulsion UUV is expressed as where is the generalized pose vector, is the generalized velocity vector, is the linear velocity coordinate transformation matrix, is the angular velocity coordinate transformation matrix; the six - degree - of - freedom agile dynamic model of the vector propulsion UUV is expressed as M where M is the body inertia matrix, C and D are the Coriolis force matrices, g is the gravity and buoyancy vector, is the thrust vector, is the thruster deflection angle vector, is the environmental disturbance force vector.

[0095] Considering the target - oriented principle, the state - space dimension is 15, including the deviation between the actual pose and the desired pose of the current UUV, the actual velocity of the current UUV, and the actual deflection angle of the vector thruster, which is expressed as Considering the inertia of the vector thruster, the action - space dimension is 3, including the thrust of the vector thruster and the desired deflection angle, which is expressed as normalize(T,θ d1 ,θ d2 ), where T represents the magnitude of the thrust, and the subscript d represents the desired value. Considering the target - oriented principle, the energy - saving requirement, and the chattering phenomenon of the vector thruster, the reward function includes an error - related term, an action - related term, a smoothing term, and a survival term, which is expressed as R(S t ,A t ,S t+1 )=c1r error +c2r action +c3r smooth +c4r alive , that is, the weighted sum of four rewards, where c i represents the weights of the four rewards. In actual operation, c can be set either as a fixed quantity or as a variable quantity.

[0096] Deep neural networks such as Figure 3 , the evaluation network Critic such as Figure 4 and the policy network Actor such as Figure 5 As shown, both the deep neural network and the evaluation network Critic are composed of 3 fully connected layers, the hidden layer dimension is 512, and the activation function is Relu. The input layer dimension of the deep neural network is 18, the output layer dimension is 16, and the activation function is Tanh. The input layer of the evaluation network Critic is 18, the output layer dimension is 1, and the activation function is the identity function. The policy network Actor is composed of 2 RNN layers and 2 fully connected layers, the hidden layer dimension is 512, the activation function is Relu, the input layer dimension is 15, the output layer dimension is 3, and the activation function is Tanh. The input of the deep neural network is the current state S t , the current action A t , and the output state differential δS t ; the evaluation network Critic inputs the current state S t , the current action A t , and outputs the state-action value; the policy network Actor inputs the current state S t , and outputs the current action A t , and the entropy H of the policy network Actor.

[0097] The policy initialization stage is as Figure 6 shown, where the parameters of the vector propulsion UUV agile motion model are fixed values. The data structure of the sub-optimal experience buffer is a double-ended queue structure, its capacity is expressed as bufferSize, and its element type is an experience tuple type, including the current state S t , the current action A t , the state S at the next moment t+1 and the reward value R t+1 .

[0098] Interact with the vector propulsion UUV six-degree-of-freedom model through the backup controller to obtain sub-optimal experience tuples, and save them in the sub-optimal experience buffer. Collect batchSize sub-optimal experience tuples from the sub-optimal experience buffer, and initialize the deep neural network model by minimizing the mean square error (MSE) to approximate the vector propulsion UUV six-degree-of-freedom model and the policy network Actor to approximate the backup controller.

[0099] The policy optimization stage is as Figure 7 shown. During the simulation process, the actual environment model is represented by the vector propulsion UUV agile motion model, where the model parameters are a fixed value plus a random quantity obeying a truncated normal distribution with a mean of zero, expressed as c = c0 + ε, ε ∼ N(0, σ 2 , a, b).

[0100] During specific operations, the sub-optimal experience buffer and the real experience buffer are actually one buffer. The reason for distinguishing them into two names is that both the buffer usage stage and the data source stored are different. Specifically, the data stored in this buffer during the policy initialization stage comes from the interaction between the backup controller and the six-degree-of-freedom model of the vector propulsion UUV. At this time, it is named the sub-optimal experience buffer; while the data stored during the policy optimization stage comes from the interaction between the policy network Actor and the actual environment model. At this time, it is named the real experience buffer. The data structure of the imagined experience buffer is still a deque, and its element types include the current state S t , the current action A t , the state S' at the nth step t+n and the cumulative value R of n rewards t+1,...,t+n , and its capacity is expressed as imgBufferSize. Due to the frequent update of the deep neural network model, there are large deviations in the past imagined experiences. Therefore, the capacity of the imagined experience buffer is much smaller than that of the real experience buffer, that is, imgBufferSize << bufferSize.

[0101] In each interaction between the policy network Actor and the actual environment model, first randomly initialize the state of the actual environment model as S t , input S t into the policy network Actor and then output the action A t , input A t into the actual environment model and then output the state value S t+1 at the next moment and the reward value R t+1 . Judge the termination condition done. If done is false, then add the real experience tuple (S t , A t , S t+1 , R t+1 ) to the real experience buffer and update the time step to the next moment, otherwise terminate this round of interaction. When the number of tuples in the real experience buffer exceeds the batch size batchSize, then update the parameters of the deep neural network model, the evaluation network Critic, and the policy network Actor at each time step in the subsequent interactions. The termination condition is that the interaction reaches the maximum number of time steps or the pose deviation reaches the maximum threshold.

[0102] Randomly sample batchSize experience tuples from the real experience buffer to update the parameters of the deep neural network model. The process is the same as the update process in the policy initialization stage. Collect batchSize real experience tuples from the real experience buffer and optimize the deep neural network model by minimizing the mean square error (MSE).

[0103] The loss value Loss calculated by the parameter update process of the real experience tuple input into the deep neural network model is output as the sampling step number n after the sampling step number function. The loss function is expressed as where bs is the batchSize, and δS t is the output value of the deep neural network model; the sampling step number function is expressed as where k loss and c loss are scaling coefficients, and n max is the maximum sampling step number. When n = 0, the difference between the deep neural network model and the actual environment model is too large. At this time, the real experience is used to calculate the 1-step time difference target, and the evaluation network Critic is updated using the 1-step TD algorithm. The loss function is expressed as where bs is the batchSize; Q w is the state-action value; γ is the reward discount rate; represents the sampling data calculated by the policy network Actor rather than from the imagined experience; H is the entropy of the policy network Actor; β is a regularization coefficient used to control the importance of the entropy.

[0104] When n > 0, after inputting the batchSize real experience tuples of the sampling into the deep neural network model, the n-step interaction process between the policy network Actor and the deep neural network model is started. In addition to calculating the cumulative reward value within each time step and storing the imagined experience tuple (S t , A t , S’ t+n , R t+1,...,t+n ) into the imagined experience buffer at the nth time step, other processes are the same as the interaction process between the policy network Actor and the actual environment model. Among them, there are batchSize imagined experience tuples stored in the imagined buffer at the nth time step.

[0105] Sample imgBatchSize imagined experience tuples from the imagined experience buffer, input them into the evaluation network Critic, and then output the predicted state-action value Q w (S t , A t ). Update the evaluation network Critic according to the n-step TD algorithm. Among them, the time difference target is expressed as The loss function is expressed as where ibs is the imgBatchSize.

[0106] After inputting the imgBatchSize imagined experience tuples of the sampling into the policy network Actor, output and the entropy value Update the policy network Actor according to the SAC algorithm. Among them, the loss function of the policy is expressed as

[0107]

[0108] In the policy deployment stage, as Figure 8 shown, judge the degree of environmental change according to the difference between the deep neural network model and the actual environment model. When the environmental change is small, the difference between the deep neural network model and the actual environment model is small, that is, the loss value Loss is less than the selection threshold T switch , at this time, select the policy network Actor for actual control output; when the environmental change is large, the difference between the deep neural network model and the actual environment model is large, that is, the loss value Loss is greater than or equal to the selection threshold T switch , at this time, select the backup controller for actual control output, and continue to optimize the policy network Actor.

[0109] Particularly, in some preferred embodiments of the present invention, a computer device is further provided, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the steps of the vector propulsion UUV agile motion control method based on improved maximum entropy reinforcement learning described in any of the above embodiments are implemented.

[0110] In some other preferred embodiments of the present invention, a computer-readable storage medium is further provided, on which computer programs / instructions are stored. When the computer program is executed by the processor, the steps of the vector propulsion UUV agile motion control method based on improved maximum entropy reinforcement learning described in any of the above embodiments are implemented.

[0111] In summary, the present invention takes the application scenario of UUV performing complex operation tasks in a large-scale motion state space as the background, and takes the UUV configured with vector thrusters but without fins and rudders as the object. A six-degree-of-freedom agile motion model of the vector propulsion UUV is established. The deep neural network model is combined with the Soft Actor-Critic (SAC) algorithm framework to accelerate the training speed and adapt to environmental changes. The traditional non-linear controller is used as a backup controller and combined with the trained reinforcement learning strategy to improve control safety, so as to realize the three-dimensional space agile motion control of the vector propulsion UUV under parameter perturbation or completely unknown, unknown environmental interference, and state constraints. In the six-degree-of-freedom agile motion model of the vector propulsion UUV, the center of gravity of the UUV is close to the center of buoyancy, that is, it has the characteristics of weak underwater stability, which is beneficial to the agile motion control of the vector propulsion UUV; the deflection angle response of the vector thruster has a certain delay and inertia, and a first-order inertial system is described to improve the modeling accuracy; the model parameters (such as hydrodynamic coefficients, body inertia tensors, etc.) are not constant values, but a fixed value plus a random quantity, which is used to simulate the influence of parameter perturbation and unknown environmental interference in the actual environment. A deep neural network is established to approximate the actual environment within a finite time and finite error range to reduce the actual sampling cost and improve the training efficiency; in the policy initialization stage, the backup controller interacts with the six-degree-of-freedom model of the vector propulsion UUV to obtain sub-optimal experience data, and the deep neural network model is trained to approximate the six-degree-of-freedom model of the vector propulsion UUV by minimizing the mean square error (MSE), and the policy network Actor is trained to approximate the backup controller; in the policy optimization stage, the sampling step number n is selected according to the difference between the deep neural network model and the actual environment model, and the imaginary experience data is obtained through n-step interaction between the policy network (Actor) and the deep neural network model, and the n-step TD algorithm is used to train the evaluation network (Critic) with the imaginary experience data to balance the variance and bias of the state-action value estimation. In the policy deployment stage, the degree of environmental change is judged according to the difference between the deep neural network model and the actual environment model, and then the policy network (Actor) or the backup controller is selected for the actual control output.

[0112] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above embodiments of the agile motion control method of the vector propulsion UUV based on improved maximum entropy reinforcement learning, which will not be repeated here.

[0113] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above-mentioned embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the method of using a bionic robotic fish to identify and track aquatic biological communities as described above.

[0114] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or N embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0115] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0116] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of executable instructions including one or more steps for implementing a customized logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0117] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of the computer-readable medium include the following: an electrical connection part (electronic device) having one or N wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing when necessary, and then storing it in a computer memory.

[0118] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0119] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0120] In addition, each functional unit in various embodiments of the present invention may be integrated into a processing module, may exist physically separately for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0121] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A vector propulsion UUV intelligent motion control method based on improved maximum entropy reinforcement learning, characterized in that Including: Step 1: Establish a six-degree-of-freedom agile motion model of a vector propulsion UUV; Step 2: Transform the agile motion control problem of the vector propulsion UUV into a Markov decision process; Step 3: Use the improved SAC algorithm to train the agile motion control task of the vector propulsion UUV; Among them, the network model based on the improved SAC algorithm includes a deep neural network, an evaluation network Critic, and a policy network Actor; the sampling step number n is selected according to the error between the deep neural network and the actual environment model, and the evaluation network Critic is trained by the n-step TD algorithm; the policy network Actor is trained according to the state-action value output by the evaluation network Critic and the entropy of the policy network Actor; Step 4: After training, the state output by the actual environment model is input into the controller. According to the difference between the deep neural network and the actual environment model, the controller selects the control action to be output in real time; when the environmental change is small, the difference between the deep neural network and the actual environment model is small, that is, the loss value Loss of the deep neural network is less than the selection threshold T switch , at this time the controller selects the policy network Actor for actual control output; when the environmental change is large, the difference between the deep neural network and the actual environment model is large, that is, the loss value Loss of the deep neural network is greater than the selection threshold T switch , at this time the controller selects the backup controller for actual control output.

2. The vector propulsion UUV intelligent motion control method based on improved maximum entropy reinforcement learning according to claim 1, characterized in that, In the said Step 2, the six-degree-of-freedom agile motion model of the vector propulsion UUV includes a kinematic model and a dynamic model; the kinematic model: Among them, is the generalized pose vector; is the generalized velocity vector; is the linear velocity coordinate transformation matrix; is the angular velocity coordinate transformation matrix; The dynamic model: where M is the body inertia matrix; C and D are the Coriolis force matrices; g is the gravity and buoyancy vector; is the thrust vector; is the thruster deflection angle vector; is the environmental disturbance force vector.

3. The vector propulsion UUV intelligent motion control method based on improved maximum entropy reinforcement learning according to claim 1, characterized in that In the said Step 2, the Markov decision process consists of an action space, a state space, and a reward function; the state space includes the deviation between the current actual pose and the desired pose of the UUV, the current actual speed of the UUV, and the actual deflection angle of the vector thruster; the action space includes the thrust of the vector thruster and the desired deflection angle; the reward function R is designed as follows: R(S t ,A t ,S t+1 ) = c1r error + c2r action + c3r smooth + c4r alive where r error is the error-related term; r action is the action-related term; r smooth is the smoothing term; r alive is the survival term; c i is the weight of the four rewards.

4. The vector propulsion UUV intelligent motion control method based on improved maximum entropy reinforcement learning according to claim 1, characterized in that, In the said Step 3, the specific structures of the deep neural network, the policy network Actor, and the evaluation network Critic are: The deep neural network consists of 3 fully connected layers, the activation function of the hidden layer is Relu, and the activation function of the output layer is Tanh; The evaluation network Critic consists of 3 fully connected layers, the activation function of the hidden layer is Relu, and the activation function of the output layer is the identity function; The policy network Actor consists of 2 RNN layers and 2 fully connected layers, the activation functions of the hidden layers are all Relu, and the activation function of the output layer is Tanh.

5. The vector propulsion UUV intelligent motion control method based on improved maximum entropy reinforcement learning according to claim 1, characterized in that, The said Step 3 specifically includes the following steps: Step 3.1: Initialize the policy network Actor and the deep neural network; interact the backup controller with the six-degree-of-freedom model of the vector UUV to obtain sub-optimal experience data, and save it in the sub-optimal experience buffer; collect batchSize sub-optimal experience tuples from the sub-optimal experience buffer, and train the deep neural network model to approximate the six-degree-of-freedom model of the vector propulsion UUV and train the policy network Actor to approximate the backup controller by minimizing the mean square error; the backup controller is a traditional non-linear controller; Step 3.2: Interact with the actual environment model through the policy network Actor to obtain real experience data and store it in the real experience buffer; the parameters of the actual environment model consist of the six-degree-of-freedom agile motion model parameters of the vector propulsion UUV and random quantities; design the interaction termination condition done. If the termination condition done is true, stop the interaction process between the policy network Actor and the actual environment model, and store the obtained real experience tuple (S t , A t , S t+1 , R t+1 ) in the real experience buffer; if the termination condition done is false, add the real experience tuple (S t , A t , S t+1 , R t+1 ) to the real experience buffer and update the time step to the next moment; when the number of tuples in the real experience buffer exceeds the batch size batchSize, then update the parameters of the deep neural network, the evaluation network Critic, and the policy network Actor at each time step in each subsequent round of interaction; Step 3.3: Update the deep neural network: Collect batchSize real experience tuples from the real experience buffer, train the deep neural network model to approximate the actual environment model by calculating the loss value Loss, and select the sampling step number n according to the error between the deep neural network and the actual environment model; Step 3.4: Update the evaluation network Critic: If the sampling step number n = 0, use the real experience tuple (S t , A t , S t+1 , R t+1 ) in the real experience buffer to calculate the 1-step temporal difference target and use the 1-step TD algorithm to update the evaluation network Critic; if n > 0, after inputting the sampled batchSize real experience tuples into the deep neural network model, start the n-step interaction between the policy network Actor and the deep neural network model, calculate the cumulative reward value at each time step, and store the imaginary experience tuple (S t , A t , S’ t+n , R t+1,...,t+n ) in the imaginary experience buffer at the nth time step; input the imgBatchSize imaginary experience tuples in the imaginary experience buffer into the evaluation network Critic and then output the predicted state-action value Q w (S t , A t ), calculate the n-step temporal difference target and use the n-step TD algorithm to update the evaluation network Critic; Step 3.5: Update the policy network Actor: After inputting the sampled imgBatchSize imaginary experience tuples into the policy network Actor, output and the entropy value Update the policy network Actor according to the SAC algorithm.

6. The vector propulsion UUV intelligent motion control method based on improved maximum entropy reinforcement learning according to claim 5, wherein Step 3.2 specifically means randomly initializing the state of the actual environment model as S t , input S t to the policy network Actor and then outputting the action A t , input A t to the actual environment model and then outputting the next moment state value S t+1 and the reward value R t+1 , judging the termination condition done; the termination condition includes that the interaction reaches the maximum number of time steps or the pose deviation reaches the maximum threshold.

7. The vector propulsion UUV intelligent motion control method based on improved maximum entropy reinforcement learning according to claim 5, characterized in that In the said Step 3.3, randomly sample batchSize experience tuples from the real buffer and input them into the deep neural network, and calculate the loss value Loss by minimizing the mean square error: where bs is the batchSize; δS t is the output value of the deep neural network model; The sampling step number n: where k loss and c loss are scaling factors; n max is the maximum number of sampling steps.

8. The vector propulsion UUV agile motion control method based on improved maximum entropy reinforcement learning according to claim 5, characterized in that In step 3.4, the one-step temporal difference target Among them, Q w is the state-action value; γ is the reward discount rate; is calculated through the policy network Actor; H is the entropy of the policy network Actor; β is a regularization coefficient used to control the importance of entropy; The loss function L of the 1-step TD algorithm Q1 (w): The n-step temporal difference target The loss function L of the n-step TD algorithm Qn (w): Among them, ibs is imgBatchSize.

9. The vector propulsion UUV intelligent motion control method based on improved maximum entropy reinforcement learning according to claim 5, wherein The loss function L π (θ) of the said strategy:

10. The vector propulsion UUV intelligent motion control method based on improved maximum entropy reinforcement learning according to claim 1, characterized in that, In step 4, when selecting a backup controller for actual control output, continue to optimize the deep neural network, the policy network Actor, and the evaluation network Critic for the next control movement.