Autonomous underwater vehicle trajectory tracking method and system based on deep reinforcement learning
By adopting the strategy-evaluation network and empirical buffer pool method in the deep reinforcement learning framework of autonomous underwater vehicles, the problem of insufficient convergence speed when the environment changes rapidly is solved, and more efficient trajectory tracking control is achieved.
Patent Information
- Application Number
- CN202510086851.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The existing deep reinforcement learning algorithms have insufficient convergence speed when the environment changes rapidly, resulting in low trajectory tracking efficiency of autonomous underwater vehicles in complex and changing environments, which cannot meet the needs of practical applications.
A deep reinforcement learning framework based on strategy-evaluation network is adopted to construct control targets on the state information and action information of autonomous underwater vehicles, and randomly sample and screen using transfer tuples in the experience buffer pool to form the final training data set to accelerate the training process of strategy-evaluation network.
Under the limited data set, the convergence speed of the autonomous underwater vehicle trajectory tracking controller is significantly improved, reducing the waste of time and resources, and improving the trajectory tracking efficiency.
Smart Images

Figure CN119937604A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of autonomous underwater vehicle control, and in particular to a method and system for autonomous underwater vehicle trajectory tracking based on deep reinforcement learning. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Autonomous underwater vehicles (AUVs) can replace manual work to complete a variety of tasks such as marine environmental exploration, underwater autonomous operations, marine hydrological environment detection, and seabed terrain search. In addition to the civilian field, autonomous underwater robots have great application potential in underwater early warning, underwater survey, formation escort, mine detection, intelligence collection, and other aspects. In order to ensure that AUV can successfully complete the assigned tasks, the motion control system is of great significance. Dynamic positioning, path tracking, and trajectory tracking are the three basic motion control problems faced by AUV in actual tasks. Trajectory tracking is one of the most important research contents, ensuring that the AUV reaches the specified location along the desired trajectory. Trajectory tracking involves environmental perception, decision making, action execution, etc., and plays a significant role in scientific exploration (such as seabed topography mapping) and commercial activities (such as inspection and maintenance of underwater facilities).
[0004] In the actual underwater mission execution, such as environmental detection, shipwreck search, etc., environmental factors (such as water flow, obstacles, underwater terrain, etc.) are highly dynamic and unpredictable. Although the current deep reinforcement learning method shows a certain adaptability in the absence of a model, the convergence speed when the environment changes rapidly cannot meet the actual needs. That is, due to the problem of sample efficiency, the existing algorithm needs to be further improved.
[0005] Algorithms that can quickly adapt to environmental changes have important research and application value for improving the autonomy and mission success rate of AUVs in complex and changing environments. At present, the mainstream reinforcement learning algorithms are designed and optimized based on the Actor-Critic framework. Today, the Soft Actor-Critic (SAC) algorithm is widely used for its excellent performance and has achieved good results. The existing SAC algorithm has a slow convergence speed due to sample efficiency issues. Low sample efficiency means that the algorithm requires a large amount of interaction data to learn effective strategies, which is not feasible in many practical applications, especially in scenarios where data acquisition is expensive or time-sensitive. AUVs are expensive, so the cost of acquiring data is high, so the control algorithm that requires trajectory tracking needs to converge as quickly as possible during the training phase to reduce resource waste. Summary of the invention
[0006] In order to solve the above technical problems, the present invention provides an autonomous underwater vehicle trajectory tracking method and system based on deep reinforcement learning, which can converge faster when training the network under a limited data set, and can reduce the waste of time and resources in actual engineering applications.
[0007] In order to achieve the above object, the present invention adopts the following technical solution:
[0008] A first aspect of the present invention provides a trajectory tracking method for an autonomous underwater vehicle based on deep reinforcement learning.
[0009] In one or more embodiments, a method for tracking trajectory of an autonomous underwater vehicle based on deep reinforcement learning is provided, comprising:
[0010] Under the autonomous underwater system, the state information of the autonomous underwater vehicle is taken as input, the thrust of the propeller and the vertical rudder angle are taken as output, and the control target of the vehicle trajectory tracking controller is constructed;
[0011] Based on the predefined state vector, action vector and reward function of the autonomous underwater vehicle, the control target of the vehicle trajectory tracking controller is converted into the autonomous underwater vehicle trajectory tracking control target under the deep reinforcement learning framework based on the policy-evaluation network;
[0012] According to the current state vector of the autonomous underwater vehicle, the current action vector and its corresponding reward function value and the next state vector are sampled to form a transfer tuple and store it in the experience buffer pool D;
[0013] Randomly sample the transfer tuples in the experience buffer pool twice to obtain two initial data sets of set size, and then select a number of top transfer tuples from one of the initial data sets according to the score of the current complete exploration process to obtain a screened data set, and randomly replace the same number of transfer tuples in the other initial data set with the screened data set to obtain the final training data set for iterative training of the policy-evaluation network;
[0014] The iteratively trained policy-evaluation network is used as the control network to control the autonomous underwater vehicle.
[0015] As an implementation method, the control goal of the vehicle trajectory tracking controller is to solve the optimal control strategy so that the objective function R t (τ) is maximized;
[0016] R t (τ)=∑ i≥t γ i-t r;
[0017] Where γ is the discount factor, r is the reward function, and τ is the system output, namely the thrust of the propeller and the vertical rudder angle; R t (τ) is a function related to τ; t is the time.
[0018] As an implementation method, the reward function at time t is defined as r t , the reward function represents taking action a at time t t The reward obtained; according to the current position, expected position, current heading angle, expected heading angle, and output action of the autonomous underwater vehicle, the reward function is set as:
[0019]
[0020] in, represents the reward close to the expected trajectory, Indicates the reward close to the expected heading angle, Represents the action reward.
[0021] As an implementation method, the objective function J of the policy network in the policy-evaluation network is π The expression of (φ) is:
[0022]
[0023] Among them, π φ (a t |s t ) is the policy network, φ is the network parameter; Q θ (s t ,a t ) is the evaluation network, θ is the evaluation network parameter; α is the temperature coefficient; S t is the state vector; a t is the action vector; a t ∽π φ Indicates action a t According to the current strategy π φ , given the state s t The distribution of sampling at time; s t ∽D represents state s t is sampled from the experience buffer pool D.
[0024] As an implementation method, the objective function J of the evaluation network in the strategy-evaluation network is Q The expression of (θ) is:
[0025]
[0026] Among them, Q θ (s t ,a t ) is the evaluation network, θ is the evaluation network parameter; πφ (a t |s t ) is the strategy network, α is the temperature coefficient; represents the value function, represents the target evaluation network parameters; γ is the discount factor; s t is the state vector; a t is the action vector; (s t , a t )∽D represents the state (s t , a t ) is sampled from the experience buffer pool D; a t ∽π represents action a t According to the current strategy π φ Sampling from a distribution given in a given state; s t+1 ∽P represents the observed state at the next moment; r(s t ,a t ) represents the immediate reward obtained after the AUV takes action at time t; Represents the value output by the target evaluation network based on the state at time t and the action taken.
[0027] As an implementation method, the objective function J of the temperature coefficient α α The expression is:
[0028]
[0029] in, is the entropy value of the action distribution of the policy network; It means that the expectation is about actions, and these actions are based on the current policy π φ , given the state s t The distribution of sampling.
[0030] As an implementation method, the calculation process of the score of the current complete exploration process is: accumulating the reward value obtained from each step of the current exploration.
[0031] A second aspect of the present invention provides an autonomous underwater vehicle trajectory tracking system based on deep reinforcement learning.
[0032] In one or more embodiments, a deep reinforcement learning-based autonomous underwater vehicle trajectory tracking system includes:
[0033] A control target building module is used to build a control target of the vehicle trajectory tracking controller by taking the state information of the autonomous underwater vehicle as input and the thrust of the propeller and the vertical rudder angle as output under the autonomous underwater positioning system;
[0034] A control target conversion module, which is used to convert the control target of the vehicle trajectory tracking controller into the autonomous underwater vehicle trajectory tracking control target under the deep reinforcement learning framework based on the policy-evaluation network based on the state vector, action vector and reward function of the autonomous underwater vehicle in advance;
[0035] A transfer tuple building module is used to sample the current action vector and its corresponding reward function value and the next state vector according to the current state vector of the autonomous underwater vehicle, form a transfer tuple and store it in the experience buffer pool;
[0036] The strategy-evaluation network training module is used to randomly sample the transfer tuples in the experience buffer pool twice to obtain two initial data sets of set size, and then select a number of top transfer tuples from one of the initial data sets according to the score of the current complete exploration process to obtain a screened data set, and randomly replace the same number of transfer tuples in the other initial data set with the screened data set to obtain the final training data set for iterative training of the strategy-evaluation network;
[0037] The vehicle control module is used to control the autonomous underwater vehicle by using the iteratively trained policy-evaluation network as a control network.
[0038] A third aspect of the present invention provides a computer-readable storage medium.
[0039] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the above-mentioned method for tracking trajectory of an autonomous underwater vehicle based on deep reinforcement learning.
[0040] A fourth aspect of the present invention provides an electronic device.
[0041] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the above-mentioned method for tracking trajectory of an autonomous underwater vehicle based on deep reinforcement learning are implemented.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] The present invention converts the control target of a vehicle trajectory tracking controller into an autonomous underwater vehicle trajectory tracking control target under a deep reinforcement learning framework based on a policy-evaluation network. It does not require complex modeling. It randomly samples transfer tuples in an experience buffer pool twice to obtain two initial data sets of set sizes. Then, according to the score of the current complete exploration process, a number of top transfer tuples are screened out from one of the initial data sets to obtain a screened data set. The screened data set is randomly replaced with the same number of transfer tuples in another initial data set to obtain a final training data set for iterative training of a policy-evaluation network. The excellence of the exploration round in which different transfer tuples obtained by random sampling are located is taken into account for iterative training of the policy-evaluation network. The priority experience set is increased, and the convergence speed of the policy-evaluation network during training is accelerated. In actual engineering applications, the waste of time and resources can be reduced, thereby improving the trajectory tracking efficiency of autonomous underwater vehicles. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0045] Figure 1 is a schematic diagram of the trajectory tracking control principle of an autonomous underwater vehicle according to an embodiment of the present invention;
[0046] FIG. 2( a ) is an environmental exploration process of the improved SAC algorithm according to an embodiment of the present invention;
[0047] FIG2( b ) is an improved specific network update process of an embodiment of the present invention;
[0048] Figure 3 is the SAC algorithm network structure of an embodiment of the present invention;
[0049] Figure 4 This is the Actor network update process of an embodiment of the present invention;
[0050] Figure 5 This is the Critic network update process of an embodiment of the present invention;
[0051] Figure 6 is a flow chart of a method for tracking an autonomous underwater vehicle trajectory based on deep reinforcement learning according to an embodiment of the present invention;
[0052] Figure 7 is a schematic diagram of the structure of an autonomous underwater vehicle trajectory tracking system based on deep reinforcement learning according to an embodiment of the present invention;
[0053] Figure 8 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0055] It should be noted that the following detailed descriptions are all illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0056] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0057] Figure 1 is a schematic diagram of the trajectory tracking control principle of an autonomous underwater vehicle according to an embodiment of the present invention; Figure 6 It is a flow chart of an autonomous underwater vehicle trajectory tracking method based on deep reinforcement learning according to an embodiment of the present invention.
[0058] Combination Figure 1 and Figure 6 , the autonomous underwater vehicle trajectory tracking method based on deep reinforcement learning in this embodiment may include:
[0059] S601, in the autonomous underwater system, taking the state information of the autonomous underwater vehicle as input, and the thrust of the propeller and the vertical rudder angle as output, to construct a control target of the vehicle trajectory tracking controller;
[0060] S602, based on the predefined state vector, action vector and reward function of the autonomous underwater vehicle, converting the control target of the vehicle trajectory tracking controller into the autonomous underwater vehicle trajectory tracking control target under the deep reinforcement learning framework based on the policy-evaluation network;
[0061] S603, according to the current state vector of the autonomous underwater vehicle, sampling to obtain the current action vector and its corresponding reward function value and the next state vector, forming a transfer tuple and storing it in the experience buffer pool;
[0062] S604, randomly sampling the transfer tuples in the experience buffer pool twice to obtain two initial data sets of set size, and then filtering out a number of top transfer tuples from one of the initial data sets according to the score of the current complete exploration process to obtain a filtered data set, and randomly replacing the same number of transfer tuples in the other initial data set with the filtered data set to obtain a final training data set for iterative training of the policy-evaluation network;
[0063] S605, using the iteratively trained policy-evaluation network as a control network to control the autonomous underwater vehicle.
[0064] This embodiment converts the control target of the vehicle trajectory tracking controller into an autonomous underwater vehicle trajectory tracking control target under a deep reinforcement learning framework based on a policy-evaluation network. It does not require complex modeling, randomly samples the transfer tuples in the experience buffer pool twice, and obtains two initial data sets of set size, and then selects a number of top transfer tuples from one of the initial data sets according to the score of the current complete exploration process to obtain a screened data set, and randomly replaces the screened data set with the same number of transfer tuples in the other initial data set to obtain a final training data set for iterative training of the policy-evaluation network. The excellence of the exploration round in which the different transfer tuples obtained by random sampling are located is taken into account for iterative training of the policy-evaluation network, and the priority experience set is increased, which speeds up the convergence speed of the policy-evaluation network training. In actual engineering applications, it can reduce the waste of time and resources, thereby improving the efficiency of trajectory tracking of autonomous underwater vehicles.
[0065] In step S601, in the autonomous underwater positioning system, the state information of the autonomous underwater vehicle includes position and speed; position [x t ,y t ,z t , θ t ,ψ t ], where x t ,y t ,z t represents the position of the AUV (autonomous underwater vehicle) at time t, θ t ,ψ t They represent the roll angle, pitch angle and yaw angle of the AUV in the fixed system, and the velocity [u t ,v t ,w t ,p t ,q t ,r t ],u t 、v t 、w trepresents the linear velocity of the AUV in three directions at time t under the dynamic system, p t ,q t ,r t They represent the roll angular velocity, pitch angular velocity and yaw angular velocity respectively.
[0066] This example is a two-dimensional AUV trajectory tracking problem, so the vertical motion of the AUV is not considered. The final input of the system becomes: position [x t ,y t ,ψ t ], speed [u t ,v t ,r t ].
[0067] The system output is τ t =[f t ,δ t ,], including the propeller thrust f t and the vertical rudder angle δ t .
[0068] Determine the position error:
[0069] Position error in the x-axis direction at time t Position error in y-axis direction
[0070] The position error at time t is expressed as: It is the difference between the sensor position information and the reference position information.
[0071] The control goal of the vehicle trajectory tracking controller is to solve the optimal control strategy so that the objective function R t (τ) is maximized;
[0072] R t (τ)=∑ i≥t γ i-t r;
[0073] Where γ is the discount factor, r is the reward function, and τ is the system output, namely the thrust of the propeller and the vertical rudder angle; R t (τ) is a function related to τ; t is the time.
[0074] In step S602, the state vector is defined:
[0075] The state vector is
[0076] The current position information of AUV x t ,y t ,ψ t ; The current speed information u of AUV t,v t ,r t ; The expected position information of AUV at the current moment The error between the current AUV and the expected position e t ; The output of the AUV actuator at the last moment τ t-1 , including propeller thrust and vertical rudder angle.
[0077] Define the action vector:
[0078] Define the action vector at time t as the system output a at time t t =τ t , where τ t =[f t ,δ t ,].
[0079] Define the reward function:
[0080] Define the reward function at time t as r t , the reward function represents taking action a at time t t The reward obtained; according to the current position, expected position, current heading angle, expected heading angle, and output action of the autonomous underwater vehicle, the reward function is set as:
[0081]
[0082] in, represents the reward close to the expected trajectory, Indicates the reward close to the expected heading angle, Represents the action reward.
[0083] Objective function of the policy network in the policy-critic network J π The expression of (φ) is:
[0084]
[0085] Among them, π φ (a t |s t ) is the policy network, φ is the network parameter; Q θ (s t ,a t ) is the evaluation network, θ is the evaluation network parameter; α is the temperature coefficient; s t is the state vector; a t is the action vector; a t ∽π φ Indicates action a t According to the current strategy π φ , given state a t The distribution of sampling at time; s t∽D represents state s t is sampled from the experience buffer pool D.
[0086] Among them, the objective function J of the evaluation network in the strategy-evaluation network is Q The expression of (θ) is:
[0087]
[0088] Among them, Q θ (s t ,a t ) is the evaluation network, θ is the evaluation network parameter; π φ (a t |s t ) is the strategy network, α is the temperature coefficient; represents the value function, represents the target evaluation network parameters; γ is the discount factor; s t is the state vector; a t is the action vector; (s t , a t )∽D represents the state (s t , a t ) is sampled from the experience buffer pool D; a t ∽π represents action a t According to the current strategy π φ Sampling from a distribution given in a given state; s t+1 ∽P represents the observed state at the next moment; r(s t ,a t ) represents the immediate reward obtained after the AUV takes action at time t; Represents the value output by the target evaluation network based on the state at time t and the action taken.
[0089] The objective function J of the temperature coefficient α α The expression is:
[0090]
[0091] in, is the entropy value of the action distribution of the policy network; It means that the expectation is about actions, and these actions are based on the current policy π φ , given the state s t The distribution of sampling.
[0092] As shown in Figure 2(a), Figure 2(b) and Figure 3 As shown in the figure, the construction process of the strategy-criticism network (SAC network) is:
[0093] Build the Actor network:
[0094] By building an Actor network, the current action is output according to the current input state. In order to make the output strategy more stable, the clip function is used to limit the range of change of the output action of the new strategy.
[0095] like Figure 4 As shown in Figure 1, the Actor network consists of 1 input layer, 2 hidden layers, and 2 output layers, and each layer is a fully connected layer. The input layer input is the observation s t , the number of neurons is the observation quantity s t The dimension of the hidden layer is 300 neurons; the output layer outputs action a t The mean and standard deviation of , and by establishing a normal distribution based on this, and performing re-parameter sampling to obtain the final output action.
[0096] The activation function of the input layer and the hidden layer uses the Relu function, and the output layer uses the tanh function to limit the output range.
[0097] Constructing the Critic network:
[0098] By constructing a Critic network to evaluate the actions output by the current Actor network, the value Q is obtained.
[0099] like Figure 5 As shown in the figure, the Critic network consists of four networks, namely Q1 network, Q2 network, Q1_target network and Q2_target network. The structure of each network is the same.
[0100] Each Critic network consists of an input layer, two hidden layers, and an output layer. Each layer of the neural network is a fully connected layer. The input layer input is the current state s t and the action taken t , the number of neurons is state s t And action a t The total dimension is 300. The number of neurons in the hidden layer is 300. The output layer outputs the evaluation of the current action, with a dimension of 1.
[0101] The activation functions of the input layer and the hidden layer are sampled with the Relu function, and the output layer uses uniform distribution to initialize the weights and bias parameters of the fully connected layer.
[0102] Determine your target strategy:
[0103] When the average reward of each round tends to be stable starting from a certain time step t, it can be considered that the target strategy is obtained, and the strategy learned after time step t is used as the output.
[0104] In combination with step S603 and step S604, the process of iterative training strategy-evaluation network is as follows:
[0105] (1) Parameter settings:
[0106] Optimizer learning rate lr, discount factor γ, the size of the experience buffer pool D is M, the size of the randomly sampled data set for each training is b, the size of the advantage data set is k, the target Q network update parameter f, the maximum number of steps during training T, and the number of steps that meet the training start conditions T_train.
[0107] (2) Initialize the Soft Actor-Critic network parameters:
[0108] Randomly initialize the Actor network parameters φ, the parameters of the two Critic networks Q1 and Q2 networks θ1 and θ2, and the parameters of the two target networks Q1_target and Q2_target networks Initializing the experience buffer pool means creating a new empty training set.
[0109] (3) Obtain the state transfer tuple to the experience buffer pool:
[0110] Initialization state t , the iteration begins. First, input the current state s t Go to the Actor network and sample to get action a t And the reward value r(s t ,a t ), and get the next state s t+1 , the obtained transfer tuple {s t , a t ,r(s t ,a t ), s t+1} is stored in the experience buffer pool D; In addition, this method introduces a new parameter ρ based on the original state transfer array e , which represents the score of the e-th exploration process. This is a complete exploration rather than the score of each step. Therefore, ρ in the transfer tuple during an exploration process e The values are the same; therefore, the new transfer tuple is as follows {s t , a t ,r(s t ,a t ), s t+1 , ρ e}. Repeat the above process until the set number of steps T is met.
[0111] Among them, the score of the current complete exploration process is ρ e The calculation process is: accumulate the reward value obtained from each exploration step in the current step.
[0112] (4) Obtain training data set:
[0113] When the number of iteration steps meets the set number of steps T_train, it indicates that the experience buffer pool has collected a certain number of transfer tuples, and training can begin. Sample from the experience buffer pool D to form a training set for the neural network to update. The specific operation process is: first, randomly sample the transfer tuples in the experience buffer pool to obtain a b-dataset B1; then randomly sample the transfer tuples in the experience buffer pool to obtain a b-dataset B2; according to the parameter ρ e Prioritize the data in B2 and select ρ from the B2 data set e The highest k The transferred tuple data are combined into a new data set B2_prior; the obtained data set B2_prior is merged with the data set B1 to form a new data set B for subsequent training.
[0114] (5) Critic Network Update:
[0115] When the training set is ready, the neural network is updated. Q (θ) function, and perform gradient descent to update the network parameters θ1 and θ2 of the two Critic networks Q1 and Q2.
[0116]
[0117] (6) Actor network update:
[0118] Then according to J π (φ) function, performs gradient descent to update the parameters φ of the Actor network.
[0119]
[0120] (7) Update of temperature coefficient α:
[0121] The temperature coefficient α is then updated.
[0122]
[0123] (8) Target network update:
[0124] Finally, the parameters of the two target networks Q1_target and Q2_target are soft-updated.
[0125]
[0126] After the neural network is completely updated, it enters the next iteration process and repeats steps (7) to (9) until the iteration ends.
[0127] The iteration is stopped, and the improved SAC algorithm of the embodiment of the present invention is updated and the final θ1, θ2, and φ are output.
[0128] The learned Actor and Critic networks are used as the control network to realize trajectory tracking control of the autonomous underwater vehicle.
[0129] Figure 7 is a schematic diagram of the structure of an autonomous underwater vehicle trajectory tracking system based on deep reinforcement learning in an embodiment of the present invention. Figure 6 The trajectory tracking method of autonomous underwater vehicles based on deep reinforcement learning corresponds to Figure 7 As shown, the autonomous underwater vehicle trajectory tracking system based on deep reinforcement learning in this embodiment may include:
[0130] The control target building module 701 is used to build the control target of the vehicle trajectory tracking controller by taking the position, speed and other state information of the autonomous underwater vehicle as input and the thrust of the propeller and the vertical rudder angle as output under the autonomous underwater positioning system;
[0131] A control target conversion module 702 is used to convert the control target of the vehicle trajectory tracking controller into the autonomous underwater vehicle trajectory tracking control target under the deep reinforcement learning framework based on the policy-evaluation network based on the state vector, action vector and reward function of the autonomous underwater vehicle in advance;
[0132] A transfer tuple construction module 703 is used to sample the current action vector and its corresponding reward function value and the next state vector according to the current state vector of the autonomous underwater vehicle, form a transfer tuple and store it in the experience buffer pool;
[0133] The strategy-evaluation network training module 704 is used to randomly sample the transfer tuples in the experience buffer pool twice to obtain two initial data sets of set size, and then select a number of top transfer tuples from one of the initial data sets according to the score of the current complete exploration process to obtain a screened data set, and randomly replace the same number of transfer tuples in the other initial data set with the screened data set to obtain a final training data set for iterative training of the strategy-evaluation network;
[0134] The vehicle control module 705 is used to control the autonomous underwater vehicle by using the iteratively trained policy-evaluation network as a control network.
[0135] It should be noted here that Figure 7 The various modules in the deep reinforcement learning-based autonomous underwater vehicle trajectory tracking system in Figure 6The various steps in the deep reinforcement learning-based autonomous underwater vehicle trajectory tracking method in correspondence one to one, and the specific implementation process is the same, which will not be repeated here.
[0136] Reference Figure 8 , a schematic diagram of an electronic device is given. It should be noted that, Figure 8 The electronic device 800 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0137] like Figure 8 As shown, electronic device 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage part 808 into a random access memory (RAM) 803. In RAM 803, various programs and data required for system operation are also stored. Central processing unit 801, ROM 802 and RAM 803 are connected to each other via a bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0138] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a local area network (LAN) card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed, so that a computer program read therefrom is installed into the storage section 808 as needed.
[0139] When the central processing unit 801 in the electronic device of this embodiment executes the program, the following is achieved: Figure 6 The steps in the deep reinforcement learning based autonomous underwater vehicle trajectory tracking method are shown.
[0140] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, the computer program including a computer program for executing Figure 6In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 809, and / or installed from a removable medium 811. When the computer program is executed by the central processing unit 801, various functions defined in the apparatus of the present application are executed.
[0141] in, Figure 6 The computer program instructions corresponding to the method shown may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0142] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0143] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A trajectory tracking method for an autonomous underwater vehicle based on deep reinforcement learning, characterized in that: include: Under the autonomous underwater system, the state information of the autonomous underwater vehicle is taken as input, the thrust of the propeller and the vertical rudder angle are taken as output, and the control target of the vehicle trajectory tracking controller is constructed; Based on the predefined state vector, action vector and reward function of the autonomous underwater vehicle, the control target of the vehicle trajectory tracking controller is converted into the autonomous underwater vehicle trajectory tracking control target under the deep reinforcement learning framework based on the policy-evaluation network; According to the current state vector of the autonomous underwater vehicle, the current action vector and its corresponding reward function value and the next state vector are sampled to form a transfer tuple and store it in the experience buffer pool; Randomly sample the transfer tuples in the experience buffer pool twice to obtain two initial data sets of set size, and then select a number of top transfer tuples from one of the initial data sets according to the score of the current complete exploration process to obtain a screened data set, and randomly replace the same number of transfer tuples in the other initial data set with the screened data set to obtain the final training data set for iterative training of the policy-evaluation network; The iteratively trained policy-evaluation network is used as the control network to control the autonomous underwater vehicle.
2. The method for tracking trajectory of an autonomous underwater vehicle based on deep reinforcement learning according to claim 1, characterized in that: The control goal of the vehicle trajectory tracking controller is to solve the optimal control strategy so that the objective function R t (τ) maximized; R t (τ)=∑ i≥t c i-t r; Where γ is the discount factor, r is the reward function, and τ is the system output, namely the thrust of the propeller and the vertical rudder angle; R t (τ) is a function related to τ; t is the time.
3. The autonomous underwater vehicle trajectory tracking method based on deep reinforcement learning according to claim 1, characterized in that: Define the reward function at time t as r t , the reward function represents the state s t Take action a t The reward obtained; according to the current position, expected position, current heading angle, expected heading angle, and output action of the autonomous underwater vehicle, the reward function is set as: in, represents the reward close to the expected trajectory, Indicates the reward close to the expected heading angle, Represents the action reward.
4. The method for tracking trajectory of an autonomous underwater vehicle based on deep reinforcement learning according to claim 1, characterized in that: Objective function of the policy network in the policy-critic network J π The expression of (φ) is: Among them, π φ (a t |s t ) is the policy network, φ is the network parameter; Q θ (s t ,a t ) is the evaluation network, θ is the evaluation network parameter; α is the temperature coefficient; s t is the state vector; a t is the action vector; Indicates action a t According to the current strategy π φ , given the state s t The distribution of sampling when Indicates state s t is sampled from the experience buffer pool D.
5. The autonomous underwater vehicle trajectory tracking method based on deep reinforcement learning according to claim 1, characterized in that: Objective function of the evaluation network in the policy-evaluation network J Q The expression of (θ) is: Among them, Q θ (s t ,a t ) is the evaluation network, θ is the evaluation network parameter; π φ (a t |s t ) is the strategy network, α is the temperature coefficient; represents the value function, represents the target evaluation network parameters; γ is the discount factor; s t is the state vector; a t is the action vector; Indicates the status (s t , a t ) is sampled from the experience buffer pool D; Indicates action a t According to the current strategy π φ Sampling from a distribution given in a given state; represents the observed state at the next moment; r(s t ,a t ) represents the immediate reward obtained after the AUV takes action at time t; Represents the value output by the target evaluation network based on the state at time t and the action taken.
6. The autonomous underwater vehicle trajectory tracking method based on deep reinforcement learning according to claim 4 or 5, characterized in that: The objective function J of the temperature coefficient α α The expression is: in, is the entropy value of the action distribution of the policy network; It means that the expectation is about actions, and these actions are based on the current policy π φ , given the state s t The distribution of sampling.
7. The method for tracking trajectory of an autonomous underwater vehicle based on deep reinforcement learning according to claim 1, characterized in that: The calculation process of the score of the current complete exploration process is: accumulate the reward values obtained from each step of the current exploration.
8. An autonomous underwater vehicle trajectory tracking system based on deep reinforcement learning, characterized in that: include: A control target building module is used to build a control target of the vehicle trajectory tracking controller by taking the state information of the autonomous underwater vehicle as input and the thrust of the propeller and the vertical rudder angle as output under the autonomous underwater positioning system; A control target conversion module, which is used to convert the control target of the vehicle trajectory tracking controller into the autonomous underwater vehicle trajectory tracking control target under the deep reinforcement learning framework based on the policy-evaluation network based on the state vector, action vector and reward function of the autonomous underwater vehicle in advance; A transfer tuple building module is used to sample the current action vector and its corresponding reward function value and the next state vector according to the current state vector of the autonomous underwater vehicle, form a transfer tuple and store it in the experience buffer pool; The strategy-evaluation network training module is used to randomly sample the transfer tuples in the experience buffer pool twice to obtain two initial data sets of set size, and then select a number of top transfer tuples from one of the initial data sets according to the score of the current complete exploration process to obtain a screened data set, and randomly replace the same number of transfer tuples in the other initial data set with the screened data set to obtain the final training data set for iterative training of the strategy-evaluation network; The vehicle control module is used to control the autonomous underwater vehicle by using the iteratively trained policy-evaluation network as a control network.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the trajectory tracking method of an autonomous underwater vehicle based on deep reinforcement learning as described in any one of claims 1 to 7 are implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the autonomous underwater vehicle trajectory tracking method based on deep reinforcement learning as described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Autonomous underwater vehicle trajectory tracking control method based on deep reinforcement learning
CN108803321A
Aircraft route tracking method based on deep reinforcement learning
CN110806759A
Underwater vehicle bottom layer control method and system based on deep reinforcement learning
CN114839884A
Target tracking control method for autonomous underwater vehicle based on trajectory prediction
CN115657689A
A Safety Control Method for Trajectory Tracking of the Unmanned Underwater Vehicle Based on the Slew Rate Constraint
LU500243B1
Cited By
Spacecraft low-thrust trajectory data generation method and system
CN120621718A
Robot generative adversarial self-imitation learning method based on large language model feedback
CN121706839A