Unmanned submersible vehicle path tracking method based on DDPG-DKNN
By adopting the DDPG-DKNN method in unmanned submarine path tracking, combining deep neural networks and dynamic KNN algorithms, the problems of insolent path tracking and high computational complexity in the existing technology are solved, and high-precision, stable and adaptive path tracking is achieved.
Patent Information
- Application Number
- CN202510170922.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-23
AI Technical Summary
The existing unmanned submarine path tracking methods are difficult to generate robust paths when dealing with complex submarine terrain and dynamic currents, and have high computational complexity and difficult to output control strategies in real time, resulting in increased navigation risks.
The unmanned submarine path tracking method based on DDPG-DKNN is adopted, and the path tracking strategy is output in real time through the DDPG algorithm combined with deep neural network and deterministic strategy gradient, and the K value is dynamically adjusted through the DKNN algorithm to adapt to changes in the sea area environment.
It improves the accuracy and stability of unmanned submarine path tracking, reduces the computational complexity, realizes adaptive response to complex sea environments, and enhances the generalization ability and real-time nature of the algorithm.
Smart Images

Figure CN120029289A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of marine intelligent systems, and in particular to an unmanned submersible path tracking method based on DDPG-DKNN. Background Art
[0002] With the rapid development of unmanned systems, unmanned submersibles have been widely used in dangerous target identification, resource exploration, and environmental monitoring due to their unique advantages in ocean exploration. Path tracking technology is a key technology for realizing autonomous underwater navigation of unmanned submersibles, ensuring their precise movement along a preset path. However, considering the complex seabed topography, dynamic ocean currents, tides and other disturbance factors, existing path tracking methods are often difficult to cope with changes in the external environment and cannot generate robust path tracking strategies. In addition, traditional methods have high computational complexity when processing large-scale time-varying ocean data, and it is difficult to output control strategies in real time, thereby reducing the underwater mission execution efficiency of unmanned submersibles and increasing navigation risks. Therefore, a DDPG-DKNN-based unmanned submersible path tracking method is needed to effectively improve the path tracking accuracy of unmanned submersibles. Summary of the invention
[0003] In view of this, the purpose of the present invention is to propose an unmanned submersible path tracking method based on DDPG-DKNN to solve the problem that the existing technology has high computational complexity and difficulty in real-time output of control strategies when processing large-scale time-varying ocean data.
[0004] Based on the above purpose, the present invention provides an unmanned submersible path tracking method based on DDPG-DKNN, comprising the following steps:
[0005] Step 100, collect the navigation speed, heading angle, longitude, latitude and current water depth of the unmanned submersible as the state space of the DDPG (Deep Deterministic Policy Gradient) algorithm, and use the DDPG algorithm to solve the path tracking problem in the continuous action space by combining the deep neural network and the deterministic policy gradient to obtain the state s t 、Action a t , reward function r t and the state s at the next moment t+1 ;
[0006] Step 200, based on the experience playback technology, the state s t 、Action a t , reward function r t and the state s at the next moment t+1The state s is stored in the replay buffer, and the parameter k of the DKNN (Dynamic K-Nearest Neighbor) algorithm is dynamically optimized, and the nearest neighbor sampling is performed in the replay buffer. t Quickly find the most similar K samples and dynamically adjust the action strategy based on similar situations;
[0007] Step 300: Output the path tracking strategy of the unmanned underwater vehicle in real time based on the DDPG algorithm.
[0008] Preferably, DDPG adopts a two-layer network architecture, including a prediction network and a target network. The prediction network calculates the loss function and updates the network parameters by calculating the value function, and the target network is used to enhance the stability and convergence speed of the DDPG algorithm.
[0009] Preferably, the value function Q(s t ,a t ) is calculated as follows:
[0010]
[0011] in, represents the Bellman equation. γ represents the discount factor, Q(s t+1 ,μ(s t+1 )) indicates that the unmanned submersible is in state s t+1 When executing the action strategy μ(s t+1 ) is the value function obtained.
[0012] Preferably, step 200 includes the following steps:
[0013] Step 201, based on the number of samples M in the buffer, the K value is adjusted to perform nonlinear dynamic adjustment;
[0014] Step 202, the Actor network maximizes the value function by outputting the action strategy and calculates the loss function of the network based on the mean square error;
[0015] In step 203, the Critic network evaluates the quality of the action strategy based on the value function, and the Actor network updates the prediction network based on the stochastic gradient descent method.
[0016] Preferably, in step 200, the K samples are K samples closest to the query point selected based on the Euclidean distance, and a sample point s in the buffer zone m and query point s t The Euclidean distance between them is p(s m ,s t ):
[0017]
[0018] Where m represents the mth sample, and i represents five vectors in the state space, namely, navigation speed, heading angle, longitude, latitude, and current water depth. mi and ti The i-th vector corresponding to a sample and query point respectively.
[0019] Preferably, in step 201, the calculation formula of the K value is:
[0020] K = log(c·M+b)
[0021] Among them, c and b represent adjustment factors.
[0022] Preferably, in step 202, the calculation formula of the loss function L is:
[0023]
[0024] Among them, N is the number of samples in the small batch sampling, Q'(s t+1 ,μ'(s t+1 ) is the target network in state s t+1 Execute action strategy μ'(s t+1 ) is the value function obtained, and rt represents the reward value obtained at time t.
[0025] Preferably, in step 203, the formula for updating the prediction network of the Actor network based on the stochastic gradient descent method is:
[0026]
[0027] in, θ μ Represents the parameters of the Actor network, To predict the Actor network in the network θ μ gradient. represents the gradient obtained by the target network when executing action a in state s, and a=μ(s) represents the action value selected from the strategy μ(s).
[0028] Beneficial effects of the present invention:
[0029] A. This invention is based on the DDPG algorithm. By combining deep neural networks and deterministic policy gradient methods, the algorithm performance is gradually improved through a large amount of offline training, so that it can directly output action values in the continuous action space to adapt to different sea environments and path tracking requirements.
[0030] B. Based on the DKKN algorithm, the present invention introduces an adjustment factor to dynamically adjust the k value according to the time-varying characteristics of the samples, so that it can adaptively change in complex sea environments, while reducing the computational complexity and effectively improving the generalization ability and stability of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0032] Figure 1 A schematic diagram of a flow chart of a method for tracking a path of an unmanned submersible according to an embodiment of the present invention;
[0033] Figure 2 Schematic diagram of the DKNN algorithm flow in an embodiment of the present invention. DETAILED DESCRIPTION
[0034] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments.
[0035] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0036] This embodiment provides a path tracking method for an unmanned submersible based on DDPG-DKNN. Figure 1 .
[0037] The unmanned submersible path tracking method based on DDPG-DKNN comprises the following steps:
[0038] Step 100, the DDPG algorithm solves the path tracking problem in the continuous action space by combining deep neural networks and deterministic policy gradients. The navigation speed, heading angle, longitude, latitude and current water depth of the unmanned submersible are collected as the state space of the DDPG algorithm. The navigation speed and heading angle of the unmanned submersible are set as the action space. The average path tracking rate of the unmanned submersible is selected as the reward function.
[0039] At time t, the unmanned submersible is in state s t Next, perform action a t , get the state s t+1 , get the instant reward value r(s) of the current stage t ,a t DDPG uses a two-layer network architecture, including a prediction network and a target network. The former calculates the value function Q(s t ,a t ), thereby calculating the loss function and updating the network parameters. The latter is to enhance the stability and convergence speed of the DDPG algorithm. The value function Q(s t ,a t ) is calculated as follows:
[0040]
[0041] in, represents the Bellman equation. γ represents the discount factor, Q(s t+1 ,μ(s t+1 )) indicates that the unmanned submersible is in state s t+1 When executing the action strategy μ(s t+1 ) is the value function obtained.
[0042] Step 200, based on the experience playback technology, the state s t 、Action a t , reward function r t and the state s at the next moment t+1 Stored in the replay buffer. In deep reinforcement learning, the sample distribution in the experience replay buffer is usually complex. If directly sampled randomly, it may contain many irrelevant or worthless samples. By dynamically optimizing the parameter k of the KNN algorithm, an improved KNN algorithm, namely the DDKN algorithm, is proposed. This algorithm performs nearest neighbor sampling in the replay buffer and uses the current state s t Quickly find the most similar K samples, dynamically adjust the action strategy according to similar situations, improve the generalization ability of the algorithm, and enable the unmanned underwater vehicle to adapt to different environments and tasks.
[0043] Step 201, set the total number of samples in the replay buffer to M, and set the current state s tSet as the query point, the DKNN algorithm selects the K samples closest to the query point based on the Euclidean distance, and a sample point s in the buffer m and query point s t The Euclidean distance between them is p(s m ,s t ):
[0044]
[0045] Among them, i is five vectors in the state space, namely, navigation speed, heading angle, longitude, latitude and current water depth. mi and ti The i-th vector corresponding to a sample and query point respectively.
[0046] Considering the complex marine environment, a smaller K value will cause the KNN algorithm to overfit, while a larger K value may lead to underfitting. Therefore, the traditional KNN algorithm using a fixed K value cannot adapt to time-varying navigation data. Based on the number of samples M in the buffer, the K value is adjusted nonlinearly and dynamically to improve the adaptability and robustness of the KNN algorithm, so as to better cope with the time-varying and uncertainty of navigation data.
[0047] K = log(c·M+b)
[0048] Among them, c and b represent adjustment factors.
[0049] Step 202, the two-layer network of the DDPG algorithm includes an Actor network and a Critic network. The Actor network maximizes the value function by outputting action strategies to reduce the correlation between samples. The loss function L of the network is calculated based on the mean square error:
[0050]
[0051] Among them, N is the number of samples in the small batch sampling, Q'(s t+1 ,μ'(s t+1 ) is the target network in state s t+1 Execute action strategy μ'(s t+1 ) is the value function obtained, and rt represents the reward value obtained at time t.
[0052] Step 203, by iteratively calculating the loss function, the Critic network evaluates the quality of the action strategy according to the value function and feeds it back to the Actor network. The Actor network updates the prediction network based on the stochastic gradient descent method.
[0053]
[0054] in, θμ Represents the parameters of the Actor network, To predict the Actor network in the network θ μ gradient. represents the gradient obtained by the target network when executing action a in state s, and a=μ(s) represents the action value selected from the strategy μ(s).
[0055] Step 300, using the soft update method, regularly update the target network with the parameters of the prediction network, thereby improving the stability and convergence of the algorithm. Based on the KNN algorithm, the size of the K value is dynamically adjusted to adapt to different navigation conditions. By combining the DDPG algorithm to output the path tracking strategy of the unmanned submersible in real time, it can efficiently complete different underwater tasks, thereby improving the path tracking accuracy of the unmanned submersible.
[0056] Those skilled in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present invention is limited to these examples; under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes in different aspects of the present invention as described above, which are not provided in detail for the sake of simplicity. Any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for unmanned submersible path tracking based on DDPG-DKNN, characterized in that: The following steps are involved: Step 100, collect the navigation speed, heading angle, longitude, latitude and current water depth of the unmanned submersible as the state space of the DDPG algorithm, and use the DDPG algorithm to solve the path tracking problem in the continuous action space by combining the deep neural network and the deterministic policy gradient to obtain the state s t 、Action a t , reward function r t and the state s at the next moment t+1 ; Step 200: Based on the experience playback technology, the state st and action a t , reward function r t and the state s at the next moment t+1 The parameters k of the DKNN algorithm are dynamically optimized, and the nearest neighbor sampling is performed in the replay buffer. The current state s is used to t Quickly find the most similar K samples and dynamically adjust the action strategy based on similar situations; Step 300: Output the path tracking strategy of the unmanned underwater vehicle in real time based on the DDPG algorithm.
2. The unmanned submersible path tracking method based on DDPG-DKNN according to claim 1, characterized in that: DDPG adopts a two-layer network architecture, including a prediction network and a target network. The prediction network calculates the loss function and updates the network parameters by calculating the value function, and the target network is used to enhance the stability and convergence speed of the DDPG algorithm.
3. The unmanned submersible path tracking method based on DDPG-DKNN according to claim 2 is characterized in that: Predict the value function Q(s t ,a t ) is calculated as follows: in, represents the Bellman equation, γ represents the discount factor, Q(s t+1 ,μ(s t+1 )) indicates that the unmanned submersible is in state s t+1 When executing the action strategy μ(s t+1 ) is the value function obtained.
4. The unmanned submersible path tracking method based on DDPG-DKNN according to claim 3 is characterized in that: Step 200 includes the following steps: Step 201, based on the number of samples M in the buffer, the K value is adjusted nonlinearly and dynamically, where the K value is the number of K samples; Step 202, the Actor network maximizes the value function by outputting the action strategy and calculates the loss function of the network based on the mean square error; In step 203, the Critic network evaluates the quality of the action strategy based on the value function, and the Actor network updates the prediction network based on the stochastic gradient descent method.
5. The unmanned submersible path tracking method based on DDPG-DKNN according to claim 3 is characterized in that: In step 200, K samples are K samples closest to the query point selected based on the Euclidean distance. A sample point s in the buffer zone m and query point s t The Euclidean distance between them is p(s m ,s t ): Among them, m represents the mth sample, i represents five vectors in the state space, namely, navigation speed, heading angle, longitude, latitude and current water depth, s mi and ti The i-th vector corresponding to a sample and query point respectively.
6. The unmanned submersible path tracking method based on DDPG-DKNN according to claim 3, characterized in that: In step 201, the calculation formula of K value is: K = log(c·M+b) Among them, c and b represent adjustment factors.
7. The unmanned submersible path tracking method based on DDPG-DKNN according to claim 3, characterized in that: In step 202, the calculation formula of the loss function L is: Among them, N is the number of samples in the small batch sampling, Q'(s t+1 ,μ'(s t+1 ) is the target network in state s t+1 Execute action strategy μ'(s t+1 ) is the value function obtained, and rt represents the reward value obtained at time t.
8. The unmanned submersible path tracking method based on DDPG-DKNN according to claim 1, characterized in that: In step 203, the formula for updating the prediction network of the Actor network based on the stochastic gradient descent method is: Among them, θ μ Represents the parameters of the Actor network, To predict the Actor network θ in the network μ The gradient of represents the gradient obtained by the target network when executing action a in state s, and a=μ(s) represents the action value selected from the strategy μ(s).