A UAV dynamic target tracking control method based on deep reinforcement learning
Through deep reinforcement learning and SAC algorithms, end-to-end controllers are designed to solve the problem of drones' difficult response in dynamic target tracking, and efficient and robust target tracking control is achieved.
Patent Information
- Application Number
- CN202211404733.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-11-10
AI Technical Summary
Drones lack prior knowledge of target motion patterns during dynamic target tracking, resulting in uncertain changes in the difficulty in responding to targets accurately and quickly.
Using a drone dynamic target tracking and control method based on deep reinforcement learning, a deep neural network is extracted using SAC algorithm and multiple features, an end-to-end integrated controller is designed, and the speed control instructions are directly output through image information to realize a perception-control closed loop.
This method simplifies the dynamic target tracking process of drone, has strong robustness, fast real-time response speed and strong adaptability to different target motion modes.
Smart Images

Figure CN115686065B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicle (UAV) flight control, and particularly relates to a method for dynamically tracking and controlling a target of a UAV based on deep reinforcement learning. Background Art
[0002] With the continuous expansion of UAV application scenarios and the continuous improvement of its autonomous intelligence level, UAVs are applied in various fields. In the civilian field, UAVs are widely used in aerial photography, security patrol, agricultural inspection, earthquake rescue, field positioning, etc.; in the military field, UAVs are used for aerial reconnaissance and monitoring, and can realize the surveillance, positioning and precise strike of dynamic targets within a region. In the above application scenarios, UAVs are required to track specified targets. Therefore, UAV target tracking has important research significance.
[0003] A UAV is a typical perception-control system, where perception and control serve as the input end and output end of the system respectively. External information is obtained through perception as the input of the system, and after a series of computational processes, a control signal is output to drive the UAV to complete the movement under a specific task scenario. UAV target tracking control, as a flight task, also requires the UAV to achieve a full-process closed loop from perception to control, with strong systematic and multidisciplinary cross characteristics. The motion form of the target often changes continuously, presenting randomness, diversity and complexity, which pose great challenges to the perception and control system of the UAV. Since the UAV lacks prior knowledge of the motion pattern of the target to be tracked, how to ensure that the UAV can accurately and quickly respond to the uncertain changes of the target has become an urgent problem to be solved.
[0004] The combination of reinforcement learning and deep neural networks has produced an interdisciplinary field called deep reinforcement learning. Through a multi-layer network structure and non-linear transformation, the deep neural network combines low-level features to form abstract and easily distinguishable high-level representations, and can discover the distributed feature representations of data, focusing on the perception and expression of things. Therefore, the deep reinforcement learning method has the ability to both perceive complex inputs and make decisions, can well adapt to the end-to-end perception-control system, and has strong versatility.
[0005] Traditional UAV control schemes often need to be realized through a series of modules in series, such as sensor signal processing, mapping, pose estimation, planning, trajectory tracking, and low-level control. This form of being separated layer by layer is very likely to cause the accumulation of delays and errors in each module, and the interface links between each module are often designed manually, which may lead to weak applicability of these interfaces in the face of different task scenarios. Summary of the Invention
[0006] Based on the deficiencies of the prior art, the present invention proposes a method for controlling the dynamic target tracking of an unmanned aerial vehicle (UAV) based on deep reinforcement learning. This method is based on the Soft Actor-Critic (SAC) algorithm and uses an end-to-end integrated controller, which can simplify the dynamic target tracking process of the UAV and has the characteristics of strong robustness, fast real-time response speed, and strong adaptability to different target motion patterns.
[0007] The complete technical solution of the present invention is as follows:
[0008] A method for controlling the dynamic target tracking of an unmanned aerial vehicle (UAV) based on deep reinforcement learning, comprising the following steps:
[0009] Step S1: Establish a UAV target tracking model based on the Markov decision process;
[0010] Step S2: Design a UAV target tracking reward function r
[0011] Design the reward function r according to three factors: the relative distance, relative azimuth, and episode termination condition between the UAV and the target in the horizontal direction;
[0012] Step S3: Construct a deep neural network for multi-feature extraction;
[0013] Step S4: Train the deep neural network based on the SAC algorithm;
[0014] Step S5: Use the deep neural network trained in Step S4 to control the dynamic target tracking of the UAV.
[0015] Preferably, the reward function r in Step S2 is designed as follows:
[0016] S201: Design a relative distance reward function r1
[0017]
[0018] Where, d r is the horizontal distance between the current UAV and the target; d r_last is the horizontal distance between the previous UAV and the target; the approaching step number n approach is cleared when the UAV is moving away from the target and incremented by 1 when the UAV is approaching the target;
[0019] S202: Design a relative azimuth reward function r2
[0020] Determine the actual azimuth angle of the target's current position relative to the UAV. Let a represent the UAV's action direction vector, and a θr represent the actual azimuth direction vector. The angle between a and a θr is θ error , and let θ error be less than a threshold θ thresh;
[0021]
[0022]
[0023] S203: Design the round termination condition reward function r3,
[0024] (1) When the relative distances between the UAV and the target in the x and y directions are respectively greater than the geographical boundaries constrained by the camera's field of view, x lim and y lim respectively, it is determined that the task of this round fails, this round terminates, and a negative reward r out is directly given to the UAV, and the influences of the two rewards r1 and r2 are blocked;
[0025] (2) When the horizontal distance between the UAV and the target is less than the threshold d r_thresh , it is considered that the UAV has successfully completed the task of reaching above the target. This round terminates, and the UAV obtains a positive reward r success weighted by the number of consecutive successful times n success , where the number of successful times n success only counts the number of steps that continuously satisfy the UAV reaching above the target in several consecutive steps, otherwise it is cleared;
[0026] Then,
[0027]
[0028] S304: Combine r1, r2, and r3 to obtain the UAV target tracking reward function r;
[0029]
[0030] where, w1 and w2 represent weight coefficients.
[0031] Preferably, the step S1 includes:
[0032] The camera image and the self - state of the UAV at the next moment only depend on the control instructions generated and executed by the UAV according to the current camera image. Regard the camera image of the UAV as the observable state s t , and the control instruction as the action a t . The alternation between s t and a t within a finite time domain constitutes a sequence of state - action in time series, denoted as the trajectory τ = s0, a0, …, s t-1 , a t-1 , s t ,..., a T-1 , s T, where s0 is the initial state and T is the termination time of the finite time domain.
[0033] Preferably, the multi-feature extraction deep neural network in step S3 includes an Actor network and a Critic network. The Actor network structure has seven layers, including the input layer of the Actor network structure, the first convolutional layer of the Actor network structure, the second convolutional layer of the Actor network structure, the third convolutional layer of the Actor network structure, the spatial exponential normalization layer of the Actor network structure, the first fully connected layer of the Actor network structure, the second fully connected layer of the Actor network structure, and the output layer of the Actor network structure. Taking the preprocessed RGB image of 120×120×3 as the input, the convolutional kernel sizes used in the convolutional layers of the Actor network structure are 7×7, 5×5, and 5×5 respectively, the number of nodes in the fully connected layers are 16 and 8 in sequence. The three convolutional layers of the Actor network structure add the Rule activation function during transmission. The spatial exponential normalization layer determines the positions of the image spatial points with the maximum activation value in each channel of the output feature map of the previous layer network with the help of the exponential normalization function. The two fully connected layers of the Actor network structure add the Leaky Rule activation function during transmission. The output layer of the Actor network structure has two branches, which are respectively used to calculate the mean and logarithmic variance of the generated random action Gaussian distribution; The Critic network structure has seven layers, including the input layer of the Critic network structure, the first convolutional layer of the Critic network structure, the second convolutional layer of the Critic network structure, the third convolutional layer of the Critic network structure, the spatial exponential normalization layer of the Critic network structure, the first fully connected layer of the Critic network structure, the second fully connected layer of the Critic network structure, and the output layer of the Critic network structure. Taking the preprocessed RGB image of 120×120×3 as the input, the convolutional kernel sizes used in the convolutional layers of the Critic network structure are 7×7, 5×5, and 5×5 respectively, the number of nodes in the fully connected layers are 16 and 8 in sequence. The three convolutional layers of the Critic network structure add the Rule activation function during transmission. The spatial exponential normalization layer determines the positions of the image spatial points with the maximum activation value in each channel of the output feature map of the previous layer network with the help of the exponential normalization function. The two fully connected layers of the Critic network structure add the Leaky Rule activation function during transmission. The output layer of the Critic network structure outputs the Q value and the V value for the Q value network and the V value network respectively.
[0034] Preferably, step S4 specifically includes:
[0035] S401: Initialize the starting position of the drone in the simulation environment and reset the parameters in the interaction environment;
[0036] S402: After obtaining the UAV camera image, send it as the state to the agent.
[0037] S403: Use the current state as the input of the Actor network, and generate an action after action sampling.
[0038] S404: After normalizing the action, convert it into a speed control instruction and send it to the UAV model in the simulation environment.
[0039] S405: Through simulation, the UAV executes the speed control instruction, and its various physical state quantities are updated; calculate the reward at the current moment through the reward function, and at the same time update the UAV camera image to obtain a new state.
[0040] S406: Send the current state, current action, new state, and current moment reward to the agent.
[0041] S407: Store the current state, current action, new state, and current moment reward as an experience data into the experience pool.
[0042] S408: Sample replay experiences from the experience pool and give them to the optimizers of each network as training data.
[0043] S409: Update the policy network, two Q-value networks, and the behavior V-value network once; perform a soft update on the target V-value network at a certain number of steps interval.
[0044] Preferably, the step S5 specifically includes: The Actor network trained through step S4 is directly used as the target tracking controller of the UAV.
[0045] Compared with the prior art, the present invention has the following advantages:
[0046] 1. The UAV dynamic target tracking control method based on deep reinforcement learning proposed by the present invention is an end-to-end control method. Integrate a series of modules such as traditional UAV target tracking sensor signal processing, mapping, pose estimation, planning, trajectory tracking, and low-level control through a neural network, and directly obtain speed control instructions using image information, simplifying the UAV dynamic target tracking process.
[0047] 2. In the UAV dynamic target tracking control method based on deep reinforcement learning proposed by the present invention, the controller trained by the SAC algorithm is a deep neural network, which has good robustness and can track targets under different motion conditions.
[0048] 3. The trained deep neural network of the present invention has high operation efficiency during actual deployment, which is beneficial to improving the real-time response speed of the UAV to dynamic targets and the adaptability to different target motion patterns. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. By referring to the drawings, the features and advantages of the present invention will be more clearly understood. The drawings are schematic and should not be construed as limiting the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 It is a schematic diagram of the Markov decision process for UAV target tracking in the present invention.
[0051] Figure 2a It is a design diagram of the Actor network structure of the present invention;
[0052] Figure 2b It is a design diagram of the Critic network structure of the present invention;
[0053] Figure 3 It is a schematic diagram of the UAV target tracking training framework based on SAC of the present invention;
[0054] Figure 4 It is a schematic diagram of the use of the UAV target tracking neural network after training in the present invention;
[0055] Figure 5 It is a 100-time static target tracking trajectory diagram of Embodiment 1 of the present invention;
[0056] Figure 6a It is a 3D trajectory diagram of the UAV tracking a moving target along a square trajectory in Embodiment 1 of the present invention;
[0057] Figure 6b It is an X-Y plane trajectory diagram of the UAV tracking a moving target along a square trajectory in Embodiment 1 of the present invention;
[0058] Figure 6c It is an x-axis position curve of the UAV tracking a moving target along a square trajectory in Embodiment 1 of the present invention;
[0059] Figure 6d It is a y-axis position curve of the UAV tracking a moving target along a square trajectory in Embodiment 1 of the present invention;
[0060] Figure 7a It is a 3D trajectory diagram of the UAV tracking a moving target along a broken line trajectory in Embodiment 1 of the present invention;
[0061] Figure 7b It is the X-Y plane trajectory diagram of the drone in Embodiment 1 of the present invention tracking a moving target along a broken line trajectory;
[0062] Figure 7c It is the x-axis position curve of the drone in Embodiment 1 of the present invention tracking a moving target along a broken line trajectory;
[0063] Figure 7d It is the y-axis position curve of the drone in Embodiment 1 of the present invention tracking a moving target along a broken line trajectory;
[0064] Figure 8a It is the 3D trajectory diagram of the drone in Embodiment 1 of the present invention tracking a moving target along a lemniscate trajectory;
[0065] Figure 8b It is the X-Y plane trajectory diagram of the drone in Embodiment 1 of the present invention tracking a moving target along a lemniscate trajectory;
[0066] Figure 8c It is the x-axis position curve of the drone in Embodiment 1 of the present invention tracking a moving target along a lemniscate trajectory;
[0067] Figure 8d It is the y-axis position curve of the drone in Embodiment 1 of the present invention tracking a moving target along a lemniscate trajectory. Detailed implementation manners
[0068] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0069] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0070] In order to solve the problems of time delay, error accumulation in discrete modules and weak applicability of module interfaces to different task scenarios in traditional drone target tracking control methods, the present invention proposes a drone dynamic target tracking control method based on deep reinforcement learning, which directly outputs speed control commands from the original image during the interaction between the drone target tracking agent and the simulation environment. Through anti-normalization processing of the network output, the speed control commands can be directly obtained as the input of the subsequent controller, thereby realizing the perception-control closed loop of drone dynamic target tracking. The technical solutions adopted include the following steps:
[0071] Step S1: Design a Markov decision process for drone target tracking;
[0072] Through the analysis of the UAV dynamic target tracking task, it can be seen that the camera image and state of the UAV at the next moment only depend on the control instructions generated and executed according to the current image. If the camera image of the UAV is regarded as the observable state s t , and regard the control instruction as action a t , then the alternation between the two in a finite time domain will constitute a set of state-action sequences in time sequence, recorded as trajectory τ=s0,a0,…,s t-1 ,a t-1 ,s t ,...,a T-1 ,s T , where s0 is the initial state and T is the end time of the finite time domain. Figure 1 The Markov decision process for target tracking control of UAV is shown.
[0073] The Markov decision process of drone target tracking can be described by the tuple {S, A, P, R, γ}, where S is the state space, i.e., the set of states. Considering that the original image of the drone camera is large in size, the image after size compression and pixel value normalization to [0, 1] is defined as the state, so A is the action space, that is, the set of actions. According to the previous analysis, the action is defined as the normalized desired speed control instruction of the drone in the horizontal direction, so A={a=(v cmd_x ,v cmd_y ) T |v cmd_x ,v cmd_y ∈[-1,1]}; P is the state transfer function, which describes the dynamic characteristics of the Markov decision process and can be recorded as P:S×A×S→[0,1]. The meaning of this function is the probability p(s'|s,a) that the drone will obtain image s' after perceiving image s and taking action a; R is the reward function, which is generally a function of the current state and action and can be recorded as It is used to evaluate the quality of the speed control command generated according to the image of the drone's onboard camera at the current moment; γ is the discount factor used to calculate the cumulative reward, γ∈(0,1).
[0074] Step S2: Design a drone target tracking reward function;
[0075] The designed reward function is mainly composed of three items: the first item r1 is related to the relative distance between the drone and the target in the horizontal direction at the current moment, the second item r2 is related to the action direction calculated by the agent at the current moment and the relative position of the drone and the target, and the third item r3 is related to the round termination condition.
[0076] S2-1: Design the relative distance reward function r1;
[0077] The design of r1 aims to encourage the UAV to approach the target and punish the UAV for moving away from the target. Moreover, the closer the UAV is to the target during the approach process, the greater the positive reward value it obtains, and the farther the UAV is from the target during the departure process, the greater the absolute value of the negative reward it obtains. Since the return in one episode is the cumulative reward value over a period of time, when the UAV changes from approaching in the previous step to moving away in the current step, the absolute value of the punishment value in the current step should be greater than the reward in the previous step to cover the reward effect of the previous step; when the UAV approaches the target for several consecutive steps, the number of consecutive approach steps is used for weighting. The calculation formula of r1 is as follows:
[0078]
[0079] where the number of approach steps n approach is cleared when the UAV moves away from the target and accumulates 1 when the UAV approaches the target.
[0080] S2-2: Design the relative azimuth reward function r2;
[0081] The design of r2 aims to reward and punish the action direction. First, calculate the actual azimuth angle of the target relative to the UAV based on the current positions of the UAV and the target. Secondly, use the cosine theorem to calculate the angle θ between the action direction vector a and the actual relative azimuth direction vector error . If the angle between the two is less than a threshold θ thresh , a positive reward inversely proportional to this angle is given; otherwise, this item is set as a negative reward and the greater the angle, the greater the absolute value of the negative reward. In addition, to prevent excessive values caused by θ error approaching 0, it is necessary to limit the amplitude of r2 when it is a positive reward. The calculation formula of r2 is as follows:
[0082]
[0083]
[0084] S2-3: Design the relative azimuth reward function r3;
[0085] When designing r3, mainly consider the judgment of the episode termination conditions. Assume that there are 3 conditions for triggering the episode termination, namely, the failure of the current episode task due to the loss of the target in the field of view, the success of the current episode task due to the UAV moving to a certain area range directly above the target and meeting certain conditions, and reaching the maximum number of steps in the episode. Only set reward functions for the first two conditions. When the relative distances between the UAV and the target in the x and y directions are respectively greater than the geographical boundaries x constrained by the camera field of viewlim , y lim , the task of this round is determined to fail, and a negative reward r with a relatively large absolute value is directly given to the UAV out and the influence of the other two rewards is blocked; when the horizontal distance between the UAV and the target is less than a certain threshold d r_thresh , it is considered that the UAV has successfully completed the task of reaching above the target. At this time, a positive reward r weighted by the consecutive success times n success is added on the basis of the first two rewards success , where the success times n success only counts the number of steps that continuously satisfy the UAV within the threshold d above the target for several steps, otherwise it is cleared r_thresh .
[0086]
[0087] In summary, the reward function designed by the present invention for the deep reinforcement learning problem of UAV target tracking is as follows:
[0088]
[0089] where w1 and w2 are corresponding weight coefficients
[0090] Step S3: Design a targeted deep neural network structure
[0091] Since the state space of the UAV target tracking control problem is a high-dimensional space, in order to maintain a good extraction ability for complex high-dimensional state features, the design idea of a multi-feature extraction network is adopted when designing the policy network. A multi-layer convolutional neural network is selected as the first half of the policy network π(a|s; θ), and a hidden layer with spatial feature extraction function is added before the fully connected layer in the second half to enhance the expression of the position information of the target in the image. Considering the stability and convergence effect of neural network training, it is also necessary to normalize the state input of the agent's policy
[0092] Design as Figure 2a - 2bThe shown multi-feature extraction network structure, as the Actor network and the Critic network, is used for the end-to-end learning of speed commands. The above multi-feature extraction network takes a preprocessed RGB image of 120×120×3 as the input and first passes through 3 convolutional layers in sequence to extract the visual features of the target image. Compared with the conventional fully connected network, in the convolutional neural network, convolutional operations are used to replace the general matrix multiplication operation. The convolutional kernel traverses the input layer in a sliding manner and obtains the result after the convolutional operation through weighted summation in the local area. Usually, the size of the convolutional kernel is smaller than the input size, which makes the convolutional neural network have the characteristics of sparse interaction and parameter sharing, and thus can extract features with fewer parameters and higher efficiency. Here, convolutional kernels of sizes 7×7 and 5×5 are used for operations respectively, and the common ReLU is selected as the activation function for each layer. Since the target logo is a black-and-white image, the VALID mode is adopted during convolutional operation padding to avoid the interference of zero-padding operation in the SAME mode on the target feature extraction.
[0093] After the convolutional layer and before the fully connected layer, a spatial exponential normalization layer is added to extract the spatial features of the target, corresponding to the "SS" (Spatial Softmax) layer in Figure 2a and Figure 2b The spatial exponential normalization layer determines the positions of the image spatial points with the maximum activation values in each channel of the output feature map of the previous layer network by means of the exponential normalization function, specifically including 2 calculation processes:
[0094] (1) Let the activation value at the (i, j) position on the c-th channel in the output feature map of the previous layer network be a cij , then the spatial exponential normalization value s cij is calculated by the following formula:
[0095]
[0096] where α is the parameter to be trained.
[0097] (2) The expectation of the 2D coordinates of the Softmax probability distribution of each channel (i.e., the average value of the spatial exponential normalization values) is given by the following formula:
[0098]
[0099] It describes the positions of the image spatial points with the maximum activation values in each channel.
[0100] The Actor network has two branches in the final output layer, which are used to calculate the mean and log variance of the Gaussian distribution of the generated random actions respectively. After the backbone network of the Critic network, for the Q-value network, the upper-layer feature vector needs to be concatenated with the current action vector and used as the input of the subsequent fully connected layer. The final output is a scalar, that is, the estimated Q-value. For the V-value network, there is no need for the action vector as an additional input. The feature vector output by the spatial exponential normalization layer can be directly used as the input of the subsequent fully connected layer, and the final output is the V-value.
[0101] Step S4: Training of the speed command perception controller based on the SAC algorithm;
[0102] The framework of end-to-end learning of speed command perception based on SAC is as Figure 3 shown, where the left side corresponds to the modules in the designed interaction environment, and the right side corresponds to the deployment of the agent in the training mode.
[0103] The training process of the SAC agent is as follows: In each episode,
[0104] (1) Initialize the starting position of the drone in the simulation environment and reset the parameters in the interaction environment;
[0105] (2) After obtaining the drone camera image, send it as the state to the agent;
[0106] (3) Use the current state as the input of the Actor network, and generate an action after action sampling;
[0107] (4) After normalizing the action, convert it into a speed control command and send it to the drone model in the simulation environment;
[0108] (5) Through simulation, the drone executes the speed control command, and its various physical state quantities are updated; and calculate the reward at the current moment through the reward function, and at the same time update the drone camera image to obtain a new state;
[0109] (6) Send the current state, the current action, the new state, and the reward at the current moment to the agent;
[0110] (7) Store the current state, the current action, the new state, and the reward at the current moment as an experience data into the experience pool;
[0111] (8) Sample replay experiences from the experience pool and give them to the optimizers of each network as training data;
[0112] (9) Update the policy network, the two Q-value networks, and the behavior V-value network once; Soft update the target V-value network once at a certain number of steps interval.
[0113] Step S5: Use the UAV dynamic target tracking controller;
[0114] As Figure 4 shown, in practical applications, the Actor network structure obtained through training in Step S4 is directly used as the target tracking controller of the UAV. The input is the image information after normalization processing, and the output is the UAV speed control amount after normalization processing. By performing inverse normalization processing on the network output, a speed control command can be generated for interaction with the environment.
[0115] To facilitate the understanding of the above technical solution of the present invention, the above technical solution of the present invention will be described in detail through the following specific examples.
[0116] Embodiment 1
[0117] Assume that there is a moving target with a changing movement speed within the reconnaissance range of the UAV. Through the tracking control method of the present invention, a UAV dynamic target tracking controller is designed and trained so that it can stably track the target. The entire controller design and verification process are completed in a simulation environment.
[0118] Set the parameters of the training algorithm as follows: the size of the experience pool is 10,000, the size of the experience replay batch is 128, the discount factor is 0.99, the maximum number of steps per episode is 50, and the learning rate of each network is 0.0003.
[0119] Apply the controller trained by the present invention to the following four different scenarios to verify the tracking effect of the trained controller and verify the effectiveness of the method of the present invention.
[0120] A. Tracking a target in a stationary state
[0121] Initially, the UAV is at an arbitrary position in space, and the target remains stationary and is located at the center position of space. Use the end-to-end controller to make the UAV track the stationary target. Conduct 100 random starting point tests. The UAV trajectory during the test is as Figure 5 shown. It can be seen from the figure that the UAV can fly above the fixed target from any starting position, and the task success rate is 100%, indicating that the obtained policy network is feasible for completing the equivalent task of UAV dynamic target tracking. Subsequent simulation tests will directly use this policy network to test tracking dynamic targets.
[0122] B. Tracking a target moving along a square trajectory
[0123] Starting from the target at (0, 0, 0) m, it moves along a square trajectory with a side length of 8 m. It moves at a constant linear speed of 0.5 m / s on the sides of the square. When it reaches the vertices of the square, the direction of its speed changes by 90°, and then it continues to move in a straight line in the new direction. The UAV hovers directly above the target at the starting moment, that is, at (0, 0, 5) m. The results of the simulation experiment are as Figure 6a - 6d shown. The solid line in the figure is the trajectory of the UAV, and the dashed line is the trajectory of the target. It can be seen from the figure that the UAV can stably track the target along the square trajectory.
[0124] C. Tracking a target moving along a broken-line trajectory
[0125] The target starts from (0, 0, 0) m and moves along a broken-line trajectory. Specifically, it first moves along the straight line y = x for a certain distance, then the direction of its speed undergoes a sudden change of -135° and it moves along the negative y-axis for a certain distance, and then the direction of its speed undergoes a sudden change of +135° again, and so on. After repeating several times, the target stops moving. The resultant speed of the target is set to 0.5 m / s during the movement process. The UAV hovers directly above the target at the starting moment of the target movement, that is, at (0, 0, 5) m. The results of the simulation experiment are as Figure 7a - 7d shown. The solid line in the figure is the trajectory of the UAV, and the dashed line is the trajectory of the target. It can be seen from the figure that the UAV can stably track the target along the broken-line trajectory.
[0126] D. Tracking a target moving along a lemniscate trajectory
[0127] It is set that the target starts from (0, 0, 0) m and moves along a lemniscate trajectory. The UAV hovers directly above the target at the starting moment of the target movement. The movement speed of the target is slower at the places where the curvature of the lemniscate is smaller and faster at the places where the curvature is larger, and generally varies within the range of 0.5 m / s to 1.2 m / s. The results of the simulation experiment are as Figure 8a - 8d shown. The solid line in the figure is the trajectory of the UAV, and the dashed line is the trajectory of the target. It can be seen from the figure that the UAV can stably track the target along the lemniscate trajectory.
[0128] In the present invention, unless otherwise clearly specified and defined, the terms "installation", "connection", "connection", "fixation" and other terms should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral body; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.
[0129] In the present invention, unless otherwise clearly specified or limited, the first feature being “on” or “under” the second feature may include direct contact between the first and second features, or may also include indirect contact between the first and second features through additional features therebetween. Moreover, the first feature being “above”, “over” and “on top of” the second feature includes the first feature being directly above and obliquely above the second feature, or merely indicating that the horizontal height of the first feature is higher than that of the second feature. The first feature being “under”, “beneath” and “underneath” the second feature includes the first feature being directly below and obliquely below the second feature, or merely indicating that the horizontal height of the first feature is lower than that of the second feature.
[0130] In the present invention, the terms “first”, “second”, “third” and “fourth” are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. The term “plurality” means two or more, unless otherwise clearly defined.
[0131] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A dynamic target tracking control method for unmanned aerial vehicles based on deep reinforcement learning, characterized in that, It includes the following steps: Step S1: Establish a UAV target tracking model based on the Markov decision process; Step S2: Design a UAV target tracking reward function r, and design the reward function r according to three factors: the relative distance, relative azimuth, and episode termination condition between the UAV and the target in the horizontal direction; Step S3: Construct a multi-feature extraction deep neural network; Step S4: Train the deep neural network based on the SAC algorithm; Step S5: Use the deep neural network trained in Step S4 to control the dynamic target tracking of the UAV; The reward function r in Step S2 is designed as follows: S201: Design the relative distance reward function r1 Among them, d r is the horizontal distance between the current UAV and the target; d r_last is the horizontal distance between the UAV and the target in the previous step; the approaching step number n approach is cleared when the UAV moves away from the target and increments by 1 when the UAV approaches the target; S202: Design the relative azimuth reward function r2 Determine the actual azimuth angle of the target's current position relative to the drone. Let a represent the drone's action direction vector, represent the actual azimuth direction vector. The included angle between a and is θ error . Let θ error be less than a threshold value θ thresh ; S203: Design the episode termination condition reward function r3, (1) When the relative distances between the drone and the target in the x and y directions are respectively greater than the geographical boundaries x lim and y lim at this time, it is determined that the mission of this round fails, this round terminates and a negative reward r is directly assigned to the drone out and the influence of the two rewards r1 and r2 is blocked; (2) When the horizontal distance between the UAV and the target is less than the threshold d r_thresh When the drone successfully completes the task of reaching the target, the round ends and the drone obtains the number of consecutive successes n success Weighted positive reward r success , where the number of successes n success Only the number of steps that the drone reaches the target in a certain number of consecutive steps is counted, otherwise it is reset to zero; Then, S204: Combine r1, r2, and r3 to obtain the UAV target tracking reward function r; where w1 and w2 represent weight coefficients.
2. The method for controlling dynamic target tracking of an unmanned aerial vehicle according to claim 1, wherein Step S1 includes: The camera image and its own state of the UAV at the next moment only depend on the control instructions generated and executed by the UAV based on the current camera image. The camera image of the UAV is regarded as the observable state s t , and the control instruction is regarded as the action a t . Within a finite time domain, the alternation between s t and a t constitutes a sequence of state-action in time series, denoted as the trajectory τ = s0, a0, L, s t-1 , a t-1 , s t ,..., a T-1 , s T , where s0 is the initial state and T is the termination moment of the finite time domain.
3. The drone dynamic target tracking control method according to claim 1, wherein The multi-feature extraction deep neural network in Step S3 includes an Actor network and a Critic network.
4. The method for controlling dynamic target tracking of an unmanned aerial vehicle according to claim 3, wherein, The Actor network structure has eight layers, including the input layer of the Actor network structure, the first convolutional layer of the Actor network structure, the second convolutional layer of the Actor network structure, the third convolutional layer of the Actor network structure, the spatial exponential normalization layer of the Actor network structure, the first fully connected layer of the Actor network structure, the second fully connected layer of the Actor network structure, and the output layer of the Actor network structure; Taking the preprocessed RGB image of 120×120×3 as the input, the sizes of the convolutional kernels adopted by the convolutional layers of the Actor network structure are 7×7, 5×5, and 5×5 in sequence, the number of nodes in the fully connected layers are 16 and 8 in sequence. The ReLU activation function is added during the transfer of the three convolutional layers of the Actor network structure. The spatial exponential normalization layer determines the positions of the image spatial points with the maximum activation values in each channel of the output feature map of the previous layer network by means of the exponential normalization function. The Leaky ReLU activation function is added during the transfer of the two fully connected layers of the Actor network structure.
5. The method for dynamically tracking and controlling a target of an unmanned aerial vehicle according to claim 3, wherein, The output layer of the Actor network structure has two branches, which are respectively used to calculate the mean and logarithmic variance of the Gaussian distribution of the generated random actions; the Critic network structure has eight layers, including the input layer of the Critic network structure, the first convolutional layer of the Critic network structure, the second convolutional layer of the Critic network structure, the third convolutional layer of the Critic network structure, the spatial exponential normalization layer of the Critic network structure, the first fully connected layer of the Critic network structure, the second fully connected layer of the Critic network structure, and the output layer of the Critic network structure; using a preprocessed RGB image of 120×120×3 as the input, the convolutional kernel sizes adopted by the convolutional layers of the Critic network structure are 7×7, 5×5, and 5×5 in sequence, the number of nodes in the fully connected layers are 16 and 8 in sequence, ReLU activation functions are added during the transfer of the three convolutional layers of the Critic network structure, the spatial exponential normalization layer determines the positions of the image spatial points with the maximum activation values in each channel of the output feature map of the previous layer network by means of the exponential normalization function, Leaky ReLU activation functions are added during the transfer of the two fully connected layers of the Critic network structure, and the output layer of the Critic network structure outputs Q values and V values for the Q-value network and the V-value network respectively.
6. The method for dynamically tracking and controlling a target of a drone according to claim 1, characterized in that The specific steps of step S4 include: S401: Initialize the starting position of the drone in the simulation environment and reset each parameter in the interaction environment; S402: After obtaining the drone camera image, send it as the state to the agent; S403: Use the current state as the input of the Actor network, and generate an action after action sampling; S405: Convert the action into a speed control command after normalization processing and send it to the drone model in the simulation environment; S406: Simulate the drone to execute the speed control command through simulation, and its various physical state quantities are updated; calculate the reward at the current moment through the reward function, and at the same time update the drone camera image to obtain a new state; S407: Send the current state, current action, new state, and current moment reward to the agent; S408: Store the current state, current action, new state, and current moment reward as an experience data in the experience pool; S409: Sample the replay experience from the experience pool and give it to the optimizers of each network as training data; S409: Update the policy network, two Q-value networks, and the behavior V-value network once; perform a soft update on the target V-value network at a certain number of steps interval.
7. The method for controlling dynamic target tracking of an unmanned aerial vehicle according to claim 1, wherein The specific steps of step S5 include: The Actor network trained through step S4 is directly used as the target tracking controller of the drone.
Citation Information
Patent Citations
Unmanned aerial vehicle end-to-end control method based on deep reinforcement learning
CN111460650A
Unmanned surface vehicle path tracking method based on deep reinforcement learning
CN115016496A