A dynamic target tracking method for unmanned boats

Through the DDPG reinforcement learning algorithm and the three-degree-of-freedom parallel mechanism motion platform, the dynamic target tracking method of unmanned boats is optimized, which solves the problems of low accuracy and high training complexity in complex environments of traditional methods, and achieves efficient and accurate dynamic target tracking, which is suitable for unmanned boats, drones and mobile robots.

CN117058547BActive Publication Date: 2025-08-22SHANGHAI UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311046258.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-18
Publication Date
2025-08-22
Estimated Expiration
2043-08-18

AI Technical Summary

Technical Problem

Traditional unmanned boat dynamic target tracking methods are susceptible to light changes, background interference and rapid changes in target movement in complex environments, resulting in low tracking accuracy, high training complexity, large energy consumption, and difficult to design reward function.

Method used

The dynamic target tracking network model using DDPG reinforcement learning algorithm is used to obtain environmental images through the camera gimbal, identify dynamic targets, build state amounts and calculate the gimbal motor action. Combined with priority experience playback and noise addition, the reward function is optimized, and the three-degree of freedom parallel mechanism motion platform is used to improve tracking accuracy and efficiency.

Benefits of technology

It improves the accuracy and robustness of dynamic target tracking of unmanned boats in complex environments, reduces training costs and energy consumption, adapts to various complex scenarios, and has a wide range of application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058547B_ABST
    Figure CN117058547B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer vision technology, and discloses a method for tracking dynamic targets in an unmanned boat. The method comprises the following steps: 1) using a camera carried by the unmanned boat to acquire an environmental image and find the image of a dynamic target in the environmental image; 2) determining the center coordinates of the environmental image and the dynamic target image, and constructing a state quantity based on the center coordinates of the environmental image and the dynamic target image; 3) inputting the state quantity into a dynamic target tracking network model to calculate the rotational attitude angle information of the camera gimbal; 4) calculating the movement of the gimbal motor based on the rotational attitude angle information, and controlling the gimbal motor to move based on the movement of the gimbal motor; and 5) repeating the above steps 1-4) to enable the camera to follow the movement of the dynamic target and complete the tracking of the dynamic target. The dynamic target tracking algorithm of the present invention can adaptively, accurately, and stably track various complex targets, and has broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and in particular to a method for tracking dynamic targets of unmanned boats. Background Art

[0002] Unmanned vessels (UVs) are maritime robotic systems and key nodes in networked marine unmanned systems. Equipped with various sensor equipment, they are widely used in marine transportation and environmental surveys, marine archaeology, underwater search and rescue, intelligence gathering, reconnaissance and evidence collection, surveillance and patrol, fire strikes, ship escort, mine countermeasures, and anti-submarine warfare. UUVs are a system that revolutionizes traditional naval warfare, spawning a new marine equipment system and holding significant significance for the development of marine resources and the protection of national maritime rights and interests. Consequently, they are highly valued by maritime powers such as the United States and Europe. Dynamic target tracking technology for UUVs is a key research area within their application. With the continuous advancement of UUV technology, an increasing number of UUVs are being used for aerial photography, monitoring, detection, search and rescue, and other missions. In these missions, UUVs are equipped with cameras to capture real-time image or video information, making accurate tracking of dynamic targets crucial.

[0003] Traditional methods for tracking dynamic targets on unmanned aerial vehicles (UAVs) primarily rely on rule-based control strategies or tracking algorithms based on visual features. However, these methods often perform poorly in complex environments and are susceptible to factors such as lighting variations, background interference, and rapid changes in target motion, resulting in low tracking accuracy or even failure.

[0004] To overcome the limitations of traditional methods, reinforcement learning has gained widespread attention in the field of dynamic target tracking in recent years. Reinforcement learning is a machine learning method that uses an intelligent agent to interactively learn from its environment to achieve a specific goal. In the case of unmanned vehicle dynamic target tracking, the agent can optimize its tracking strategy by continuously trying different motion adjustments, observing environmental feedback, and using the resulting rewards.

[0005] Reinforcement learning-based dynamic target tracking for unmanned underwater vehicles (UAVs) offers a promising solution due to its ability to adaptively adjust its strategy based on the environment and target state. By interactively learning with the environment, the agent can gradually improve the accuracy and robustness of its camera-based dynamic target tracking, enabling it to more reliably complete tasks in complex scenarios.

[0006] However, in practical applications, reinforcement learning-based methods for tracking dynamic targets with cameras on unmanned vehicles (UAVs) still face several challenges. These include: 1) Tracking accuracy: Current UAV dynamic target tracking methods primarily focus on the overall motion control of the UAV, without specifically addressing the camera gimbal. This directly leads to insufficient accuracy and precision during dynamic tracking. 2) Training complexity: Reinforcement learning algorithms typically require extensive interactive training. In the case of UAV dynamic target tracking, each interaction consumes time and resources, so reducing training complexity is a critical issue. 3) Reward function design: Designing an effective reward function is a key challenge in reinforcement learning tasks. The reward function must accurately reflect the performance of camera dynamic target tracking and possess good continuity and learnability. 4) Energy consumption: In UAV applications, energy is a precious resource. How to rationally plan the camera movement frequency and amplitude, as well as the UAV's navigation path, to reduce energy consumption while ensuring tracking effectiveness remains a challenge. Summary of the Invention

[0007] In view of the problems and shortcomings in the prior art, the present invention aims to provide a method for dynamic target tracking of an unmanned boat.

[0008] To achieve the purpose of the invention, the technical solution adopted by the present invention is as follows:

[0009] A first aspect of the present invention provides a training method for an unmanned vehicle dynamic target tracking network model, comprising the following steps:

[0010] S1: Use the camera on the camera gimbal of the unmanned boat to obtain the current environment image, identify and analyze the environment image, and find the image of the dynamic target that needs to be tracked in the environment image.

[0011] S2: Determine whether the size of the dynamic target image is within a preset image threshold range. If the size of the dynamic target image is not within the preset image threshold range, the unmanned boat moves and then executes step S1; if the size of the dynamic target image is within the preset image threshold range, executes step S3;

[0012] S3: Determine the center coordinates of the environment image and the dynamic target image, and construct the current state quantity s based on the center coordinates of the environment image and the dynamic target image. t ;

[0013] S4: The state quantity s t Input the dynamic target tracking network model to calculate the rotation angle information of the camera gimbal; calculate the action a of the gimbal motor that controls the movement of the camera gimbal according to the rotation angle information of the camera gimbal. t , according to the calculated motion a of the pan / tilt motort Control the pan / tilt motor to move; at the same time, based on the state quantity s t Constructing the reward function r of the dynamic target tracking network model t (s t ); wherein the dynamic target tracking network model is a DDPG (Deep Deterministic Policy Gradient) reinforcement learning algorithm neural network;

[0014] S5: Repeat steps S1-S3 to obtain the state quantity s at the next moment t ′, the current state quantity s t 、Action a t , reward r t , the state quantity s at the next moment t ′ as a set of empirical data (s t ,a t ,r t ,s t ′);

[0015] S6: Repeat steps S1-S5 to obtain multiple sets of experience data, and store all the experience data in the experience pool; then determine whether the number of repetitions of steps S1-S5 is equal to the number of network parameter update steps M of the dynamic target tracking network model; if the number of repetitions of steps S1-S5 is less than M, execute step S1; if the number of repetitions of steps S1-S5 is equal to M, execute step S7;

[0016] S7: using a priority experience replay algorithm to filter out n sets of experience data from the experience pool, and using the filtered n sets of experience data to sequentially update the network parameters of the dynamic target tracking network model;

[0017] S8: Repeat steps S1-S7 and count the number of repetitions of steps S1-S7. If the number of repetitions of steps S1-S7 is greater than or equal to a preset number of cycles N, the loop ends, the dynamic target tracking network model training is completed, and the optimal network parameters of the dynamic target tracking network model are obtained; the preset number of cycles N is greater than the number of repetitions of steps S1-S7 when the dynamic target tracking network model converges;

[0018] S9: Using the optimal network parameters obtained in step S8 to update the network parameters of the dynamic target tracking network model, to obtain a trained dynamic target tracking network model.

[0019] According to the above-mentioned training method of the unmanned boat dynamic target tracking network model, preferably, the state quantity s t =(x f ,y f , x c ,yc ), the reward function r t (s t ) is calculated as follows:

[0020]

[0021] Among them, (x f ,y f ) is the center coordinate of the dynamic target image; (x c ,y c ) is the center coordinate of the environment image; R max is the maximum reward value returned when the center coordinates of the dynamic target and the center coordinates of the environment image coincide; k, α>0 is the correction coefficient; R max , α and k are all adjustable parameters, T is the pan / tilt motor receiving action a t After completing the action a t total time.

[0022] According to the above-mentioned training method of the unmanned boat dynamic target tracking network model, preferably, the dynamic target tracking network model is composed of an Actor network (strategy network), a Critic network (state value function estimation network), a TargetActor network (target strategy network) and a Target Critic network (target state value function estimation network); the Actor network has the same network architecture as the Target Actor network, and the Critic network has the same network architecture as the Target Critic network.

[0023] According to the training method of the unmanned boat dynamic target tracking network model, preferably, the input of the Actor network is the state quantity s t , the output is the attitude angle information of the gimbal camera; the input of the Target Actor network is the state quantity s at the next moment t ′, the output is the attitude angle information of the gimbal camera at the next moment; the input of the Critic network is the state-action pair (s t , a t ), the output is a state-action pair (s t , a t ) state value evaluation value; the input of the TargetCritic network is the state action pair (s t ′,a t ′), the output is the state-action pair (s t ′,a t ′)’s state value target value.

[0024] According to the training method of the unmanned boat dynamic target tracking network model, preferably, the Actor network consists of an input layer, two linear hidden layers and an output layer; wherein the input of the input layer is the state quantity s t , state quantity s t The first and second linear hidden layers are sequentially fully connected and normalized to obtain the strategic hidden layer data. The strategic hidden layer data obtained through the second linear hidden layer is input into the output layer for full connection processing to obtain the attitude angle information of the gimbal camera. More preferably, the input layer of the actor network contains four nodes, and the four nodes of the input layer are connected to the first linear hidden layer via a full connection. The first linear hidden layer contains 64 nodes. After normalization and activation, the nodes in the first linear hidden layer are fully connected to the second linear hidden layer. The second linear hidden layer contains 32 nodes, and the nodes in the second linear hidden layer are fully connected to the output layer.

[0025] According to the training method of the unmanned boat dynamic target tracking network model, preferably, the Critic network includes two branches, the first branch includes a state input layer and two linear hidden layers, the input of the state input layer is the state quantity s t , state quantity s t The two linear hidden layers of the first branch are fully connected and normalized to obtain the state evaluation hidden layer data. The second branch includes an action input layer and a linear hidden layer. The input of the action input layer is the action a of the pan / tilt motor. t , the action of the pan / tilt motor a t The linear hidden layer of the second branch is fully connected and normalized to obtain action evaluation hidden layer data. The state evaluation hidden layer data obtained by the first branch and the action evaluation hidden layer data obtained by the second branch are normalized and added to obtain the state value evaluation value and output it. More preferably, in the first branch of the Critic network, the state input layer contains four nodes, and the four nodes of the state input layer are connected to the first linear hidden layer via a full connection. The first linear hidden layer contains 64 nodes, and the nodes of the first linear hidden layer are fully connected to the second linear hidden layer after normalization activation; in the second branch of the Critic network, the action input layer contains three nodes, and the action input layer is connected to the linear hidden layer via a full connection.

[0026] According to the above-mentioned training method of the unmanned boat dynamic target tracking network model, preferably, in step S4, in order to enhance the exploration ability of the dynamic target tracking network model, during the training of the dynamic target tracking network model, the action a of the pan-tilt motor is calculated. t When noise is added, the action of the pan / tilt motor a t The calculation formula is as follows:

[0027] a t=μ(s t |θ μ )+N t

[0028]

[0029] Where s t Represents the current state; θ μ represents the Actor network parameters; μ(·) represents the Actor network; N t represents the random noise parameter; θ is the parameter, θ>0; σ is the parameter, σ>0; B τ is the standard Brownian motion; τ is the standard time; t is the actual time.

[0030] According to the above-mentioned training method of the unmanned watercraft dynamic target tracking network model, preferably, in step S7, the formula for selecting the experience data from the experience pool using the priority experience replay algorithm is as follows:

[0031]

[0032] Where, P(c) represents a set of experience data (s t , a t , r t , s t ′); c represents the sequence number of the extracted empirical data, p k represents the priority of the extracted empirical data, and α represents the preset parameter used to adjust the priority sampling degree of data samples.

[0033] According to the training method of the unmanned boat dynamic target tracking network model, preferably, the action a of the pan-tilt motor that controls the movement of the camera pan-tilt is calculated according to the rotation attitude angle information of the camera pan-tilt t The method is to inversely solve the rotation angle information of the camera gimbal to obtain the action a of the gimbal motor. t .

[0034] According to the above-mentioned training method for the unmanned vehicle dynamic target tracking network model, preferably, in order to improve training efficiency and reduce training costs, in step S1, an unmanned vehicle simulation environment model can be first established based on the actual environment, and the unmanned vehicle simulation environment model can be used to obtain the current simulation environment image. The unmanned vehicle simulation environment model includes an unmanned vehicle model, an unmanned vehicle motion model, a camera model, a camera gimbal model, a dynamic target model, and a sea level environment model.

[0035] According to the training method of the unmanned boat dynamic target tracking network model, preferably, the camera gimbal is a three-degree-of-freedom parallel mechanism motion platform, which can effectively correspond to more complex marine environment dynamic tracking. More preferably, when the camera gimbal is a three-degree-of-freedom parallel mechanism motion platform, the rotation attitude angle information of the camera gimbal includes the up and down pitch angles. Horizontal rotation angle and axial rotation angle The pan / tilt motor action a t (θ1, θ2, θ3), θ1, θ2, θ3 are the pitch, horizontal, and spin motor angles.

[0036] According to the above-mentioned training method of the unmanned boat dynamic target tracking network model, preferably, in step S9, the optimal network parameters obtained in step S8 are used to update the network parameters of the dynamic target tracking network model. The specific process is:

[0037] (1) Update the network parameters of the Actor network:

[0038] Use the Actor network to calculate the state s t The pan / tilt motor action a t , PTZ motor action a t The calculation formula is as follows:

[0039] a t =μ(s t |θ μ )

[0040] Among them, θ μ is the network parameter of the Actor network, μ(·) represents the Actor network;

[0041] Use the Critic network to calculate the state action pair (s t , a t )’s state value evaluation value q t , state value evaluation value q t The calculation formula is as follows:

[0042] q t =Q(s t , a t |θ Q )

[0043] Among them, θ Q is the network parameter of the Critic network, Q(·) represents the Critic network;

[0044] Use the gradient ascent algorithm to maximize the state value evaluation value q t , thus adjusting the parameters θ in the Actor networkμ Make updates;

[0045] (2) Update of network parameters of the Critic network:

[0046] Use the Target Actor network to calculate the next state s t ′The pan / tilt motor moves a t ′, PTZ motor action a t The calculation formula of ′ is as follows:

[0047] a t ′=μ′(s t ′|θ μ )

[0048] Among them, θ μ′ is the network parameter of the Target Actor network, μ′(·) represents the Target Actor network;

[0049] Use Target Critic network to calculate the state action pair (s t , a t )’s state value target value y t , state value target value y t The calculation formula is as follows:

[0050] y t =r t +γ(1-done)Q′(s t ′,a t ′|θ Q )

[0051] Among them, r t is the reward function, γ is the discounted return rate, done indicates whether the authority correction is completed, θ Q′ is the network parameter of the TargetCritic network, Q′(·) represents the Target Critic network

[0052] Use the gradient descent algorithm to minimize the state value evaluation value q t and state value target value y t The difference between L c , thereby updating the parameters in the Critic network, the difference L c The calculation formula is as follows:

[0053] L c =(y t -q t ) 2

[0054] (3) Network parameter update of Target Actor network and Target Critic network

[0055] To update the target network, the DDPG reinforcement learning algorithm uses a soft update method, also known as exponential moving average. That is, a learning rate (or momentum) τ is introduced to take the weighted average of the old target network parameters and the new corresponding network parameters, and then assign them to the target network.

[0056] Among them, the formula for the network parameter update process of the Target Actor network is as follows:

[0057] θ μ′ =τθ μ +(1-τ)θ μ′

[0058] The formula for the Target Critic network update process is as follows:

[0059] θQ ′ =τθ Q +(1-τ)θQ ′

[0060] θ μ is the network parameter of the Actor network; θ μ′ is the network parameter of the Target Actor network; θ Q is the network parameter of the Critic network; θ Q′ is the network parameter of the Target Critic network; τ is the learning rate (momentum), τ∈(0, 1), usually taking the value of 0.005.

[0061] A second aspect of the present invention provides a method for tracking a dynamic target on an unmanned vehicle, comprising the following steps:

[0062] S11: Using the camera on the unmanned boat to obtain an environmental image, perform recognition and analysis on the environmental image, and find the image of the dynamic target in the environmental image;

[0063] S22: Determine whether the size of the dynamic target image is within a preset image threshold range. If the size of the dynamic target image is not within the preset image threshold range, the unmanned boat moves and then executes step S11; if the size of the dynamic target image is within the preset image threshold range, executes step S33;

[0064] S33: Determine the center coordinates of the environment image and the dynamic target image, and construct a state quantity according to the center coordinates of the environment image and the dynamic target image;

[0065] S44: Input the state quantity into the dynamic target tracking network model to calculate the rotation attitude angle information of the camera gimbal; calculate the action of the gimbal motor that controls the movement of the camera gimbal according to the rotation attitude angle information of the camera gimbal, and control the gimbal motor to move according to the calculated action of the gimbal motor; the dynamic target tracking network model is the trained dynamic target tracking network model described in the first aspect above.

[0066] S55: Repeat steps S11-S44 to enable the camera to follow the dynamic target and complete the tracking of the dynamic target.

[0067] The third aspect of the present invention provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the training method for the unmanned boat dynamic target tracking network model as described in the first aspect above, or the unmanned boat dynamic target tracking method as described in the second aspect above.

[0068] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and is characterized in that when the computer program is executed by a processor, it implements the training method of the unmanned boat dynamic target tracking network model as described in the first aspect above, or the unmanned boat dynamic target tracking method as described in the second aspect above.

[0069] Compared with the prior art, the technical solution of the present invention has the following positive and beneficial effects:

[0070] (1) The dynamic target tracking network model of the present invention is a DDPG reinforcement learning algorithm neural network. It adopts the DDPG reinforcement learning algorithm and has a high degree of self-learning ability. It can be applied in practice after high-level training in a virtual environment, greatly reducing training costs and improving training efficiency.

[0071] (2) The goal of the dynamic target tracking network model training of the present invention is to make the camera lens center approach the dynamic target center. Therefore, based on the principle that the farther the distance between the camera lens center of the unmanned boat and the dynamic target center on the image, the lower the reward, the stability of the reward parameters is ensured by setting the correction coefficient. At the same time, in order to ensure the high efficiency of the camera movement, the total reward is deducted from the corrected time to design a reward function that can be flexibly adjusted according to the actual situation. The reward function is used to control the camera gimbal of the unmanned boat, which greatly improves the accuracy and efficiency of the camera in tracking dynamic targets.

[0072] (3) When training the dynamic target tracking network model, the present invention is the action a of the pan / tilt motor t Noise is added. The addition of noise can greatly enhance the exploration ability of the dynamic target tracking network model and improve the accuracy and precision of dynamic target tracking.

[0073] (4) When training a dynamic target tracking network model, the present invention adopts a priority experience replay method to filter out experience data from the experience pool, and updates the network parameters of the dynamic target tracking network model based on the filtered experience data. The priority experience replay algorithm can select experience data with higher priority according to the current state, efficiently utilize the experience data and update it, thereby greatly improving the accuracy and efficiency of the network parameters.

[0074] (5) The present invention adopts a three-degree-of-freedom parallel mechanism motion platform as a camera gimbal, which not only has more accurate positioning and higher precision, but is also easier to operate, can adapt to various complex scenes, and has strong robustness.

[0075] (6) The unmanned boat dynamic target tracking method of the present invention can adaptively, accurately and stably track various complex targets, which has brought a significant technological breakthrough to the unmanned boat dynamic target tracking task; and the method of the present invention has strong universality, which can be applied not only to unmanned boats, but also to platforms such as drones and mobile robots after simple modification, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 This is a flow chart of a training method for a dynamic target tracking network model of an unmanned boat according to the present invention;

[0077] Figure 2 Schematic diagram of the structure of the three-degree-of-freedom parallel mechanism motion platform in the present invention; in the figure, 1 is the motor bracket, 2 is the lower connecting rod, 3 is the upper connecting rod, 4 is the camera, 5 is the camera bracket, 6 is the upper platform, 7 is the motor, and 8 is the base;

[0078] Figure 3 A schematic diagram is provided for establishing coordinate parameters of the three-degree-of-freedom parallel mechanism motion platform in the present invention;

[0079] Figure 4 Schematic diagram of the network architecture of the Actor network in the dynamic target tracking network model of the present invention;

[0080] Figure 5 Schematic diagram of the network architecture of the Critic network in the dynamic target tracking network model of the present invention. DETAILED DESCRIPTION

[0081] The present invention is further described in detail below through specific examples, but the scope of the present invention is not limited thereto.

[0082] Example 1:

[0083] A training method for a dynamic target tracking network model of an unmanned boat, such as Figure 1 The specific steps are as follows:

[0084] S1: Use the camera mounted on the camera gimbal of the unmanned boat to obtain the current environmental image, use the YOLO algorithm to identify and analyze the environmental image, find the image of the dynamic target to be tracked in the environmental image, determine the position information of the dynamic target to be tracked in the environmental image, and select the dynamic target.

[0085] To improve training efficiency and reduce training costs, in step S1, a UAV simulation environment model can be established based on the actual environment. This model is then used to obtain the current image of the simulated environment. The UAV simulation environment model includes the UAV model, the UAV motion model, the UAV camera model, the camera gimbal model, the dynamic target model, and the sea level environment model.

[0086] Furthermore, in order to effectively cope with the dynamic target tracking in the complex ocean environment and make the camera platform have a larger range of motion and flexibility, the present invention adopts a three-degree-of-freedom parallel mechanism motion platform as the camera platform to control the camera to obtain the environment image. The structural diagram of the three-degree-of-freedom parallel mechanism motion platform used in the present invention is shown in FIG. Figure 2 As shown. Figure 2 It can be seen that the three-degree-of-freedom parallel mechanism motion platform is composed of an upper platform, a base and three pairs of branches consisting of lower connecting rods and upper connecting rods. The motor fixed to the base and the lower connecting rod, the lower connecting rod and the upper connecting rod, the upper platform and the upper connecting rod are all connected by a revolute pair. The 9 revolute pair axes intersect at point O. The point O is called the rotation center of the three-degree-of-freedom parallel mechanism motion platform (i.e., the center of the sphere). The plane passing through point O and parallel to the base is called the center of the sphere plane. Driven by the three motors, the upper platform can achieve three-degree-of-freedom rotation relative to the base. In order to facilitate the description of the motor action of the camera gimbal, the present invention constructs a coordinate system for the three-degree-of-freedom parallel mechanism motion platform, that is: according to the DH method, a fixed coordinate system O-X0Y0Z0, a moving coordinate system O-XYZ, and a connecting rod coordinate system OX are established with the center of the sphere O as the origin. ij Y ij Z ij , i, j = 1, 2, 3, indicating the jth revolute pair of the i-th branch. The Z0 axis is the line connecting the center H of the base and the origin O, with the positive direction pointing to the upper platform; the Z axis is the line connecting the origin O and the center of the upper platform, with the positive direction pointing to the upper platform; the Z axis of the connecting rod coordinate system is the line connecting the origin O and the center of the upper platform, with the positive direction pointing to the upper platform; ij The axis is the jth rotation sub-direction of the i-th branch, and the positive direction points to the outside of the sphere; the direction of the X0 axis is the direction of the Z0 axis and the Z 11 The normal direction of the plane formed by the X-axis is the Z0 axis and the Z 13 The normal direction of the plane formed by the axis; X ij The axis direction is Z ij Axis and Z i(j+1) The normal direction of the plane formed by the axes; the Y-axis direction of each system is determined by the right-hand rule (the coordinate construction diagram of the three-degree-of-freedom parallel mechanism motion platform is as follows Figure 3 As shown, Figure 3 In order to highlight the coordinate system, the structure of the three-degree-of-freedom parallel mechanism motion platform is simplified).

[0087] Based on the coordinate system of the constructed three-degree-of-freedom parallel mechanism motion platform, the structural parameters of the three-degree-of-freedom parallel mechanism motion platform are: α1, α2, β1, β2, η, where α1 and α2 are the structural angles between the lower link and the upper link, respectively; β1 and β2 are the semi-cone angles between the base and the upper platform, respectively; and η is the structural angle between the upper platform and the base in the initial state, i.e., the initial torsion angle. These structural parameters can then be used to calculate the rotational quantities of the three motors in the three-degree-of-freedom parallel mechanism motion platform. The specific calculation method is as follows:

[0088] The gimbal can only rotate relative to the base, and its posture can be expressed by the direction cosine matrix R of the moving coordinate system O-XYZ relative to the fixed coordinate system O-X0Y0Z0 XYZ To express. R XYZ The calculation formula is shown in formula 1:

[0089]

[0090] Where, Respectively represent the up and down pitch angle, horizontal rotation angle and axial rotation angle of the gimbal, Respectively

[0091] The gimbal's rotation angle is known Using the middle hinge direction vector v i The hinge direction vector w of the upper platform i The angle between them is α2, and the constraint equation is established. According to the constraint equation, the inverse solution of the gimbal's rotation attitude angle is obtained, and the movements of the three motors can be obtained.

[0092] The constraint equation is shown in Equation 2:

[0093]

[0094] Where T represents the matrix transpose.

[0095] S2: Compare the length and width pixel values ​​of the dynamic target image extracted in step S1 with a preset image threshold to determine whether the length and width pixel values ​​of the dynamic target image are within the preset image threshold range. If the length and width pixel values ​​of the dynamic target image are not within the preset image threshold range, it indicates that the current unmanned vehicle is too far or too close to the dynamic target, which is not conducive to subsequent dynamic target tracking. The unmanned vehicle needs to move its position and then execute step S1 to reacquire the environment image and dynamic target image. If the length and width pixel values ​​of the dynamic target image are within the preset image threshold range, execute step S3.

[0096] S3: Calculate the center coordinates O of the environment image c (x c ,y c ), the center coordinate O of the dynamic target image t (x f ,y f ), construct the current state quantity s according to the center coordinates of the environment image and the dynamic target image t , s t =(x f ,y f , x c ,y c ).

[0097] S4: The state quantity s t Input the dynamic target tracking network model to calculate the rotation angle information of the camera platform (preferably, the rotation angle information of the camera platform is the pitch angle). Horizontal rotation angle and axial rotation angle ); According to the constraint equation (calculation formula as shown in Equation 2), the attitude inverse solution of the gimbal's rotation attitude angle is obtained to obtain the action a of the gimbal motor that controls the camera gimbal movement t , and according to the calculated pan-tilt motor action a t Preferably, in order to enhance the exploration capability of the dynamic target tracking network model, during the training of the dynamic target tracking network model, the action a of the pan-tilt motor is calculated. t When noise is added, the action of the pan / tilt motor a t The calculation formula is shown in Equation 3 below:

[0098] a t =μ(s t |θ μ )+N t Formula 3

[0099] Where s t Represents the current state; θ μ represents the Actor network parameters; μ(·) represents the Actor network; N t represents the random noise parameter; θ is a parameter, θ>0.

[0100] Among them, the random noise parameter N t The calculation formula of the table is shown in Formula 4:

[0101]

[0102] Where, θ is a parameter, θ>0; σ is a parameter, σ>0; B τis the standard Brownian motion; τ is the standard time; t is the actual time.

[0103] Since the training goal of the dynamic target tracking network model is to make the camera lens center close to the center of the dynamic target being tracked, the farther the distance between the camera lens center of the unmanned boat and the center of the dynamic target on the image, the lower the reward is. A correction coefficient is set to ensure the stability of the reward parameters. Moreover, to ensure the high efficiency of the camera movement, the total reward will be subtracted from the corrected time. Therefore, based on the state quantity s t Constructing the reward function r of the dynamic target tracking network model t (s t ). The reward function r t (s t ) is calculated as shown in Equation 5 below:

[0104]

[0105] In the formula, (x f ,y f ) is the center coordinate of the dynamic target image; (x c ,y c ) is the center coordinate of the environment image; R max is the maximum reward value returned when the center coordinates of the dynamic target and the center coordinates of the environment image coincide; k, α>0 is the correction coefficient; R max , α and k are all adjustable parameters, T is the pan / tilt motor receiving action a t After completing the action a t total time.

[0106] The dynamic target tracking network model described in the present invention is a DDPG (Deep Deterministic Policy Gradient) reinforcement learning algorithm neural network. The dynamic target tracking network model is composed of an Actor network (policy network), a Critic network (state value function estimation network), a Target Actor network (target policy network) and a TargetCritic network (target state value function estimation network); the Actor network has the same network architecture as the Target Actor network, and the Critic network has the same network architecture as the Target Critic network. The input of the Actor network is the state quantity s t , the output is the attitude angle information of the gimbal camera; the input of the Target Actor network is the state quantity s at the next moment t ′, the output is the attitude angle information of the gimbal camera at the next moment; the input of the Critic network is the state-action pair (s t , a t), the output is a state-action pair (s t , a t ) state value evaluation value; the input of the Target Critic network is the state action pair (s t ′,a t ′), the output is the state-action pair (s t ′,a t ′)’s state value target value.

[0107] The Actor network (such as Figure 4 As shown in Figure 2, it consists of an input layer, two linear hidden layers, and an output layer; the input of the input layer is the state quantity s t , state quantity s t The first and second linear hidden layers are sequentially fully connected and normalized to obtain the strategic hidden layer data. The strategic hidden layer data obtained through the second linear hidden layer is input into the output layer for full connection processing to obtain the attitude angle information of the gimbal camera. More preferably, the input layer of the actor network contains four nodes, and the four nodes of the input layer are connected to the first linear hidden layer via a full connection. The first linear hidden layer contains 64 nodes. After normalization and activation, the nodes in the first linear hidden layer are fully connected to the second linear hidden layer. The second linear hidden layer contains 32 nodes, and the nodes in the second linear hidden layer are fully connected to the output layer.

[0108] The Critic network (such as Figure 5 The first branch includes a state input layer and two linear hidden layers. The input of the state input layer is the state quantity s t , state quantity s t The two linear hidden layers of the first branch are fully connected and normalized to obtain the state evaluation hidden layer data. The second branch includes an action input layer and a linear hidden layer. The input of the action input layer is the action a of the pan / tilt motor. t , the action of the pan / tilt motor a t The linear hidden layer of the second branch is fully connected and normalized to obtain action evaluation hidden layer data. The state evaluation hidden layer data obtained by the first branch and the action evaluation hidden layer data obtained by the second branch are normalized and added to obtain the state value evaluation value and output it. More preferably, in the first branch of the Critic network, the state input layer contains four nodes, and the four nodes of the state input layer are connected to the first linear hidden layer via a full connection. The first linear hidden layer contains 64 nodes, and the nodes of the first linear hidden layer are fully connected to the second linear hidden layer after normalization activation; in the second branch of the Critic network, the action input layer contains three nodes, and the action input layer is connected to the linear hidden layer via a full connection.

[0109] S5: Repeat steps S1-S3 to obtain the state quantity s at the next moment t ′, the current state quantity s t 、Action a t , reward r t , the state quantity s at the next moment t ′ as a set of empirical data (s t , a t , r t , s t ′).

[0110] S6: Repeat steps S1-S5 to obtain multiple sets of experience data, and store all the experience data in the experience pool; then determine whether the number of repeated cycles of steps S1-S5 is full of the network parameter update steps M of the dynamic target tracking network model. If the number of repeated cycles of steps S1-S5 is less than M, execute step S1; if the number of repeated cycles of steps S1-S5 is equal to M, execute step S7.

[0111] S7: Using the priority experience replay algorithm to filter out n sets of experience data from the experience pool, and using the filtered n sets of experience data to update the network parameters of the dynamic target tracking network model in sequence. The formula for filtering experience data from the experience pool using the priority experience replay algorithm is shown in the following formula 6:

[0112]

[0113] Where, P(c) represents a set of experience data (s t , a t , r t , s t ′); c represents the sequence number of the extracted empirical data, p k represents the priority of the extracted empirical data, and α represents the preset parameter used to adjust the priority sampling degree of data samples.

[0114] S8: Repeat steps S1-S7 and count the number of repetitions of steps S1-S7. If the number of repetitions of steps S1-S7 is greater than or equal to a preset number of cycles N, the loop ends, the dynamic target tracking network model training is completed, and the optimal network parameters of the dynamic target tracking network model are obtained. The preset number of cycles N is greater than the number of repetitions of steps S1-S7 when the dynamic target tracking network model converges.

[0115] The method for determining whether the dynamic target tracking network model has converged is as follows: inputting the current state quantity into the dynamic target tracking network model, calculating the rotation attitude angle information of the camera gimbal, and calculating the action a of the gimbal motor that controls the movement of the camera gimbal according to the rotation attitude angle information of the camera gimbal. t, according to the calculated motion a of the pan / tilt motor t After the gimbal motor is controlled to move, the deviation between the center of the dynamic target and the center of the environment image in the new image obtained by the camera is less than the preset maximum value D. If the deviation is less than the preset maximum value D, it means that no matter how the scene changes, the current network parameters can ensure that the center of the target is always in the center field of view of the image, and the model converges.

[0116] S9: Using the optimal network parameters obtained in step S8 to update the network parameters of the dynamic target tracking network model, to obtain a trained dynamic target tracking network model.

[0117] The specific process of using the optimal network parameters obtained in step S8 to update the network parameters of the dynamic target tracking network model is as follows:

[0118] (1) Update the network parameters of the Actor network:

[0119] Use the Actor network to calculate the state s t The pan / tilt motor action a t , PTZ motor action a t The calculation formula is as follows:

[0120] a t =μ(s t |θ μ )

[0121] Among them, θ μ is the network parameter of the Actor network, μ(·) represents the Actor network;

[0122] Use the Critic network to calculate the state action pair (s t , a t )’s state value evaluation value q t , state value evaluation value q t The calculation formula is as follows:

[0123] q t =Q(s t , a t |θ Q )

[0124] Among them, θ Q is the network parameter of the Critic network, Q(·) represents the Critic network;

[0125] Use the gradient ascent algorithm to maximize the state value evaluation value q t , thus adjusting the parameters θ in the Actor network μ Make updates;

[0126] (2) Update of network parameters of the Critic network:

[0127] Use the TargetActor network to calculate the next state s t ′The pan / tilt motor moves a t ′, PTZ motor action a t The calculation formula of ′ is as follows:

[0128] a t ′=μ′(s t ′|θ μ′ )

[0129] Among them, θ μ′ is the network parameter of the TargetActor network, μ′(·) represents the Target Actor network;

[0130] Use Target Critic network to calculate the state action pair (s t , a t )’s state value target value y t , state value target value y t The calculation formula is as follows:

[0131] y t =r t +γ(1-done)Q′(s t ′,a t ′|θ Q )

[0132] Among them, r t is the reward function, γ is the discounted return rate, done indicates whether the authority correction is completed, θ Q′ is the network parameter of the TargetCritic network, Q′(·) represents the Target Critic network

[0133] Use the gradient descent algorithm to minimize the state value evaluation value q t and state value target value y t The difference between L c , thereby updating the parameters in the Critic network, the difference L c The calculation formula is as follows:

[0134] L c =(y t -q t ) 2

[0135] (3) Network parameter update of Target Actor network and Target Critic network

[0136] To update the target network, the DDPG reinforcement learning algorithm uses a soft update method, also known as exponential moving average. That is, a learning rate (or momentum) τ is introduced to take the weighted average of the old target network parameters and the new corresponding network parameters, and then assign them to the target network.

[0137] Among them, the formula for the network parameter update process of the Target Actor network is as follows:

[0138] θ μ′ =τθ μ +(1-τ)θ μ′

[0139] The formula for the Target Critic network update process is as follows:

[0140] θ Q′ =rθ Q +(1-τ)θ Q′

[0141] θ μ is the network parameter of the Actor network; θ μ′ is the network parameter of the Target Actor network; θ Q is the network parameter of the Critic network; θ Q′ is the network parameter of the Target Critic network; τ is the learning rate (momentum), τ∈(0, 1), usually taking the value of 0.005.

[0142] Example 2:

[0143] A method for tracking dynamic targets of an unmanned boat, the specific steps are as follows:

[0144] S11: Using the camera on the unmanned boat to obtain an environmental image, perform recognition and analysis on the environmental image, and find the image of the dynamic target in the environmental image;

[0145] S22: Determine whether the size of the dynamic target image is within a preset image threshold range. If the size of the dynamic target image is not within the preset image threshold range, the unmanned boat moves and then executes step S11; if the size of the dynamic target image is within the preset image threshold range, executes step S33;

[0146] S33: Determine the center coordinates of the environment image and the dynamic target image, and construct a state quantity according to the center coordinates of the environment image and the dynamic target image;

[0147] S44: Input the state quantity into the dynamic target tracking network model to calculate the rotation attitude angle information of the camera gimbal; calculate the action of the gimbal motor that controls the movement of the camera gimbal based on the rotation attitude angle information of the camera gimbal, and control the gimbal motor to move based on the calculated action of the gimbal motor; the dynamic target tracking network model is a trained dynamic target tracking network model obtained by training using the training method described in the above embodiment 1.

[0148] S55: Repeat steps S11-S44 to enable the camera to follow the dynamic target and complete the tracking of the dynamic target.

[0149] Example 3:

[0150] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the training method for the unmanned boat dynamic target tracking network model as described in the above embodiment 1, or the unmanned boat dynamic target tracking method as described in the above embodiment 2.

[0151] Example 4:

[0152] A computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the training method for the unmanned boat dynamic target tracking network model as described in the above-mentioned embodiment 1 or the unmanned boat dynamic target tracking method as described in the above-mentioned embodiment 2 is implemented.

[0153] The above description is only a preferred embodiment of the present invention, but is not limited to the above examples. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A training method for a dynamic target tracking network model of an unmanned vehicle, characterized in that: The following steps are involved: S1: Use the camera on the camera gimbal of the unmanned boat to obtain the current environment image, identify and analyze the environment image, and find the image of the dynamic target that needs to be tracked in the environment image. S2: Determine whether the size of the dynamic target image is within a preset image threshold range. If the size of the dynamic target image is not within the preset image threshold range, the unmanned boat moves and then executes step S1; if the size of the dynamic target image is within the preset image threshold range, executes step S3; S3: Determine the center coordinates of the environment image and the dynamic target image, and construct the current state quantity s based on the center coordinates of the environment image and the dynamic target image. t ; S4: The state quantity s t Input the dynamic target tracking network model to calculate the rotation angle information of the camera gimbal; calculate the action a of the gimbal motor that controls the movement of the camera gimbal according to the rotation angle information of the camera gimbal. t , according to the calculated motion a of the pan / tilt motor t Control the pan / tilt motor to move; at the same time, based on the state quantity s t Constructing the reward function r of the dynamic target tracking network model t (s t ); wherein the dynamic target tracking network model is a DDPG reinforcement learning algorithm neural network; S5: Repeat steps S1-S3 to obtain the state quantity s at the next moment t ′, the current state quantity s t 、Action a t , reward r t , the state quantity s at the next moment t ′ as a set of empirical data (s t , a t , r t , s t ′); S6: Repeat steps S1-S5 to obtain multiple sets of experience data, and store all the experience data in the experience pool; then determine whether the number of repetitions of steps S1-S5 is equal to the number of network parameter update steps M of the dynamic target tracking network model; if the number of repetitions of steps S1-S5 is less than M, execute step S1; if the number of repetitions of steps S1-S5 is equal to M, execute step S7; S7: using a priority experience replay algorithm to filter out n sets of experience data from the experience pool, and using the filtered n sets of experience data to sequentially update the network parameters of the dynamic target tracking network model; S8: Repeat steps S1-S7 and count the number of repetitions of steps S1-S7. If the number of repetitions of steps S1-S7 is greater than or equal to a preset number of cycles N, the loop ends, the dynamic target tracking network model training is completed, and the optimal network parameters of the dynamic target tracking network model are obtained; the preset number of cycles N is greater than the number of repetitions of steps S1-S7 when the dynamic target tracking network model converges; S9: Using the optimal network parameters obtained in step S8 to update the network parameters of the dynamic target tracking network model, to obtain a trained dynamic target tracking network model.

2. The training method according to claim 1, characterized in that The state quantity s t =(x f ,y f , x c ,y c ), the reward function r t (s t ) is calculated as follows: Among them, (x f ,y f ) is the center coordinate of the dynamic target image; (x c ,y c ) is the center coordinate of the environment image; R max is the maximum reward value returned when the center coordinates of the dynamic target and the center coordinates of the environment image coincide; k, α>0 is the correction coefficient; T is the pan-tilt motor receiving action a t After completing the action a t total time.

3. The training method according to claim 2, characterized in that The dynamic target tracking network model is composed of an Actor network, a Critic network, a Target Actor network, and a Target Critic network. The Actor network has the same network architecture as the Target Actor network, and the Critic network has the same network architecture as the Target Critic network. The input of the Actor network is the state quantity s t , the output is the attitude angle information of the gimbal camera; the input of the TargetActor network is the state quantity s at the next moment t ′, the output is the attitude angle information of the gimbal camera at the next moment; The input of the Critic network is the state-action pair (s t , a t ), the output is a state-action pair (s t , a t ) state value evaluation value; the input of the Target Critic network is the state action pair (s t ′,a t ′), the output is the state-action pair (s t ′,a t ′)’s state value target value.

4. The training method according to claim 3, characterized in that The Actor network consists of an input layer, two linear hidden layers and an output layer; the input of the input layer is the state quantity s t , state quantity s t The first linear hidden layer and the second linear hidden layer are fully connected and normalized in sequence to obtain the strategy hidden layer data. The strategy hidden layer data obtained by the second linear hidden layer processing is input into the output layer for full connection processing to obtain the attitude angle information of the gimbal camera. The Critic network includes two branches. The first branch includes a state input layer and two linear hidden layers. The input of the state input layer is the state quantity s t , state quantity s t The two linear hidden layers of the first branch are fully connected and normalized to obtain the state evaluation hidden layer data. The second branch includes an action input layer and a linear hidden layer. The input of the action input layer is the action a of the pan / tilt motor. t , the action of the pan / tilt motor a t The linear hidden layer of the second branch is fully connected and normalized to obtain the action evaluation hidden layer data. The state evaluation hidden layer data obtained by the first branch and the action evaluation hidden layer data obtained by the second branch are normalized and added to obtain the state value evaluation value and output it.

5. The training method according to claim 3, characterized in that: In step S4, the action a of the pan / tilt motor is calculated. t When adding noise, the action of the pan / tilt motor a t The calculation formula is as follows: a t =μ(s t |θ μ )+N t Where s t Represents the current state; θ μ represents the Actor network parameters; μ(·) represents the Actor network; N t represents the random noise parameter; θ is the parameter, θ>0; σ is the parameter, σ>0; B τ is the standard Brownian motion; τ is the standard time; t is the actual time.

6. The training method according to any one of claims 1 to 5, characterized in that: The formula for selecting experience data from the experience pool using the priority experience replay algorithm in step S7 is as follows: Where, P(c) represents a set of experience data (s t , a t , r t , s t ′); c represents the sequence number of the extracted empirical data, p k represents the priority of the extracted empirical data, and α represents the preset parameter used to adjust the priority sampling degree of data samples.

7. The training method according to claim 6, characterized in that The camera gimbal is a three-degree-of-freedom parallel mechanism motion platform.

8. A method for tracking dynamic targets of an unmanned boat, characterized in that: The following steps are involved: S11: Using the camera on the unmanned boat to obtain an environmental image, perform recognition and analysis on the environmental image, and find the image of the dynamic target in the environmental image; S22: Determine whether the size of the dynamic target image is within a preset image threshold range. If the size of the dynamic target image is not within the preset image threshold range, the unmanned boat moves and then executes step S11; if the size of the dynamic target image is within the preset image threshold range, executes step S33; S33: Determine the center coordinates of the environment image and the dynamic target image, and construct a state quantity according to the center coordinates of the environment image and the dynamic target image; S44: Inputting the state quantity into a dynamic target tracking network model to calculate rotational attitude angle information of the camera gimbal; calculating the movement of a gimbal motor that controls the movement of the camera gimbal based on the rotational attitude angle information of the camera gimbal, and controlling the gimbal motor to move based on the calculated movement of the gimbal motor; the dynamic target tracking network model is the trained dynamic target tracking network model according to any one of claims 1 to 7; S55: Repeat steps S11-S44 to enable the camera to follow the dynamic target and complete the tracking of the dynamic target.

9. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and when the processor executes the computer program, it implements the training method of the unmanned boat dynamic target tracking network model according to any one of claims 1 to 7, or the unmanned boat dynamic target tracking method according to claim 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the training method for the unmanned boat dynamic target tracking network model according to any one of claims 1 to 7 or the unmanned boat dynamic target tracking method according to claim 8 is implemented.

Citation Information

Patent Citations

  • Visual tracking method and device based on continuous movements under guidance of deep reinforcement learning

    CN108549928A