Visual Active Tracking and Aiming Method for Unmanned Weapon Platform Based on Deep Meta-Reinforcement Learning

By building a virtual simulation environment in the UE4 simulation engine and designing reward functions suitable for unmanned platforms and weapon gimbals, using deep meta reinforcement learning model to improve the generalization ability and tracking accuracy of the visual active tracking and targeting system of the unmanned weapon platform, solving the problems of poor generalization ability and inappropriate reward functions in the existing technology.

CN115187631BActive Publication Date: 2025-06-10BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210722984.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2025-06-10
Estimated Expiration
2042-06-24

AI Technical Summary

Technical Problem

In the prior art, deep reinforcement learning is prone to overfitting a single task, resulting in poor generalization capabilities of active visual tracking and aiming systems. At the same time, the existing reward function is not suitable for the problem of different resolutions of unmanned platforms and weapon gimbals.

Method used

The unmanned weapon platform visual active tracking and aiming method is adopted based on deep meta reinforcement learning. By building a virtual simulation environment in the UE4 simulation engine, a rich training set and test set are generated, a deep meta reinforcement learning model is built using PPO algorithm and LSTM network, and a reward function suitable for unmanned platforms and weapon gimbals is designed to improve the generalization ability of the model and tracking and aiming accuracy.

Benefits of technology

It realizes end-to-end visual active tracking and targeting targets of the unmanned weapon platform in the real environment, improves the system's generalization ability and tracking and targeting accuracy, and is suitable for military robots and biometric collection and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115187631B_ABST
    Figure CN115187631B_ABST
Patent Text Reader

Abstract

The present invention discloses a visual active tracking and aiming method for an unmanned weapon platform based on deep meta-reinforcement learning. A virtual simulation environment is built based on the UE4 simulation engine to generate a series of visual active tracking and aiming tasks for the unmanned weapon platform to form a training set and a test set. The proximal policy optimization algorithm of deep reinforcement learning is used to construct a deep meta-reinforcement learning model. A reward function for the tracking and aiming task is designed, and the deep meta-reinforcement learning model is trained in the training set until the model converges. The converged model is tested in the test set. The model is deployed on the unmanned weapon platform to verify the performance of the model in the real environment, achieve rapid adaptation to new tasks, and improve the generalization ability and tracking and aiming accuracy of the visual active tracking and aiming system of the unmanned weapon platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a visual active tracking and aiming method for an unmanned weapon platform based on deep meta-reinforcement learning, belonging to the fields of computer vision, deep meta-reinforcement learning, and robotics, and can be used in military robots, automatic face or iris acquisition systems, etc. Background Technique

[0002] Visual tracking refers to detecting and tracking a target in consecutive frames collected by a visual sensor, and can be applied in multiple fields such as autonomous driving, monitoring, robot control, and human assistance. Visual tracking is divided into passive visual tracking and active visual tracking. Passive visual tracking refers to detecting and tracking a target in a video with the tracker's position unchanged. The passive method is limited to the situation where the target is within the camera's field of view. Active visual tracking means that the tracker has to actively search for and track the target, and the tracker's position changes. This technology requires the tracker to position itself relative to the target and plan an optimal trajectory to track the target in real time.

[0003] Deep meta-reinforcement learning has good generalization ability for new tasks. After the network weights are fixed upon training completion, it can dynamically adjust the hidden layer through an internal recurrent neural network to quickly adapt to new tasks. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: overcoming the deficiencies of the prior art, and proposing a visual active tracking and aiming method for an unmanned weapon platform based on deep meta-reinforcement learning. The purpose of this method is to solve the problems that deep reinforcement learning is prone to overfitting for a single task, resulting in poor generalization ability of the active visual tracking and aiming system based on deep reinforcement learning, and the existing reward function is not suitable for the different action control resolutions of the unmanned platform and the weapon pan-tilt. This method uses a training set with rich task types, a good reward function, and a deep meta-reinforcement learning model to improve the generalization ability and tracking and aiming accuracy of the visual active tracking and aiming system for the unmanned weapon platform.

[0005] The technical solution of the present invention is:

[0006] A visual active tracking and aiming method for an unmanned weapon platform based on deep meta-reinforcement learning, and the specific steps of this method include:

[0007] Step S1, build a virtual simulation environment based on the UE4 simulation engine, and generate a series of visual active tracking and aiming tasks for the unmanned weapon platform to form a training set and a test set;

[0008] Step S2, use the deep reinforcement learning PPO algorithm and the LSTM network to build a deep meta-reinforcement learning model;

[0009] Step S3, design a reward function for the tracking and aiming task, and train the deep meta-reinforcement learning model in the training set until the model converges;

[0010] Step S4, test the model trained to convergence in step S3 in the test set;

[0011] Step S5, deploy the model tested in step S4 on the unmanned weapon platform to verify the performance of the model in the real environment.

[0012] In the described step S1, when building a virtual simulation environment based on the UE4 simulation engine, randomly change environmental parameters such as lighting, terrain, obstacles, color, and texture, unmanned weapon platform parameters such as friction coefficient, mass, centroid position, inertia matrix, and shape size, and target parameters such as type, initial position, initial velocity, color, and texture, etc. Generate a series of visual active tracking and aiming tasks for the unmanned weapon platform. These tasks form the training set and the test set, and the tasks in the test set are not included in the training set;

[0013] In the described step S3, the unmanned weapon platform includes two parts: an unmanned platform and a weapon pan-tilt. Since the motion control resolutions of these two parts are different, the reward function of the tracking and aiming task is divided into the unmanned platform reward r agent and the weapon pan-tilt reward r weapon into two parts;

[0014] r = r agent + μr weapon , (if w agent_yaw < w precision : μ = 1 else: μ = 0)

[0015]

[0016] where, Δx and Δy represent the offsets of the x-axis and y-axis between the unmanned weapon platform and the target in the world coordinate system; w agent_yaw is the yaw angle required for the center line of the unmanned platform to align with the target; w precision is the motion control resolution of the unmanned platform; w weapon_yaw and w weapon_pitch are the yaw angle and pitch angle required for the center line of the weapon on the weapon pan-tilt to aim at the center point of the target.

[0017] Beneficial effects

[0018] (1) The method of the present invention applies deep meta-reinforcement learning to the tracking and aiming system of the unmanned weapon platform, and utilizes the characteristics of strong generalization ability and fast adaptation to new tasks of deep meta-reinforcement learning to achieve end-to-end visual active tracking and aiming at the target by the unmanned weapon platform in the real environment.

[0019] (2) The method of the present invention generates a training set with rich tasks based on the UE4 engine; designs a reward function for the tracking and aiming task according to the different motion control resolutions of different parts of the unmanned weapon platform, improving the tracking and aiming accuracy of the model; uses the deep meta-reinforcement learning model to slowly learn and adjust the weights of the LSTM network during training, and dynamically adjusts the hidden layer state through the LSTM network during deployment to quickly adapt to new tasks and improve the generalization ability of the model.

[0020] (3) The method of the present invention designs a reward function for the tracking and aiming task, which can accurately and quickly track and aim at the target, and can be applied to fields such as military robot tracking and aiming, biometric collection and identification.

[0021] (4) The present invention discloses a visual active tracking and aiming method for an unmanned weapon platform based on deep meta-reinforcement learning, including a deep meta-reinforcement learning model, a visual active tracking and aiming task training set and a test set for the unmanned weapon platform.

[0022] (5) The method of the present invention includes the following steps: S1: Build a virtual simulation environment based on the UE4 simulation engine to generate a series of visual active tracking and aiming tasks for the unmanned weapon platform to form a training set and a test set; S2: Use the Proximal Policy Optimization (PPO) algorithm of deep reinforcement learning and the Long Short-Term Memory (LSTM) network to construct a deep meta-reinforcement learning model; S3: Design a reward function for the tracking and aiming task, and train the deep meta-reinforcement learning model in the training set until the model converges; S4: Test the model trained to convergence in step S3 in the test set; S5: Deploy the model on the unmanned weapon platform to verify the performance of the model in the real environment.

[0023] (6) The method of the present invention realizes the visual active tracking and aiming task in an end-to-end manner, taking the stacked video sequence at the current moment, the action at the previous moment and the reward value at the current moment as inputs, and outputting the action of the unmanned weapon platform.

[0024] (7) The method of the present invention uses a training set with rich task types, a well-designed reward function and a deep meta-reinforcement learning model, slowly learns and adjusts the weights of the LSTM network during training, and dynamically adjusts the hidden state through the LSTM network during deployment to quickly adapt to new tasks and improve the generalization ability and tracking and aiming accuracy of the visual active tracking and aiming system for the unmanned weapon platform. Brief Description of the Drawings

[0025] Figure 1 It is a schematic diagram of the composition of the unmanned weapon platform in the present invention;

[0026] Figure 2Schematic diagram of the method flow of the present invention;

[0027] Figure 3 Schematic diagram of the simulation parameter composition of the visual active tracking and aiming task training set and test set of the unmanned weapon platform in the present invention;

[0028] Figure 4 Schematic diagram of the deep meta-reinforcement learning model of the present invention;

[0029] Figure 5 For calculating the reward r of the unmanned platform in the present invention agent Schematic diagram for illustration;

[0030] Figure 6 For calculating the reward r of the weapon pan-tilt in the present invention weapon Schematic diagram for illustration. Specific implementation manners

[0031] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0032] The present invention includes a deep meta-reinforcement learning model, a visual active tracking and aiming task training set and a test set of an unmanned weapon platform. The unmanned weapon platform used in the present invention includes an unmanned platform, a weapon pan-tilt and a camera, as Figure 1 shown. The present invention trains the deep meta-reinforcement learning model in the training set, and tests the generalization ability and tracking and aiming performance of the model in the test set and the real environment. The present invention proposes a visual active tracking and aiming method for an unmanned weapon platform based on deep meta-reinforcement learning, as Figure 2 shown, including the following steps:

[0033] Step S1: Build a virtual simulation environment based on the UE4 simulation engine, and generate a series of visual active tracking and aiming tasks of the unmanned weapon platform to form a training set and a test set;

[0034] The specific content of step S1 is: Since directly training the deep meta-reinforcement learning model in the real environment has problems such as long training time and wear of the unmanned weapon platform, the algorithm model is usually trained in the simulation environment. The present invention builds a virtual simulation environment based on the UE4 simulation engine, as Figure 3As shown, randomly change environmental parameters such as lighting, terrain, obstacles, color, and texture, unmanned weapon platform parameters such as friction coefficient, mass, centroid position, inertia matrix, and shape and size, and target parameters such as type, initial position, initial velocity, color, and texture, etc. to generate a series of visual active tracking and aiming tasks for the unmanned weapon platform. These tasks form a training set and a test set, and the tasks in the test set are not included in the training set. In the present invention, the targets include ground targets and air targets, enabling the present invention to have the ability to track and aim at various targets.

[0035] Step S2: Use the deep reinforcement learning PPO algorithm and the LSTM network to construct a deep meta-reinforcement learning model;

[0036] The specific content of step S2 is as follows: The PPO algorithm is divided into an evaluation network and a policy network, and both networks are composed of LSTM networks, as Figure 4 shown. Take the stacked video sequence at the current moment, the action at the previous moment, and the reward value at the current moment as inputs, extract features through a convolutional neural network and a fully connected network, input the features into the policy network and the evaluation network, the policy network outputs the action of the unmanned weapon platform, and the evaluation network outputs the state-action value.

[0037] Step S3: Design a reward function for the tracking and aiming task, and train the deep meta-reinforcement learning model in the training set until the model converges;

[0038] The specific content of step S3 is as follows: The deep meta-reinforcement learning model realizes the tracking and aiming task by maximizing the final reward. Design a reward function for the tracking and aiming task, and train the deep meta-reinforcement learning model in the training set until the model converges. Since the unmanned weapon platform includes two parts, an unmanned platform and a weapon gimbal, and due to the different action control resolutions of these two parts, the reward function of the tracking and aiming task is divided into the unmanned platform reward r agent and the weapon gimbal reward r weapon two parts, as Figure 5 and Figure 6 shown;

[0039] r = r agent + μr weapon , (if w agent_yaw < w precision : μ = 1 else: μ = 0)

[0040]

[0041] Among them, Δx and Δy represent the offsets of the x-axis and y-axis between the unmanned weapon platform and the target in the world coordinate system; w agent_yaw is the yaw angle required for the center line of the unmanned platform to align with the target; w precision is the motion control resolution of the unmanned platform; wweapon_yaw and w weapon_pitch are the yaw angle and pitch angle required for the weapon center line on the weapon turret to aim at the target center point.

[0042] Step S4: Test the model trained to convergence in step S3 in the test set;

[0043] The specific content of step S4 is: Test the model trained to convergence in step S3 in the test set of the visual active tracking and aiming task of the unmanned weapon platform. Replace the weapon on the weapon turret with a camera for testing. When the model meets the following indicators in all tasks of the test set, it is determined that the model can track the target stably and accurately.

[0044] d target -d safe ≤Δd

[0045]

[0046] t aim ≥Δt

[0047] where d target represents the distance between the unmanned weapon platform and the target; d safe represents the safe distance that needs to be maintained between the unmanned weapon platform and the target; Δd represents the allowable distance error; Δx target and Δy target represent the offsets of the center point of the image of the camera on the weapon turret to the center point of the target on the x-axis and y-axis in the pixel coordinate system; Δσ represents the allowable aiming accuracy error. t aim represents the tracking and aiming duration when the first two indicators are continuously met; Δt represents the minimum tracking and aiming duration.

[0048] Step S5: Deploy the model on the unmanned weapon platform and verify the performance of the model in the real environment.

[0049] The specific content of step S5 is: Deploy the model on the unmanned weapon platform and verify the performance of the model in the real environment according to the indicators given in step 4, that is, the unmanned weapon platform can track and aim at the target stably and accurately in the real environment and meet the indicators given in step 4.

[0050] The application examples of the present invention are listed as follows:

[0051] Application example 1: Application of the visual active tracking and aiming system of the unmanned weapon platform based on deep meta-reinforcement learning on the military unmanned weapon platform.

[0052] The present invention can be applied to military unmanned weapon platforms to perform tracking tasks, reconnaissance tasks, and aiming and attacking tasks. By inputting the video sequence collected by the camera at the current moment, the reward at the current moment, and the action at the previous moment, it outputs the action instruction of the unmanned weapon platform to autonomously track, aim at, and attack the target.

[0053] Application Example 2: The vision active tracking and aiming system of the unmanned weapon platform based on deep meta-reinforcement learning is applied to the system for collecting iris or face information.

[0054] The present invention can be applied to a system that can actively approach the person to be collected without the cooperation of the person to be collected to collect iris or face information. By replacing the weapon gimbal in the present invention with a camera gimbal, the present invention can actively move in front of the person to be collected, and the camera on the camera gimbal can scan and identify the iris or face of the person to be collected, which is quite convenient for the biological information collection work.

[0055] It should be noted that the technical solutions not detailed in this application adopt well-known technologies. The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. Visual Active Tracking and Aiming Method for Unmanned Weapon Platform Based on Deep Meta-Reinforcement Learning, characterized in that the steps of this method include: Step S1, build a virtual simulation environment based on the UE4 simulation engine, and generate a series of visual active tracking and aiming tasks for the unmanned weapon platform to form a training set and a test set; Step S2, use the deep reinforcement learning PPO algorithm and the LSTM network to construct a deep meta-reinforcement learning model; Step S3, design a reward function for the tracking and aiming task, and train the deep meta-reinforcement learning model in the training set until the model converges; Step S4, test the model trained to convergence in the test set according to Step S3; Step S5, deploy the model tested in Step S4 on the unmanned weapon platform to achieve visual active tracking and aiming of the unmanned weapon platform; In the said Step S3, the deep meta-reinforcement learning model realizes the tracking and aiming task by maximizing the final reward, designs a reward function for the tracking and aiming task, and trains the deep meta-reinforcement learning model in the training set until the model converges; The reward function for the tracking and aiming task is divided into the reward r of the unmanned platform agent and the reward r of the weapon gimbal weapon in two parts: r = r agent + μr weapon ,(if w agent_yaw < w precision : μ = 1 else: μ = 0) where, Δx and Δy represent the offsets of the x-axis and y-axis between the unmanned weapon platform and the target in the world coordinate system; w agent_yaw is the yaw angle required for the center line of the unmanned platform to align with the target; w precision is the motion control resolution of the unmanned platform; w weapon_yaw and w weapon_pitch are the yaw angle and pitch angle required for the center line of the weapon on the weapon turret to aim at the center point of the target.

2. The visual active tracking and aiming method for the unmanned weapon platform based on deep meta-reinforcement learning according to Claim 1, characterized in that: In the said Step S1, when building the virtual simulation environment based on the UE4 simulation engine, randomly change the environment parameters, randomly change the unmanned weapon platform parameters, and randomly change the target parameters.

3. The visual active tracking and aiming method for the unmanned weapon platform based on deep meta-reinforcement learning according to Claim 2, characterized in that: The said environment parameters include illumination, terrain, obstacles, color, and texture; The said unmanned weapon platform parameters include friction coefficient, mass, centroid position, inertia matrix, and shape and size; The said target parameters are target type, initial position, initial velocity, color, and texture.

4. The visual active tracking and aiming method for the unmanned weapon platform based on deep meta-reinforcement learning according to any one of Claims 1-3, characterized in that: In the said Step S1, the tasks in the test set are not included in the training set, and the targets include ground targets and air targets.

5. The visual active tracking and aiming method for the unmanned weapon platform based on deep meta-reinforcement learning according to Claim 1, characterized in that: In the said Step S2, the deep reinforcement learning PPO algorithm includes an evaluation network and a policy network, and both the evaluation network and the policy network are composed of LSTM networks.

6. The visual active tracking and aiming method for the unmanned weapon platform based on deep meta-reinforcement learning according to Claim 5, characterized in that: The detailed steps of the deep reinforcement learning PPO algorithm are: take the stacked video sequence at the current moment, the action at the previous moment, and the reward value at the current moment as inputs, extract features through a convolutional neural network and a fully connected network, and input the extracted features into the policy network and the evaluation network. The policy network outputs the action of the unmanned weapon platform, and the evaluation network outputs the state-action value.

7. The visual active tracking and aiming method for the unmanned weapon platform based on deep meta-reinforcement learning according to Claim 1, characterized in that: In the said Step S4, the method for testing the model trained to convergence in the test set is: Replace the weapon on the weapon turret with a camera for testing. The model is determined to be able to stably and accurately track the target when it meets the following metrics in all tasks of the test set; d target -d safe ≤Δd t aim ≥Δt Among them, d target represents the distance between the unmanned weapon platform and the target; d safe represents the safe distance that needs to be maintained between the unmanned weapon platform and the target; Δd represents the allowable distance error; Δx target and Δy target represent the offsets of the center point of the image of the camera on the weapon turret to the center point of the target on the x-axis and y-axis in the pixel coordinate system; Δσ represents the allowable aiming accuracy error, t aim represents the tracking and aiming duration when the first two indicators are continuously satisfied; Δt represents the minimum tracking and aiming duration.

8. The visual active tracking and aiming method for an unmanned weapon platform based on deep meta-reinforcement learning according to claim 1, characterized in that: Deploy the model after the test in step S4 on a military unmanned weapon platform for application, and perform tracking tasks, reconnaissance tasks, and aiming and attacking tasks. Input the video sequence collected by the camera at the current moment, the reward at the current moment, and the action at the previous moment, and output the action instruction of the unmanned weapon platform to autonomously track, aim at, and attack the target, or deploy the model after the test in step S4 on a system for collecting iris or face information.