Dynamic single-target long-time tracking method and system
By using a lightweight dynamic single-target tracking method, combined with a deep twin tracking network and a Kalman filter, the problems of high computational complexity and poor stability of single-target tracking on edge computing devices are solved, and stable tracking is achieved in occlusion and complex environments.
Patent Information
- Application Number
- CN202511207276.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-11-18
AI Technical Summary
Existing single-target tracking technologies suffer from high computational complexity on edge computing devices, poor stability under occlusion conditions, and weak environmental adaptability, making it difficult to meet real-time tracking requirements.
A lightweight, dynamic, long-term single-target tracking method is adopted, which combines a deep twin tracking network and a Kalman filter to achieve stable target tracking through feature extraction, state determination, and policy switching.
It achieves efficient and robust single-target tracking on edge computing devices, capable of handling occlusion and complex environments, maintaining stability and meeting real-time tracking requirements.
Smart Images

Figure CN120976836A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of computer vision and artificial intelligence, and specifically relates to a dynamic single-target long-time tracking method and system applied to an edge computing device. BACKGROUND
[0002] Target recognition, tracking and gimbal servo are core research directions in modern artificial intelligence, computer vision and unmanned system engineering, aiming to realize the automatic detection, recognition and continuous tracking of specific targets in video or image sequences. This technology is widely used in video surveillance, autonomous driving, human-computer interaction, robot navigation and many other fields, greatly promoting the automation and intelligent development of the industry.
[0003] Existing single-target tracking technologies are mainly divided into traditional discriminative tracking algorithms and generative tracking algorithms based on deep learning. Traditional algorithms such as CSK and KCF based on coherent filtering are fast but lack accuracy and robustness. Deep learning-based algorithms, especially those represented by Siamese Network, extract target features through deep neural networks and have made significant breakthroughs in tracking accuracy.
[0004] However, existing technologies still face serious challenges in engineering applications, especially on edge computing devices (such as unmanned aerial computers) with limited computing power: High computational complexity: Advanced deep learning models, especially those with complex structures such as Transformers, have huge parameter quantities and computational loads. This results in extremely slow inference speed on edge devices, making it difficult to meet the real-time tracking requirements and reducing the value of engineering applications.
[0005] Poor tracking stability: When the target encounters long-time or complete occlusion, rapid motion or severe deformation, most tracking algorithms will fail due to loss of target features. They are difficult to continuously predict the target's trajectory during occlusion and quickly re-lock after the target reappears, resulting in interruption of the tracking task.
[0006] Weak environmental adaptability: Complex and variable environments, such as changes in light, background clutter and similar object interference, can severely affect the accuracy of feature extraction, thereby reducing the robustness of tracking. Artificially designed features in traditional algorithms are difficult to cope with variable scenarios, and deep learning models are also susceptible to interference, leading to tracking drift or failure.
[0007] Therefore, how to design a lightweight computing method that can effectively deal with occlusion and maintain stable tracking in complex environments is a problem that needs to be solved in the current technical field. SUMMARY
[0008] The application aims to solve the problems of high computational complexity, poor stability under occlusion and weak environmental adaptability of the single target tracking method in the prior art, and provide a lightweight dynamic single target long-time tracking method and system with high efficiency and strong robustness.
[0009] The dynamic single target long-time tracking method comprises the following steps: S1: obtaining a video stream for shooting a target, and determining a tracking target as a template image in a frame of the video stream; S2: initializing a deep twin tracking network and a trajectory prediction model based on the template image; S3: using the deep twin tracking network to perform feature extraction and correlation calculation on the template image and a search area of a subsequent frame of the video stream, to generate a response map containing target position information, and determine a preliminary tracking result therefrom; S4: determining the real-time tracking state of the target based on the response map, to distinguish whether the target is in a normal tracking state or an occlusion or loss state; S5: performing strategy switching, when the target is determined to be in the normal tracking state, taking the preliminary tracking result as a final tracking result of the current frame, and using the final tracking result to update the state of the trajectory prediction model; when the target is determined to be in the occlusion or loss state, activating the trajectory prediction model, predicting the position of the target in the current frame according to the historical motion information of the target, and taking the predicted position as the final tracking result of the current frame.
[0010] Further to the above technical solution, the S2 comprises: determining a target region as the template image by frame selection in any frame of the video stream; initializing the deep twin tracking network according to the template image; and assigning the target position information determined by the frame selection to a Kalman filter as the trajectory prediction model, to set the initial state thereof.
[0011] Further, the deep twin tracking network decomposes target tracking into a classification task and a regression task; the classification task generates a foreground probability map by convolution operation on the response map, to determine the position of the target; the regression task generates a target frame size map by convolution operation on the response map, to determine the size of the target frame; The candidate target frame generated according to the foreground probability map and the target frame size map is subjected to aspect ratio penalty scoring and scale penalty scoring, and the target frame with the highest comprehensive score is selected as the preliminary tracking result.
[0012] Furthermore, the deep twin tracking network includes a first branch and a second branch. The first branch extracts features from the template image, and the second branch extracts features from the search region. The search region is calculated based on the template image.
[0013] Furthermore, the determination of the occlusion or loss state includes: the occlusion or loss state is determined by a combination of average peak correlation energy (APCE) and color histogram similarity, wherein the APCE determines the occlusion state of the target based on the fluctuation state of the response map, and the color histogram determines the existence of the target by calculating the cosine similarity between the predicted region and the initial target region.
[0014] Furthermore, the average peak correlation energy (APCE) is calculated using the following formula: Where f is the APCE value and A is the response. , These are the maximum and minimum values of the response, respectively. The calculated APCE value is compared with a preset threshold as a preliminary basis for determining whether the target is occluded. The calculation of the color histogram similarity includes: resampling the currently predicted target region to the same size as the template image; statistically analyzing the color distribution of the R, G, and B channels for the template image and the resampled predicted region to generate their respective color histogram vectors; and quantifying the degree of appearance similarity between the two by calculating the cosine similarity between the two color histogram vectors.
[0015] Furthermore, the trajectory prediction model employs a Kalman filter that incorporates depth features to predict the target's position in the current frame, and uses the depth feature vector extracted by the deep twin tracking network for the current target ( ), and the position observation at time t ( The data are combined to form an extended observation vector. The observation matrix (H) and observation noise covariance matrix (R) of the Kalman filter are adjusted accordingly to process the enhanced extended observation vector.
[0016] Furthermore, the method also includes an expanded search strategy: when the trajectory prediction model fails to restore the tracking state to normal tracking within a preset number of consecutive frames, an expanded search strategy is activated, wherein the expanded search strategy is expressed by the formula... Dynamically calculate the search region radius for the next frame ,in, For the target instantaneous velocity, SL represents the average movement speed, and SL represents the target side length. is an attenuation coefficient, when the tracking target is not detected for more than a preset time, the tracking is terminated.
[0017] Further, the method further comprises a gimbal servo control step: according to the position coordinates of the target in the image frame determined by the tracking target information, the deviation of the target from the image center is calculated, and a corresponding control signal is generated to drive the gimbal to rotate, so as to keep the target in the center of the picture.
[0018] A lightweight dynamic single-target long-time tracking system for implementing the above method, comprising: an onboard computer, a gimbal connected with the onboard computer, and a ground display device in communication connection with the onboard computer. The onboard computer is configured with: a main tracking module configured to track the target in the video stream using a deep twin tracking network to generate a preliminary tracking result; a state determination module connected to the main tracking module and configured to determine the real-time tracking state of the target based on the preliminary tracking result and output a determination signal; a trajectory prediction module configured to predict the current position of the target according to the historical motion information of the target when the determination signal indicating that the target is in an occluded or lost state is received; a control unit connected to the state determination module, the main tracking module and the trajectory prediction module, configured to select the preliminary tracking result or the position predicted by the trajectory prediction module as the final tracking result according to the determination signal a communication unit for pushing the tracking result to the ground display device through a data link.
[0019] Advantages: Compared with the prior art, the advantages of the present application are: 1. Lightweight tracking method based on deep learning: A cross-platform general FFmpeg streaming media acceleration framework is built, and hardware encoding and decoding and rate adaptation strategy are used to realize cross-platform efficient push-pull streaming, which significantly reduces the transmission delay. The dynamic rate control algorithm analyzes the complexity parameters of the video frame (such as motion amplitude, color richness, etc.) through real-time network bandwidth fluctuations, balances the code rate allocation and visual quality under bandwidth constraints, and implements differentiated bit allocation for different complexity scenes through hierarchical coding technology, prioritizing key frames and motion area image quantity, thereby maintaining the stability of the tracking features in a high compression environment. In addition, it supports multi-stream parallel processing, optimizes resource scheduling through thread pool, and meets the needs of multi-target tracking scenarios. The proposed lightweight twin tracking network uses depth separable convolution to reconstruct the feature extraction module, which greatly reduces the number of parameters. Using a multi-thread asynchronous strategy for feature extraction under the NPU architecture, one branch of the twin tracking network is used to extract the features of the template image, and the other branch is used to extract the features of the search area. Thus, a cross-platform lightweight general efficient target tracking method is formed.
[0020] 2. Deep feature comparison mechanism: Based on the multi-level deep features extracted by the training model, the similarity between the target area and the initial template is measured in the feature space. The correlation between the features obtained by the two branches is calculated. In addition, considering that the deep twin tracking network ignores the initial features of the tracking target, and the color features of the target remain basically unchanged during the entire tracking process. Therefore, the color histogram is used as one of the indicators for target loss or target occlusion. Through the feature comparison mechanism, the incorrect target tracking process is stopped in time, the target mis-tracking rate is reduced, and the interference of complex environments on single target tracking is addressed.
[0021] 3. Deep feature and Kalman filter fusion prediction mechanism: Because the target features in the search area are completely lost when the target is completely occluded, the twin network has weak anti-occlusion ability. For short-term and completely occluded scenes, a new anti-occlusion strategy is proposed. A Kalman filter that combines KCF and deep features is used to compensate for the shortcomings of the deep twin tracking network in dealing with occlusions. Using Kalman filtering can predict the target position in a completely occluded state, but traditional Kalman filtering performs poorly in nonlinear and complex environments or when the system model is not completely known. Therefore, deep features and Kalman filtering are combined to improve the state prediction and update steps of Kalman filtering using the powerful feature extraction capability of deep learning, and to improve the stability of continuous target tracking. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 The figure is a schematic diagram of the overall architecture of the dynamic single target long-time tracking system in the embodiment of the application.
[0023] Figure 2 A detailed flowchart of the dynamic single-target long-time tracking method in the embodiments of the present application. DETAILED DESCRIPTION
[0024] The technical solutions of the present application will be described in detail below with reference to the accompanying drawings, but the protection scope of the present application is not limited to the described embodiments.
[0025] Embodiment 1: The embodiments of the present application provide a dynamic single-target long-time tracking system, which comprises an airborne computer, a holder connected with the airborne computer, and a ground display device in communication connection with the airborne computer, for realizing the method shown in the figure. Figure 1 The airborne computer is preferably a domestic edge computing device containing a neural network processing unit (NPU) to meet the real-time computing requirements. The holder can be a three-axis or two-axis holder, which carries a visible light camera.
[0026] Based on the system, the embodiments of the present application provide a dynamic single-target long-time tracking method, the flow of which is shown in the figure, and the specific steps are as follows: Figure 2 S1: Start the holder camera and initialize it, use FFmpeg to perform real-time compression and encoding on the captured video stream, upload the encoded video stream to the airborne computer with artificial intelligence computing power; S2: Frame the tracking target in any frame of the video stream as a template image, initialize the lightweight deep twin tracking network according to the template image, and assign the position information of the target to the Kalman filter, while modifying the initial value and initial prediction of the Kalman filter; at the same time, save the template image.
[0027] S3: Perform feature extraction under the NPU architecture using a multi-thread asynchronous strategy, one branch of the twin tracking network is used to extract the features of the template image, and the other branch extracts the features of the search region; the search region is calculated according to the template region. The size of the tracking region is processed as 127*127, and the size of the search region is processed as 255*255.
[0028] S4: Calculate the correlation of the features from the two branches, including the shape, color, and texture features of the target, to obtain a response map; simultaneously predict the tracking target and the position region from the response map; S5: Perform aspect ratio penalty scoring and scale penalty scoring on the obtained possible target region to obtain the target region with the highest score as the tracking region of the predicted target.
[0029] S6: Occlusion judgment is performed on the obtained region, and color histogram feature similarity calculation is performed with the initialization region. When it is determined that the target is not occluded and the target exists in the video stream, tracking is performed using a depth twin network, and the obtained image target region is re-assigned to the Kalman filter.
[0030] S7: When the target is deformed, occluded, or interfered by an object, a Kalman prediction model of fused depth features is started, and the tracking region is predicted using the prediction model; and the color and running characteristics of the predicted target region are analyzed; if the predicted region and the initialization target region meet the expectation through analysis and comparison, it is determined that the prediction result is valid, and the next tracking is performed. The prediction result of the Kalman filter is updated to the template region of the depth twin tracking network, and when the template region is updated, the target region image is re-extracted using the state parameters (bounding box coordinates) of the Kalman filter as a new template of the twin network to adaptively handle occlusion or deformation.
[0031]
[0032] S9: When one of the above depth twin tracking network or Kalman prediction tracking result meets the expectation, the relative position of the target and the current field of view of the gimbal camera is calculated according to the position of the target in the image, and the gimbal servo drive controller is set according to the adaptive difference in the unmanned aerial vehicle gimbal coordinate system according to the actual target position and the expected target position, and the servo actuator is driven to adjust the rotation of the gimbal in the horizontal and vertical directions, so that the target is basically kept in the middle of the picture.
[0033] Specifically, in step S2, a tracking target is determined in any frame of the video stream by manual framing or an upstream target detection algorithm, and the target region is used as a template image T. Based on the template image, the following initialization operations are performed: initializing the depth twin tracking network: sending the template image T into the network, pre-computing and caching its depth features; assigning the target position information determined by framing to the Kalman filter as a trajectory prediction model to set its initial state value and initial prediction; saving the initial template image T for subsequent appearance similarity comparison.
[0034] In step S3, for each subsequent frame of the video stream, the following main tracking process is performed based on the depth twin network: Search area generation: Calculate the target center point coordinates according to the target position determined in the last frame, and generate a rectangular search area based on the center point coordinates of the target in the last frame, with the point as the center. The area covers the maximum possible displacement range of the target as much as possible. In particular, the algorithm dynamically adjusts the search area size in real time according to the target size change through the bilinear interpolation algorithm, ensuring that the best search range is maintained when the target is scaled.
[0035] Feature extraction: The deep twin network is used as the backbone network for feature extraction. The network includes two branches with shared weights: template branch: the input of template T is [1,3,127,127], and after passing through the backbone network, the template feature output1 [1,48,8,8] is obtained; search branch: the input of search area X is [1,3,255,255], and after passing through the backbone network feature extraction network, the image feature output2 [1,48,16,16] is obtained.
[0036] To realize the lightweight of the model, the deep separable convolution is used instead of the standard convolution in the feature extraction module to reduce the model parameter quantity and the calculation complexity; at the same time, the INT8 hybrid quantization is used to compress the model volume to several MB levels, which greatly reduces the memory occupation and the inference delay while ensuring the feature extraction capability. When deployed on an airborne computer, multiple independent RKNN instances are initialized, each of which is bound to a different NPU core (such as NPU_CORE_0, NPU_CORE_1, NPU_CORE_2), ensuring no resource competition between threads. This multi-thread cooperative inference architecture fully utilizes the parallel computing characteristics of NPU, and through task fragmentation and dynamic load balancing strategy of the template branch and the search branch, the feature extraction efficiency is significantly improved. Response map generation: The features obtained by the two branches are correlated, that is, each position on the feature map is multiplied by a simple element level to generate a new response map. The position of each pixel on the response map reflects the similarity of the template and the search area feature at that position, indicating the possible position of the target in the search area. Correlation operation is performed on it, and feature fusion is performed. As shown in formula (1): (1) In the formula, each position (i,j) in R can be mapped to the corresponding point (x,y) in the search area.
[0037] Regression and classification tasks: The deep twin network decomposes the target tracking into two tasks. The classification branch outputs the feature as a probability map of foreground and background by convolution operation on the response map, where each value in the foreground probability map represents the probability of the current coordinate appearing the tracking target, which is used to predict the class of each position in the response and output together with the center point to determine the position of the target. The regression branch outputs the feature as a target box scale map of the corresponding position by convolution operation, where each element has four scale information, which are the horizontal and vertical coordinates of the upper left corner of the target box and its length and width, which are used to determine the size of the target box. In order to prevent the target scale from changing too much, the algorithm uses a learning rate to smooth the target scale of adjacent frames. By combining the results of the two tasks, the most likely position of the target in the current frame is obtained through the foreground probability map, and then the corresponding position is located in the target box scale map to obtain the predicted target box scale information, and finally the target box is drawn to complete the tracking task of the target.
[0038] Candidate target box screening: The aspect ratio penalty score and the scale penalty score of the candidate target box are calculated, and the target box with the highest comprehensive score is selected as the prediction box of the twin tracking network, and the prediction box replaces the original template image T. This iteration is repeated to continuously predict the position and area of the target.
[0039] When the deep feature tracking fails due to long-time occlusion, the Kalman prediction based on the deep feature is used to track the target.
[0040] When the target is occluded by natural and man-made obstacles such as trees and buildings, and there is no target in the search area, since the new search area is generated based on the last tracking area, the original features of the target initialization are ignored, and the deep twin network cannot determine whether the original target exists and its position according to the response map A, resulting in the search area being more and more deviated, and the tracking fails. In order to overcome the above defects, an occlusion judgment module is used to determine whether to start the trajectory prediction module and update the template image feature, i.e. using the average peak correlation energy APCE and the color histogram hybrid mechanism to judge the occlusion state of the target. The trajectory prediction module collects the position information of the target during tracking and completes the position prediction of the target. The specific implementation method is: Average peak correlation energy APCE: During tracking, the deep twin tracking network outputs the response through the classification branch, the maximum value of the response can determine the position of the target, and the main peak state and fluctuation of the response can be used as the confidence of the current tracking result. When the target is not occluded, the response is a single peak, which represents good tracking effect. When the target is partially occluded, the side lobe of the response will increase. When the target is completely occluded, not only the side lobe of the response will increase, but also it will show a multi-peak form. Therefore, the fluctuation state of the response, i.e. the average peak correlation energy APCE of the response, is used as the preliminary basis for determining whether the tracking state of the target changes, and APCE is represented by formula (2): (2) In formula (2), A is the response, and A is the APCE value. , The maximum and minimum values of the response, respectively. When the target is not occluded, the APCE value of the response is larger; when the target is partially occluded until completely occluded, the fluctuation of the response gradually increases or even appears "multi-peak", and the corresponding APCE value gradually decreases. Therefore, the APCE is used to establish a criterion, when the APCE is greater than the threshold value , it is preliminarily determined that the target is not occluded; when the APCE is less than the threshold value , it is preliminarily determined that the target is occluded.
[0041] Color histogram determination: considering that the depth twin tracking network ignores the initial features of the tracked target, and the color features of the target remain basically unchanged during the entire tracking process. Therefore, the color histogram can be used as an important indicator of target loss or whether the target is occluded. The calculation method is as follows: First, the predicted region size is resampled to the initial target tracking region size, then the number of times each R, G, B primary color appears in the initial target tracking region and the resampled tracking region is counted, and a histogram of the frequency of the appearance of the three primary colors in the region is constructed, that is, the histogram feature of the target tracking region, then the cosine similarity of each channel is calculated, and the similarity of the two regions is obtained by weighting. First, the number of times each pixel value of 0-255 of each primary color appears is obtained, obtaining a 255-dimensional vector, which is the fingerprint of the histogram. After obtaining the fingerprints of the initial target tracking region and the resampled predicted tracking region, the similarity between the two vectors is calculated to obtain the similarity of the two regions. Assuming that and are two n-dimensional vectors, the similarity of P and Q is represented by formula (3), and finally the similarity of each primary color is weighted to obtain the similarity score of the two regions. If the similarity score is greater than 1, it is determined that the target still exists in the search region, and if it is less than 1, it is determined that the target is occluded or lost (3) The target similarity coefficient is obtained by weighting the average peak value correlation energy and the color histogram, which is represented by formula (4). When the coefficient is less than a certain threshold value, it is determined that the target is occluded and enters the next step.
[0042] (4) In the formula, represents the weight of the average peak value correlation energy, represents the color histogram similarity score.
[0043] Trajectory prediction and update based on Kalman filter When the target is completely occluded, the target features in the search region are completely lost. Due to the estimation of the search region in the t-th frame according to the target position in the t-1-th frame, the tracker is difficult to recapture the target due to the loss of target features caused by occlusion. Considering that the time interval between adjacent frames in the tracking process is in the order of milliseconds, it can be considered that the target maintains a uniform straight-line motion in the horizontal and vertical directions in such a short time; therefore, when the target is completely occluded (i.e. APCE < a and color feature histogram < 1), the prediction of the target position in the completely occluded state can be realized by using Kalman filtering.
[0044] The traditional Kalman filter is the optimal estimator of a linear system, which relies on an accurate system model and noise statistical information. However, its performance will decrease in a nonlinear and complex environment or when the system model is not completely known. Depth features usually have good representation ability for objects in images, and can capture the shape, texture and other information of objects. By combining depth features with Kalman filtering, the powerful feature extraction capability of deep learning can be used to improve the state prediction and update steps of Kalman filtering, thereby improving the accuracy and robustness of target tracking.
[0045] In the specific tracking process, the implementation process of Kalman filtering is as follows: a. Depth feature extraction Let be the state vector of the target, be the observation value at time t (corresponding to the position in the image). The depth feature extracted by an L-layer depth network can be represented as: (5) where is the depth feature vector, is the transformation of the layer (such as convolution, full connection, attention mechanism, etc.).
[0046] b. State prediction State prediction follows the traditional Kalman filter, which uses a dynamic model to estimate the state at the next time point. Let be the prior state estimate based on the observation at time t-1, then: (6) F is the state transition matrix, corresponding to the motion mode of the target; B is the control input matrix, is the control vector. At the same time, the predicted state covariance matrix , that is: (7) where Q is the process noise covariance matrix.
[0047] c. Measurement update When a new observation is available, first calculate the measurement residual and the corresponding covariance matrix: (8) (9) where H is the observation matrix, representing the correspondence between the target real state value and the observation value; R is the observation noise covariance matrix. Calculate the Kalman gain : (10) Update the state using the Kalman gain: (11) Update the state covariance matrix: (12) d. Depth feature fusion In order to fuse the depth feature, modify the observation model H, which not only contains the position information, but also contains the depth feature , this method extends the observation vector including depth features as shown in equation (13), and adjust the observation matrix H equation (14) and the observation noise covariance matrix R equation (15) accordingly, so as to be able to process the enhanced observation vector.
[0048] (13) (14) where, Through back propagation online update, is the kernel mapping function.
[0049] (15) where, is the position noise, dynamic adjustment, update according to the depth feature matching quality.
[0050] Through the above operation, the Kalman filter predicts the position information of the target in the occlusion state according to the previous motion model; calculate the color feature histogram of the predicted region and the original frame target region respectively, and get the color similarity score. When the score is greater than a certain threshold, the next round of tracking is continued. When the Kalman filter is used continuously to predict the target for more than 10 frames, the next step is started.
[0051] Expansion search strategy: when the target moves fast, is long-term occluded, or has large scale variation in the video frame, the depth feature tracking and Kalman prediction may both fail due to lack of effective observation information (i.e. no target in the picture for a long time), and the robustness and accuracy of the tracking algorithm can be improved through the expansion search strategy.
[0052] Define the state vector of the target , which contains the state information of the target at time step k, including position, velocity, direction, etc. For target tracking in a two-dimensional plane, the state vector can be simply represented as: (16) wherein, x and y represent the position coordinates of the target, vx and vy are the velocity components along the x-axis and the y-axis.
[0053] When the t-th frame is determined to be occluded and the target is still not recaptured in the t+m-th frame, the expansion search strategy is entered. Expansion is performed according to the moving speed of the target before the target is lost and the pixel length occupied by the target in the image. The dynamic expansion search radius is shown in formula (17). When the target is lost, the search area is expanded according to the speed and direction of the target before the target is lost and the number of lost frames, until the target appears again in the search frame. If the target is not recaptured for more than 3s, the expansion search strategy is exited and it is considered that the tracking is ended.
[0054] (17) wherein, R represents the search area radius of the next frame (unit: pixel), v represents the average moving speed of the target, L represents the length of the target at the t-th moment as the basis for the target expansion radius, a decay coefficient, a weight parameter for smoothing the speed change. The greater the value, the smaller the influence of the historical speed.
[0055] The present application also includes gimbal servo control: initialization and calibration, setting the basic parameters of the gimbal to ensure that the hardware can work normally. Before starting tracking each time, the gimbal is automatically calibrated to zero to ensure that its pointing direction in the static state is preset.
[0056] Target detection and positioning: the pixel coordinates of the upper left and lower right corner points of the target frame obtained according to the above tracking algorithm are obtained. The target pixel area is calculated through the frame corner points , and the pixel coordinates of the center point of the target detection frame are calculated.
[0057] The angle offset of the target relative to the gimbal is calculated. The initial position of the gimbal is considered, and is compared with the initial tracking frame area When , it is considered that the target distance is far from the camera plane distance, and a smaller angle difference value is adopted according to the motion of the target detection box center point pixel coordinates When , it is considered that the target distance is close to the camera plane distance, and the angle difference value is adaptively adjusted by formula (16) . Wherein, is an adaptive control coefficient, which is set according to the specific application scene.
[0058] (18) Output control signal: according to the calculated angle error, send angle control, drive the holder to rotate. After receiving the signal, the holder adjusts the posture, real-time follows the target and basically keeps the target in the center area of the picture.
[0059] As described above, although the present application has been shown and described with reference to certain preferred embodiments thereof, it is to be understood that such is by way of illustration and not of limitation. Various substitutions and changes can be made without departing from the spirit of the application defined in the following claims.
Claims
1. A dynamic single-target long-time tracking method, characterized in that, The method comprises the following steps: S1: obtaining a video stream for shooting a target, and determining a tracking target in any frame of the video stream as a template image; S2: initializing a deep twin tracking network and a trajectory prediction model based on the template image; S3: using the deep twin tracking network to perform feature extraction and correlation calculation on the template image and a search region of a subsequent frame of the video stream, to generate a response map containing target position information, and to determine a preliminary tracking result therefrom; S4: determining a real-time tracking state of the target based on the preliminary tracking result, to distinguish whether the target is in a normal tracking state or an occlusion or loss state; S5: when it is determined that the target is in the normal tracking state, taking the preliminary tracking result as tracking target information, and updating a state of the trajectory prediction model using target position information in the tracking target information; when it is determined that the target is in the occlusion or loss state, starting the trajectory prediction model, predicting a position of the target in a current frame according to historical motion information of the target, taking the predicted position as the tracking target information, and updating a template region of the deep twin tracking network.
2. The dynamic single-target long-time tracking method according to claim 1, characterized in that, The S2 comprises: determining a target region as the template image by frame selection in any frame of the video stream; initializing the deep twin tracking network according to the template image; and assigning target position information determined by the frame selection to a Kalman filter as the trajectory prediction model, to set an initial state thereof.
3. The dynamic single-target long-time tracking method according to claim 1, wherein, The deep twin tracking network decomposes target tracking into a classification task and a regression task; the classification task generates a foreground probability map by performing convolution operation on the response map, to determine a position of the target; and the regression task generates a target frame size map by performing convolution operation on the response map, to determine a size of a target frame. A candidate target frame generated according to the foreground probability map and the target frame size map is subjected to aspect ratio penalty scoring and size penalty scoring, and a target frame with the highest comprehensive score is selected as the preliminary tracking result.
4. The dynamic single-target long-time tracking method according to claim 1, wherein, The deep twin tracking network comprises a first branch and a second branch; the first branch extracts features of the template image; and the second branch extracts features of a search region, which is calculated according to the template image.
5. The dynamic single-target long-time tracking method according to claim 1, wherein, The determination of the occlusion or loss state comprises: the occlusion or loss state is determined comprehensively by average peak correlation energy (APCE) and similarity of a color histogram; the APCE is used to determine an occlusion state of the target according to fluctuation state of a response map; and the color histogram is used to determine whether the target exists by calculating cosine similarity between a predicted region and an initial target region.
6. The dynamic single-target long-time tracking method according to claim 5, wherein, The average peak correlation energy (APCE) is calculated by the following formula: wherein f is the APCE value, A is the response, , is the maximum and minimum value of the response, respectively. The calculated APCE value is compared with a preset threshold value, to serve as a preliminary basis for determining whether the target is occluded; The color histogram similarity calculation includes: resampling the current predicted target region to be consistent with the size of the template image; for the template image and the resampled predicted region, respectively, count the color distribution of the R, G, and B channels to generate respective color histogram vectors; and by calculating the cosine similarity between the two color histogram vectors, the appearance similarity of the two is quantified.
7. The dynamic single-target long-time tracking method according to claim 1, wherein, The trajectory prediction model adopts a Kalman filter fused with depth features to predict the position of the target in the current frame, and extracts a depth feature vector of the target in the current frame by the depth twin tracking network , and a position observation value at time t are combined to form an extended observation vector , and the observation matrix H and the observation noise covariance matrix R of the Kalman filter are adjusted accordingly to process the enhanced extended observation vector.
8. The dynamic single-target long-time tracking method according to claim 7, characterized in that, The method further comprises an expansion search strategy: when the trajectory prediction model still fails to restore the tracking state to normal tracking within a preset number of consecutive frames, the expansion search strategy is started, and the expansion search strategy is calculated by a formula dynamically calculating a search area radius of a next frame wherein, is a target instantaneous speed, is an average moving speed, SL is a target side length, is a decay coefficient, and when a tracking target is not detected for more than a preset time, the tracking is terminated.
9. The dynamic single-target long-time tracking method according to claim 1, wherein, The method further includes a gimbal servo control step: according to the position coordinates of the target in the image frame determined by the tracking target information, the deviation of the target from the image center is calculated, and a corresponding control signal is generated to drive the gimbal to rotate, so as to keep the target in the center of the picture.
10. A system for implementing the dynamic single-target long-time tracking method of claim 1, characterized by, The method comprises: an airborne computer, a gimbal connected to the airborne computer, and a ground display device in communication connection with the airborne computer; wherein the airborne computer is configured with: a main tracking module configured to track a target in a video stream using a deep twin tracking network to generate a preliminary tracking result; a state determination module connected to the main tracking module and configured to determine the real-time tracking state of the target based on the preliminary tracking result and output a determination signal; a trajectory prediction module configured to predict the current position of the target according to its historical motion information when the determination signal indicating that the target is in an occluded or lost state is received; a control unit connected to the state determination module, the main tracking module, and the trajectory prediction module, and configured to select the preliminary tracking result or the position predicted by the trajectory prediction module as the final tracking result according to the determination signal a communication unit for pushing the tracking result to the ground display device through a data link.
Citation Information
Cited By
Method and system for continuously measuring target sight angle of quadruped robot
CN121746454A
A method and system for continuous measurement of target line-of-sight angle of a quadruped robot
CN121746454B
Unmanned aerial vehicle tracking method based on continuous local frame matching
CN122335909A