Environment-adaptive gimbal target tracking method and device, gimbal and storage medium
Patent Information
- Application Number
- CN202611007893.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-07-08
AI Technical Summary
[0003]现有基于云台的目标跟踪方法通常采用模板匹配或相关滤波框架:以初始帧中的目标区域作为参考模板,在当前帧的候选区域内搜索与模板最相似的图像块作为跟踪结果,但在实际动态场景中,当目标发生快速形变、尺度变化、旋转或受到局部遮挡时,固定的模板特征难以准确表征目标的实时外观状态,尤其在云台主动转动导致图像边缘区域产生运动模糊、光学畸变或背景干扰加剧的情况下,在线更新机制极易将包含背景噪声的不可靠帧引入模板,容易导致图像特征丢失,从而导致跟踪灵敏度下降、准确率不足
1、通过引入退化风险分布图,将环境对目标特征的干扰进行空间量化,使跟踪系统具备环境感知能力。通过在当前帧图像中确定退化风险分布图,该分布图能够表征目标所在区域各位置的特征丢失风险,本质是将图像空间划分为“高可靠性区域”与“高风险区域”,为后续的位置偏差计算和模板更新决策提供了先验的环境信息,从而避免算法在未知干扰下盲目工作;
Smart Images

Figure CN122510307B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of gimbal technology, and in particular to an environment-adaptive gimbal target tracking method, device, gimbal, and storage medium. Background Technology
[0002] As an electromechanical device that can flexibly control the camera's posture, a pan-tilt head, when combined with a target tracking algorithm, can actively adjust the field of view to keep the target in the center of the frame.
[0003] Existing gimbal-based target tracking methods typically employ template matching or correlation filtering frameworks: using the target region in the initial frame as a reference template, and searching for the most similar image patch within the candidate region of the current frame as the tracking result. However, in real-world dynamic scenes, when the target undergoes rapid deformation, scale changes, rotation, or is partially occluded, fixed template features are difficult to accurately represent the real-time appearance of the target. Especially when the gimbal actively rotates, causing motion blur, optical distortion, or increased background interference in the image edge region, the online update mechanism is prone to introducing unreliable frames containing background noise into the template, which can easily lead to the loss of image features, resulting in decreased tracking sensitivity and insufficient accuracy. Summary of the Invention
[0004] To overcome the shortcomings of the prior art, this application provides an environment-adaptive gimbal target tracking method, device, gimbal, and storage medium, which improves the robustness and stability of target tracking in complex scenes through image processing algorithm optimization.
[0005] The first aspect of this application provides an environment-adaptive gimbal target tracking method, the method comprising:
[0006] Acquire the first frame image and the user-specified initial target bounding box, crop the image region corresponding to the initial target bounding box from the first frame image as the target reference template, and extract the initial feature vector of the target reference template; Acquire the current frame image, and determine a degradation risk distribution map based on the current frame image to characterize the risk of feature loss in the target's region; In the current frame image, a candidate image block of the same size as the target reference template is cropped with the target position of the previous frame as the center, and the current frame feature vector of the candidate image block is extracted. The target's current predicted position in the current frame is obtained by matching the current frame feature vector with the initial feature vector. Analyze the degradation risk distribution within a preset range near the current predicted location of the target in the degradation risk distribution map, and determine the location with the lowest risk within the preset range; Calculate the first deviation between the current predicted position of the target and the geometric center of the current frame image, and calculate the second deviation between the current predicted position of the target and the position with the lowest risk; A gimbal control signal is generated based on the first deviation and the second deviation, and the gimbal is driven to track the target based on the gimbal control signal. Calculate the confidence index of the current frame matching result, and determine whether the current tracking state is a reliable tracking state or a tracking failure state based on the consistency between the confidence index and the target motion trajectory. When the tracking is determined to be in a failure state, the target re-acquisition mechanism is activated to retrieve the target; when the tracking is determined to be in a reliable state, it is determined whether the preset update conditions are met. If the preset update conditions are met, the reference feature template is updated using the tracking results of the current frame; otherwise, the reference feature template is not updated.
[0007] In an optional implementation, the step of matching the current frame feature vector with the initial feature vector to obtain the target's current predicted position in the current frame includes: Calculate the correlation between the current frame feature vector and the initial feature vector to generate a two-dimensional response map; The peak position of the two-dimensional response map is taken as the current predicted position of the target.
[0008] In an optional implementation, determining whether the current tracking state is a reliable tracking state or a tracking failure state based on the consistency between the confidence index and the target motion trajectory includes: The peak-to-side-lobe ratio of the two-dimensional response plot is calculated as the confidence index; The target position in the current frame is predicted using a Kalman filter to obtain the predicted target position; The Euclidean distance between the current predicted position of the target and the predicted position of the target is calculated as a consistency measure of the target's trajectory. When the peak sidelobe ratio is greater than a first preset threshold and the Euclidean distance is less than a second preset threshold, it is determined to be a reliable tracking state; When the peak sidelobe ratio is less than or equal to a first preset threshold, or the Euclidean distance is greater than or equal to a second preset threshold, the tracking is determined to be in a failure state.
[0009] In an optional implementation, when the tracking is determined to be in a failure state, the method further includes: By fitting a motion model using the historical target position data of the most recent N frames, the candidate positions of the target within a preset search area are predicted. Multi-scale template matching is then performed using the target reference template within the preset search area to find the best matching position as the new target position, and the tracking is reset to a reliable state. When the tracking status is determined to be reliable, it is determined whether the preset update conditions are met at the same time. The preset update conditions include a first update condition and a second update condition. The first update condition is that the peak-to-sidelobe ratio is greater than a third preset threshold, and the second update condition is that the ratio of the main peak to the secondary peak of the two-dimensional response graph is greater than a fourth preset threshold. When both the first update condition and the second update condition are met, the initial feature vector is updated with a preset learning rate.
[0010] In an optional implementation, determining the degradation risk distribution map characterizing the risk of feature loss in the target's region based on the current frame image includes: Perform semantic segmentation on the current frame image to obtain the semantic segmentation result; Using the target position of the previous frame as the center, a region of interest is generated, and the local variance of each pixel in the region of interest is calculated as a texture richness index. Based on the semantic segmentation results and the texture richness index, a degradation risk distribution map of the same size as the region of interest is generated; each pixel value in the degradation risk distribution map represents the probability that the current pixel location will lose features within a preset number of frames in the future.
[0011] In an optional implementation, extracting the current frame feature vector of the candidate image patch includes: Multiple preset types of features are extracted from the candidate image blocks to obtain multiple candidate feature components; Extract the center point value of the degradation risk distribution map as the degradation risk value of the target's current location; The fusion weights corresponding to the multiple candidate feature components are dynamically determined based on the degradation risk value. The multiple candidate feature components are weighted and summed according to their respective fusion weights to obtain the current frame feature vector.
[0012] In an optional implementation, the multiple preset types of features include histogram of oriented gradients features, color naming features, and shallow convolutional neural network features. The step of dynamically determining the fusion weights corresponding to the multiple candidate feature components based on the degradation risk value includes: When the degradation risk value is greater than the first risk threshold, the weight of the directional gradient histogram feature is reduced, and the weights of the color naming feature and the shallow convolutional neural network feature are increased. When the degradation risk value is less than the second risk threshold, the balanced weight of each preset type feature is maintained; When the degradation risk value is between the second risk threshold and the first risk threshold, the fusion weight is determined by linear interpolation. Wherein, the first risk threshold is greater than the second risk threshold.
[0013] A second aspect of this application provides an environment-adaptive gimbal target tracking device, the device comprising: An initialization module is used to acquire a first frame image and a user-specified initial target bounding box, crop the image region corresponding to the initial target bounding box from the first frame image as a target reference template, and extract the initial feature vector of the target reference template. The degradation risk perception module is used to acquire the current frame image and determine a degradation risk distribution map based on the current frame image to characterize the risk of feature loss in the target's region. The candidate extraction module is used to crop out a candidate image block of the same size as the target reference template from the target position of the previous frame in the current frame image, and extract the current frame feature vector of the candidate image block. The position prediction module is used to perform matching calculations based on the current frame feature vector and the initial feature vector to obtain the target's current predicted position in the current frame; The risk analysis module is used to analyze the degradation risk distribution within a preset range near the current predicted location of the target in the degradation risk distribution map, and to determine the location with the lowest risk within the preset range; The deviation calculation module is used to calculate the first deviation between the current predicted position of the target and the geometric center of the current frame image, and to calculate the second deviation between the current predicted position of the target and the position with the lowest risk. The gimbal control module is used to generate a gimbal control signal based on the first deviation and the second deviation, and drive the gimbal to track the target based on the gimbal control signal; The status determination module is used to calculate the confidence index of the matching result of the current frame, and determine whether the current tracking status is a reliable tracking status or a tracking failure status based on the consistency between the confidence index and the target motion trajectory. The template update module is used to initiate a target re-acquisition mechanism to retrieve the target when the tracking is determined to be in a failed state; when the tracking is determined to be reliable, it determines whether a preset update condition is met. If the preset update condition is met, the reference feature template is updated using the tracking result of the current frame; otherwise, the reference feature template is not updated.
[0014] A third aspect of this application provides a gimbal, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the environment-adaptive gimbal target tracking method.
[0015] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described environment-adaptive gimbal target tracking method.
[0016] In summary, the environment-adaptive gimbal target tracking method, apparatus, gimbal, and storage medium provided in this application have at least one of the following beneficial effects: 1. By introducing a degradation risk distribution map, the interference of the environment on target features is spatially quantified, enabling the tracking system to have environmental awareness. By determining the degradation risk distribution map in the current frame image, this distribution map can characterize the feature loss risk at each location in the target area. Essentially, it divides the image space into "high reliability areas" and "high risk areas," providing prior environmental information for subsequent position deviation calculation and template update decisions, thereby avoiding the algorithm working blindly under unknown interference. 2. By jointly considering the first deviation between the target's current predicted position and the image's geometric center, and the second deviation between the target's current predicted position and the position with the lowest risk, a gimbal control signal is generated, achieving a dynamic balance between tracking accuracy and environmental adaptability. Specifically, after obtaining the target's current predicted position, the degradation risk distribution within a preset range near that position is further analyzed, and the position with the lowest risk within that range is identified. Simultaneously, the deviation between the predicted position and the geometric center (driving the target to center) and the deviation between the predicted position and the position with the lowest risk (driving the target away from interference) are calculated. By combining these two deviations to generate the gimbal control signal, the gimbal's trajectory is no longer passively following the target, but actively balancing "maintaining the target's center" and "avoiding highly degraded areas." This proactively guides the target to remain in image areas with stable features and high signal-to-noise ratios during tracking, fundamentally reducing the risk of feature loss and template contamination caused by environmental degradation. 3. By constructing a dual judgment mechanism based on confidence index and motion trajectory consistency, a strict distinction is made between reliable tracking and tracking failure states, and the baseline feature template is conditionally updated only in the reliable tracking state. Specifically, the confidence index of the current frame matching result is calculated, which directly reflects the credibility of the current frame feature matching; at the same time, the consistency analysis of the target motion trajectory is combined to determine whether the target position change conforms to the laws of physical motion, and the two are used to comprehensively determine the tracking status. When the tracking is determined to be a failure, the recapture mechanism is immediately activated to avoid erroneous results contaminating the template; when the tracking is determined to be reliable, the template is not updated immediately, but further judgment is made as to whether the preset update conditions are met (for example, the average degradation risk of the target area in the current frame is lower than a threshold, or the cumulative change in the target appearance exceeds a certain level). Only when both "reliable tracking" and "update conditions" are met is the baseline feature template updated using the current frame result. This progressive judgment logic ensures that only high-quality, low-risk frames that truly reflect the new appearance of the target can participate in template evolution, thereby strictly excluding background noise and degradation interference from the template update process, maintaining the purity and stability of template features, and thus improving the sensitivity and accuracy of long-term tracking. 4. A closed loop of "perception-avoidance-judgment-update / recovery" is formed. The degradation risk distribution map guides the gimbal control signal to direct the target to a low-risk area, reducing the probability of feature loss. At the same time, the confidence and motion consistency judgment ensures that the template is only updated when the target is in a low-risk area and the match is good, further consolidating the template quality. Once the target enters a tracking failure state due to unavoidable reasons such as sudden occlusion, the re-acquisition mechanism can independently retrieve the target from the existing template, avoiding permanent tracking loss. Thus, this application significantly improves the adaptability and long-term stability of the gimbal target tracking system in complex dynamic scenarios from three aspects: spatial avoidance of environmental interference, quality admission of template updates, and robust recovery after failure. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating an environment-adaptive gimbal target tracking method according to an embodiment of this application; Figure 2 This is a functional block diagram of an environment-adaptive gimbal target tracking device shown in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a gimbal shown in an embodiment of this application. Detailed Implementation
[0018] The present application will be further described below with reference to the accompanying drawings and embodiments.
[0019] The following will clearly and completely describe the concept, specific structure, and resulting technical effects of this application in conjunction with embodiments and accompanying drawings, so as to fully understand the purpose, features, and effects of this application. Obviously, the described embodiments are only a part of the embodiments of this application, not all of them. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the scope of protection of this application. Furthermore, all connections / linkages involved in the patent do not simply refer to direct contact between components, but rather to the ability to form a better connection structure by adding or reducing connecting accessories according to specific implementation conditions. The various technical features in this application can be combined interactively without contradicting each other.
[0020] To facilitate understanding of the inventive concept of this application, the following embodiments illustrate a gimbal-based target tracking method that adapts to the environment, using a gimbal as the executing entity.
[0021] Reference Figure 1 The diagram shown is a flowchart illustrating an environment-adaptive gimbal target tracking method according to an embodiment of this application. The environment-adaptive gimbal target tracking method includes the following steps.
[0022] S11, acquire the first frame image and the user-specified initial target bounding box, crop the image region corresponding to the initial target bounding box from the first frame image as the target reference template, and extract the initial feature vector of the target reference template.
[0023] In some embodiments, before tracking begins, a video sequence transmitted back from the PTZ is acquired, and the first frame image I1 of the video sequence is obtained. The user specifies the initial target bounding box in I1 through a human-computer interaction interface. This box precisely selects the target object to be tracked (such as pedestrians, vehicles, etc.) and its initial position. .in, The coordinates of the center of the target bounding box and These are the width and height of the target bounding box, respectively. Cropped from the current frame image I1. The corresponding image region is used as the target reference model T.
[0024] Furthermore, various preset types of features are extracted from the target baseline model T. In this embodiment, the gimbal extracts three complementary features in parallel: HOG features (Histogram of Oriented Gradients), CN features (Color Naming Features), and shallow CNN features.
[0025] HOG features are used to describe the shape and edge information of the target. During extraction, the template image is divided into several cell units, a gradient orientation histogram is calculated within each cell unit, and adjacent cell units are grouped together and normalized to obtain the final HOG feature vector. .
[0026] CN features are used to describe the color distribution information of the target. The template image is mapped from the RGB space to an 11-dimensional color name space, with each pixel corresponding to a probability distribution of a color name, resulting in the CN feature vector. This feature exhibits a certain degree of robustness to changes in illumination.
[0027] Shallow CNN features are used to describe the semantic texture information of the target. A pre-trained convolutional neural network (such as the first three convolutional layers of VGG-16) is employed. The template image is input into the network, and the feature maps output from the intermediate layers are extracted as the shallow CNN feature vectors. .
[0028] Next, the HOG features, CN features, and shallow CNN features are concatenated or weighted to obtain the initial fused feature vector. (Also known as the initial feature vector). In this embodiment, the gimbal can calculate the initial fused feature vector using a weighted summation method. : In the initial stage, the weights of the three features can be taken as a balanced value, for example... =0.4, =0.3, =0.3.
[0029] S12, acquire the current frame image, and determine a degradation risk distribution map based on the current frame image to characterize the risk of feature loss in the target area.
[0030] In some embodiments, during target tracking, the gimbal acquires the current frame image It (t≥2) and uses this current frame image as the analysis object to perform environmental perception and degradation risk analysis. The gimbal performs scene understanding on the current frame image, extracts texture distribution information and illumination distribution information of the area surrounding the target, and comprehensively evaluates the stability of available visual features at each image location in the near future, thereby generating a degradation risk distribution map corresponding to the target's location. This degradation risk distribution map serves as the environmental basis for subsequent dynamic feature routing and gimbal look-ahead control, enabling the pre-emptive perception of risk areas that the target will face before feature loss actually occurs.
[0031] In an optional implementation, determining the degradation risk distribution map characterizing the risk of feature loss in the target's region based on the current frame image includes: Perform semantic segmentation on the current frame image to obtain the semantic segmentation result; Using the target position of the previous frame as the center, a region of interest is generated, and the local variance of each pixel in the region of interest is calculated as a texture richness index. Based on the semantic segmentation results and the texture richness index, a degradation risk distribution map of the same size as the region of interest is generated; each pixel value in the degradation risk distribution map represents the probability that the current pixel location will lose features within a preset number of frames in the future.
[0032] In some embodiments, the gimbal performs fast semantic segmentation on It, which can employ a lightweight semantic segmentation network (such as ENet, ICNet, etc.) to divide the image into different region types, including but not limited to: high-texture regions (such as grass, trees, building textures), low-texture regions (such as sky, water surface, smooth walls), abnormal lighting regions (overexposed areas, shadow areas), and dynamic background interference regions (such as moving non-target objects). The target position in the previous frame... Generate dimensions centered on The Region of Interest (ROI) is the area within which the target may be present in the current frame.
[0033] The local variance of each pixel within the ROI is calculated as a texture richness index. For pixel (i,j), a k×k neighborhood window centered on it is taken (k=5 in this embodiment), and the variance of the pixel grayscale values within the window is calculated. The larger the local variance, the richer the texture of the region, and the less likely the features are to be lost; the smaller the local variance, the sparser the texture of the region, and the higher the risk of feature loss.
[0034] Combining semantic segmentation results and texture richness metrics, a degradation risk heatmap Rt of the same size as the ROI is generated. Specifically, for each pixel (i,j) in the ROI, its corresponding degradation risk value is determined according to the following rules. : If the pixel belongs to the "low-texture area" or "abnormal lighting area" in the semantic segmentation result, its base risk value is increased.
[0035] Normalization is performed based on local variance; the lower the local variance, the higher the risk value.
[0036] Taking into account the two factors mentioned above, the degradation risk value of each pixel is... A larger value indicates a higher probability of feature loss at that location within the next 2-3 frames. The degradation risk heatmap Rt can be represented as a matrix: .
[0037] Through the aforementioned optional implementation methods, the distribution of visual feature degradation risk in the target area can be characterized at the pixel-level granularity, providing accurate and quantitative environmental perception data for subsequent dynamic feature routing and gimbal look-ahead control. Compared to traditional methods that rely solely on the confidence level of a single frame response map for post-event judgment, this approach enables early perception and quantitative assessment of environmental changes, obtaining risk warning information before feature loss occurs. This creates conditions for proactively adjusting feature fusion strategies and gimbal motion decisions in subsequent steps.
[0038] S13, in the current frame image, a candidate image block of the same size as the target reference template is cropped with the target position of the previous frame as the center, and the current frame feature vector of the candidate image block is extracted.
[0039] In some embodiments, in the current frame image It, the target position of the previous frame Centered on the target reference template T, a candidate image block Pt of the same size as the target reference template T is cropped. Using the same implementation method as step S11 above, various preset types of features are extracted from Pt to obtain multiple candidate feature components, namely HOG features, CN features, and shallow CNN features. , and .
[0040] Extract the center point value of the degradation risk heatmap Rt as the degradation risk value of the target's current location. Because the center of the ROI is the same as the target position in the previous frame. Alignment, therefore the center point of Rt corresponds to the position of the target in the current frame.
[0041] Based on degradation risk value The fusion weights of the three candidate feature components are dynamically calculated. In this embodiment, the gimbal employs a piecewise linear weight allocation strategy: when > When the first risk threshold (e.g., set to 0.7) is reached (high risk: target is in a low-texture or abnormally lit region), actively reduce the dependence on texture gradient features (HOG) and enhance the dependence on color features (CN) and semantic features (CNN). =0.1, =0.6, =0.3.
[0042] when When the risk threshold is less than the second risk threshold (e.g., set to 0.3) (low risk: target is in a textured area), maintain a balanced weighting and take... =0.4, =0.3, =0.3.
[0043] When 0.3≤ When the weights are ≤0.7 (medium risk), the weights of the three candidate feature components are determined by linear interpolation, that is, a linear transition is made between the two sets of weights mentioned above.
[0044] The three features are weighted and summed according to dynamically calculated fusion weights to obtain the fused feature vector of the current frame. (Also known as the current frame feature vector): .
[0045] By using a dynamic weight allocation method, "predictive switching" of features is achieved, which proactively adjusts the feature fusion strategy before environmental degradation occurs, avoiding the use of feature channels that are about to degrade.
[0046] S14, perform matching calculations based on the current frame feature vector and the initial feature vector to obtain the target's current predicted position in the current frame.
[0047] In some embodiments, after extracting the feature vector of the current frame, the gimbal performs a matching calculation between the current frame feature vector and the initial feature vector saved in step S11, determining the target's position in the current frame by measuring the similarity between the two. Specifically, the gimbal determines a search area centered on the target's position in the previous frame, and performs a sliding window-style correlation calculation between the current frame feature vector and the initial feature vector within this area to generate a response map reflecting the degree of matching between each candidate position and the target template. The peak position of this response map is the optimal predicted position of the target in the current frame. Through the above matching calculation, the gimbal can achieve frame-by-frame target localization in a continuous video stream, providing a precise position reference for subsequent gimbal control.
[0048] In an optional implementation, the step of matching the current frame feature vector with the initial feature vector to obtain the target's current predicted position in the current frame includes: Calculate the correlation between the current frame feature vector and the initial feature vector to generate a two-dimensional response map; The peak position of the two-dimensional response map is taken as the current predicted position of the target.
[0049] In some embodiments, the gimbal uses a preset correlation filter (e.g., a kernelized correlation filter (KCF)) for filtering operations. The core idea of the KCF algorithm is to perform dense sampling of the target region through cyclic shifting, and to utilize the property that the cyclic matrix is diagonalizable in the Fourier domain to transform the training and detection process from the time domain to the frequency domain, significantly reducing computational complexity.
[0050] Initial feature vectors based on the target baseline template T, either in the initial frame or after each model update. Ridge regression is used to train the correlation filter. Let the training sample be x, and the desired output be a Gaussian response y (with the peak located at the center of the target). Then the objective function of ridge regression is: ; in, λ is the regression function, and λ is the regularization parameter (in this embodiment, λ=0.0001) used to prevent overfitting.
[0051] By performing dense sampling of the base samples through cyclic shifting, a cyclic matrix X is generated. Utilizing the property that the cyclic matrix can be diagonalized through Fourier transform, the frequency domain closed-form solution to the above optimization problem is: ; in and Let x and y be the discrete Fourier transforms, respectively, and ⊙ represent the Hadamard product (element-wise product). Indicates complex conjugation.
[0052] For nonlinear regression problems, a kernel function is introduced. The features are mapped to a high-dimensional space. In this embodiment, the gimbal uses a Gaussian kernel: ; Where σ is the bandwidth parameter of the Gaussian kernel. After introducing the kernel trick, the solution in the dual space is: ; in, This is the Fourier transform of the first row of the kernel matrix.
[0053] For the current frame, the target position of the previous frame Centered on the training samples, candidate regions z are extracted, and their kernel correlation with the training samples is calculated. Then the responses of all candidate regions are: ; By using inverse Fourier transform Transforming back to the time domain yields a two-dimensional response map St of the same size as the candidate region. The peak position of the response map St is the predicted position of the target in the current frame. For ease of distinction, this is referred to as the target's current predicted position.
[0054] Through the aforementioned optional implementation methods, the gimbal utilizes a kernel correlation filter to efficiently calculate the correlation between the current frame feature vector and the initial feature vector in the frequency domain, achieving high-precision target localization results while ensuring real-time tracking. The KCF algorithm achieves dense sampling through cyclic shifting, avoiding the huge computational overhead of exhaustive search, enabling this method to be deployed on embedded platforms with limited computing power. Simultaneously, by introducing a Gaussian kernel function to map linear features to a nonlinear high-dimensional space, the filter's ability to discriminate changes in target appearance is enhanced, providing reliable position input for subsequent tracking status verification and gimbal control.
[0055] S15, Analyze the degradation risk distribution within a preset range near the current predicted position of the target in the degradation risk distribution map, and determine the position with the lowest risk within the preset range.
[0056] In some embodiments, the gimbal analyzes the current predicted position of the target in the degradation risk heatmap Rt. The distribution of degradation risk within a preset range. The preset range is defined as follows: Centered on, with dimensions of A rectangular region is defined. Within this region, the coordinates of the point with the minimum degradation risk are searched. This point is the location with the lowest risk of feature loss in the nearby area within the preset range, and is called the lowest risk location.
[0057] S16, calculate the first deviation between the current predicted position of the target and the geometric center of the current frame image, and calculate the second deviation between the current predicted position of the target and the position with the lowest risk.
[0058] S17, Generate a gimbal control signal based on the first deviation and the second deviation, and drive the gimbal to track the target based on the gimbal control signal.
[0059] In some embodiments, the gimbal then calculates a first deviation Δfollow between the predicted target position and the geometric center of the current frame image, and a second deviation Δavoid between the predicted target position and the position of lowest risk. For an image with a resolution of W×H, the coordinates of the geometric center of the image are (W / 2, H / 2), which corresponds to the position of the gimbal's current optical axis in the image coordinate system.
[0060] The first deviation, Δfollow, causes the gimbal to pull the target back to the center of the frame, achieving basic tracking; the second deviation, Δavoid, causes the gimbal to pull the target towards a lower-risk area, achieving proactive avoidance. The final gimbal control signal is a weighted sum of these two deviations. ; Here, α=0.7 and β=0.3 are chosen, which means that basic following is given priority, while also taking into account forward-looking avoidance.
[0061] The gimbal generates commands for the rotation speed and angle of the gimbal motor in the horizontal and vertical directions based on the Control signal, thereby driving the gimbal to move.
[0062] S18, calculate the confidence index of the current frame matching result, and determine whether the current tracking state is a reliable tracking state or a tracking failure state based on the consistency of the confidence index and the target motion trajectory.
[0063] In some embodiments, after obtaining the target's current predicted position in step S14, the gimbal also needs to evaluate the reliability of the prediction result to determine whether the current tracking is reliable. Due to factors such as occlusion, rapid target movement, or background interference in complex scenes, the response map of the correlation filter may exhibit phenomena such as indistinct peaks or multi-peak competition. In this case, directly using the peak position as the target position will lead to the accumulation of tracking errors or even target loss. Therefore, the gimbal calculates the confidence index of the current frame matching result and performs consistency verification in conjunction with the target's historical motion trajectory to comprehensively determine whether the current tracking state is reliable tracking or tracking failure, providing a decision-making basis for subsequent model update strategies and target re-acquisition mechanisms.
[0064] In an optional implementation, determining whether the current tracking state is a reliable tracking state or a tracking failure state based on the consistency between the confidence index and the target motion trajectory includes: The peak-to-side-lobe ratio of the two-dimensional response plot is calculated as the confidence index; The target position in the current frame is predicted using a Kalman filter to obtain the predicted target position; The Euclidean distance between the current predicted position of the target and the predicted position of the target is calculated as a consistency measure of the target's trajectory. When the peak-to-sidelobe ratio is greater than a first preset threshold and the Euclidean distance is less than a second preset threshold, the system is determined to be in a reliable tracking state; otherwise, it is determined to be in a tracking failure state.
[0065] In some embodiments, the peak-to-sidelobe ratio (PSR) of the gimbal's calculated response map St is used as a confidence metric: ; Where max(St) is the peak value of the response map, and μ and σ are the mean and standard deviation of the sidelobe region of the response map (the region excluding a certain range around the peak value), respectively. The higher the PSR value, the more prominent the peak value of the response map, and the more reliable the tracking results.
[0066] Next, the gimbal uses a Kalman filter to predict the target position in the current frame. For ease of distinction, this is referred to as the target predicted position. The Kalman filter uses a linear motion model to make predictions based on the target's historical trajectory (position and velocity). Specifically, the state vector is set as... ,in and These represent the velocity components in the horizontal and vertical directions, respectively. The state transition matrix and observation matrix are set according to the uniform motion model.
[0067] Furthermore, calculate the predicted location of the target. With the peak value of the response map (i.e., the current predicted position of the target) Euclidean distance of ) .
[0068] In addition, the gimbal can also preset a first preset threshold and a second preset threshold. When the PSR is greater than the first preset threshold (3.0 in this embodiment) and the Euclidean distance d is less than the second preset threshold (0.5 times the length of the target box diagonal in this embodiment), it is determined to be a reliable tracking state; otherwise, it is determined to be a tracking failure state (caused by severe occlusion or complete loss of features).
[0069] Through the aforementioned optional implementation methods, the gimbal combines response map confidence with motion trajectory consistency to form a dual verification mechanism. A single confidence index is insufficient to distinguish between two scenarios: "the target undergoes drastic deformation" and "the target is occluded." The former requires updating the model to adapt to the appearance change, while the latter requires keeping the model unchanged and initiating reacquisition. Using a single index for judgment can easily lead to incorrect decisions. This implementation introduces the motion prediction results of a Kalman filter as an independence verification in the spatial domain and uses the peak-to-sidelobe ratio of the response map as a reliability measure in the confidence domain. The two mutually verify each other, accurately distinguishing different states such as successful tracking, target deformation, and complete loss, effectively avoiding model contamination or reacquisition delays caused by misjudgment.
[0070] S19, when the tracking failure state is determined, the target re-acquisition mechanism is started to retrieve the target; when the reliable tracking state is determined, it is determined whether the preset update condition is met. If the preset update condition is met, the reference feature template is updated using the tracking result of the current frame; otherwise, the reference feature template is not updated.
[0071] In some embodiments, after determining the tracking status, the gimbal executes differentiated follow-up processing strategies based on different determination results. If the tracking of the current frame is determined to be in a failed state, it indicates that the target has been lost or the response is unreliable. If tracking or model updates are blindly continued at this time, not only will the correct target position not be obtained, but background noise will also be introduced into the model, accelerating the overall failure of the tracking system. Therefore, the gimbal needs to pause the regular tracking process, start the target re-acquisition mechanism, and use the target's historical motion information and appearance template to search within the prediction area to find the target as soon as possible and resume normal tracking. If the tracking of the current frame is determined to be in a reliable state, it indicates that the current positioning result is reliable. At this time, it is necessary to further determine whether to update the reference feature template. During the tracking process, the appearance of the target will gradually change due to factors such as pose changes and illumination changes. If the template is not updated, the difference between the template and the current appearance of the target will gradually accumulate, eventually leading to tracking failure. However, if updates are performed in every frame, the model will be contaminated and irreversible drift will occur once there is a brief interference or slight occlusion. Therefore, the gimbal only performs model updates when the preset strict update conditions are met, and skips updates when the conditions are not met to maintain the stability of the model. Through the aforementioned branching processing mechanism of "failure recapture + reliable selective update", the gimbal can balance the robustness and adaptability of tracking in complex scenarios.
[0072] In an optional implementation, when the tracking is determined to be in a failure state, the method further includes: By fitting a motion model using the historical target position data of the most recent N frames, the candidate positions of the target within a preset search area are predicted. Multi-scale template matching is then performed using the target reference template within the preset search area to find the best matching position as the new target position, and the tracking is reset to a reliable state. When the reliable tracking state is determined, it is determined whether the first update condition and the second update condition are met at the same time. The first update condition is that the peak-to-sidelobe ratio is greater than the third preset threshold, and the second update condition is that the ratio of the main peak to the secondary peak of the two-dimensional response graph is greater than the fourth preset threshold. When both the first update condition and the second update condition are met, the initial feature vector is updated with a preset learning rate.
[0073] In some embodiments, when a tracking failure is determined, a target re-acquisition mechanism is initiated. The gimbal can utilize the target position historical data of the most recent N frames (where N is an integer greater than or equal to 0, and N=5 in this embodiment) to fit a quadratic motion model and predict the approximate area where the target may appear in the current frame (i.e., the preset search area). Specifically, based on the historical position sequence, the least squares method is used to fit a quadratic polynomial relationship between the position and the frame number, and the predicted position of the current frame is extrapolated to obtain the predicted position of the current frame.
[0074] Centered on the predicted location, the preset search area is set to Within this region, multi-scale sliding window matching (normalized cross-correlation method) is performed using the target reference template T to find the optimal matching position as the new target position, and the tracking state is reset to a reliable tracking state.
[0075] When a reliable tracking state is determined, it is checked whether the preset update conditions are met. This embodiment employs a dual-threshold triggering mechanism: First update condition: Current PSR > third preset threshold (e.g., set to 0.75) to ensure high confidence in the tracking results; The second update condition is: the ratio of the main peak to the secondary peak in the response graph is greater than the fourth preset threshold (for example, set to 1.5), to ensure that the response graph is clean and unambiguous.
[0076] Model updates (including baseline feature template T updates and classifier coefficient updates) are triggered only when both of the above update conditions are met simultaneously. (Update) to avoid introducing occlusions or background noise into the model and prevent model drift.
[0077] When the baseline feature template needs to be updated, it is updated with a preset learning rate η in the following manner: ; In this embodiment, the preset learning rate η = 0.025.
[0078] At the same time, update the classifier coefficients of the relevant filters: ; If the update conditions are not met, this update will be skipped.
[0079] Through the aforementioned optional implementation methods, two differentiated processing strategies—model updating and target recapture—are executed for the two distinct states of reliable tracking and tracking failure, forming a complete tracking closed loop. On the one hand, in the reliable tracking state, a dual-threshold joint judgment mechanism composed of peak-to-sidelobe ratio and main-to-secondary-peak ratio is used to perform model updates only when the tracking result is highly reliable and the response map is clean and unambiguous. This allows the model to smoothly adapt to the natural gradations of the target's appearance while effectively suppressing noise introduced by abnormal situations such as occlusion and sudden changes in illumination. On the other hand, in the tracking failure state, a two-step recapture strategy using trajectory prediction to narrow the search range and multi-scale template matching for precise positioning achieves rapid target retrieval without excessively increasing computational overhead, avoiding the problem of long-term system failure after tracking failure. The conservative update strategy under reliable tracking and the active recapture strategy under tracking failure work together to ensure that this embodiment can maintain good tracking sensitivity and positioning accuracy during long-term tracking.
[0080] Reference Figure 2 The diagram shown is a functional block diagram of an environment-adaptive gimbal target tracking device according to an embodiment of this application.
[0081] In some embodiments, the environment-adaptive gimbal target tracking device 20 may include multiple functional modules composed of computer program segments. The computer programs for each program segment of the environment-adaptive gimbal target tracking device 20 may be stored in the memory of the gimbal and executed by at least one processor to perform (see details). Figure 1 (Description) The gimbal target tracking function is environmentally adaptive. Based on its functions, it can be divided into multiple functional modules. These modules may include: an initialization module 201, a degradation risk perception module 202, a candidate extraction module 203, a position prediction module 204, a risk analysis module 205, a deviation calculation module 206, a gimbal control module 207, a status determination module 208, and a template update module 209. The term "module" in this application refers to a series of computer program segments that can be executed by at least one processor and perform a fixed function, stored in memory. In this embodiment, the functions of each module will be detailed in subsequent embodiments.
[0082] The initialization module 201 is used to acquire a first frame image and a user-specified initial target bounding box, crop the image region corresponding to the initial target bounding box from the first frame image as a target reference template, and extract the initial feature vector of the target reference template.
[0083] The degradation risk perception module 202 is used to acquire the current frame image and determine a degradation risk distribution map based on the current frame image to characterize the risk of feature loss in the target's region.
[0084] The candidate extraction module 203 is used to crop out a candidate image block of the same size as the target reference template from the target position of the previous frame in the current frame image, and extract the current frame feature vector of the candidate image block.
[0085] The position prediction module 204 is used to perform matching calculations based on the current frame feature vector and the initial feature vector to obtain the target's current predicted position in the current frame.
[0086] The risk analysis module 205 is used to analyze the degradation risk distribution within a preset range near the current predicted position of the target in the degradation risk distribution map, and to determine the lowest risk position within the preset range.
[0087] The deviation calculation module 206 is used to calculate the first deviation between the current predicted position of the target and the geometric center of the current frame image, and to calculate the second deviation between the current predicted position of the target and the position with the lowest risk.
[0088] The gimbal control module 207 is used to generate a gimbal control signal based on the first deviation and the second deviation, and drive the gimbal to track the target based on the gimbal control signal.
[0089] The state determination module 208 is used to calculate the confidence index of the current frame matching result, and determine whether the current tracking state is a reliable tracking state or a tracking failure state based on the consistency between the confidence index and the target motion trajectory.
[0090] The template update module 209 is used to initiate a target re-acquisition mechanism to retrieve the target when the tracking failure state is determined; when the tracking is determined to be reliable, it determines whether a preset update condition is met. If the preset update condition is met, the reference feature template is updated using the tracking result of the current frame; otherwise, the reference feature template is not updated.
[0091] It should be understood that the various variations and specific embodiments of the environment-adaptive gimbal target tracking method provided in the above embodiments are also applicable to the environment-adaptive gimbal target tracking device of this embodiment. Through the foregoing detailed description of the environment-adaptive gimbal target tracking method, those skilled in the art can clearly understand the implementation method of the environment-adaptive gimbal target tracking device of this embodiment. For the sake of brevity, it will not be described in detail here.
[0092] See Figure 3 The diagram shown is a schematic representation of the structure of a gimbal in an embodiment of this application. In a preferred embodiment of this application, the gimbal 3 includes a memory 31, at least one processor 32, and at least one communication bus 33.
[0093] Those skilled in the art should understand that Figure 3 The structure of the gimbal shown does not constitute a limitation of the embodiments of this application. It can be a bus structure or a star structure. The gimbal 3 may also include more or fewer other hardware or software than shown, or different component arrangements.
[0094] In some embodiments, the gimbal 3 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits, programmable gate arrays, digital processors, and embedded devices. The gimbal 3 may also include user equipment, which includes, but is not limited to, any electronic product that allows human-computer interaction with a user via a keyboard, mouse, remote control, touchpad, or voice control device, such as personal computers, tablet computers, smartphones, and digital cameras.
[0095] In the embodiments provided in this application, it should be understood that the disclosed methods, apparatus, computer-readable storage media, and gimbals can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple components or modules may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be indirect couplings or communication connections between devices, components, or modules through some interfaces, and may be electrical, mechanical, or other forms.
[0096] The components described as separate parts may or may not be physically separate. The components shown as components may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the components can be selected to achieve the purpose of this embodiment according to actual needs.
[0097] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each component can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0098] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0099] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0100] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0101] The above is a detailed description of the preferred embodiments of this application. However, the invention of this application is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. An environment-adaptive gimbal target tracking method, characterized in that, The method includes: Acquire the first frame image and the user-specified initial target bounding box, crop the image region corresponding to the initial target bounding box from the first frame image as the target reference template, and extract the initial feature vector of the target reference template; The current frame image is acquired, and semantic segmentation is performed on the current frame image to obtain the semantic segmentation result. A region of interest is generated with the target position of the previous frame as the center, and the local variance of each pixel in the region of interest is calculated as a texture richness index. Based on the semantic segmentation result and the texture richness index, a degradation risk distribution map of the same size as the region of interest is generated. Each pixel value in the degradation risk distribution map represents the probability that the current pixel position will lose features within a preset number of future frames. In the current frame image, a candidate image block of the same size as the target reference template is cropped with the target position of the previous frame as the center, and the current frame feature vector of the candidate image block is extracted. The target's current predicted position in the current frame is obtained by matching the current frame feature vector with the initial feature vector. Analyze the degradation risk distribution within a preset range near the current predicted location of the target in the degradation risk distribution map, and determine the location with the lowest risk within the preset range; Calculate the first deviation between the current predicted position of the target and the geometric center of the current frame image, and calculate the second deviation between the current predicted position of the target and the position with the lowest risk; A gimbal control signal is generated based on the first deviation and the second deviation, and the gimbal is driven to track the target based on the gimbal control signal. Calculate the confidence index of the current frame matching result, and determine whether the current tracking state is a reliable tracking state or a tracking failure state based on the consistency between the confidence index and the target motion trajectory. When the tracking is determined to be in a failure state, the target re-acquisition mechanism is activated to retrieve the target; when the tracking is determined to be in a reliable state, it is determined whether the preset update conditions are met. If the preset update conditions are met, the target reference template is updated using the tracking results of the current frame; otherwise, the target reference template is not updated.
2. The environment-adaptive gimbal target tracking method according to claim 1, characterized in that, The step of matching the current frame feature vector with the initial feature vector to obtain the target's current predicted position in the current frame includes: Calculate the correlation between the current frame feature vector and the initial feature vector to generate a two-dimensional response map; The peak position of the two-dimensional response map is taken as the current predicted position of the target.
3. The environment-adaptive gimbal target tracking method according to claim 2, characterized in that, The step of determining whether the current tracking status is a reliable tracking status or a tracking failure status based on the consistency between the confidence index and the target motion trajectory includes: The peak-to-side-lobe ratio of the two-dimensional response plot is calculated as the confidence index; The target position in the current frame is predicted using a Kalman filter to obtain the predicted target position; The Euclidean distance between the current predicted position of the target and the predicted position of the target is calculated as a consistency measure of the target's trajectory. When the peak sidelobe ratio is greater than a first preset threshold and the Euclidean distance is less than a second preset threshold, it is determined to be a reliable tracking state; When the peak sidelobe ratio is less than or equal to a first preset threshold, or the Euclidean distance is greater than or equal to a second preset threshold, the tracking is determined to be in a failure state.
4. The environment-adaptive gimbal target tracking method according to claim 3, characterized in that, When the tracking is determined to be in a failure state, the method further includes: By fitting a motion model using the historical target position data of the most recent N frames, the candidate positions of the target within a preset search area are predicted. Multi-scale template matching is then performed using the target reference template within the preset search area to find the best matching position as the new target position, and the tracking is reset to a reliable state. When the tracking status is determined to be reliable, it is determined whether the preset update conditions are met at the same time. The preset update conditions include a first update condition and a second update condition. The first update condition is that the peak-to-sidelobe ratio is greater than a third preset threshold, and the second update condition is that the ratio of the main peak to the secondary peak of the two-dimensional response graph is greater than a fourth preset threshold. When both the first update condition and the second update condition are met, the initial feature vector is updated with a preset learning rate.
5. The environment-adaptive gimbal target tracking method according to claim 1, characterized in that, The extraction of the current frame feature vector of the candidate image patch includes: Multiple preset types of features are extracted from the candidate image blocks to obtain multiple candidate feature components; Extract the center point value of the degradation risk distribution map as the degradation risk value of the target's current location; The fusion weights corresponding to the multiple candidate feature components are dynamically determined based on the degradation risk value. The multiple candidate feature components are weighted and summed according to their respective fusion weights to obtain the current frame feature vector.
6. The environment-adaptive gimbal target tracking method according to claim 5, characterized in that, The various preset feature types include histogram of oriented gradients features, color naming features, and shallow convolutional neural network features. The dynamic determination of the fusion weights corresponding to the multiple candidate feature components based on the degradation risk value includes: When the degradation risk value is greater than the first risk threshold, the weight of the directional gradient histogram feature is reduced, and the weights of the color naming feature and the shallow convolutional neural network feature are increased. When the degradation risk value is less than the second risk threshold, the balanced weight of each preset type feature is maintained; When the degradation risk value is between the second risk threshold and the first risk threshold, the fusion weight is determined by linear interpolation. Wherein, the first risk threshold is greater than the second risk threshold.
7. An environment-adaptive gimbal target tracking device, characterized in that, The device includes: An initialization module is used to acquire a first frame image and a user-specified initial target bounding box, crop the image region corresponding to the initial target bounding box from the first frame image as a target reference template, and extract the initial feature vector of the target reference template. The degradation risk perception module is used to acquire the current frame image, perform semantic segmentation on the current frame image, and obtain semantic segmentation results; generate a region of interest centered on the target position of the previous frame, and calculate the local variance of each pixel in the region of interest as a texture richness index; generate a degradation risk distribution map of the same size as the region of interest based on the semantic segmentation results and the texture richness index; each pixel value in the degradation risk distribution map represents the probability of feature loss at the current pixel position within a preset number of future frames; The candidate extraction module is used to crop out a candidate image block of the same size as the target reference template from the target position of the previous frame in the current frame image, and extract the current frame feature vector of the candidate image block. The position prediction module is used to perform matching calculations based on the current frame feature vector and the initial feature vector to obtain the target's current predicted position in the current frame; The risk analysis module is used to analyze the degradation risk distribution within a preset range near the current predicted location of the target in the degradation risk distribution map, and to determine the location with the lowest risk within the preset range; The deviation calculation module is used to calculate the first deviation between the current predicted position of the target and the geometric center of the current frame image, and to calculate the second deviation between the current predicted position of the target and the position with the lowest risk. The gimbal control module is used to generate a gimbal control signal based on the first deviation and the second deviation, and drive the gimbal to track the target based on the gimbal control signal; The status determination module is used to calculate the confidence index of the matching result of the current frame, and determine whether the current tracking status is a reliable tracking status or a tracking failure status based on the consistency between the confidence index and the target motion trajectory. The template update module is used to initiate a target re-acquisition mechanism to retrieve the target when the tracking is determined to be in a failed state; when the tracking is determined to be reliable, it determines whether a preset update condition is met. If the preset update condition is met, the target reference template is updated using the tracking result of the current frame; otherwise, the target reference template is not updated.
8. A gimbal, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the environment-adaptive gimbal target tracking method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the environment-adaptive gimbal target tracking method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Target real-time tracking method and device, computer device and storage medium
CN108985162A
Spatial-temporal semantic association-based moving ship adaptive tracking method
CN121353327A