Holder tracking method and device based on binocular camera, and storage medium

By using the depth information and three-dimensional coordinate transformation of the binocular camera, combined with the Kalman filter, the problem of distinguishing between the real motion and apparent deformation of the target in pan-tilt tracking is solved, and accurate tracking of the target in complex scenes is achieved.

CN120780031AActive Publication Date: 2025-10-14SHENZHEN EMEET TECH CO LTD

Patent Information

Application Number
CN202511286656.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-10-14
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

Existing pan-tilt tracking technology lacks a mechanism to compensate for depth changes on characteristic scales and cannot distinguish between the target's real motion and apparent deformation, resulting in poor target tracking results.

Method used

Based on the image data output by the binocular camera, the target is located and tracked through feature information matching, depth information is calculated, and three-dimensional coordinates are generated. Combined with the offset from the camera optical center to the gimbal rotation center, the three-dimensional gimbal coordinates are generated through coordinate transformation. The feature information is used to match historical features to determine the historical trajectory. The state is updated through the Kalman filter, and the gimbal motor is driven to rotate to center the target.

Benefits of technology

Through depth information and three-dimensional coordinate transformation, the stability and robustness of target tracking are improved, characteristic scale distortion and motion blur are overcome, and accurate tracking of targets in complex scenes is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780031A_ABST
    Figure CN120780031A_ABST
Patent Text Reader

Abstract

The invention discloses a cradle head tracking method and device based on a binocular camera and a storage medium, and relates to the technical field of computer vision, and the method comprises the steps: processing image data based on a binocular parallax principle, and generating a three-dimensional coordinate of a center point of a tracking target; based on the three-dimensional coordinates of the camera coordinate system and the offset from the optical center of the camera to the rotation center of the holder, generating three-dimensional holder coordinates through coordinate transformation solution; determining a historical track based on the tracking target feature information and a historical target feature matching result, and outputting an identifier and a three-dimensional position observation value through correlation verification of a three-dimensional holder coordinate and the historical track; inputting an observation updating equation correction state through the identifier and the three-dimensional position observation value, and outputting a three-dimensional prediction position; based on the three-dimensional prediction position and a deviation formula, calculating the angle deviation with the camera image center under the holder coordinate system, and driving the holder to center the target in the picture center according to the angle deviation. The problem that the target tracking effect is poor is solved, and the robustness of target tracking in a complex scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, and particularly relates to a gimbal tracking method based on a binocular camera, a device and a storage medium. BACKGROUND

[0002] The gimbal tracking technology is a complex system technology integrating mechanical and electrical control, sensor fusion, real-time image processing, target detection and tracking algorithm, and its development is based on the cross-fusion of multiple disciplines. Related technologies maintain target locking by updating the 2D image feature template when tracking the target before and after the movement, but due to the lack of a depth change compensation mechanism for feature scale, it cannot distinguish between the real movement of the target and the apparent deformation, thereby leading to poor target tracking effect.

[0003] The above content is only used to assist in understanding the technical solutions of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY

[0004] The main purpose of the present application is to provide a gimbal tracking method based on a binocular camera, a device and a storage medium, which aims to solve the technical problem of poor target tracking effect caused by the lack of a depth change compensation mechanism for feature scale, which cannot distinguish between the real movement of the target and the apparent deformation.

[0005] To achieve the above purpose, the present application provides a gimbal tracking method based on a binocular camera, which comprises: Based on the image data output by the binocular camera, the tracking target is located and tracked through feature information matching, and the depth information is calculated according to the difference of the corresponding pixel positions of the tracking target in the image data, and the three-dimensional coordinates of the tracking target center point in the current frame camera coordinate system are calculated and generated; Based on the three-dimensional coordinates and the offset amount of the camera optical center to the gimbal rotation center, the three-dimensional gimbal coordinates of the tracking target are calculated and generated through the coordinate transformation relationship; Based on the matching result of the feature information corresponding to the tracking target and the historical target feature, the historical trajectory of the tracking target is determined, and the identifier of the tracking target and the three-dimensional position observation value of the tracking target are output through the correlation verification of the three-dimensional gimbal coordinates corresponding to the tracking target and the historical trajectory corresponding to the tracking target; By inputting the identifier of the tracking target and the three-dimensional position observation value of the tracking target into the observation update equation to correct the state, the three-dimensional predicted position of the tracking target is output; Based on the three-dimensional predicted position corresponding to the tracking target and the bias formula, the angle deviation of the three-dimensional predicted position in the gimbal coordinate system from the center of the camera image is calculated, and the gimbal motor is driven to rotate according to the angle deviation, so that the tracking target is centered in the center of the camera image.

[0006] In an embodiment, the tracking target is identified and the pixel coordinates of the tracking target in the image data are obtained through the feature information extracted in the image data; Based on the same pixel position of the left and right images corresponding to the image data, the corresponding pixel coordinates are stereomatched to generate an initial disparity map; Based on the initial disparity map, the camera focal length and the baseline distance, the depth information of the pixel coordinates relative to the gimbal is generated through geometric conversion formula calculation; According to the pixel coordinates and the depth information of the tracking target, the three-dimensional coordinates of the center point of the tracking target in the camera coordinate system are generated.

[0007] In an embodiment, based on the camera images under different gimbal attitude angles, the camera optical center is calibrated, and the coordinate difference value is calculated with the gimbal rotation center, and the offset is output; The three-dimensional coordinates in the camera coordinate system are superimposed with the offset through the coordinate transformation relationship to calculate and generate the three-dimensional gimbal coordinates.

[0008] In an embodiment, based on the feature information corresponding to the tracking target and the matching failure of the historical target features, the tracking target is taken as a new target of the current frame; According to the three-dimensional gimbal coordinates of the new target, a state vector is constructed to update the historical trajectory corresponding to the new target.

[0009] In an embodiment, according to the matching result of the feature information corresponding to the tracking target and the historical target features, the tracking target is taken as an index to determine the historical trajectory of the tracking target in the database; Based on the historical trajectory, the three-dimensional gimbal coordinates are associated and verified to generate a verification parameter; According to the verification parameter, the three-dimensional gimbal coordinates and the corresponding historical trajectory are updated, and the identifier and the three-dimensional position observation value of the tracking target are output.

[0010] In an embodiment, based on the identifier of the tracking target, the Kalman filter state vector and the state covariance matrix corresponding to the identifier are retrieved in the database; The three-dimensional position observation value, the Kalman filter state vector and the state covariance matrix are input into the observation update equation to generate the real-time three-dimensional position and the real-time velocity of the tracking target; According to the real-time three-dimensional position and the real-time velocity, the three-dimensional predicted position of the tracking target at the next time stamp is predicted according to the time sequence.

[0011] In an embodiment, the three-dimensional predicted position corresponding to the tracking target is input into a bias formula, and horizontal angle bias and pitch angle bias of the tracking target in the gimbal coordinate system are output; Based on the horizontal angle bias and the pitch angle bias, a bias correction strategy is generated according to a preset bias adjustment rule; Based on the device parameters of the gimbal and the bias correction strategy, a correction parameter is generated through a cooperative mechanism of feedforward compensation and feedback adjustment; According to the correction parameter, the rotation of the gimbal motor is driven, and the tracking target is centered on the center position of the camera picture.

[0012] In an embodiment, the integrity of the initial disparity map is verified based on a disparity map integrity rule, and an integrity judgment result is generated; If the integrity judgment result shows that there is a hole in the initial disparity map, valid pixels adjacent to the hole are collected, and a pixel mean value corresponding to the valid pixels is calculated; Based on a mean value filling strategy, the pixel mean value is introduced to fill the hole in the initial disparity map, and an optimized depth map is generated; According to the depth map, the camera focal length and the baseline distance, the pixel coordinates are input into a geometric conversion formula for calculation, and the depth information of the pixel coordinates relative to the gimbal is generated.

[0013] In addition, to achieve the above-mentioned purposes, the present application also proposes a gimbal tracking device, which comprises a memory, a processor and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the above-mentioned gimbal tracking method based on the binocular camera.

[0014] In addition, to achieve the above-mentioned purposes, the present application also proposes a storage medium, which is a computer readable storage medium, and the storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the above-mentioned gimbal tracking method based on the binocular camera.

[0015] The application provides a gimbal tracking method based on a binocular camera, which comprises target space positioning processing of image data based on the parallax principle of the binocular camera, generation of three-dimensional coordinates of a current frame tracking target center point in a camera coordinate system, combination of a pre-calibrated offset of a camera optical center to a gimbal rotation center, solving of three-dimensional coordinates of the target in a gimbal coordinate system through rigid body coordinate transformation, correlation of a historical motion trajectory of the target by using a matching result of target feature information and a historical feature library, and output of a target unique identifier and a three-dimensional position observation value through spatial consistency verification of three-dimensional gimbal coordinates and a trajectory prediction point, input of the identifier and the observation value into a Kalman filter observation update equation to correct a state vector, output of a three-dimensional prediction position of the target at a next moment, calculation of horizontal angle deviation and pitch angle deviation of the target from a camera image center according to the position, generation of motor angular velocity instructions by an incremental controller to drive the gimbal to rotate, and realization of target dynamic optical axis alignment, which effectively overcomes feature scale distortion and motion blur when the target moves.

[0016] In summary, the application obtains target depth information through the binocular parallax principle, converts 2D coordinates of a person in an image into accurate 3D coordinates in a gimbal coordinate system, solves horizontal angle and pitch angle that the gimbal needs to rotate through geometric mapping, significantly improves tracking stability, simultaneously fuses motion information in the depth dimension, constructs a 3D Kalman filter model to predict a target trajectory, and dynamically adjusts gimbal response through a depth adaptive control strategy, thereby solving the technical problem that a feature scale compensation mechanism for depth change is lacking, a target real motion and apparent deformation cannot be distinguished, and target tracking effect is poor, and comprehensively improving the robustness of target tracking in a complex scene. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate an embodiment consistent with the present application and, together with the description, serve to explain the principles of the application.

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows, and obviously, other drawings can also be obtained by those skilled in the art without any creative labor.

[0019] Figure 1 A flowchart of a first embodiment of the gimbal tracking method based on the binocular camera of the present application; Figure 2 A flowchart of a second embodiment of the gimbal tracking method based on the binocular camera of the present application; Figure 3 A flowchart of the gimbal tracking method based on the binocular camera of the present application; Figure 4A structure schematic diagram of a gimbal tracking device of the present application.

[0020] The object realization, functional features and advantages of the present application will be further explained in conjunction with the embodiments, with reference to the accompanying drawings. DETAILED DESCRIPTION

[0021] It should be understood that the specific embodiments described herein are merely intended to explain the technical solutions of the present application, and are not intended to limit the present application.

[0022] In related technologies, when tracking a target moving in front and behind, a target lock is maintained by updating a 2D image feature template, but due to the lack of a depth change compensation mechanism for feature scale, it is unable to distinguish between real target motion and apparent deformation, thus leading to poor target tracking effect.

[0023] The present application provides a solution: first, based on image data output by a binocular camera, a tracking target is located and tracked through feature information matching, and depth information is calculated according to the difference in pixel position of the tracking target in the image data, to solve and generate a three-dimensional coordinate of a center point of the tracking target in a camera coordinate system of a current frame, then based on the three-dimensional coordinate and an offset amount of a camera optical center to a gimbal rotation center, a three-dimensional gimbal coordinate of the tracking target is solved and generated through a coordinate transformation relationship, and based on the matching result of the feature information corresponding to the tracking target and historical target features, a historical trajectory of the tracking target is determined, and the three-dimensional gimbal coordinate corresponding to the tracking target is associated and verified with the historical trajectory corresponding to the tracking target, to output an identifier of the tracking target and a three-dimensional position observation value of the tracking target, then the identifier of the tracking target and the three-dimensional position observation value of the tracking target are input to a state of a measurement update equation to correct the state, to output a three-dimensional predicted position of the tracking target, and finally, based on the three-dimensional predicted position corresponding to the tracking target and a bias formula, an angle bias of the three-dimensional predicted position in a gimbal coordinate system with respect to a camera image center is solved, and a gimbal motor is driven to rotate according to the angle bias, so as to center the tracking target in the camera picture center.

[0024] It should be understood that the execution subject of the present embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a gimbal tracking device, etc. capable of realizing the above functions. The present embodiment and the following embodiments will be described below taking the gimbal tracking device as an example.

[0025] In order to better understand the technical solutions of the present application, the following will be described in detail in conjunction with the drawings and specific embodiments of the present application.

[0026] The present embodiment provides a gimbal tracking method based on a binocular camera, which will be described in detail with reference to the accompanying drawings. Figure 1, Figure 1 The flowchart of the gimbal tracking method based on the binocular camera according to the first embodiment of the application is shown in the figure.

[0027] In this embodiment, the gimbal tracking method based on the binocular camera comprises steps S10-S50: In step S10, the tracking target is located and tracked by matching the feature information based on the image data output by the binocular camera, and the depth information is calculated according to the difference of the pixel positions corresponding to the tracking target in the image data, so as to solve and generate the three-dimensional coordinates of the center point of the tracking target in the camera coordinate system of the current frame.

[0028] In this embodiment, the feature information refers to the features extracted from the target in the image, including the shape features and motion features of the target. The depth information refers to the straight-line distance of the center point of the target relative to the rotation center of the gimbal. The binocular camera refers to a visual system composed of two optical cameras separated in space, which simulates human stereovision by synchronously collecting left and right view images. The image data refers to the left and right view original image frames output by the binocular camera. The three-dimensional coordinates refer to the physical position of the target in the camera coordinate system, which is calculated by parallax and obtained by inverse projection of the pixel coordinates.

[0029] As an optional implementation, the left and right images of the current frame are synchronously collected by the mounted binocular camera, and the left and right images are input into a target detection network, such as a YOLOv5 model. The feature information is compared in the database to identify the tracking target, and the target detection network locates and frames the tracking target to output the bounding box of the tracking target. The edge and corner coordinates of the bounding box are determined to calculate the pixel coordinates of the center point of the bounding box. Based on the binocular baseline distance and the camera focal length, the target parallax value is calculated by the parallax formula, and the three-dimensional coordinates of the target in the camera coordinate system are calculated in combination with the camera intrinsic matrix. Through the pre-trained target detection network, the identification of the tracking target and the determination of the coordinates are directly performed, which has low resource occupancy and is suitable for embedded devices.

[0030] As another optional implementation, the left and right images are subjected to full-image dense parallax calculation by a stereo matching algorithm to generate an initial parallax map. All valid parallax points in the target detection frame region are extracted, and a parallax distribution histogram is counted to take the highest frequency peak as the overall depth estimation parallax of the target. The tracking target center pixel coordinates are subjected to distortion correction to obtain normalized coordinates. The three-dimensional coordinates of the tracking target in the camera coordinate system are calculated by the inverse projection transformation formula. Through the generation of the parallax distribution histogram, the histogram statistics are used to resist local matching noise, so that the generated parallax is more accurate.

[0031] Step S20, based on the three-dimensional coordinates, and the offset amount of the camera optical center to the gimbal rotation center, the three-dimensional gimbal coordinates of the tracking target are calculated through a coordinate transformation relationship.

[0032] In this embodiment, the camera optical center refers to the optical center point of the camera lens, the imaging reference position. The gimbal rotation center refers to the physical intersection of the gimbal pitch and azimuth axes. The offset amount refers to the rigid displacement of the camera optical center to the gimbal rotation center. The coordinate transformation relationship refers to mapping the coordinates from one coordinate system to another coordinate system through rotation and translation operations. The three-dimensional gimbal coordinates refer to the position of the target in the gimbal coordinate system, with the gimbal rotation center as the origin.

[0033] As an optional implementation, based on the three-dimensional coordinates in the camera coordinate system, the pre-calibrated relationship parameters between the camera and the gimbal space are loaded, the camera coordinates are rotated and aligned to the gimbal coordinate system direction through rigid body transformation, the intermediate coordinates are obtained, and the three-dimensional gimbal coordinates are output by superimposing the translation offset, completing the coordinate system conversion. Through this coordinate transformation, the transformation calculation is efficient, and various open source frameworks are supported.

[0034] As another optional implementation, a quaternion is used to represent the rotation relationship, the three-dimensional coordinates in the camera coordinate system are input into the quaternion rotation formula, the three-dimensional coordinates are converted into quaternion form, the rotation operation is realized through quaternion multiplication, and the three-dimensional gimbal coordinates are generated by converting the Euler coordinates and superimposing the translation offset. Through the conversion formula of the quaternion, the attitude is gradually changed, the smoothness of the interpolation is maintained, and the smoothness of the motion trajectory is improved.

[0035] Step S30, based on the matching result of the feature information corresponding to the tracking target and the historical target feature, the historical trajectory of the tracking target is determined, and the three-dimensional gimbal coordinates corresponding to the tracking target are associated and verified with the historical trajectory corresponding to the tracking target, to output the identifier of the tracking target and the three-dimensional position observation value of the tracking target.

[0036] In this embodiment, the feature information corresponding to the tracking target refers to the features extracted from the target in the image, including the shape features and motion features of the target. The historical target feature refers to the past frame target feature data stored in the feature database. The matching result refers to the similarity score and corresponding relationship of the current feature and the historical feature. The historical trajectory refers to the historical position sequence of the target in the gimbal coordinate system. The association verification refers to verifying the spatial consistency of the current coordinates and the predicted point of the historical trajectory. The identifier refers to the unique identity number of the target. The three-dimensional position observation value refers to the spatial coordinates of the target output after verification.

[0037] As an optional implementation, the feature information of the tracking target in the current frame is extracted, cosine similarity matching is performed with the historical feature library, if the highest similarity exceeds the similarity threshold, it is determined as the same target and the historical trajectory point sequence thereof is indexed. The Kalman filter is used to extrapolate the trajectory prediction position in the current frame, the Euclidean distance error of the three-dimensional gimbal coordinates is calculated, if the error is less than the preset error threshold, the spatial consistency check is passed, the unique identifier of the target is output, and the three-dimensional gimbal coordinates corresponding to the target and the historical three-dimensional gimbal coordinates are output as three-dimensional position observations. Through the matching of the feature information and the setting of the threshold, only the dot product operation of the feature vector is required to complete the matching process, and the Kalman filter is used to suppress single-frame jitter, thereby ensuring the continuity of the trajectory.

[0038] As another optional implementation, a pre-trained convolutional neural network is used to extract feature information of a target with a preset dimension, the nearest neighbor search of the historical feature vector is performed through the database, and the historical trajectory data is loaded after successful matching. The current frame position is predicted by polynomial fitting, the speed direction consistency and position deviation are calculated in combination with the three-dimensional gimbal coordinates, and the three-dimensional position observations are output after the identifier is bound after passing the double check. The feature comparison is performed through the pre-trained neural network, the efficiency of feature matching is improved, and the resource investment is saved.

[0039] In step S40, the identifier of the tracking target and the three-dimensional position observations of the tracking target are input into an observation update equation to correct the state, and the three-dimensional prediction position of the tracking target is output.

[0040] In the embodiment, the observation update equation refers to a state correction formula for fusing the observation value and the prediction value in the Kalman filter. The corrected state refers to the update of the state vector and the covariance matrix of the target. The three-dimensional prediction position refers to the prediction of the spatial coordinates of the target at the next preset time.

[0041] As an optional implementation, the exclusive Kalman filter instance of the target identifier is retrieved, the three-dimensional position observations are input into the observation update equation, the Kalman gain matrix is calculated and the prediction error is fused, the position and speed components in the state vector are updated, the state covariance matrix is adjusted to reflect the estimation uncertainty, and finally the three-dimensional prediction position of the target at the next preset time is extrapolated through the state transition equation. Through the fixed gain matrix operation, the construction of the matrix does not need to be performed again, the calculation efficiency is improved, and the resource occupancy rate is reduced.

[0042] As another optional implementation, an adaptive Kalman filter is used to dynamically adjust the process noise covariance matrix according to the observation value residual, and the state covariance weight is increased when the observation value deviates from the predicted value by more than a threshold value, so as to suppress abnormal observation interference and output an accurate three-dimensional predicted position. Through the adaptive Kalman filter, the prediction stability is improved when the motion is suddenly changed / occluded, and the parameters are dynamically adjusted when there is noise, without the need for manual parameter adjustment, thereby improving the prediction accuracy.

[0043] In step S50, the angle deviation of the three-dimensional predicted position in the gimbal coordinate system from the center of the camera image is calculated based on the three-dimensional predicted position corresponding to the tracking target and a deviation formula, and the gimbal motor is driven to rotate according to the angle deviation, so that the tracking target is centered in the camera image center.

[0044] In this embodiment, the deviation formula refers to a geometric relationship formula for calculating the angle deviation of the target position from the center of the camera optical axis. The center of the camera image refers to the projection point of the optical center of the camera imaging plane. The angle deviation refers to the horizontal angle and the pitch angle of the target deviating from the optical axis. The gimbal motor rotation refers to driving the motor of the gimbal to rotate to perform the action of rotating the azimuth axis and the pitch axis. The centering in the camera image center refers to the state of the target imaging point coinciding with the image center. This way has simple control structure and small calculation amount, and the disadvantage is that the anti-external disturbance ability is weak, and it is suitable for indoor stable scenes.

[0045] As an optional implementation, based on the three-dimensional predicted position, the horizontal angle deviation and the pitch angle deviation in the gimbal coordinate system are calculated through a geometric relationship combined with a deviation formula, the deviation values are input into an incremental controller to generate azimuth axis angular velocity instructions and pitch axis angular velocity instructions, which are converted into motor driving voltages through pulse width modulation, and the gimbal motor is driven to rotate until the angle deviation is zero, so that the tracking target is dynamically locked in the image center.

[0046] As another optional implementation, a feedforward compensation mechanism is introduced, the angle deviation of the three-dimensional predicted position in the gimbal coordinate system from the center of the camera image is combined with the target depth to dynamically adjust the proportional gain of the controller. At the same time, the angular acceleration is predicted based on the target motion speed, and the feedforward torque instruction is superimposed to the motor driver, and the position residual is eliminated through closed-loop feedback, so that the tracking target is centered in the camera image center. This way has the advantages of high tracking accuracy and strong anti-interference, and the disadvantages are complex algorithm and the need for pre-calibration of system inertia parameters, and is suitable for dynamic scenes such as outdoor unmanned aerial vehicle tracking.

[0047] Exemplarily, in a film shooting scene, a target actor runs along an S-shaped route, and a binocular camera synchronously collects 1080p@60fps images. A target positioning processing module detects the actor's bounding box and extracts the center point pixel coordinates by YOLOv5, and calculates its camera coordinate system three-dimensional coordinates based on the parallax principle, such as the kth frame: X_c=0.3m, Y_c=-0.1m, Z_c=8.2m. Load the pre-calibrated offset between the camera and the gimbal, T=[0.05, -0.02, 0.1]^T, R is the rotation matrix with a pitch angle of 10°, and the gimbal coordinates are output after coordinate transformation, X_p=0.28m, Y_p=-0.08m, Z_p=8.25m. The feature matching module extracts the actor's clothing texture feature vector through ResNet-18, and matches the cosine similarity with the historical feature library, with a similarity of 0.92 greater than the similarity threshold 0.85, associates its historical trajectory and checks that the spatial error between the current coordinates and the Kalman prediction point is 0.15m, which is less than the preset error threshold 0.3m, and outputs the identifier ID=007 and the observation value, X_obs=0.28m, Y_obs=-0.08m, Z_obs=8.25m. The observation value is input into the Kalman filter corresponding to ID=007, and the next frame prediction position is output after the state is corrected by the observation update equation, X_p=0.35m, Y_p=-0.12m, Z_p=8.30m. Based on the bias formula, the angle deviation is calculated, Δθ=arctan(0.35 / 8.30)≈2.4°, Δφ=arctan(-0.12 / 8.30)≈-0.8°, through the deep adaptive PID controller, K_p=1.2×5 / 8.3≈0.72, the motor control instruction is generated, and the gimbal is driven to rotate to a new position in the center of the tracking target within 0.1s.

[0048] Due to binocular parallax depth perception, the edge target angle calculation distortion problem caused by unknown distance in traditional monocular gimbal tracking is solved, and the axial tracking loss defect when the target approaches / away is overcome by fusing depth motion information through three-dimensional Kalman prediction, thereby improving the accuracy, robustness and adaptability of dynamic target tracking.

[0049] Based on any of the above embodiments, in the second embodiment of the application, refer to Figure 2 , Figure 2 is a flowchart of the second embodiment of the gimbal tracking method based on the binocular camera of the application. The step S10 includes steps A11-A14: Step A11, the tracking target is identified and the pixel coordinates of the tracking target in the image data are obtained through the feature information extracted from the image data.

[0050] In this embodiment, the target detection network refers to a deep learning-based image recognition model used to locate and classify specific targets in images. The image data output by the binocular camera refers to the synchronously captured left and right view image frames. Recognizing the tracking target refers to the detection network outputting target category labels and confidence levels. The pixel coordinates in the image data refer to the location coordinates of the target in the two-dimensional image plane, measured in pixels.

[0051] As an optional implementation, the left and right image frames output synchronously by the binocular camera are processed by a pre-trained target detection network such as YOLOv5. The input images are normalized and scaled to a fixed size, and the network inference layer extracts convolutional features and predicts anchor boxes. The output includes target category probabilities and bounding box coordinates. The geometric center point coordinates of the bounding box are analyzed as the pixel coordinates of the target in the image. This approach has the advantages of extremely fast inference speed and low resource consumption, but the center point of the detected target box may be offset, making it suitable for embedded gimbal devices with limited computing power.

[0052] As another optional implementation, a visual self-attention transformer is used as the detection network backbone. The left and right images are segmented into blocks and embedded. The self-attention is calculated through a pre-set stage level shift window. The left and right image features are fused through a cross-window attention mechanism at the feature layer, and the target detection bounding box is output. The geometric center point of the bounding box is taken as the pixel coordinates. This approach has the advantages of strong global context modeling capability and small center point positioning error, but has high computational complexity, making it suitable for high-precision video-level gimbal tracking systems.

[0053] Step A12: Based on the same pixel position of the left and right images corresponding to the image data, the corresponding pixel coordinates are stereo matched to generate an initial disparity map.

[0054] In this embodiment, the binocular camera disparity principle refers to a geometric method for calculating target depth based on the position difference between left and right lenses. Stereo matching refers to finding the corresponding pixel positions of the same spatial point in left and right images. The left and right images corresponding to the image data refer to the synchronously captured view difference image pair by the binocular camera. The initial disparity map is a two-dimensional matrix that records the horizontal pixel displacement of corresponding points in left and right images.

[0055] As an optional implementation, the left and right images output by the binocular camera are preprocessed to remove noise data in the images. A stereo matching algorithm is used to calculate the matching cost of corresponding pixel coordinates in the left and right images within a pre-set disparity range. The disparity result is optimized through multi-path cost aggregation. The initial disparity map is generated through left-right consistency verification and sub-pixel interpolation. This approach has the advantages of low computational resource consumption and no need for training data, making it suitable for indoor structured scenes.

[0056] As another optional implementation, an end-to-end binocular stereo matching deep learning model, such as a GC-Net end-to-end deep learning model, is used to input left and right images into a convolutional network sharing weight features, to predict an initial disparity map through a cost volume construction and 3D convolution regularization, and finally to output continuous disparity values by applying a disparity regression mechanism. An initial disparity map is generated based on the continuous disparity values. This approach has the advantages of high matching accuracy in weak texture areas and strong anti-light change capability, but has the disadvantages of high computational load and dependence on synthetic data sets for training, and is suitable for outdoor complex environments.

[0057] Step A13, based on the initial disparity map, camera focal length and baseline distance, depth information of the pixel coordinates relative to the gimbal is generated by a geometric conversion formula.

[0058] In this embodiment, the camera focal length refers to the distance from the lens optical center to the imaging plane. The baseline distance refers to the physical distance between the left and right lens optical centers of the binocular camera. The geometric conversion formula refers to the disparity-depth mapping relationship based on the principle of triangulation.

[0059] As an optional implementation, pre-calibrated camera focal length and baseline distance are loaded, and for the disparity value corresponding to the target center pixel coordinates in the initial disparity map, depth information is calculated through geometric conversion relationship analysis, and depth information of the target relative to the center of rotation of the gimbal is output. This approach has the advantages of efficient calculation and does not require additional training, but has the disadvantage of relying on disparity map accuracy, and is suitable for scenes with high disparity confidence.

[0060] As another optional implementation, a neural network model is used to perform dense depth regression on the initial disparity map, and the target region disparity block and camera intrinsic parameters are input to learn the non-linear mapping between disparity and depth through convolutional layers, and to directly output sub-pixel level depth information. This approach has the advantages of strong anti-disparity noise capability and continuous depth value output, but has the disadvantage of requiring 10,000 sets of calibration data for training and high inference delay, and is suitable for weak texture or occluded targets.

[0061] As another optional implementation, the target detection network and the depth map are used to obtain the distance of the person region: the target detection network is used to obtain the region of the person in the image. The center point coordinates of the person region are calculated. The distance of the person is obtained at the corresponding position of the depth map through the center point coordinates. This approach has the advantages of simple process and compatibility with any depth map generation method, but has the disadvantage of relying on the accuracy of the detection frame, and is suitable for scenes where the person target is prominent and the motion is smooth.

[0062] Step A14, according to the pixel coordinates and the depth information of the tracking target, a three-dimensional coordinate of the center point of the tracking target in the camera coordinate system is generated.

[0063] In this embodiment, the camera coordinate system refers to a three-dimensional rectangular coordinate system with the camera optical center as the origin and the optical axis as the Z-axis.

[0064] As an optional implementation, based on the target center pixel coordinates and depth information, and introducing the camera intrinsic parameters, the three-dimensional coordinates in the camera coordinate system are calculated by back projection to output the three-dimensional coordinates of the tracking target center point. This method has the advantages of fewer calculation steps and high real-time performance, and the disadvantage is that the edge target coordinate error increases due to uncorrected lens distortion, which is suitable for medium and long focal length lens scenes.

[0065] As another optional implementation, load the camera distortion correction parameters, and perform distortion correction on the pixel coordinates to obtain normalized coordinates, then calculate the three-dimensional coordinates combined with the depth information, and output the three-dimensional coordinates of the tracking target center point. This method has the advantages of eliminating the influence of barrel / pincushion distortion, and the disadvantage is high calculation load, which is suitable for wide-angle lens or high-precision measurement scenes.

[0066] Exemplarily, in the scene of tracking the lead singer in a concert, the binocular camera synchronously collects left and right view image data of the stage area. The lead singer target under the spotlight is identified by the YOLOv7 target detection network, and its pixel coordinates (u_L, v_L) in the left image are output. Based on the principle of parallax, SGBM stereo matching is performed on the left and right images to generate an initial disparity map and extract the target region disparity value d. Combined with the camera focal length f_x=1200 pixels and the baseline distance B=0.2 meters, the depth information Z of the lead singer is calculated by the geometric conversion formula Z=B·f_x / d, which is Z=8.5 meters. According to the pixel coordinates (u_L, v_L) and the depth Z, the three-dimensional coordinates of the lead singer center point in the camera coordinate system are calculated by the back projection formula X=(u_L-u_0)·Z / f_x, Y=(v_L-v_0)·Z / f_y, which are X_c=0.3m, Y_c=-0.1m, Z_c=8.5m.

[0067] Further, the depth map obtained by binocular parallax is divided into the following steps: the binocular intrinsic matrix is obtained by using the chessboard calibration through Zhang Zhengyou calibration method. The essential matrix E and the external parameter are solved through feature point comparison and intrinsic matrix. The epipolar correction and stereo matching are performed through the external parameter, and the disparity map is obtained through the SGBM algorithm. The nearby average is used to fill the hollow area of the binocular disparity map. The depth map is generated through the disparity map.

[0068] Since the depth measurement accuracy is improved by stereo matching and geometric conversion fusion, and the target space position is positioned at millimeter level by pixel and three-dimensional coordinate back projection, the stability and positioning real-time performance of tracking in dynamic scenes are improved.

[0069] Based on any of the above embodiments, in embodiment three of the present application, the step S20 includes steps B11-B12: Step B11, based on camera images under different gimbal attitude angles, calibrate the camera optical center, and calculate the coordinate difference with the gimbal rotation center, output the offset.

[0070] In this embodiment, the gimbal attitude angle refers to the real-time angle value of the azimuth and elevation of the gimbal. Joint calibration refers to synchronously collecting multiple groups of data to solve the spatial relationship between the camera and the gimbal. The gimbal rotation center refers to the physical intersection of the two axes of the gimbal.

[0071] As an optional implementation, images of the checkerboard calibration board are taken under different attitudes of the gimbal, the corner pixel coordinates are extracted and the camera extrinsic parameters are solved, the gimbal encoder angles are recorded synchronously, and the translation vector and rotation matrix of the camera optical center relative to the gimbal rotation center are optimized and solved by minimizing the re-projection error, and the offset corresponding to the translation and rotation is output. The advantage of this method is low equipment cost, only a checkerboard is needed and the calibration speed is fast, and the disadvantage is that the precision depends on the stability of the corner detection, which is suitable for consumer-grade gimbal systems.

[0072] As another optional implementation, the gimbal is controlled to rotate to a preset angle sequence, and a laser tracker synchronously measures the physical coordinates of the camera optical center and the gimbal rotation center, and the optimal rigid body transformation parameters are calculated by singular value decomposition, and a high-precision offset is output. The advantage of this method is sub-millimeter level precision and anti-environmental interference, and the disadvantage is that it depends on high-end measuring equipment, which is suitable for military / medical high-precision gimbals.

[0073] Step B12, superimpose the three-dimensional coordinates in the camera coordinate system and the offset through the coordinate transformation relationship to solve and generate the three-dimensional gimbal coordinates.

[0074] As an optional implementation, the pre-calibrated rotation matrix and translation vector are loaded, the three-dimensional coordinates in the camera coordinate system are multiplied by the rotation matrix to obtain intermediate coordinates, and then the translation vector is superimposed to output the final three-dimensional gimbal coordinates. The advantage of this method is high computational efficiency and strong industrial compatibility, which is suitable for scenes where the gimbal elevation angle is limited.

[0075] As another optional implementation, a quaternion is used to represent the rotation relationship, the camera coordinates are converted to quaternion form, the rotation transformation is realized through quaternion multiplication, and then the Euler coordinates are converted back, and the translation offset is added to generate the three-dimensional gimbal coordinates. The advantage of this method is that there is no singularity problem and the motion interpolation is smooth, and the disadvantage is that the calculation is complex, which is suitable for panoramic shooting or mechanical arm large-angle motion scenes.

[0076] Exemplarily, in a sports event live broadcast scene, the sprint process of an athlete in track and field is shot by a binocular camera, and the attitude angles of a holder are recorded synchronously: an azimuth angle of 30° and a pitch angle of -5°. The holder is controlled to rotate in an azimuth angle interval of 0°~90° at a step of 10°, an image of a chessboard calibration board is shot, and corner point pixel coordinates are extracted. The offset of the camera optical center to the holder rotation center is solved by re-projection error minimization optimization based on the holder attitude angle data recorded by the encoder, and a translation vector T=[0.05, -0.02, 0.15] meters and a rotation matrix R are obtained. The three-dimensional coordinates of the athlete in the camera coordinate system, X_c=1.2 m, Y_c=0.3 m, and Z_c=25 m, are input into a coordinate transformation model with the offset, and the three-dimensional holder coordinates in the holder coordinate system, X_p=1.18 m, Y_p=0.28 m, and Z_p=25.1 m, are output through rotation matrix multiplication and translation vector superposition.

[0077] Further, a target tracking algorithm is performed based on the target detection position and depth: the pixel coordinates of the person are converted into the coordinates in the camera coordinate system: (u, v): pixel coordinates of target detection (the upper left corner of the image is the origin (0, 0)); (cx, cy): camera principal point (image center); (fx, fy): camera focal length (pixel unit); Z: depth value (distance from the target to the camera optical center) obtained by binocular distance measurement.

[0078] (X_c, Y_c, Z_c) is the position of the target in the camera coordinate system. The coordinate system is defined as follows: origin: camera optical center; Z axis: pointing to the target direction along the optical axis; X axis: horizontally to the right (consistent with the image coordinate system); Y axis: vertically downward (consistent with the image coordinate system).

[0079] The coordinates in the camera coordinate system are converted into the coordinates in the holder coordinate system: (Tx, Ty, Tz): offset of the camera optical center to the holder rotation center (in the holder coordinate system); (Xp, Yp, Zp): coordinates of the target in the holder base coordinate system.

[0080] Since the non-common origin error of the coordinate system caused by the mechanical installation deviation of the camera and the holder is solved through multi-attitude joint calibration, the mapping of the target position from the camera system to the holder system is realized through rigid body transformation solution, the coordinate projection distortion caused by the rotation of the visual angle is eliminated, and the positioning consistency and motion stability of the target tracking are improved.

[0081] Based on any of the above embodiments, in the fourth embodiment of the present application, the step S30 further comprises steps C11-C12 before the step S30: Step C11, based on the feature information corresponding to the tracking target, if the matching with the historical target feature fails, the tracking target is taken as a new target of the current frame.

[0082] In the present embodiment, the matching failure means that the similarity between the current feature and the historical feature is lower than a set threshold. The new target means a new tracking target instance that does not exist in the historical database. The current frame means the latest image frame collected by the binocular camera.

[0083] As an optional implementation, the feature vector of the tracking target of the current frame is extracted, and cosine similarity calculation is performed on all features in the historical feature library. If the highest similarity value is lower than a preset similarity threshold, it is determined that the matching fails, a unique identifier in a preset format is assigned to the target, the Kalman filter state vector and the covariance matrix are initialized, and the current feature vector is added to the historical feature library, and the target of the current frame is marked as a new target. This method has the advantages of lightweight calculation and globally unique identifier, and the disadvantage is that the similarity is distorted under severe light changes, and it is suitable for indoor light stable scenes.

[0084] As another optional implementation, the feature information is compared with the fast approximate nearest neighbor search library, and if the optimal matching Hamming distance is greater than a preset distance threshold, it is determined that the matching fails, a new target identifier is generated, an empty historical trajectory record is created, an independent Kalman filter is started to track the target, and the tracking target of the current frame is taken as a new target. This method has the advantages of anti-feature scale change and fast search speed, and the disadvantage is that GPU acceleration index construction is required, and it is suitable for outdoor light dynamic scenes.

[0085] Step C12, according to the three-dimensional PTZ coordinates of the new target, a state vector is constructed, and the historical trajectory corresponding to the new target is updated.

[0086] In the present embodiment, constructing a state vector means initializing a state matrix containing target position and velocity components. Updating the historical trajectory means adding the new target coordinates to the exclusive motion trajectory sequence thereof.

[0087] As an optional implementation, a state vector of a preset dimension is created for the new target, the position of the state vector is the current three-dimensional PTZ coordinates, and the velocity is initialized to zero. The covariance matrix is initialized to an identity matrix, and the coordinates are added to the historical trajectory point sequence of the new target, a time stamp index is established, and the historical trajectory corresponding to the new target is generated. This method has the advantages of zero-delay calculation and low resource occupation, and the disadvantage is that the prediction of sudden maneuvering targets is lagging, and it is suitable for uniform speed start scenes.

[0088] As another optional implementation, a long short-term memory network trajectory prediction model is adopted, three-dimensional gimbal coordinates of the new target are input into a pre-trained network to generate an initial state vector, a trajectory cache queue is created to store continuous coordinate points, and a sliding window is used to update a trajectory feature embedding vector. This method has the advantages of modeling complex motion patterns, but has the disadvantages of high GPU inference and high model storage occupation, and is suitable for tracking targets with high-speed changes in direction.

[0089] For example, in a concert spotlight scene, when a new singer suddenly appears on stage under the spotlight, a binocular camera captures image data of the singer, a hierarchical window self-attention visual architecture extracts a singer costume texture feature vector, and a cosine similarity matching is performed with all singer features in the historical feature library. If the highest similarity is 0.7, which is lower than the threshold value 0.8, it is determined that the matching fails. The target is marked as a new target ID_New, and its three-dimensional gimbal coordinates X_p=2.1m, Y_p=-0.3m, and Z_p=15m are obtained. A 6-dimensional state vector [2.1, -0.3, 15, 0, 0, 0] (position + velocity) is constructed, the Kalman filter parameters are initialized, and the coordinates are added to the ID_New exclusive historical trajectory sequence, and an independent tracking thread is started.

[0090] Due to the dynamic feature matching threshold mechanism, the problem of response lag of traditional tracking systems to suddenly appearing targets is solved, the real-time initialization of the state vector is realized, the zero-delay trajectory modeling of the new target is realized, the independent thread management avoids the interference of the new target with the existing target tracking, and the system robustness and tracking coverage of the multi-target emergence are improved.

[0091] Based on any of the above embodiments, in Embodiment Five of the present application, the step S30 includes steps D11-D13: Step D11, according to the matching result of the feature information corresponding to the tracking target and the historical target feature, the tracking target is taken as an index to determine the historical trajectory of the tracking target in the database.

[0092] In this embodiment, the index refers to locating the associated data in the database through a unique identifier. The database refers to a time series data warehouse that stores target historical trajectories.

[0093] As an optional implementation, the feature information of the current tracking target is extracted, and a cosine similarity calculation is performed with the historical feature values in the database. If the highest similarity exceeds the similarity threshold, the unique identifier corresponding to the target is taken as an index key to retrieve the complete historical trajectory point sequence in the database. This method has the advantages of simple query logic and compatibility with lightweight embedded databases, and is suitable for systems with thousands of target scales.

[0094] As another optional implementation, a graph database is used to store the target trajectory relationship network, and a feature vector that matches successfully is used as a node query condition to return a data stream of the historical trajectory of the associated tracking target through a query statement. This mode has the advantages of fast complex relationship retrieval and support for trajectory topology analysis, and is suitable for a smart city security system with a scale of ten thousand targets.

[0095] Step D12, generating a verification parameter based on the historical trajectory and the three-dimensional gimbal coordinate association verification.

[0096] In this embodiment, the association verification refers to verifying the spatial consistency of the current coordinate and the historical trajectory prediction point. The verification parameter refers to an index quantifying the association confidence.

[0097] As an optional implementation, a cubic spline curve is fitted based on the historical trajectory points to generate a current position prediction value, and the Euclidean distance residual of the current position prediction value and the three-dimensional gimbal coordinate is calculated. If the Euclidean distance residual is less than or equal to a preset distance, the verification parameter is output as a first confidence, otherwise the confidence is attenuated according to a negative exponential function. This mode has the advantages of strict position consistency verification, and the disadvantage is that sudden maneuvering targets are easy to trigger confidence attenuation, and is suitable for indoor service robot scenarios with smooth motion trajectories.

[0098] As another optional implementation, the target average speed vector calculated through the historical trajectory points is combined with the current speed vector generated from the current coordinate and the previous point position. According to the target average speed vector and the current speed vector, the cosine value of the included angle between the two is calculated according to the cosine formula, and the cosine value of the motion direction consistency parameter is output as the verification parameter. This mode has the advantages of capturing motion trend mutations, and the disadvantage is that it cannot identify small position shifts of uniform linear motion, and is suitable for direction-sensitive scenarios such as highway vehicle tracking.

[0099] Step D13, updating the three-dimensional gimbal coordinate and its corresponding historical trajectory according to the verification parameter, and outputting the identifier and the three-dimensional position observation value of the tracking target.

[0100] In this embodiment, updating refers to correcting the coordinate or trajectory data based on the verification result.

[0101] As an optional implementation, the three-dimensional gimbal coordinate and the historical trajectory prediction value are weighted and fused according to the confidence of the verification parameter. And the historical trajectory is updated according to the fusion result, and the target identifier and the three-dimensional position observation value are output. This mode has the advantages of suppressing single-frame abnormal jitter and retaining historical motion trend, and the disadvantage is that the trajectory deviates from the true position when the confidence is low, and is suitable for medium and low speed vibration scenarios.

[0102] As another optional implementation, when the check parameter is the direction consistency coefficient, if the check parameter is greater than or equal to the preset parameter threshold, the three-dimensional gimbal coordinates are directly output as the three-dimensional position observation value. Otherwise, the historical trajectory moving average value is used as the observation value, the trajectory database is updated synchronously, the target identifier and the corrected three-dimensional position observation value are output. This method has the advantages of fast response speed and enabling the historical average value to prevent drift when the direction suddenly changes. The disadvantage is that the sliding window average value blurs the instantaneous position, and it is suitable for tracking high-speed maneuvering targets.

[0103] Exemplarily, in a concert spotlight scene, the lead singer suddenly appears from the stage riser, and the binocular camera captures the image thereof, a 1024-dimensional lead singer costume texture feature vector is extracted through the hierarchical window self-attention visual architecture, cosine similarity matching is performed on all singer features in the historical feature library, the highest similarity 0.75 is lower than the threshold 0.8, it is determined that the matching fails, the target is assigned an identifier ID_New as a new target, and it is confirmed that there is no historical trajectory in the database. Based on the current frame three-dimensional gimbal coordinates: X_p=1.2m, Y_p=-0.5m, Z_p=10m, an empty trajectory sequence is constructed. A predicted position is generated through a cubic spline curve fitting: X_pred=1.18m, Y_pred=-0.48m, Z_pred=10.1m, the Euclidean distance residual Δd=0.05m is calculated from the current coordinates, because Δd<0.3m threshold, the check parameter confidence is 0.98. The current coordinates and the predicted value are fused with a weight of 0.98:0.02, the three-dimensional gimbal coordinates are updated to X_obs=1.199m, Y_obs=-0.499m, Z_obs=10.01m, the historical trajectory database of ID_New is updated synchronously, and the identifier ID_New and the three-dimensional position observation value X_obs, Y_obs, Z_obs are output.

[0104] Due to the binding mechanism of dynamic features and trajectories, the trajectory prediction cold start problem when the historical data of a new target is missing is solved, the target motion state is seamlessly connected through real-time trajectory updating, and the tracking smoothness and positioning accuracy in a sudden target surge scene are improved.

[0105] Based on any of the above embodiments, in the sixth embodiment of the present application, the step S40 includes steps E11-E13: Step E11, based on the identifier of the tracking target, searching the Kalman filter state vector and state covariance matrix corresponding to the identifier in the database.

[0106] In this embodiment, searching refers to the operation of querying data records through an index key. The Kalman filter state vector refers to a matrix containing state variables such as target position and velocity. The state covariance matrix refers to an error covariance matrix that describes the uncertainty of state estimation.

[0107] As an optional implementation, a retrieval operation is performed in the key-value database according to the target identifier, a pre-stored structure is retrieved, a state vector field and a covariance matrix field therein are parsed, and the state vector and the covariance matrix are output to a Kalman filter update process. This mode has the advantages of intuitive data structure and compatibility with cross-platform systems, and is suitable for small and medium-sized tracking systems.

[0108] As another optional implementation, the latest timestamp record is queried in the time series database based on the target identifier, and a binary encoding stream of the state vector and the covariance matrix is extracted. This mode has the advantages of high transmission efficiency and compact storage, and the disadvantage of strict data alignment, and is suitable for high-speed real-time systems.

[0109] Step E12, inputting the three-dimensional position observation value, the Kalman filter state vector and the state covariance matrix into an observation update equation to generate a real-time three-dimensional position and a real-time velocity of the tracking target.

[0110] In this embodiment, the real-time three-dimensional position refers to the optimal position estimation of the target after correction in the gimbal coordinate system. The real-time velocity refers to the velocity component of the target in the XYZ direction after correction.

[0111] As an optional implementation, the three-dimensional position observation value is input as an observation value, and the Kalman filter current state vector and the state covariance matrix are input into the observation update equation to calculate the Kalman gain matrix. The corrected real-time three-dimensional position and real-time velocity are output by weighted fusion of the observation value and the predicted value, and the state covariance matrix is updated synchronously. This mode has the advantages of standardized calculation process and low delay, and the disadvantage of fixed observation noise parameter leading to estimation divergence in the presence of sudden interference, and is suitable for tracking scenarios with stable sensor noise.

[0112] As another optional implementation, an adaptive observation noise mechanism is adopted to dynamically adjust the observation noise covariance matrix according to the observation residual. When the residual exceeds a preset residual threshold, the value of the noise covariance matrix is increased, the observation weight is reduced, and the real-time three-dimensional position and real-time velocity are output with anti-interference. This mode has the advantages of strong robustness and no need for manual parameter tuning, and the disadvantage of increased calculation amount for real-time residual statistics, and is suitable for outdoor security scenarios with frequent occlusions.

[0113] Step E13, predicting the three-dimensional predicted position of the tracking target at the next frame timestamp according to the real-time three-dimensional position and the real-time velocity in time sequence.

[0114] In this embodiment, time series prediction refers to extrapolating future states in time sequence based on historical states. The next frame time refers to the acquisition time of the subsequent image frame.

[0115] As an optional implementation, based on real-time three-dimensional position and real-time speed, the displacement increment of the next frame timestamp is calculated by the uniform motion model, and the three-dimensional prediction position is output by superimposing the current position. The advantage of this method is that the calculation is simple and the delay is low, and the disadvantage is that the high-speed target prediction is lagged due to the neglect of acceleration, which is suitable for low-speed uniform speed scene.

[0116] As another optional implementation, the state transition matrix multiplication is adopted, the state vector containing position and speed is input into the predefined transition matrix, and the next frame state vector is directly output by matrix multiplication, and the position component is extracted as the three-dimensional prediction position. The advantage of this method is that it strictly conforms to the Kalman filter theoretical framework and the state is updated synchronously, and the disadvantage is that the calculation amount is high, which is suitable for high-speed and high-maneuvering target.

[0117] Exemplarily, in the live scene of a car race, the tracking target is car No. 7, and the Kalman filter state vector of the tracking target is retrieved through the identifier ID_007 in the Redis database: [X=102.3m, Y=-5.2m, Z=150m, V_x=60m / s, V_y=0.3m / s, V_z=0] and the state covariance matrix. The three-dimensional position observation value output by the binocular system: X_obs=102.5m, Y_obs=-5.1m, Z_obs=149.8m is input into the observation update equation, and the real-time three-dimensional position after fusion is output: X_est=102.4m, Y_est=-5.15m, Z_est=149.9m and the real-time speed: V_x=60.2m / s, V_y=0.28m / s, V_z=0.01m / s. Based on the real-time position and speed, the next frame three-dimensional prediction position is predicted by the uniform model according to the frame interval Δt=0.033s: X_pred=102.4+60.2×0.033≈104.4m, Y_pred=-5.15+0.28×0.033≈-5.14m, Z_pred=149.9m, and the gimbal is driven to turn ahead to lock the car trajectory.

[0118] Since the dynamic weighted fusion of observation and prediction is adopted, the single-frame measurement noise interference is overcome, the high-speed target is tracked with zero lag through kinematic extrapolation prediction, and the tracking smoothness and trajectory prediction accuracy in high-speed dynamic scene are improved.

[0119] Based on any of the above embodiments, in the seventh embodiment of the present application, the step S50 comprises steps F11-F14: Step F11, inputting the three-dimensional prediction position corresponding to the tracking target into the bias formula, outputting the horizontal angle bias and the pitch angle bias of the tracking target in the gimbal coordinate system.

[0120] In the embodiment, the horizontal angle deviation refers to the angle of the tracking target deviating from the horizontal direction of the optical axis. The pitch angle deviation refers to the angle of the tracking target deviating from the vertical direction of the optical axis.

[0121] As an optional implementation, based on the three-dimensional predicted position, the horizontal angle deviation and the pitch angle deviation are calculated through a geometric projection relationship, and two-axis angle deviation values are output. This mode has the advantages of efficient calculation and no additional parameters, and the disadvantage is that the angle error of the target at the edge of the wide-angle lens increases due to uncorrected lens distortion, and is suitable for medium and long focal gimbal systems.

[0122] As another optional implementation, camera distortion correction parameters are introduced, the three-dimensional predicted position is converted into normalized camera coordinates, and the angle deviation is calculated through the inverse tangent function. This mode has the advantages of eliminating barrel / pincushion distortion, and the disadvantage is high computational load, which is suitable for wide-angle / fisheye lenses.

[0123] Step F12, based on the horizontal angle deviation and the pitch angle deviation, a deviation correction strategy is generated according to a preset deviation adjustment rule.

[0124] In the embodiment, the preset deviation adjustment rule refers to a pre-defined mapping strategy of deviation and control quantity. The deviation correction strategy refers to a control instruction for driving the gimbal motor to move.

[0125] As an optional implementation, based on the horizontal angle deviation and the pitch angle deviation, a preset parameter table is queried, and azimuth axis angular velocity instructions and pitch axis angular velocity instructions are generated through independent feedback controllers. According to the azimuth axis angular velocity instructions and the pitch axis angular velocity instructions, a deviation correction strategy is generated and output according to the preset deviation adjustment rule. This mode has the advantages of fast response and adjustable parameters online, and the disadvantage is that fixed parameters are difficult to adapt to nonlinear disturbances, which is suitable for indoor shooting scenes with constant gimbal load.

[0126] As another optional implementation, a fuzzy control rule base is used, the values and rates of change of the horizontal angle deviation and the pitch angle deviation are input into a fuzzy inference machine, continuous angular velocity instructions are generated by de-fuzzing through a rule matrix, and a high-robustness deviation correction strategy is output after dead zone compensation. This mode has the advantages of dynamically adapting to complex disturbances and not requiring an accurate model, and the disadvantage is that the rule base design is complex, which is suitable for outdoor unmanned aerial vehicle follow-up shooting and other strong interference scenes.

[0127] Step F13, based on the device parameters of the gimbal and the deviation correction strategy, a correction parameter is generated through a cooperative mechanism of feedforward compensation and feedback adjustment.

[0128] In this embodiment, the gimbal device parameters refer to the inherent properties of the gimbal hardware. Feedforward compensation refers to predicting disturbances based on system models and applying control quantities in advance. Feedback regulation refers to closed-loop correction of control output based on real-time deviations. The cooperative mechanism refers to the algorithm architecture of superimposed fusion of feedforward and feedback control quantities. The correction parameter refers to the optimized control quantity of the final output.

[0129] As an optional implementation, load the gimbal device parameters, and output the angular velocity command based on the deviation correction strategy. Calculate the inertial torque through feedforward compensation, superimpose the torque output by the feedback controller, generate the total torque command, and convert it into the motor current command as the correction parameter output. This method has the advantages of fast dynamic response and high steady-state accuracy, and the disadvantage is that it depends on accurate inertia calibration, which is suitable for industrial robot gimbals with known load parameters.

[0130] As another optional implementation, adaptive feedforward and feedback are used in cooperation, and the feedforward weight is dynamically adjusted according to the real-time tracking error. When the error is large, the feedforward proportion is increased. The device parameters are mapped to the feedforward gain coefficient through a fuzzy rule base, and the correction parameter is calculated and generated according to the feedforward weight and the feedforward gain coefficient. This method has the advantages of dynamic adaptation to load changes and strong resistance to external disturbances, and is suitable for movie-level electric zoom gimbals.

[0131] Step F14, according to the correction parameter, drive the rotation of the gimbal motor, and center the tracking target in the center position of the camera image.

[0132] As an optional implementation, input the correction parameter into the motor driver, generate the corresponding torque through a three-phase inverter circuit to drive the gimbal motor to rotate, and link the camera optical axis to align the target space position. Real-time detection of the pixel coordinates of the target in the image until it coincides with the preset center point. This method has the advantages of fast torque response and strong overload capacity, and the disadvantages are that it requires a high-voltage drive circuit and high cost, and is suitable for heavy-duty movie gimbals.

[0133] As another optional implementation, generate a pulse width modulation signal based on the correction parameter to drive the motor control chip to adjust the motor winding voltage, and control the gimbal rotation angle through encoder feedback closed loop control, so that the real-time coordinates of the tracking target in the image converge to the center area. This method has the advantages of simple circuit and low power consumption, and is suitable for lightweight consumer gimbals.

[0134] Exemplarily, in a live broadcast scene of a car race, the tracking target is a car No. 9, based on the three-dimensional predicted position of the car: X_pred=215.6 m, Y_pred=-3.2 m, Z_pred=300 m, the horizontal angle deviation Δθ=arctan(215.6 / 300)≈35.7° and the pitch angle deviation Δφ=arctan(-3.2 / 300)≈-0.61° are calculated by inputting the deviation formula. According to the preset PID rule: K_p=1.5, K_i=0.02, K_d=0.4, the angular velocity correction strategy ω_θ=2.3 rad / s, ω_φ=-0.08 rad / s is generated. Combined with the gimbal device parameters, the moment of inertia J=0.015, the motor torque coefficient K_t=0.12, the feedforward compensation inertia torque τ_ff=J·dω / dt is superimposed on the feedback output to generate the corrected current command I_θ=19.2 A, I_φ=-0.67 A. The brushless motor of the gimbal is driven to rotate, and after 0.12 seconds of dynamic adjustment, the car imaging point is accurately centered in the center of the camera screen: the coordinates are (640, 360) under the resolution of 1280×720, and the error is <±5 pixels.

[0135] Further, the angle acquisition method is as follows: Calculate the Pan yaw axis angle of the gimbal. Pan is the rotation angle of the gimbal around the vertical axis (Y axis), which is in radian here: ; Calculate the Tilt pitch axis angle of the gimbal. Tilt is the rotation angle of the gimbal around the horizontal axis (Z axis or X axis), which is in radian here: ; Convert the radian value to the corresponding angle value, and send the angle to the gimbal. The gimbal adjusts the angle to center the person: ; ; Due to the joint control chain of geometry and physics, the problem of high-speed target large-angle steering lag is solved, and the motor start-stop jitter is suppressed by feedforward inertia compensation, which improves the stability and picture composition professionalism of target tracking in the super-speed scene.

[0136] Based on any of the above embodiments, in embodiment eight of the present application, the step A13 includes steps G11-G14: Step G11, based on the completeness rule of the disparity map, verify the completeness of the initial disparity map, and generate a completeness judgment result.

[0137] In this embodiment, verification refers to detecting the distribution and proportion of invalid regions in the disparity map. The completeness judgment result refers to a hierarchical index quantifying the effectiveness of the disparity map.

[0138] As an optional implementation, the initial disparity image pixels are traversed to count the area of the connected region of invalid values. If the maximum hollow area exceeds a preset invalid threshold or the total invalid pixel ratio is greater than a preset invalid ratio, a completeness judgment result is output: the completeness is determined to be at a "warning" level, otherwise, it is marked as "complete". This method has the advantages of comprehensive detection and intuitive parameters, and the disadvantage is high computational load, which is suitable for offline depth map quality evaluation scenarios.

[0139] As another optional implementation, a sliding window is used to scan the initial disparity map, and the proportion of valid disparity points in each window is calculated. If the valid proportion of a continuous preset number of windows is a preset valid proportion, it is determined to be "invalid", otherwise, the judgment result is output according to the global valid point proportion. This method has the advantages of efficient calculation and local sensitivity, and the disadvantage is that small area hollows may be missed, which is suitable for embedded real-time systems.

[0140] Step G12, if the completeness judgment result shows that there is a hollow in the initial disparity map, the valid pixels adjacent to the hollow are collected, and the pixel mean value corresponding to the valid pixels is calculated and generated.

[0141] In this embodiment, the hollow refers to the invalid value region in the disparity map. The adjacent valid pixels refer to the valid disparity value pixels adjacent to the hollow edge. The pixel mean value refers to the arithmetic mean of the valid pixel disparity values.

[0142] As an optional implementation, if the completeness judgment result is that there is a hollow, the hollow region boundary pixel points are located, all valid disparity value pixels in a preset neighborhood range are collected, the disparity mean value of these pixels is calculated as the filling value, and the repair disparity value of the hollow region is generated. This method has the advantages of efficient calculation and smooth repair edge, and the disadvantage is that it cannot be filled when the neighborhood valid pixels are missing, which is suitable for small area hollow repair scenarios.

[0143] As another optional implementation, an adaptive neighborhood expansion strategy is used to expand the annular region outward layer by layer from the center of the hollow until the number of valid pixels collected is greater than or equal to a preset number of valid pixels, the weighted mean value of the valid pixels is calculated, and the hollow filling value is output. This method has the advantages of intelligent adaptation to hollow size and resistance to isolated noise, and the disadvantage is complex calculation, which is suitable for large area hollows caused by complex occlusion.

[0144] Step G13, based on the mean filling strategy, the pixel mean value is introduced to fill the hollow in the initial disparity map, and an optimized depth map is generated.

[0145] In this embodiment, the mean filling strategy refers to an algorithm for filling the hollow region with the disparity average value of the valid pixels. The optimized depth map refers to the complete depth information matrix after the hollow is repaired.

[0146] As an optional implementation, the boundary of the hole region in the initial parallax map is located, all valid parallax pixel values in the neighborhood of the hole region are collected, an arithmetic mean value is calculated as a filling value, all invalid pixel values in the hole region are replaced by the calculated pixel mean value, and a continuous and complete optimized depth map is generated. The advantage of this method is that the calculation is simple and the depth after repair is smooth, and the disadvantage is that it cannot be filled when the neighborhood is completely invalid, and it is suitable for small area hole scenes caused by sensor noise.

[0147] As another optional implementation, a spatial weighted mean filling is used, the center of the hole is taken as the origin, the weight is assigned according to the pixel distance, the weighted mean value is calculated to fill the hole, and the preset pixel area median filter is used to smooth the edge transition area, and an anti-noise optimized depth map is generated. The advantage of this method is that it can repair large holes and the edge transition is natural, and the disadvantage is that the calculation load is high, and it is suitable for natural scene occlusion repair.

[0148] Step G14, according to the depth map, camera focal length and baseline distance, input into the geometric conversion formula to calculate, generate the depth information of the pixel coordinates relative to the gimbal.

[0149] As an optional implementation, based on the depth value corresponding to the target center pixel coordinates in the depth map, the camera focal length and the baseline distance are combined, and the depth information of the target relative to the gimbal is calculated through the principle of triangulation, and the depth information is output. The advantage of this method is that the calculation is efficient and does not require additional parameters, and the disadvantage is that it depends on the accuracy of the parallax map, and it is suitable for high-quality scenes.

[0150] As another optional implementation, a parallel computing architecture is used to perform sub-pixel interpolation optimization on the depth map, the pixel coordinates are converted into normalized coordinates by combining the camera distortion correction parameters, and then the accurate depth information under the tilted viewing angle is calculated through the geometric conversion formula. The advantage of this method is that it is anti-lens distortion and sub-pixel smooth, and the disadvantage is that it needs GPU acceleration, and it is suitable for fisheye lens or large-angle shooting scenes.

[0151] Exemplarily, in a stage smoke interference scene, a binocular camera captures a main singer target to generate an initial parallax map. Based on the parallax map integrity rule check, it is found that the hole region accounts for 8% (more than the 5% threshold), and the integrity is determined to be "invalid". The hole boundary is located, and the effective pixel values in the 8-neighborhood are collected, such as the parallax sequence [35, 37, 36], and the pixel mean value 36 is calculated. The mean filling strategy is used to replace the hole region with 36, and an optimized depth map is generated. Combined with the camera focal length f_x=1200 pixels and the baseline distance B=0.2 meters, the target depth information Z=6.7 meters is solved through the geometric conversion formula Z=B·f_x / d (d is the parallax), and output to the gimbal tracking system.

[0152] Further, with reference to Figure 3 , Figure 3The flowchart of the pan-tilt tracking method based on the binocular camera of the present application. The binocular camera synchronously collects left and right view image data, performs "target detection and depth calculation", identifies the tracking target through the target detection network and extracts the pixel coordinates, and calculates the depth value based on the binocular parallax principle to generate the three-dimensional coordinates of the target center point. Enter the "whether it is a new target" judgment link, if the feature matching result confirms that it is a new target, "initialize 3D Kalman filter target", construct the state vector and covariance matrix containing position and velocity components. If it is not a new target, perform "target association matching", associate the current detection target with the historical trajectory and output the identifier and three-dimensional position observation value. Whether it is a new target or not, enter "update Kalman filter state", retrieve the corresponding Kalman filter parameters according to the identifier, and update the state vector by fusing the three-dimensional position observation value. Based on the updated state, perform "predict the next frame target position", output the three-dimensional predicted coordinates of the target at the next time. According to the predicted coordinates, "calculate the relative angle of the pan-tilt", calculate the horizontal angle deviation Δθ = arctan (X_p / Z_p) and the pitch angle deviation Δφ = arctan (Y_p / Z_p) through the deviation formula. Finally, "the pan-tilt executes the command to center the person", generate control instructions to drive the motor to rotate according to the angle deviation, so that the target dynamically aligns with the center of the camera picture, and the process ends in a closed loop control.

[0153] Further, convert the target detection width and height to real width and height in the 3D coordinate system: ; : detection frame width and height (pixels); : camera focal length parameter; Use the 3D coordinate system coordinates and the pan-tilt coordinate system width and height of the person to construct the state vector:

[0154] Position (m): ; Size (m): (width), (height); Velocity (m / s): ; Size change rate (m / s): .

[0155] Calculate the prediction stage and the prediction covariance. State transition equation: ; Covariance prediction: ; State transition matrix: ; Process noise covariance: ; where: ; ; : time step (seconds); : acceleration noise standard deviation of position change (meters per squared second), quantifying uncertainty in target acceleration; : noise standard deviation of velocity change (meters per squared second), quantifying uncertainty in target velocity; : acceleration noise standard deviation of size change (meters per squared second), quantifying uncertainty in target size; : noise standard deviation of size change rate (meters per squared second), quantifying uncertainty in size change rate.

[0156] Compute Mahalanobis distance, observation model: ; Observation matrix: ; Mahalanobis distance computation: ; where .

[0157] Size similarity weighting: ; This scheme weighting weight: , , .

[0158] State update, observation noise covariance: ; where: ; focal length, , .

[0159] Compute Kalman gain: ; State update: ; Covariance update: .

[0160] Due to the hollow intelligent repair mechanism, the depth information fault problem caused by smoke shielding is solved, the continuity of the depth space is ensured through the neighborhood mean filling, and the robustness and measurement accuracy of the depth perception in the strong interference scene are improved.

[0161] The application provides a gimbal tracking device, which comprises at least one processor and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the binocular camera-based gimbal tracking method in the above embodiment one.

[0162] Reference will be made to the following Figure 4 , which shows a structural schematic diagram of a gimbal tracking device suitable for implementing the embodiments of the application. The gimbal tracking device in the embodiments of the application can include but is not limited to mobile terminals such as mobile phones, notebook computers, gimbal trackers, personal digital assistants (PDA), tablet computers (PAD), portable multimedia players (PMP), gimbal terminals, and the like, and fixed terminals such as vehicle-mounted gimbals and desktop computers. Figure 4 The shown gimbal tracking device is only an example, and should not bring any limitation to the functions and use range of the embodiments of the application.

[0163] As Figure 4As shown, the gimbal tracking device can include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. Various programs and data required for the operation of the gimbal tracking device are also stored in the random access memory 1004. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other by a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; the storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the gimbal tracking device to communicate with other devices wirelessly or by wire to exchange data. Although the gimbal tracking device with various systems is shown in the figure, it should be understood that all the systems shown are not required to be implemented or possessed. More or fewer systems can be alternatively implemented or possessed.

[0164] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by a communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0165] The gimbal tracking device provided by the present application adopts the gimbal tracking method based on the binocular camera in the above-mentioned embodiments, which can solve the technical problem that the target tracking effect is poor due to the lack of a depth change compensation mechanism for feature scale, which cannot distinguish between target real motion and apparent deformation. Compared with the prior art, the gimbal tracking device provided by the present application has the same beneficial effects as the gimbal tracking method based on the binocular camera provided by the above-mentioned embodiments, and the other technical features in the gimbal tracking device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.

[0166] It should be understood that various aspects of the disclosure can be implemented in hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any appropriate manner in any one or more embodiments or examples.

[0167] The above description is merely illustrative of the application and is not intended to limit the scope of the application. Any variations and modifications that can be made by any person skilled in the art within the spirit and scope of the application are intended to be encompassed by the application. The scope of the application is defined by the appended claims.

[0168] The application provides a computer readable storage medium having stored thereon computer readable program instructions (i.e., computer programs) for performing the gimbal tracking method based on a binocular camera in the above embodiments.

[0169] The computer readable storage medium provided by the application may, for example, be a U disk, but is not limited to an electric, magnetic, optical, electromagnetic, infrared, or semiconductor system, system, or device, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electric connection with one or more conductive wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination of the above. In the present embodiment, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer readable storage medium can be transmitted by any appropriate medium, including but not limited to an electric wire, an optical cable, a radio frequency (RF), and the like, or any appropriate combination of the above.

[0170] The above computer readable storage medium can be contained in the gimbal tracking device; or can exist separately and not be assembled into the gimbal tracking device.

[0171] The computer readable storage medium carries one or more programs, when the one or more programs are executed by the gimbal tracking device, the gimbal tracking device is caused to: based on image data output by a binocular camera, locate and track a target through feature information matching, and calculate depth information according to a difference of pixel positions of the target in the image data, to solve and generate three-dimensional coordinates of a center point of the target in a camera coordinate system of a current frame; based on the three-dimensional coordinates and an offset amount of a camera optical center to a gimbal rotation center, solve and generate three-dimensional gimbal coordinates of the target through a coordinate transformation relationship; based on a matching result of feature information corresponding to the target and historical target features, determine a historical trajectory of the target, and through correlation verification of the three-dimensional gimbal coordinates corresponding to the target and the historical trajectory corresponding to the target, output an identifier of the target and a three-dimensional position observation value of the target; by inputting the identifier of the target and the three-dimensional position observation value of the target to an observation update equation to correct a state, output a three-dimensional predicted position of the target; based on the three-dimensional predicted position corresponding to the target and a deviation formula, solve an angle deviation of the three-dimensional predicted position from a center of a camera image in a gimbal coordinate system, and drive a gimbal motor to rotate according to the angle deviation, so as to center the target in a center of a camera image.

[0172] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0173] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0174] The modules involved in the embodiments of the present application can be implemented in the form of software or in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself.

[0175] The computer readable storage medium provided by the present application is a computer readable storage medium, which stores computer readable program instructions (i.e. computer program) for executing the gimbal tracking method based on the binocular camera. The technical problem of poor target tracking effect caused by the lack of depth change compensation mechanism for feature scale and the inability to distinguish between target real motion and apparent deformation can be solved. Compared with the prior art, the computer readable storage medium provided by the present application has the same beneficial effects as the gimbal tracking method based on the binocular camera provided by the above-mentioned embodiments. Details are not repeated here.

[0176] The above only describes some embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the present application specification and drawings, or direct / indirect application in other related technical fields is included in the patent protection scope of the present application.

Claims

1. A pan-tilt tracking method based on a binocular camera, characterized in that: The method comprises: Based on the image data output by the binocular camera, the tracking target is located by matching feature information, and the depth information is calculated based on the difference in the corresponding pixel positions of the tracking target in the image data, and the three-dimensional coordinates of the center point of the tracking target in the current frame camera coordinate system are calculated; Based on the three-dimensional coordinates and the offset from the camera optical center to the gimbal rotation center, the three-dimensional gimbal coordinates of the tracking target are generated by solving the coordinate transformation relationship; Based on the matching result of the feature information corresponding to the tracking target and the historical target features, the historical trajectory of the tracking target is determined, and the three-dimensional pan-tilt coordinates corresponding to the tracking target are correlated and verified with the historical trajectory corresponding to the tracking target, and the identifier of the tracking target and the three-dimensional position observation value of the tracking target are output; By inputting the identifier of the tracking target and the three-dimensional position observation value of the tracking target into the observation update equation correction state, the three-dimensional predicted position of the tracking target is output; Based on the three-dimensional predicted position corresponding to the tracking target and the deviation formula, the angular deviation between the three-dimensional predicted position and the center of the camera image in the gimbal coordinate system is calculated, and the gimbal motor is driven to rotate according to the angular deviation so that the tracking target is centered in the center of the camera image.

2. The pan-tilt tracking method based on a binocular camera according to claim 1, wherein: The steps of locating the tracking target by matching feature information based on the image data output by the binocular camera, calculating depth information according to the difference in the pixel positions corresponding to the tracking target in the image data, and solving and generating the three-dimensional coordinates of the center point of the tracking target in the current frame camera coordinate system include: Identifying the tracking target and obtaining pixel coordinates of the tracking target in the image data by using the feature information extracted from the image data; Based on the same pixel position of the left and right images corresponding to the image data, stereo matching the corresponding pixel coordinates to generate an initial disparity map; Generate depth information of the pixel coordinates relative to the gimbal by calculating through a geometric transformation formula based on the initial disparity map, the camera focal length, and the baseline distance; Generate three-dimensional coordinates of a center point of the tracking target in a camera coordinate system according to the pixel coordinates and the depth information of the tracking target.

3. The pan-tilt tracking method based on a binocular camera according to claim 1, wherein: The step of generating the three-dimensional gimbal coordinates of the tracking target by solving the coordinate transformation relationship based on the three-dimensional coordinates and the offset from the camera optical center to the gimbal rotation center includes: Based on the camera images at different gimbal attitude angles, the camera optical center is calibrated, and the coordinate difference with the gimbal rotation center is solved to output the offset; The three-dimensional coordinates in the camera coordinate system and the offset are correspondingly superimposed through a coordinate transformation relationship to generate the three-dimensional gimbal coordinates.

4. The pan-tilt tracking method based on a binocular camera according to claim 1, wherein: Before the steps of determining the historical trajectory of the tracking target based on the matching result of the feature information corresponding to the tracking target and the historical target features, and outputting the identifier of the tracking target and the three-dimensional position observation value of the tracking target by correlating and verifying the three-dimensional pan-tilt coordinates corresponding to the tracking target with the historical trajectory corresponding to the tracking target, the pan-tilt tracking method based on the binocular camera further includes: If the feature information corresponding to the tracked target fails to match the feature of the historical target, the tracked target is used as a new target in the current frame; According to the three-dimensional pan-tilt coordinates of the newly added target, a state vector is constructed, and the historical trajectory corresponding to the newly added target is updated.

5. The pan-tilt tracking method based on binocular cameras according to claim 1, wherein: The steps of determining the historical trajectory of the tracking target based on the matching result of the feature information corresponding to the tracking target and the historical target features, and outputting the identifier of the tracking target and the three-dimensional position observation value of the tracking target by correlating and verifying the three-dimensional pan-tilt coordinates corresponding to the tracking target with the historical trajectory corresponding to the tracking target include: According to the matching result of the feature information corresponding to the tracking target and the feature of the historical target, the tracking target is used as an index to determine the historical trajectory of the tracking target in the database; Verify the three-dimensional gimbal coordinates based on the historical trajectory to generate verification parameters; The three-dimensional pan-tilt coordinates and their corresponding historical trajectories are updated according to the verification parameters, and the identifier and the three-dimensional position observation value of the tracking target are output.

6. The pan-tilt tracking method based on binocular cameras according to claim 1, wherein: The step of inputting the identifier of the tracking target and the three-dimensional position observation value of the tracking target into the observation update equation correction state to output the three-dimensional predicted position of the tracking target includes: Based on the identifier of the tracking target, searching a database for a Kalman filter state vector and a state covariance matrix corresponding to the identifier; Inputting the three-dimensional position observation value, the Kalman filter state vector and the state covariance matrix into the observation update equation to generate the real-time three-dimensional position and real-time velocity of the tracked target; The three-dimensional predicted position of the tracking target at the next frame timestamp is predicted according to the real-time three-dimensional position and the real-time speed in a time series.

7. The pan-tilt tracking method based on a binocular camera according to claim 1, wherein: The step of calculating the angular deviation between the three-dimensional predicted position and the camera image center in the gimbal coordinate system based on the three-dimensional predicted position corresponding to the tracked target and the deviation formula, and driving the gimbal motor to rotate according to the angular deviation so that the tracked target is centered in the camera image includes: Inputting the three-dimensional predicted position corresponding to the tracked target into the deviation formula, and outputting the horizontal angle deviation and pitch angle deviation of the tracked target in the gimbal coordinate system; Based on the horizontal angle deviation and the pitch angle deviation, generating a deviation correction strategy according to a preset deviation adjustment rule; Based on the device parameters of the gimbal and the deviation correction strategy, correction parameters are generated through a coordinated mechanism of feedforward compensation and feedback regulation; According to the correction parameters, the pan / tilt motor is driven to rotate so as to center the tracking target at the center of the camera image.

8. The pan-tilt tracking method based on binocular cameras according to claim 2, wherein: The step of generating depth information of the pixel coordinate relative to the gimbal by calculating based on the initial disparity map, the camera focal length and the baseline distance through a geometric conversion formula includes: Based on the disparity map integrity rule, verifying the integrity of the initial disparity map and generating an integrity judgment result; If the integrity judgment result shows that there is a hole in the initial disparity map, collecting valid pixels adjacent to the hole and calculating and generating pixel means corresponding to the valid pixels; Based on the mean filling strategy, the pixel mean is introduced to fill the holes in the initial disparity map to generate an optimized depth map; The depth map, camera focal length and baseline distance are input into a geometric transformation formula to generate depth information of the pixel coordinates relative to the gimbal.

9. A pan-tilt tracking device, characterized in that: The pan-tilt tracking device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the pan-tilt tracking method based on a binocular camera as described in any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the pan-tilt tracking method based on a binocular camera as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Multi-target continuous tracking system and method based on cloud deck control technology

    CN108062115A

  • Biaxial pan-tilt target tracking system and method thereof

    CN115457076A

  • Camera automatic tracking method and system based on holder control

    CN118394134A

Cited By

  • Level consistency judgment method and system, electronic equipment and computer readable storage medium

    CN121074039A

  • Panorama camera and turntable mobile measurement matching method and system

    CN121304806A

  • Panoramic camera and turntable mobile measurement calibration method and system

    CN121304806B

  • Binocular camera position adjusting method for XR equipment binocular active alignment equipment

    CN121560082A

  • A binocular camera alignment method for a binocular active alignment device of an XR device

    CN121560082B