Target tracking method and device based on rolling and pitching type holder
By employing a roll-up gimbal structure and a collaborative control algorithm, the problem of target tracking in complex environments was solved, achieving high-precision, low-latency target tracking, improving tracking accuracy and anti-interference capabilities, and reducing mechanical load.
Patent Information
- Application Number
- CN202510998039.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-20
- Publication Date
- 2025-10-28
AI Technical Summary
Existing gimbal devices suffer from problems such as decreased recognition accuracy, slow response speed, control mismatch, and excessive mechanical load when tracking targets in complex environments, making it difficult to achieve high-precision, low-latency target tracking.
Employing a roll-up gimbal structure, combined with adaptive color calibration, multi-scale target detection, template matching, and Kalman filter prediction, and through a dual-loop collaborative control algorithm of PID and FOC, the gimbal's attitude reference and mechanical transmission are optimized to achieve high-precision, low-latency target tracking.
It significantly improves the accuracy, stability, and anti-interference ability of target tracking, optimizes the performance of target tracking tasks in dynamic and complex scenarios, reduces mechanical load, and improves response speed.
Smart Images

Figure CN120852815A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automation control technology, and in particular to a target tracking method and device based on a roll-up gimbal. Background Art
[0002] In recent years, target tracking technology has developed rapidly in the field of intelligent mobile platforms and has been widely used in scenarios such as aerial photography, agricultural plant protection, power line inspection, emergency rescue, and military confrontation. With the advancement of technology, tracking stability, anti-interference ability, and device response speed in complex environments have become research focuses. Among these, the robustness of target tracking methods and the control accuracy of gimbal devices directly affect the performance of target tracking tasks.
[0003] Traditional target tracking methods often rely on single feature matching. When the target lighting changes or it encounters occlusion, the recognition accuracy drops significantly, making it difficult to quickly re-lock onto the target. This causes the gimbal to fail to accurately obtain deviation angle information, resulting in tracking offset. At the same time, most algorithms are not optimized for the motion characteristics of the gimbal, and the computational efficiency is not matched with the response speed of control commands, resulting in low consistency between the smoothness of the gimbal rotation and the target's motion trajectory.
[0004] Most existing gimbal devices use a traditional three-axis stabilization structure. While this can mitigate attitude changes to some extent, such structures are often bulky and heavy, increasing the load on the intelligent mobile platform and reducing its endurance. Furthermore, the control algorithms of traditional gimbals are relatively complex, and they are prone to response lag when rapidly tracking moving targets, making it difficult to meet the demands for low-latency, high-precision tracking.
[0005] Therefore, existing gimbal devices still have limitations in target tracking, and there is an urgent need to improve tracking accuracy and stability in complex environments. Summary of the Invention
[0006] To address the aforementioned issues, this invention provides a target tracking method and apparatus based on a roll-and-tilt gimbal. Through adaptive color calibration, multi-scale target detection, template matching, and Kalman filter prediction, it can ensure high-precision target position prediction even in scenarios where the target is temporarily occluded. In particular, the gimbal employs a roll-and-tilt structure, realizing an attitude reference mechanism. Based on a dual-loop collaborative control algorithm of PID and FOC, it compensates for disturbances in the intelligent mobile platform in real time, reducing response lag caused by mechanical transmission backlash, and achieving high-precision, low-latency target tracking. This invention significantly improves the accuracy, stability, and anti-interference capability of target tracking through collaborative optimization of the target tracking method and gimbal device, optimizing the performance of target tracking tasks in dynamic and complex scenarios.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A target tracking method based on a roll-up gimbal includes the following steps: S1. Extract the colorimetric reference target area and obtain the original spectral data through multi-frame sampling and morphological operations of the visual sensor. Preprocess the data and divide it into grids to calculate the color block features. Establish a polynomial response model. Calculate the color conversion matrix and store the calibration parameters to achieve adaptive color correction in dynamic environments. S2. Acquire images through a visual sensor, perform preprocessing, extract candidate regions, input the region into the improved YOLOv8 deep learning model, perform target detection, analyze the detection results, and determine whether they meet the preset confidence threshold. If they do, output the final detection result; otherwise, execute S3. S3. Through template matching, for candidate regions in S2 that do not meet the confidence threshold, perform secondary detection on the target to improve the recognition accuracy in scenarios where the target is partially occluded or has low confidence. Determine whether the target is detected. If yes, output the final detection result. If no, execute S2 until the target is detected. S4. Using the target coordinates of the current frame as the observation input, the target position and velocity are predicted by Kalman filtering. The predicted values are corrected by combining the actual detection results. The parameters are adaptively adjusted to cope with the scenario where the target is occluded. The rotation parameters of the gimbal device are calculated based on the prediction results. The angular acceleration is limited to achieve smooth rotation. The data is sent to the gimbal device through the serial port with a specific data protocol packet to achieve stable tracking command transmission. S5. The gimbal device collects angle sensor data in real time, establishes a precise initial attitude reference through zero-position calibration, and uses a quaternion algorithm to solve the attitude, achieving low-latency response. Combined with dynamic PID adjustment and feedforward compensation, it achieves anti-interference self-stabilizing control, ensuring the attitude accuracy and stability of the gimbal during long-term operation. S6. The gimbal device receives data protocol packets, parses the target attitude and generates error signals. After calculating the desired angular velocity through the outer loop PID, the current is decoupled and the torque component is controlled through the inner loop FOC algorithm. Through dual-loop collaborative control, the gimbal device is ensured to respond quickly and move smoothly. The SVPWM drive motor is used to smoothly adjust the attitude, achieving high-precision and low-latency target tracking.
[0008] Furthermore, S1 specifically includes: S11, Visual sensor initialization stage: Automatically capture multiple frames of the preset colorimetric reference target to obtain the original spectrum; S12. Accurately extract the colorimetric reference target region through contour analysis and morphological operations, and obtain the original spectral response data under the current ambient lighting conditions; S13. The acquired raw spectral response data is subjected to mean filtering and noise suppression. After Bayer demosaicing and white balance pre-correction preprocessing through the ISP pipeline, a standard test image in linear RGB space is generated. S14. Divide the standard test image into a 4×6 grid, calculate the mean of the RGB three channels for each color block, and obtain the original spectral response matrix. , This represents the mean value of the j-th channel of the i-th color block; S15. Establish a polynomial response model, and use the matrix... Transform into a second-order polynomial characteristic matrix Each row contains RGB values and their intersections, as follows: ; Compare the feature matrix with the standard Lab matrix pre-stored in non-volatile memory. A comparison is performed, and the color transformation matrix is solved using the least squares method. The color conversion matrix was calculated. This is used to convert RGB values to Lab values, as follows: ; in, These are the elements of the color conversion matrix, respectively corresponding to , , The polynomial coefficients of the channel; S16, Convert the color matrix Stored in non-volatile memory as calibration parameters, supporting multiple sets of calibration parameters for partitioned storage to adapt to different lighting environments, and applying appropriate calibration parameters to perform color correction on each subsequently captured frame.
[0009] Furthermore, S2 specifically includes: S21. Capture the original image containing the target object using a visual sensor to obtain visual data of the scene to be detected; S22. Preprocess the original image by performing color correction and Gaussian filtering to suppress noise interference and smooth the image. S23. In the preprocessed image, an edge detection algorithm is used to extract candidate regions that may contain target objects, narrow the detection range, reduce background interference, and improve the accuracy of subsequent target detection. S24. Input the candidate region into the improved YOLOv8 deep learning model, embed a bidirectional feature pyramid network between the model's feature extraction network and the detection head, and perform weighted fusion of feature maps at different levels through a multi-scale feature fusion path to enhance the feature representation ability of small targets and improve the robustness of detection of targets at different scales. S25. Perform non-maximum suppression processing on the detection boxes output by the model to remove redundant detection results and obtain the target's category, location coordinates (x, y, w, h) and confidence score; perform probability normalization processing on the detection results using the Softmax algorithm to determine whether they meet the preset confidence threshold. If they do, output the final detection results, including the target category, location coordinates and confidence score; if they do not meet the threshold, proceed to step S3.
[0010] Furthermore, S3 specifically includes: S31. When a target with a confidence level greater than a preset threshold is detected in S25, an image region is cropped according to the location coordinates of the target, used as a template image, stored in the cache and updated to the latest template. S32. In the current frame image, the squared difference matching algorithm is used to compare the template image stored in the buffer with the sliding window to calculate the similarity score, so as to achieve stable recognition in the scene of target partial occlusion. S33. If the maximum similarity score is greater than the matching threshold, the corresponding region is determined to have a target. The target category, location coordinates and confidence score are output. The corresponding region of the frame image is used as a template image, stored in the cache and updated with the latest template. If the similarity scores are all less than the matching threshold, it is determined that there is no target. S2 is executed again until the target is detected.
[0011] Furthermore, S4 specifically includes: S41. Obtain the target coordinates of the current frame as the observation input for the Kalman filter, and initialize the filter state parameters and covariance matrix. S42. Based on the uniform motion model, the Kalman filter algorithm predicts the target position and velocity vector in the current frame according to the target motion state of the previous frame. If the target is detected in the current frame, the predicted value is corrected using the actual detection result. The state equation is: ; in , Let be the coordinates of the target at time k. , The velocities are in the x and y directions. The frame interval time. This is process noise; S43. Based on the target's motion speed and scene complexity, the parameters of the Kalman filter are adaptively adjusted to improve prediction accuracy, achieve accurate prediction of the target's motion trajectory, and effectively cope with situations where the target is occluded for a short period of time. S44. Based on the predicted target coordinates and velocity vector, calculate the rotation angle parameters required for the gimbal device to move from the current position to the predicted target position. Based on the predicted velocity vector, limit the angular acceleration of the gimbal rotation. By controlling the rate of change of angular velocity, the gimbal rotation process can be smoothly transitioned, reducing jitter. S45. The calculated rotation parameters are sent to the gimbal device via a serial port using a specific data protocol packet to achieve stable tracking command transmission.
[0012] Furthermore, S5 specifically includes: S51. After the main controller of the gimbal device is powered on, it collects the angular velocity and acceleration data output by the angle sensor in real time through the high-speed bus interface, accurately corrects the initial angle of the gimbal device, ensures that the roll axis and pitch axis are at the preset zero position, and establishes the initial attitude reference. S52. The main controller adopts a quaternion attitude calculation algorithm, which avoids the gimbal lock-up problem of traditional Euler angles, simplifies the calculation logic, achieves low latency response, corrects the accumulated error in attitude calculation in real time, and prevents attitude deviation caused by long-term operation. S53. By dynamically adjusting the parameters of the PID controller and introducing feedforward compensation torque, the effect of device jitter is quickly offset, and the gimbal device is kept stable.
[0013] Furthermore, S6 specifically includes: S61. The gimbal device receives the data protocol packets sent by S45 through the serial port and parses them to obtain the target attitude. S62. Compare the parsed target attitude with the current attitude to generate an attitude error signal. Send the signal to the outer loop PID controller, which calculates the desired angular velocity based on the error magnitude, cumulative error, and error change rate. S63. The expected value output by the outer loop PID controller is used as the input command of the inner loop FOC algorithm. The main controller reads the magnetic encoder in real time to obtain the high-precision electrical angle of the brushless motor rotor. The FOC algorithm combines the rotor position and the input command, and decouples the stator three-phase current into excitation component and torque component through Clarke-Park transformation, and controls the Iq component to match the expected torque. S64. Utilize SVPWM technology to generate a high-frequency PWM signal, drive the FOC brushless motor driver to control the brushless motor to generate corresponding torque, and with a smooth speed curve, enable the gimbal device to adjust its attitude according to the analyzed rotation parameters, thereby achieving accurate and stable tracking of the target while maintaining the system's anti-disturbance capability.
[0014] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows: A target tracking device based on a roll-up gimbal, the device comprising a main controller, an angle sensor, a vision sensor, a magnetic encoder, and a roll-up gimbal.
[0015] The main controller is used to coordinate the work of each module, execute image processing, quaternion attitude calculation, PID control, and FOC drive algorithm to achieve stable target tracking. The angle sensor is used to collect the real-time angle of the gimbal device to establish an accurate initial attitude reference for zero-position calibration. The vision sensor is used to acquire target images and provide visual data for target tracking; The magnetic encoder is used to obtain high-precision electrical angles of the brushless motor rotor. The roll-and-tilt gimbal includes an FOC brushless motor driver and a dual-axis mechanical structure. The FOC brushless motor control board controls the brushless motor. The dual-axis mechanical structure includes a roll axis and a pitch axis. Two sets of brushless motors drive the two axes to move independently, achieving a large frame angle of 360° roll and ±90° pitch. It controls the rotation of the vision sensor to ensure that the vision sensor always tracks the target. It eliminates the yaw axis of traditional three-axis gimbals to simplify the mechanical structure, reduce structural weight, and has strong control decoupling and fast response speed. It reduces response lag caused by mechanical transmission backlash and achieves precise adjustment and dynamic compensation of the two-axis attitude.
[0016] This invention provides a target tracking method and device based on a roll-up gimbal, which has the following advantages: By synergistically optimizing adaptive color calibration, multi-scale target detection, template matching, and Kalman filter prediction, the target tracking performance in complex environments is significantly improved. This method employs an improved YOLOv8 model combined with a bidirectional feature pyramid network to enhance multi-scale detection capabilities, and effectively addresses target occlusion issues through template matching and Kalman filtering.
[0017] Employing a roll-and-tilt dual-axis structure, this invention reduces redundant degrees of freedom compared to traditional three-axis systems. Through quaternion attitude calculation, dynamic PID adjustment, and feedforward compensation technology, the motion characteristics of the gimbal device are optimized, effectively suppressing the effects of external disturbances and mechanical transmission backlash. This improves the closed-loop control bandwidth and control response speed. Combined with the PID and FOC dual-loop collaborative control algorithm, the gimbal rotation becomes smooth, achieving high-precision, low-latency gimbal response. These technological innovations significantly improve the accuracy, stability, and anti-interference capabilities of target tracking.
[0018] This invention has excellent scalability and demonstrates superior tracking stability in dynamic scenarios such as aerial photography, agricultural plant protection, power line inspection, emergency rescue, and military confrontation, providing a more efficient and reliable target tracking solution for intelligent mobile platforms. Attached Figure Description
[0019] Figure 1 A flowchart of a target tracking method based on a roll-up gimbal; Figure 2 This is a structural block diagram of a target tracking device based on a roll-up gimbal. Figure 3 A flowchart for image processing using a vision sensor; Figure 4 This is a structural block diagram of the gimbal device. DETAILED DESCRIPTION
[0020] The principles and features of the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The content mentioned is only a part of the embodiments of the present invention and is not intended to limit the scope of the claimed invention.
[0021] Please see Figure 1 and Figure 2 The embodiments of the present invention include: This invention provides a target tracking method and device based on a roll-up gimbal. Through adaptive color calibration, multi-scale target detection, template matching, and Kalman filter prediction, it can ensure high-precision target position prediction even in scenarios where the target is temporarily occluded. In particular, the gimbal adopts a roll-up structure, realizing an attitude reference mechanism. Based on a dual-loop collaborative control algorithm of PID and FOC, it compensates for disturbances of the intelligent mobile platform in real time, reduces response lag caused by mechanical transmission backlash, and achieves high-precision, low-latency target tracking. This invention significantly improves the accuracy, stability, and anti-interference capability of target tracking through the collaborative optimization of the target tracking method and gimbal device, and optimizes the performance of target tracking tasks in dynamic and complex scenarios.
[0022] Reference Figure 1 As shown, the process includes the following steps: S1. Extract the colorimetric reference target area and obtain the original spectral data through multi-frame sampling and morphological operations of the visual sensor. Preprocess the data and divide it into grids to calculate the color block features. Establish a polynomial response model. Calculate the color conversion matrix and store the calibration parameters to achieve adaptive color correction in dynamic environments. S2. Acquire images through a visual sensor, perform preprocessing, extract candidate regions, input the region into the improved YOLOv8 deep learning model, perform target detection, analyze the detection results, and determine whether they meet the preset confidence threshold. If they do, output the final detection result; otherwise, execute S3. S3. Through template matching, for candidate regions in S2 that do not meet the confidence threshold, perform secondary detection on the target to improve the recognition accuracy in scenarios where the target is partially occluded or has low confidence. Determine whether the target is detected. If yes, output the final detection result. If no, execute S2 until the target is detected. S4. Using the target coordinates of the current frame as the observation input, the target position and velocity are predicted by Kalman filtering. The predicted values are corrected by combining the actual detection results. The parameters are adaptively adjusted to cope with the scenario where the target is occluded. The rotation parameters of the gimbal device are calculated based on the prediction results. The angular acceleration is limited to achieve smooth rotation. The data is sent to the gimbal device through the serial port with a specific data protocol packet to achieve stable tracking command transmission. S5. The gimbal device collects angle sensor data in real time, establishes a precise initial attitude reference through zero-position calibration, and uses a quaternion algorithm to solve the attitude, achieving low-latency response. Combined with dynamic PID adjustment and feedforward compensation, it achieves anti-interference self-stabilizing control, ensuring the attitude accuracy and stability of the gimbal during long-term operation. S6. The gimbal device receives data protocol packets, parses the target attitude and generates error signals. After calculating the desired angular velocity through the outer loop PID, the current is decoupled and the torque component is controlled through the inner loop FOC algorithm. Through dual-loop collaborative control, the gimbal device is ensured to respond quickly and move smoothly. The SVPWM drive motor is used to smoothly adjust the attitude, achieving high-precision and low-latency target tracking.
[0023] Reference Figure 2 As shown, the structural block diagram includes the following parts: The gimbal unit consists of a main controller, an angle sensor, a vision sensor, a magnetic encoder, and a roll-up gimbal. The main controller coordinates the work of each module, performs image processing, quaternion attitude calculation, PID control, and FOC drive algorithm to achieve stable target tracking. Angle sensor is used to collect the real-time angle of the gimbal device to establish an accurate initial attitude reference for zero-position calibration; Visual sensors are used to acquire images of targets and provide visual data for target tracking; Magnetic encoders are used to obtain high-precision electrical angles of the rotor of a brushless motor. The roll-and-tilt gimbal includes an FOC brushless motor driver and a dual-axis mechanical structure. The FOC brushless motor control board controls the brushless motor, and the dual-axis mechanical structure includes a roll axis and a pitch axis. Two sets of brushless motors drive the two axes to move independently, achieving a large frame angle of 360° roll and ±90° pitch. It controls the rotation of the vision sensor to ensure that the vision sensor always tracks the target. It eliminates the yaw axis of traditional three-axis gimbals to simplify the mechanical structure, reduce structural weight, and has strong control decoupling and fast response speed. It reduces response lag caused by mechanical transmission backlash and achieves precise adjustment and dynamic compensation of the two-axis attitude.
[0024] Furthermore, referring to Figure 3 The diagram shows a flowchart of image processing using a vision sensor, specifically including the following in S1: S11, Visual sensor initialization phase, automatically performs multi-frame sampling and capture of the preset colorimetric reference target to obtain the original spectrum; S12. Accurately extract the colorimetric reference target region through contour analysis and morphological operations, and obtain the original spectral response data under the current ambient lighting conditions; S13. The acquired raw spectral response data is subjected to mean filtering and noise suppression. After Bayer demosaicing and white balance pre-correction preprocessing through the ISP pipeline, a standard test image in linear RGB space is generated. S14. Divide the standard test image into a 4×6 grid, calculate the mean of the RGB three channels for each color block, and obtain the original spectral response matrix. , This represents the mean value of the j-th channel of the i-th color block; S15. Establish a polynomial response model, and use the matrix... Transform into a second-order polynomial characteristic matrix Each row contains RGB values and their intersections, as follows: ; Compare the feature matrix with the standard Lab matrix pre-stored in non-volatile memory. A comparison is performed, and the color transformation matrix is solved using the least squares method. The color conversion matrix was calculated. This is used to convert RGB values to Lab values, as follows: ; in, These are the elements of the color conversion matrix, respectively corresponding to , , The polynomial coefficients of the channel; S16, Convert the color matrix Stored in non-volatile memory as calibration parameters, supporting multiple sets of calibration parameters for partitioned storage to adapt to different lighting environments, and applying appropriate calibration parameters to perform color correction on each subsequently captured frame.
[0025] Furthermore, referring to Figure 3 The diagram shows a flowchart of image processing using a vision sensor, specifically including the following in S2: S21. Capture the original image containing the target object using a visual sensor to obtain visual data of the scene to be detected; S22. Preprocess the original image by performing color correction and Gaussian filtering to suppress noise interference and smooth the image. S23. In the preprocessed image, an edge detection algorithm is used to extract candidate regions that may contain target objects, narrow the detection range, reduce background interference, and improve the accuracy of subsequent target detection. S24. Input the candidate region into the improved YOLOv8 deep learning model, embed a bidirectional feature pyramid network between the model's feature extraction network and the detection head, and perform weighted fusion of feature maps at different levels through a multi-scale feature fusion path to enhance the feature representation ability of small targets and improve the robustness of detection of targets at different scales. S25. Perform non-maximum suppression processing on the detection boxes output by the model to remove redundant detection results and obtain the target's category, location coordinates (x, y, w, h) and confidence score; perform probability normalization processing on the detection results using the Softmax algorithm to determine whether they meet the preset confidence threshold. If they do, output the final detection results, including the target category, location coordinates and confidence score; if they do not meet the threshold, proceed to step S3.
[0026] Furthermore, referring to Figure 3 The flowchart shown illustrates the image processing flow of a vision sensor, specifically including the following in S3: S31. When a target with a confidence level greater than a preset threshold is detected in S25, an image region is cropped according to the location coordinates of the target, used as a template image, stored in the cache and updated to the latest template. S32. In the current frame image, the squared difference matching algorithm is used to compare the template image stored in the buffer with the sliding window to calculate the similarity score, so as to achieve stable recognition in the scene of target partial occlusion. S33. If the maximum similarity score is greater than the matching threshold, the corresponding region is determined to have a target. The target category, location coordinates and confidence score are output. The corresponding region of the frame image is used as a template image, stored in the cache and updated with the latest template. If the similarity scores are all less than the matching threshold, it is determined that there is no target. S2 is executed again until the target is detected.
[0027] Furthermore, referring to Figure 3 The diagram shows a flowchart of image processing using a vision sensor, specifically including the following in S4: S41. Obtain the target coordinates of the current frame as the observation input for the Kalman filter, and initialize the filter state parameters and covariance matrix. S42. Based on the uniform motion model, the Kalman filter algorithm predicts the target position and velocity vector in the current frame according to the target motion state of the previous frame. If the target is detected in the current frame, the predicted value is corrected using the actual detection result. The state equation is: ; in , Let be the coordinates of the target at time k. , The velocities are in the x and y directions. The frame interval time. This is process noise; S43. Based on the target's motion speed and scene complexity, the parameters of the Kalman filter are adaptively adjusted to improve prediction accuracy, achieve accurate prediction of the target's motion trajectory, and effectively cope with situations where the target is occluded for a short period of time. S44. Based on the predicted target coordinates and velocity vector, calculate the rotation angle parameters required for the gimbal device to move from the current position to the predicted target position. Based on the predicted velocity vector, limit the angular acceleration of the gimbal rotation. By controlling the rate of change of angular velocity, the gimbal rotation process can be smoothly transitioned, reducing jitter. S45. The calculated rotation parameters are sent to the gimbal device via a serial port using a specific data protocol packet to achieve stable tracking command transmission.
[0028] Furthermore, referring to Figure 4 The diagram shown illustrates the structural block diagram of the gimbal device, specifically including the following components in S5: S51. After the main controller of the gimbal device is powered on, it collects the angular velocity and acceleration data output by the angle sensor in real time through the high-speed bus interface, accurately corrects the initial angle of the gimbal device, ensures that the roll axis and pitch axis are at the preset zero position, and establishes the initial attitude reference. S52. The main controller adopts a quaternion attitude calculation algorithm, which avoids the gimbal lock-up problem of traditional Euler angles, simplifies the calculation logic, achieves low latency response, corrects the accumulated error in attitude calculation in real time, and prevents attitude deviation caused by long-term operation. S53. By dynamically adjusting the parameters of the PID controller and introducing feedforward compensation torque, the effect of device jitter is quickly offset, and the gimbal device is kept stable.
[0029] Furthermore, referring to Figure 4The diagram shown illustrates the structural block diagram of the gimbal device control system, specifically including the following in S6: S61. The gimbal device receives the data protocol packets sent by S45 through the serial port and parses them to obtain the target attitude. S62. Compare the parsed target attitude with the current attitude to generate an attitude error signal. Send the signal to the outer loop PID controller, which calculates the desired angular velocity based on the error magnitude, cumulative error, and error change rate. S63. The expected value output by the outer loop PID controller is used as the input command of the inner loop FOC algorithm. The main controller reads the magnetic encoder in real time to obtain the high-precision electrical angle of the brushless motor rotor. The FOC algorithm combines the rotor position and the input command, and decouples the stator three-phase current into excitation component and torque component through Clarke-Park transformation, and controls the Iq component to match the expected torque. S64. Utilize SVPWM technology to generate a high-frequency PWM signal, drive the FOC brushless motor driver to control the brushless motor to generate corresponding torque, and with a smooth speed curve, enable the gimbal device to adjust its attitude according to the analyzed rotation parameters, thereby achieving accurate and stable tracking of the target while maintaining the system's anti-disturbance capability.
[0030] This invention provides a roll-and-tilt gimbal-based target tracking device suitable for use on mobile UAV platforms for aerial reconnaissance target tracking missions in complex environments. Adaptive color correction ensures accurate recognition under varying lighting conditions, while improved YOLOv8 multi-scale detection enhances detection capabilities for small UAVs. Template matching and Kalman filtering prediction effectively address target occlusion issues. Through dual-loop collaborative control of PID and FOC, the gimbal achieves high-precision and rapid response, enabling stable tracking of dynamic targets during aerial reconnaissance missions. This device, through collaborative optimization of algorithms and hardware, significantly improves the tracking robustness of UAVs in situations such as target loss, occlusion, or rapid movement.
[0031] Compared to traditional three-axis gimbals, the roll-and-tilt dual-axis structure offers faster response and higher control precision while reducing the load on the UAV. Through quaternion attitude calculation and dynamic PID adjustment, the system can quickly compensate for attitude disturbances during UAV flight; while SVPWM-driven FOC control ensures smooth gimbal movement, avoiding tracking jitter caused by mechanical delays. These designs enable stable, low-latency target tracking in aerial reconnaissance missions, while reducing system power consumption and extending flight time.
[0032] This invention provides target tracking in UAV aerial reconnaissance missions, including: During the initialization phase of the visual sensor, the visual sensor on the UAV gimbal automatically performs ten consecutive frames of sampling on the preset colorimetric reference target to capture the original spectral data; through contour analysis and morphological operations, it accurately extracts the target area and obtains the spectral response data under ambient lighting conditions. After mean filtering and noise suppression, a standard test image in linear RGB space is generated. This image is divided into a 4×6 grid. The RGB mean of each color block is calculated to form a spectral response matrix. Then, a second-order polynomial feature matrix is established. By comparing it with the least squares of the pre-stored standard Lab matrix, the color conversion matrix is solved and stored to achieve adaptive color correction in dynamic environments. This complete color calibration process effectively solves the problem of target recognition for UAVs under complex lighting conditions, significantly improves the accuracy of target tracking and environmental adaptability in reconnaissance missions, and can intelligently match the lighting characteristics of different time periods and weather conditions through a partitioned storage mechanism for multiple sets of calibration parameters, ensuring that UAVs maintain stable color recognition performance in long-term, large-scale aerial reconnaissance missions.
[0033] The visual sensor automatically acquires raw image data containing the target object, and eliminates lighting differences and noise interference through color correction and Gaussian filtering; Next, an edge detection algorithm is used to extract candidate target regions to reduce background interference.
[0034] To improve the detection capability of multi-scale targets, such as drones, birds, and loitering munitions, the improved YOLOv8 model embeds a bidirectional feature pyramid network between the feature extraction network and the detection head. By weighted fusion of feature maps from different levels, it significantly enhances the detection capability of small targets. During the training phase, a dataset containing 50,000 labeled samples was used, covering categories such as drones, birds, and loitering munitions. The dataset was divided into training, validation, and test sets in a 7:2:1 ratio. During model inference, a confidence threshold of 0.7 and an NMS threshold of 0.2 are set to identify and filter redundant detection boxes, and the final detection result is output through Softmax normalization. For low-confidence targets, the system automatically triggers a template matching mechanism for secondary verification to ensure tracking robustness in complex environments. It can achieve an overall average accuracy of 0.915 and a real-time detection speed of 30 FPS, meeting the high accuracy and low latency requirements of UAV reconnaissance missions.
[0035] During target tracking, when a target with a confidence level is detected, the system will store the target area as a template image in the cache. For subsequent frame images, a squared difference matching algorithm is used for sliding window comparison to effectively solve the problem of target partial occlusion; if the match is successful, the template is updated and the target information is output; otherwise, the detection process is re-executed. The system further uses the Kalman filter algorithm to predict the target's motion trajectory. Based on the state equation of the uniform motion model and the adaptive parameter adjustment mechanism, it ensures stable tracking even when the target is temporarily occluded. The prediction results are used to calculate the gimbal rotation parameters. Smooth rotation control is achieved by limiting angular acceleration. The maximum angular acceleration is limited to the range of 0.5-2 rad / s², and the maximum angular velocity is limited to the range of 1-5 rad / s. The PTZ control commands are encapsulated using a specific data protocol packet. Fixed bytes “AA” are set as the frame header identifier and “BB” as the frame tail identifier. The PTZ control commands are embedded between the frame header and frame tail to form a complete data frame.
[0036] The main controller accesses the angle sensor via a high-speed bus. While the gimbal remains stationary, it continuously reads one hundred sample values and calculates their average to obtain the static zero-point offset of the angle sensor. This offset data will be stored for real-time compensation during subsequent attitude calculations, and precise zero-point calibration of the roll and pitch axes will be performed to establish a reliable initial attitude reference. The quaternion attitude calculation algorithm effectively avoids the gimbal lock-up problem of traditional Euler angles. Through the fusion mechanism of gyroscope integral prediction and accelerometer observation correction, attitude calculation with low latency of less than 5ms and no cumulative error is achieved, which further enhances the anti-interference capability. By dynamically adjusting the PID control parameters and introducing feedforward compensation torque, airflow disturbances and mechanical vibrations during UAV flight can be quickly offset, enabling the gimbal to maintain an attitude stability accuracy of ±0.1° even under high-speed maneuvering conditions.
[0037] During the process of receiving data protocol packets, the PTZ device will verify the start and end markers of the data frames. "AA" is used as the basis for judging the start of a valid data frame, and "BB" is used as the marker for the end of data transmission of the frame. Only when a complete data protocol packet is identified is the received PTZ control command deemed valid, thereby ensuring the accuracy and integrity of the transmitted data. The parsed gimbal control commands, combined with the current gimbal attitude, generate an error signal input to the outer-loop PID controller. This controller calculates the desired angular velocity based on the error dynamic characteristics. The core task of the outer-loop PID controller is to calculate the error between the current attitude and the dynamic target attitude, and output correction commands to the FOC control loop to achieve fast and accurate target tracking. To optimize the motion trajectory and avoid shocks during start-up and shutdown and jitter during operation, the gimbal incorporates a motion trajectory planning module. When a new target angle is received, this module does not directly use the target value as the PID setpoint, but instead generates a smooth S-shaped or trapezoidal velocity curve based on the preset maximum speed and acceleration. The PID controller then tracks the instantaneous target point on this smooth curve, making the entire gimbal movement extremely smooth.
[0038] To achieve precise torque control, the system adopts an inner-loop FOC algorithm, which decouples the three-phase current through Clarke-Park transformation to precisely control the torque component Iq in order to match the PID output command; SVPWM technology generates high-frequency drive signals, enabling the brushless motor to operate smoothly according to the optimized speed curve. By proposing a target tracking method and device based on a roll-tilt gimbal, a target tracking accuracy of ±0.1° and a millisecond-level response speed are achieved on an intelligent mobile platform for unmanned aerial vehicles (UAVs). At the same time, an adaptive anti-disturbance algorithm effectively suppresses the impact of target illumination changes, occlusion, or airflow disturbances on the gimbal. The design of the roll-tilt gimbal structure reduces weight by more than 30%, effectively reducing the load on the UAV, enhancing its endurance, and making the gimbal control highly decoupled, with a fast response speed, reducing response lag caused by mechanical transmission backlash, and achieving precise adjustment and dynamic compensation of two-axis attitude.
[0039] Experiments show that even at a cruising speed of 60 km / h, the present invention can still maintain stable tracking of dynamic targets, providing reliable technical support for reconnaissance missions in complex environments.
[0040] In practical applications, a target tracking method and device based on a roll-up gimbal can be widely used in various automatic tracking scenarios that require high precision and low latency.
[0041] In drone reconnaissance missions, this invention enables stable tracking of high-speed moving targets such as flying objects, vehicles, and personnel, maintaining reliable tracking performance even under complex lighting conditions and partial obstruction. In the field of intelligent security, it can be deployed on unmanned monitoring vehicle mobile platforms for automatic identification and continuous tracking of suspicious targets. In the film and television aerial photography industry, it can accurately lock onto moving subjects and maintain image stability, effectively solving the image shake problem of traditional gimbals during rapid turns. In industrial inspection applications, it can be used in conjunction with drones to automatically track and inspect facilities such as high-voltage lines and wind turbine blades, significantly improving inspection efficiency and safety. Regardless of the application scenario, this invention ensures that the gimbal responds quickly and accurately to target movement, while possessing excellent anti-interference capabilities, meeting the tracking needs of various complex environments.
[0042] The above-described content is merely a specific embodiment of the present invention and does not constitute a limitation on the scope of patent protection of the present invention. Any equivalent structural or procedural transformations made based on the content described in the specification and drawings of this invention, or their direct or indirect application to other related technical fields, are also covered within the scope of patent protection of this invention.
Claims
1. A target tracking method based on a roll-up gimbal, characterized in that, The steps include: S1. Extract the colorimetric reference target area and obtain the original spectral data through multi-frame sampling and morphological operations of the visual sensor. Preprocess the data and divide it into grids to calculate the color block features. Establish a polynomial response model. Calculate the color conversion matrix and store the calibration parameters to achieve adaptive color correction in dynamic environments. S2. Acquire images through a visual sensor, perform preprocessing, extract candidate regions, input the region into the improved YOLOv8 deep learning model, perform target detection, analyze the detection results, and determine whether they meet the preset confidence threshold. If they do, output the final detection result; otherwise, execute S3. S3. Through template matching, for candidate regions in S2 that do not meet the confidence threshold, perform secondary detection on the target to improve the recognition accuracy in scenarios where the target is partially occluded or has low confidence. Determine whether the target is detected. If yes, output the final detection result. If no, execute S2 until the target is detected. S4. Using the target coordinates of the current frame as the observation input, the target position and velocity are predicted by Kalman filtering. The predicted values are corrected by combining the actual detection results. The parameters are adaptively adjusted to cope with the scenario where the target is occluded. The rotation parameters of the gimbal device are calculated based on the prediction results. The angular acceleration is limited to achieve smooth rotation. The data is sent to the gimbal device through the serial port with a specific data protocol packet to achieve stable tracking command transmission. S5. The gimbal device collects angle sensor data in real time, establishes a precise initial attitude reference through zero-position calibration, and uses a quaternion algorithm to solve the attitude, achieving low-latency response. Combined with dynamic PID adjustment and feedforward compensation, it achieves anti-interference self-stabilizing control, ensuring the attitude accuracy and stability of the gimbal during long-term operation. S6. The gimbal device receives data protocol packets, parses the target attitude and generates error signals. After calculating the desired angular velocity through the outer loop PID, the current is decoupled and the torque component is controlled through the inner loop FOC algorithm. Through dual-loop collaborative control, the gimbal device is ensured to respond quickly and move smoothly. The SVPWM drive motor is used to smoothly adjust the attitude, achieving high-precision and low-latency target tracking.
2. A target tracking method based on a roll-up gimbal according to claim 1, characterized in that: Step S1 specifically includes: S11, Visual sensor initialization stage: Automatically capture multiple frames of the preset colorimetric reference target to obtain the original spectrum; S12. Accurately extract the colorimetric reference target region through contour analysis and morphological operations, and obtain the original spectral response data under the current ambient lighting conditions; S13. The acquired raw spectral response data is subjected to mean filtering and noise suppression. After Bayer demosaicing and white balance pre-correction preprocessing through the ISP pipeline, a standard test image in linear RGB space is generated. S14. Divide the standard test image into a 4×6 grid, calculate the mean of the RGB three channels for each color block, and obtain the original spectral response matrix. , This represents the mean value of the j-th channel of the i-th color block; S15. Establish a polynomial response model, and use the matrix... Transform into a second-order polynomial characteristic matrix Each row contains RGB values and their intersections, as follows: ; Compare the feature matrix with the standard Lab matrix pre-stored in non-volatile memory. A comparison is performed, and the color transformation matrix is solved using the least squares method. The color conversion matrix was calculated. This is used to convert RGB values to Lab values, as follows: ; in, These are the elements of the color conversion matrix, respectively corresponding to , , The polynomial coefficients of the channel; S16, Convert the color matrix Stored in non-volatile memory as calibration parameters, supporting multiple sets of calibration parameters in partitioned storage to adapt to different lighting environments, and applying appropriate calibration parameters to perform color correction on each subsequently captured frame.
3. A target tracking method based on a roll-up gimbal according to claim 1, characterized in that: Step S2 specifically includes: S21. Capture the original image containing the target object using a visual sensor to obtain visual data of the scene to be detected; S22. Preprocess the original image by performing color correction and Gaussian filtering to suppress noise interference and smooth the image. S23. In the preprocessed image, an edge detection algorithm is used to extract candidate regions that may contain target objects, narrow the detection range, reduce background interference, and improve the accuracy of subsequent target detection. S24. Input the candidate region into the improved YOLOv8 deep learning model, embed a bidirectional feature pyramid network between the model's feature extraction network and the detection head, and perform weighted fusion of feature maps at different levels through a multi-scale feature fusion path to enhance the feature representation ability of small targets and improve the robustness of detection of targets at different scales. S25. Perform non-maximum suppression processing on the detection boxes output by the model to remove redundant detection results and obtain the target's category, location coordinates (x, y, w, h) and confidence score; perform probability normalization processing on the detection results using the Softmax algorithm to determine whether they meet the preset confidence threshold. If they do, output the final detection results, including the target category, location coordinates and confidence score; if they do not meet the threshold, proceed to step S3.
4. A target tracking method based on a roll-up gimbal according to claim 1, characterized in that: Step S3 specifically includes: S31. When a target with a confidence level greater than a preset threshold is detected in S25, an image region is cropped according to the location coordinates of the target, used as a template image, stored in the cache and updated to the latest template. S32. In the current frame image, the squared difference matching algorithm is used to compare the template image stored in the buffer with the sliding window to calculate the similarity score, so as to achieve stable recognition in the scene of target partial occlusion. S33. If the maximum similarity score is greater than the matching threshold, the corresponding region is determined to have a target. The target category, location coordinates and confidence score are output. The corresponding region of the frame image is used as a template image, stored in the cache and updated with the latest template. If the similarity scores are all less than the matching threshold, it is determined that there is no target. S2 is executed again until the target is detected.
5. A target tracking method based on a roll-up gimbal according to claim 1, characterized in that: Step S4 specifically includes: S41. Obtain the target coordinates of the current frame as the observation input for the Kalman filter, and initialize the filter state parameters and covariance matrix. S42. Based on the uniform motion model, the Kalman filter algorithm predicts the target position and velocity vector in the current frame according to the target motion state of the previous frame. If the target is detected in the current frame, the predicted value is corrected using the actual detection result. The state equation is: ; in , Let be the coordinates of the target at time k. , The velocities are in the x and y directions. The frame interval time. This is process noise; S43. Based on the target's motion speed and scene complexity, the parameters of the Kalman filter are adaptively adjusted to improve prediction accuracy, achieve accurate prediction of the target's motion trajectory, and effectively cope with situations where the target is occluded for a short period of time. S44. Based on the predicted target coordinates and velocity vector, calculate the rotation angle parameters required for the gimbal device to move from the current position to the predicted target position. Based on the predicted velocity vector, limit the angular acceleration of the gimbal rotation. By controlling the rate of change of angular velocity, the gimbal rotation process can be smoothly transitioned, reducing jitter. S45. The calculated rotation parameters are sent to the gimbal device via a serial port using a specific data protocol packet to achieve stable tracking command transmission.
6. A target tracking method based on a roll-up gimbal according to claim 1, characterized in that: Step S5 specifically includes: S51. After the main controller of the gimbal device is powered on, it collects the angular velocity and acceleration data output by the angle sensor in real time through the high-speed bus interface, accurately corrects the initial angle of the gimbal device, ensures that the roll axis and pitch axis are at the preset zero position, and establishes the initial attitude reference. S52. The main controller adopts a quaternion attitude calculation algorithm, which avoids the gimbal lock-up problem of traditional Euler angles, simplifies the calculation logic, achieves low latency response, corrects the accumulated error in attitude calculation in real time, and prevents attitude deviation caused by long-term operation. S53. By dynamically adjusting the parameters of the PID controller and introducing feedforward compensation torque, the effect of device jitter is quickly offset, and the gimbal device is kept stable.
7. A target tracking method based on a roll-up gimbal according to claim 1, characterized in that: Step S6 specifically includes: S61. The gimbal device receives the data protocol packets sent by S45 through the serial port and parses them to obtain the target attitude. S62. Compare the parsed target attitude with the current attitude to generate an attitude error signal. Send the signal to the outer loop PID controller, which calculates the desired angular velocity based on the error magnitude, cumulative error, and error change rate. S63. The expected value output by the outer loop PID controller is used as the input command of the inner loop FOC algorithm. The main controller reads the magnetic encoder in real time to obtain the high-precision electrical angle of the brushless motor rotor. The FOC algorithm combines the rotor position and the input command, and decouples the stator three-phase current into excitation component and torque component through Clarke-Park transformation, and controls the Iq component to match the expected torque. S64. Utilize SVPWM technology to generate a high-frequency PWM signal, drive the FOC brushless motor driver to control the brushless motor to generate corresponding torque, and with a smooth speed curve, enable the gimbal device to adjust its attitude according to the analyzed rotation parameters, thereby achieving accurate and stable tracking of the target while maintaining the system's anti-disturbance capability.
8. A target tracking method based on a roll-up gimbal according to claims 1-7, characterized in that: The method is applied to the following apparatus: A target tracking device based on a roll-up gimbal, the device comprising a main controller, an angle sensor, a vision sensor, a magnetic encoder, and a roll-up gimbal; The main controller is used to coordinate the work of each module, execute image processing, quaternion attitude calculation, PID control, and FOC drive algorithm to achieve stable target tracking; The angle sensor is used to collect the real-time angle of the gimbal device to establish an accurate initial attitude reference for zero-position calibration. The vision sensor is used to acquire target images and provide visual data for target tracking; The magnetic encoder is used to obtain high-precision electrical angles of the brushless motor rotor. The roll-and-tilt gimbal includes an FOC brushless motor driver and a dual-axis mechanical structure. The FOC brushless motor control board controls the brushless motor. The dual-axis mechanical structure includes a roll axis and a pitch axis. Two sets of brushless motors drive the two axes to move independently, achieving a large frame angle of 360° roll and ±90° pitch. It controls the rotation of the vision sensor to ensure that the vision sensor always tracks the target. It eliminates the yaw axis of traditional three-axis gimbals to simplify the mechanical structure, reduce structural weight, and has strong control decoupling and fast response speed. It reduces response lag caused by mechanical transmission backlash and achieves precise adjustment and dynamic compensation of the two-axis attitude.
Citation Information
Cited By
Multi-modal visual servo four-degree-of-freedom experimental recording holder device and control method
CN121364747A
A pan-tilt target tracking speed adaptive adjustment control method and a pan-tilt applying the same
CN122411087A