A multi-modal feedback based open-loop gimbal preset calibration method and system

CN122845918APending Publication Date: 2026-09-29SHENZHEN JIWEI TIMES TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202611033554.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

如果直接采用整图匹配或未筛选的局部特征进行补偿,容易将动态目标造成的图像变化误判为云台位置偏移,从而产生错误补偿

Benefits of technology

(1)针对开环步进电机丢步与回程误差累积问题,本发明在云台执行目标预置位运动之前,对步进电机脉冲序列进行S曲线加减速预补偿,并在每次换向指令前依据回程误差补偿表插入补偿脉冲Δp,从而在驱动层降低步进电机启动、停止及换向过程中的丢步误差和机械回程误差累积,减小云台到达目标预置位前的初始偏差。在一种样机验证条件下,本发明相对于无补偿步进开环方案和整图特征匹配方案,能够降低预置位偏移量并提高回位稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845918A_ABST
    Figure CN122845918A_ABST
Patent Text Reader

Abstract

This invention discloses an open-loop gimbal preset position calibration method and system based on multimodal feedback, comprising the following steps: before the gimbal performs target preset position movement, the pulse sequence of the stepper motor is pre-compensated for acceleration and deceleration by an S-curve, and compensation pulses are inserted before each commutation command according to a pre-established hysteresis error compensation table; during the gimbal's stationary period, the high-frequency component statistical data output by the image signal processor is read, the texture richness of the image sub-regions is scored, and the sub-regions that change significantly during the gimbal's stationary period are filtered out by the temporal pixel change, N static background sub-regions are selected as dynamically selected regions of interest (ROIs), and the historical reference fingerprints of the ROIs are extracted; when the gimbal rotates to the target preset position, the displacement vectors between the N ROIs and their respective historical reference fingerprints are calculated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of camera technology, and in particular to an open-loop gimbal preset position calibration method and system based on multimodal feedback. Background Technology

[0002] Pan-tilt cameras are widely used in security monitoring, intelligent surveillance, outdoor inspection, yard monitoring, and equipment inspection. Existing pan-tilt cameras typically use horizontal and vertical rotation mechanisms to rotate the camera module, enabling functions such as preset position recall, target tracking, and patrol monitoring. In cost-sensitive network cameras, the pan-tilt unit often uses a stepper motor for open-loop drive, and the controller typically estimates the current position of the pan-tilt unit by recording the number of output pulses.

[0003] However, in the open-loop stepper motor drive mode, the actual position of the pan-tilt unit is easily affected by factors such as motor step loss, gear backlash, return error of the reduction mechanism, changes in load inertia, changes in frictional resistance, and frequent reversing operations. As the pan-tilt unit operates for a long time, a cumulative deviation may occur between the pulse count position and the actual mechanical position of the pan-tilt unit, causing the image to shift when the pan-tilt unit calls the target preset position, affecting the positioning accuracy of the preset position and the target monitoring effect.

[0004] To improve the positioning accuracy of PTZ cameras, some existing solutions use encoders, Hall effect sensors, limit switches, or other position detection elements to form a closed-loop feedback. However, these solutions increase hardware costs, structural complexity, and assembly and debugging difficulties, hindering the widespread application of low-cost PTZ cameras. Other solutions correct the initial position of the stepper motor through mechanical hard stops, collision zeroing, or periodic zeroing. However, these solutions typically only reset the coordinates at the zeroing point, making it difficult to correct the remaining deviation after the PTZ returns to the target preset position in real time. If the zeroing timing is fixed, it may also interrupt the current monitoring screen or target tracking process, affecting the continuity of monitoring.

[0005] A search revealed a patent document with publication number CN115941930A, which discloses a video preset point calibration method. This method achieves automatic calibration of camera preset points through preset image acquisition, calibration triggering, and image comparison between a reference image and a preview image. While this approach can reduce the workload of manual maintenance of preset points to some extent, it primarily relies on preset point image acquisition and image comparison, resulting in relatively large overall computational and storage requirements. Furthermore, it does not provide pre-compensation for step loss and mechanical backlash errors during the start-up, stopping, and commutation processes of open-loop stepper motors, nor does it establish a lightweight ROI screening and consistency filtering mechanism for semi-static interference targets such as vehicles, pedestrians, and swaying leaves in the monitoring image.

[0006] Patent document CN111246094A discloses a gimbal offset compensation and correction method. This method reduces the backlash error during gimbal reversal by acquiring the backlash error of the gimbal in different motion directions and superimposing the backlash error in the corresponding direction when the preset motion direction is opposite to the previous motion direction. This type of solution mainly focuses on error compensation at the drive or motion control level, which can improve the mechanical backlash error problem. However, it does not combine camera image content to visually detect the actual image offset after the gimbal reaches the preset target position, nor does it use the consistency judgment of displacement vectors of multiple static background ROIs to eliminate semi-static interference areas.

[0007] US Patent 7239108B2 discloses a stepper motor position reference method, which involves using a mechanical hard stop, zero point, or reference point for position reference in an open-loop stepper motor. This approach illustrates that in stepper motor systems without position feedback, coordinate reset via a mechanical reference point is a common method. However, this type of solution typically only restores the coordinate reference at the moment of zeroing, making it difficult to correct the small residual deviations after each return of the pan-tilt unit to the target preset position in real time. Furthermore, using a fixed-period zeroing may interrupt the current monitoring or target tracking process, affecting monitoring continuity.

[0008] In addition, some existing solutions attempt to correct PTZ offset using image matching, whole-image feature comparison, or inertial sensor fusion. However, in actual monitoring scenarios, there are often semi-static or dynamic interference targets in the image, such as pedestrians standing still, vehicles moving or parked, leaves swaying, and changes in lighting. If whole-image matching or unfiltered local features are used directly for compensation, it is easy to misjudge image changes caused by dynamic targets as PTZ position offset, resulting in incorrect compensation. Furthermore, when there are drastic changes in lighting, partial occlusion, changes in image structure, or failure of historical reference images, if the system continues to perform visual compensation based on the failed reference images, it is easy to produce incorrect compensation, or even further amplify the PTZ preset position offset. If inertial devices such as IMUs are used, it will increase hardware costs and structural complexity, and cumulative drift problems may still exist.

[0009] In summary, while existing technologies cover video preset position calibration, gimbal backhaul error compensation, and open-loop stepper motor physical reference zeroing, they still have the following shortcomings: First, there is a lack of lightweight visual correction solutions suitable for low-cost IPC end-side devices; second, a progressive calibration chain is not formed by driver layer pulse pre-compensation, image layer visual correction, and event-driven physical zeroing; third, a robust consistency filtering mechanism for semi-static interference ROI is lacking; fourth, a safety degradation mechanism is not established for reference image failure or unreliable visual compensation; and fifth, the timing of physical zeroing is not dynamically selected based on visual compensation confidence and calibration urgency index. Therefore, there is an urgent need for a preset position calibration scheme suitable for open-loop stepper motor gimbals that, without adding additional hardware closed-loop feedback devices such as encoders and Hall sensors, reduces driver layer error accumulation, improves visual correction anti-interference capability, and performs timely physical zeroing calibration when visual compensation is unreliable or coordinate error risk is high, thereby improving the preset position positioning accuracy and long-term operational stability of the open-loop gimbal. Summary of the Invention

[0010] To address the problems existing in the prior art, the present invention also provides a method and system for calibrating the preset position of an open-loop gimbal based on multimodal feedback, and a computer-readable storage medium.

[0011] To achieve the above objectives, the present invention provides an open-loop gimbal preset position calibration method based on multimodal feedback, comprising the following steps: Step S0: Before the gimbal performs the target preset position movement, the pulse sequence of the stepper motor is pre-compensated for acceleration and deceleration by S curve, and a compensation pulse Δp is inserted before each commutation command according to the pre-established return error compensation table; Step S1: During the period when the gimbal is stationary, read the high-frequency component statistics output by the image signal processor (ISP), score the texture richness of the sub-regions of the image, filter out the sub-regions that have changed significantly during the period when the gimbal is stationary by the amount of pixel change in the time domain, select N static background sub-regions as dynamically selected regions of interest (ROIs), and extract the historical reference fingerprints of the ROIs. Step S2: After the gimbal rotates to the target preset position, calculate the displacement vector Vi between the N ROIs and their respective historical reference fingerprints; perform consistency checks on all displacement vectors based on the Random Sample Consensus Algorithm (RANSAC) and remove interfering ROIs that deviate from the common trend of most displacement vectors; calculate the visual compensation amount Decorr by weighted averaging of the remaining interior point displacement vectors. Step S3: Convert the visual compensation amount Dcorr into the corresponding stepper motor compensation pulse number ΔP and apply it to the gimbal drive; calculate the image correlation score Scorr between the current frame and the historical reference frame corresponding to the historical reference fingerprint extraction in real time. When Scorr is lower than the set confidence threshold Sth, suspend the visual compensation logic and enter the visual compensation untrusted state; when the visual compensation untrusted state meets the preset degradation condition, clear the historical reference fingerprint and schedule physical collision zeroing calibration. Step S4: Real-time maintenance and calibration urgency index Ical. When Ical exceeds the set urgency threshold Ith, trigger event-driven collision zeroing calibration to complete the gimbal coordinate system reset. Among them, steps S0 to S4 together constitute a three-level progressive virtual closed loop of pulse prediction, visual correction and physical zeroing.

[0012] Preferably, in step S0, when performing S-curve acceleration and deceleration pre-compensation on the drive pulse sequence of the stepper motor, the jerk (i.e., the rate of change of acceleration jerk) in the acceleration and deceleration phases is controlled to change continuously, and a compensation pulse Δp is inserted according to the return error compensation table before each commutation command, so as to suppress the accumulation of step loss error and mechanical return error of the stepper motor in the drive layer. In step S0, the S-curve acceleration / deceleration pre-compensation includes: dividing the entire acceleration phase into three sub-phases: acceleration phase, uniform acceleration phase, and deceleration phase. The duration of each sub-phase is predetermined based on the motor's rated torque and load inertia. The return error compensation table uses the mechanical reduction ratio as an index to store the number of commutation compensation pulses Δp corresponding to each reduction ratio. The Δp is written into the device firmware through the factory calibration process to suppress the accumulation of stepper motor step loss error and mechanical return error at the drive layer.

[0013] Preferably, in step S4, the calibration urgency index Ical is dynamically updated based on the time interval Tidle from the last physical calibration and the reversing operation frequency Frev within a preset statistical time window; the update rule of the calibration urgency index Ical is: Ical = α × Tidle + β × Frev, where α and β are preset weighting coefficients; When the target tracking control causes the gimbal to move near the physical limit, the limit-based zeroing path is triggered; when Ical exceeds the urgency threshold Ith and the screen is in an idle state, the idle silent zeroing path is triggered; when the cumulative duration of Ical exceeding the urgency threshold Ith exceeds the preset forced zeroing time limit Tforce, the timeout forced zeroing path is triggered. The above limit-based zeroing path, idle silent zeroing path, and timeout forced zeroing path all belong to the event-driven collision zeroing calibration in step S4.

[0014] Preferably, in the limit-following zeroing path, when the gimbal's horizontal rotation angle reaches within the preset margin angle θmargin from the physical limit, the system adds an overshoot pulse based on the current tracking command, causing the gimbal to actively touch the physical limit or the gimbal's original limit detection structure; after the collision, the system records the current pulse count and clears the absolute coordinates to zero, completing the coordinate system reset; the θmargin is dynamically adjusted according to the gimbal's horizontal rotation range and the current tracking target's movement trend; in the idle silent zeroing path, when Ical ≥ Ith and the image motion amplitude is lower than the stillness judgment threshold Mth within K consecutive frames, silent collision zeroing calibration is triggered, where K ranges from 3 to 10 frames; in the timeout forced zeroing path, regardless of whether the camera image is in a static state, collision zeroing towards the nearest physical limit is forcibly triggered, and Ical is cleared to zero after completing the coordinate system reset.

[0015] Preferably, in step S1, the high-frequency component statistics output by the ISP hardware include at least one of the autofocus evaluation function value (AF statistics) or AC (alternating current) high-frequency component energy value; the sub-region scoring rule is: the average high-frequency component of each sub-region is used as the texture richness score, and the N sub-regions with the highest scores are selected, where N ranges from 3 to 8; the temporal pixel change filtering step includes: within a continuous L frames where the gimbal is stationary, if the average inter-frame brightness difference of a certain sub-region exceeds the dynamic filtering threshold Pth, then the sub-region is removed from the ROI candidate set, where L ranges from 5 to 15 frames; During the ROI selection process, spatial arrangement constraints are also applied to the selected N ROIs to ensure that the minimum center-to-center distance between the N ROIs in the image is not less than dmin, so as to ensure that the ROIs are distributed in different regions of the image; the dmin is dynamically set according to the image resolution, and by default is not less than 15% of the image width.

[0016] Preferably, in step S2, the historical reference fingerprint is a feature descriptor of the ROI image block collected when the gimbal reaches the target preset position, including at least one of normalized grayscale histogram or gradient direction histogram; the inlier determination condition of the consistency test is: the Euclidean distance between the displacement vector Vi and the current RANSAC model estimate is less than the threshold dth, and the value of dth is in the range of 1 to 5 pixels. In step S2, the displacement vectors of the inlier points that pass the consistency test are weighted averaged when calculating the visual compensation amount Dicorr. The weight of each inlier ROI is determined based on its texture richness score and spatial location. ROIs with higher texture scores and locations closer to the center of the image are given higher weights to improve the accuracy and stability of visual compensation.

[0017] Preferably, in step S3, the image correlation score Scorr is calculated using the Normalized Cross-Correlation (NCC) algorithm, and the calculation range is the selected N ROI regions. When Scorr is lower than the confidence threshold Sth, the system marks the current visual compensation confidence as unreliable. After the unreliable state continues for more than C consecutive frames, it is determined that the visual compensation unreliable state meets the preset degradation condition. After the physical collision zeroing calibration is completed, the historical reference fingerprint is re-established. The value range of C is 2 to 8 frames. The conversion formula for converting the visual compensation amount Decorr into the number of stepper motor compensation pulses ΔP is: ΔP=Dcorr×Rpixel2step, where Rpixel2step is the conversion coefficient of the number of motor steps corresponding to a single pixel. Rpixel2step is determined by the current focal length, the image sensor pixel size, and the deceleration ratio.

[0018] This invention also provides an open-loop gimbal preset position calibration system based on multimodal feedback, including a camera module, a stepper motor-driven gimbal, and an embedded controller. The controller performs gimbal preset position calibration through the following functional modules: The motion pre-compensation module is used to reshape the target velocity sequence into an S-curve acceleration and deceleration pulse frequency sequence before the gimbal performs the target preset position motion, and inserts a compensation pulse Δp according to the return error compensation table before each reversal command to suppress the accumulation of stepper motor step loss error and mechanical return error. The dynamic ROI discovery module is used to read the high-frequency component statistics output by the image signal processor (ISP) of the camera module when the gimbal is stationary, perform texture scoring and dynamic filtering on multiple sub-regions of the image, and select N stable background sub-regions as regions of interest (ROIs) and extract their historical reference fingerprints. The visual consistency verification module is used to calculate the displacement vector between N ROIs and historical reference fingerprints after the gimbal reaches the target preset position, and after filtering out abnormal ROIs based on the Random Sample Consensus Algorithm (RANSAC), calculate the visual compensation amount Drorr based on the weighted average of the remaining inlier displacements. The compensation confidence management module is used to convert the visual compensation amount Dcorr into the corresponding number of stepper motor compensation pulses ΔP and apply it to the gimbal drive, and calculate the normalized cross-correlation score Scorr between the current frame and the historical reference frame in real time. When Scorr is lower than the confidence threshold Sth, the visual compensation logic is suspended, and the historical reference fingerprint is cleared to trigger physical collision zeroing calibration when the preset degradation conditions are met. The calibration scheduling module is used to maintain the calibration urgency index Ical and trigger event-driven collision zeroing calibration based on the relationship between Ical and the threshold Ith, including paths such as limit-based automatic zeroing, idle silent zeroing, and timeout forced zeroing.

[0019] Preferably, the calibration scheduling module includes a timeout forced zeroing subunit, which is used to force collision zeroing when the cumulative duration of Ical exceeding the threshold exceeds the forced zeroing time limit Tforce, regardless of whether the screen is in an idle state; the motion pre-compensation module stores a backhaul error compensation table indexed by the gimbal deceleration ratio; the visual consistency verification module is also used to maintain historical ROI displacement vector records within the sliding time window, and when the gimbal returns to the same target preset position multiple times in a row, the mean of the historical compensation vectors is used as the initial prior for RANSAC model estimation.

[0020] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the open-loop gimbal preset position calibration method based on multimodal feedback as described in any one of claims 1 to 7.

[0021] The technical solution of this invention has at least the following beneficial effects: (1) To address the issues of step loss and backlash error accumulation in open-loop stepper motors, this invention performs S-curve acceleration and deceleration pre-compensation on the stepper motor pulse sequence before the gimbal executes the target preset position movement. Furthermore, a compensation pulse Δp is inserted before each commutation command based on the backlash error compensation table. This reduces step loss and mechanical backlash error accumulation during stepper motor startup, shutdown, and commutation at the drive layer, thereby reducing the initial deviation of the gimbal before reaching the target preset position. Under prototype verification conditions, this invention, compared to the uncompensated open-loop stepper scheme and the whole-map feature matching scheme, can reduce the preset position offset and improve backlash stability.

[0022] (2) To address the problem that low-computing-power platforms struggle to run deep learning or complex full-image matching schemes, this invention utilizes the AF statistics and AC high-frequency component statistics natively output by the image signal processor (ISP) hardware for dynamic ROI discovery and static background reference region selection. This eliminates the need for complex feature extraction of the entire frame and avoids reliance on GPUs, NPUs, or additional AI models, thus obtaining texture-rich and relatively stable candidate regions and reducing on-device computation. Under a prototype verification condition, the on-device computational overhead of this invention is lower than that of the full-image feature matching scheme, making it suitable for deployment in cost-sensitive IPC PTZ cameras.

[0023] (3) In view of the problem that dynamic targets or semi-static interferences are easy to contaminate the visual compensation amount, the present invention filters out the sub-regions that change significantly by the temporal pixel change amount of multiple consecutive frames during the gimbal's static period, and selects the static background ROI by combining ROI spatial arrangement constraints. This can avoid misselecting moving people, vehicles, swaying leaves, flickering light, or partially occluded areas as visual reference areas, thereby improving the stability of historical reference fingerprints.

[0024] (4) To address the problem of visual compensation contamination caused by semi-static interference, this invention proposes a displacement vector consistency filtering mechanism based on the Random Sample Consensus Algorithm (RANSAC). This mechanism performs consistency checks on the displacement vectors of multiple ROIs relative to historical reference fingerprints, identifying and removing outliers that deviate from the common displacement trend of most static background ROIs. The visual compensation amount, Decorr, is then obtained through a weighted average of the interior point displacement vectors. This effectively eliminates the contamination of visual compensation by semi-static or dynamic interference areas such as parked vehicles, moving pedestrians, swaying leaves, and partial occlusion, improving the accuracy of preset position return and the stability of compensation in complex monitoring scenarios such as courtyard entrances, parking areas, and outdoor inspections.

[0025] (5) This invention uses the image correlation score Scorr to determine the credibility of visual compensation. When Scorr is lower than the credibility threshold Sth, the system suspends the visual compensation logic, clears the historical reference fingerprints after the preset degradation conditions are met in the untrusted state of visual compensation, and schedules physical collision zeroing calibration. This mechanism can prevent failed reference fingerprints from continuously participating in compensation, thereby reducing the risk of miscompensation in the case of drastic changes in illumination, scene occlusion, or failure of the reference area.

[0026] (6) To address the issue that traditional timed zeroing can easily interrupt monitoring, this invention introduces an event-driven collision zeroing mechanism. The zeroing timing is dynamically selected based on the calibration urgency index Ical, the time interval since the last physical calibration Tidle, and the reversing operation frequency Frev. Three paths are set: limit-based zeroing, idle silent zeroing, and timeout-forced zeroing. By integrating some calibration operations into the limit actions during target tracking or gimbal movement, the impact of zeroing actions on normal monitoring footage can be reduced. In scenarios with dense crowds or frequent vehicle movement where the footage is constantly busy, the timeout-forced zeroing path provides a safety net for coordinate system reset.

[0027] (7) This invention does not require additional hardware feedback devices such as encoders, Hall sensors, and IMUs. It can achieve open-loop gimbal preset position calibration through a three-level progressive virtual closed loop of "pulse prediction - visual correction - physical zeroing". This three-level progressive mechanism can take into account the driving layer error suppression, image layer visual correction and coordinate system physical reset, thereby improving the preset position positioning accuracy and stability of the open-loop gimbal during long-term operation while controlling hardware costs and computational overhead. Attached Figure Description

[0028] Figure 1 This is a schematic diagram of the system framework control of the present invention; Figure 2 This is a schematic diagram of the method flow of the present invention; Figure 3 This is a schematic diagram illustrating the dynamic ROI discovery and RANSAC consistency filtering of this invention. Figure 4 This is a timing diagram for calibrating the urgency index of this invention. Detailed Implementation

[0029] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the following embodiments are only used to illustrate the present invention and are not intended to limit the scope of protection of the present invention. In the absence of conflict, the technical features in the following embodiments can be combined with each other. In the following embodiments, Embodiment 1 is used to illustrate the method flow of the present invention, Embodiment 2 is used to illustrate the system module implementation corresponding to the method, Embodiment 3 is used to illustrate the application process of the present invention in a typical semi-static interference monitoring scenario, and Embodiment 4 is used to illustrate the implementation of a computer-readable storage medium. Each embodiment revolves around the same three-level progressive virtual closed-loop scheme of pulse prediction—visual correction—physical zeroing; the virtual closed loop refers to a software calibration closed loop formed by pre-compensation of the driving layer, visual correction of the image layer, and physical zeroing calibration without relying on hardware position closed-loop feedback such as encoders, Hall sensors, or IMUs.

[0030] Example 1: Method Example Reference Figures 1 to 4 This embodiment uses a low-cost IPC PTZ camera as an application background to illustrate an open-loop PTZ preset position calibration method based on multimodal feedback. The IPC PTZ camera uses an open-loop stepper motor to control the horizontal rotation and vertical pitch of the PTZ. The horizontal rotation range can be a preset mechanical travel range, such as 0 to 360 degrees, and the vertical pitch range can be -45 to +45 degrees. The IPC PTZ camera does not have an angle encoder, Hall sensor, or other absolute position closed-loop detection device. The image signal processor (ISP) can output the autofocus evaluation function (AF) statistics and AC component statistics for the embedded controller to execute the visual calibration algorithm. The IPC PTZ camera includes a camera module, an ISP, a stepper motor drive circuit, an embedded controller, and a memory. The embedded controller is used to perform PTZ motion control, image data statistics reading, ROI filtering, visual displacement estimation, compensation pulse conversion, and collision zeroing scheduling.

[0031] The method in this embodiment includes the following steps.

[0032] Step S0: S-curve acceleration / deceleration pre-compensation.

[0033] Before the gimbal performs the target preset position movement, the stepper motor driver chip receives the target velocity sequence V_target(t) output by the upper-level tracking controller or preset position controller. The S-curve shaping unit in the embedded controller reshapes the target velocity sequence V_target(t) into a continuous pulse frequency sequence of acceleration, thereby reducing the risk of shock and step loss during the stepper motor's start-up, stopping, and commutation processes.

[0034] Specifically, the S-curve acceleration / deceleration pre-compensation includes at least three stages: an acceleration phase, a uniform acceleration phase, and a deceleration phase. The acceleration phase corresponds to the time interval t∈[0, t1], the uniform acceleration phase to the time interval t∈[t1, t2], and the deceleration phase to the time interval t∈[t2, t3]. t1, t2, and t3 can be pre-calibrated or calculated based on the stepper motor's rated torque, the gimbal's load inertia, the mechanical reduction ratio, and the target rotation angle.

[0035] Furthermore, the commutation detection unit in the embedded controller monitors the gimbal's motion direction switching signal in real time. When a commutation of the stepper motor is detected, the commutation detection unit injects Δp compensation pulses according to a pre-established retrace error compensation table before each commutation command is executed. The retrace error compensation table can use mechanical reduction ratio, rotation direction, or gimbal model parameters as indexes to store the corresponding number of commutation compensation pulses Δp under different conditions. Δp can be written into the device firmware through the device's factory calibration process.

[0036] By injecting the aforementioned S-curve acceleration / deceleration pre-compensation and commutation compensation pulses, the accumulation of stepper motor step loss error and mechanical return error can be reduced in the drive layer, providing a smaller initial deviation range for subsequent visual correction.

[0037] Step S1: Dynamic ROI discovery and static background reference region locking.

[0038] Once the pan-tilt unit reaches the target preset position and enters a stationary state, the dynamic ROI discovery module reads the AF statistics and AC component energy values ​​output by the ISP driver interface. The embedded controller divides the current monitoring screen into M×M sub-regions. For example, when M=16, a 1080p frame can be divided into 256 sub-regions. For each sub-region i, the controller calculates the mean high-frequency component value score_i for that sub-region and uses score_i as the texture richness score.

[0039] A higher texture richness score indicates that the region has richer edge, texture, or grayscale variations, making it more suitable as a reference region for visual matching and displacement estimation. The controller selects N sub-regions with higher texture richness scores from each sub-region to enter the ROI candidate set. N can be configured to be 3 to 8, preferably 5.

[0040] Furthermore, to exclude dynamic elements in the image, such as swaying leaves, moving people, passing vehicles, and flickering light areas, the system calculates the average inter-frame brightness difference ΔBright_i for each ROI candidate region within a continuous L frames where the pan-tilt unit is stationary. L can be configured to be 5 to 15 frames, preferably 8. If the average inter-frame brightness difference ΔBright_i of a candidate region exceeds the dynamic filtering threshold Pth, where Pth represents the inter-frame brightness difference threshold, the candidate region is marked as an unstable region and removed from the ROI candidate set. Pth can be set according to the image brightness quantization bit width; for example, under 8-bit brightness quantization, Pth can be 15.

[0041] When a candidate region is removed, the system selects a candidate region with a higher score and that meets the stability condition from the remaining candidate regions to maintain the ROI set size at N. All ROIs that pass the stability filter constitute the current static background reference region set.

[0042] Furthermore, to avoid the selected ROIs being concentrated in a certain corner of the image and affecting the spatial robustness of the visual compensation calculation, this embodiment also applies spatial arrangement constraints to the ROIs. Specifically, the center-to-center distance between any two of the N ROIs in the image must meet the minimum distance constraint dmin, which can be dynamically set according to the image width. For example, dmin is not less than 15% of the image width. This allows multiple ROIs to be distributed in different areas of the image, improving the stability of subsequent displacement vector consistency judgment.

[0043] After determining the ROI set, the system extracts the feature descriptors of each ROI image block as historical reference fingerprints. The historical reference fingerprints can be normalized grayscale histograms, gradient orientation histograms, or normalized pixel block data. Preferably, normalized grayscale histograms are used as the historical reference fingerprints by default. The system associates the historical reference fingerprints of each ROI with its coordinate position in the image and a preset target identifier, storing them in the historical reference fingerprint database.

[0044] Step S2: Displacement vector consistency filtering based on RANSAC.

[0045] After the gimbal receives the target preset position reset command, the driver drives the gimbal to rotate to the target angle according to the pre-compensation pulse sequence generated in step S0. After the gimbal reaches the target preset position, the visual consistency verification module performs displacement vector calculations on the N ROIs established in step S1.

[0046] Specifically, for ROI_i, the system performs template matching in the current frame image using the historical reference coordinates of ROI_i as the center, locates the optimal matching position, and calculates the displacement vector V_i based on the offset between the optimal matching position and the historical reference position. The displacement vector V_i can be expressed as: V_i=(Δx_i, Δy_i) Where Δx_i represents the pixel offset of ROI_i in the horizontal direction of the image, and Δy_i represents the pixel offset of ROI_i in the vertical direction of the image.

[0047] After collecting the displacement vectors corresponding to N ROIs, the visual consistency verification module executes the Random Sample Consistency Algorithm (RANSAC) to filter out abnormal ROIs affected by dynamic or semi-static targets. Its core judgment criterion is that ROIs truly belonging to a static background should exhibit a relatively consistent displacement trend on the image plane after the gimbal's preset position offset; while ROIs affected by people occlusion, swaying leaves, vehicle movement, or sudden changes in local lighting will typically deviate from the common trend of most displacement vectors.

[0048] In one implementation, the RANSAC consistency filtering process includes: (1) Randomly select two displacement vectors from N displacement vectors and calculate their mean as the current model estimation vector V_model; (2) Calculate the Euclidean distance between all displacement vectors V_i and the current model estimation vector V_model; (3) If the Euclidean distance between a certain displacement vector V_i and V_model is less than the interior point threshold dth, then the displacement vector V_i is recorded as an interior point; otherwise, the displacement vector V_i is recorded as an exterior point. (4) Repeat the above random sampling and interior point statistics process, and retain the model estimation vector V_model with the largest number of interior points or the smallest residual; wherein, the number of iterations can be a preset number, such as no more than 20 times, or determined according to the number of sample combinations of N displacement vectors; (5) ROIs that deviate from the final model estimate vector by more than the inlier threshold dth are marked as disturbance outliers and are not included in the final compensation calculation.

[0049] The value of dth can be configured to be 1 to 5 pixels, and preferably, dth is 3 pixels.

[0050] After completing the RANSAC consistency filtering, the system calculates the visual compensation amount Drorr based on the final inlier set. Specifically, the system performs a weighted average of the displacement vectors in the inlier set to obtain the inlier mean vector V_inlier_mean, and uses this inlier mean vector as the basis for visual compensation. The weight w_i of each inlier ROI_i can be determined jointly based on its texture richness score score_i and its normalized distance dcenter_i from the center of the image. Preferably, ROIs with higher texture richness scores and closer locations to the center of the image are assigned higher weights.

[0051] The RANSAC consistency filtering described above can effectively eliminate abnormal ROIs corresponding to semi-static interference targets, avoiding incorrect compensation by the gimbal due to local dynamic changes.

[0052] Step S3: Visual compensation calculation and confidence management.

[0053] In step S3, the system calculates the visual compensation amount Drorr based on the inlier_mean vector V_inlier_mean obtained in step S2, and converts the pixel displacement into a motor compensation pulse. Specifically, Drorr can be the pixel displacement vector corresponding to V_inlier_mean, or the horizontal and vertical compensation amounts can be obtained from this pixel displacement vector respectively; when only the compensation amplitude is calculated, Drorr can be the magnitude of V_inlier_mean, and the direction of the compensation pulse is determined in combination with the direction of V_inlier_mean.

[0054] The system converts the visual compensation amount Drorr into the number of stepper motor compensation pulses ΔP according to the following formula: ΔP = Decorr × Rpixel2step, where Rpixel2step is the conversion factor for the number of motor steps corresponding to a single pixel; ΔP represents the number of stepper motor compensation pulses calculated based on the visual compensation amount Decorr.

[0055] In one implementation, for a fixed focal length lens, Rpixel2step can be obtained through factory calibration. Specifically, during factory calibration, the gimbal is driven to rotate precisely a preset number of steps Ncal in the horizontal or vertical direction, for example, Ncal = 200 steps; then, a calibration board or fixed reference object is photographed, and the pixel displacement Dcal of the feature points in the image is measured; based on Ncal and Dcal, the following is calculated: Rpixel2step=Ncal / Dcal The calculation results are written into the device firmware for use in subsequent preset position calibration processes.

[0056] In another implementation, Rpixel2step can also be obtained through a self-learning method upon initial power-on. Upon first startup, the device automatically executes a preset rotation and image measurement process, calculates Rpixel2step, and persists it without manual intervention.

[0057] For variable focal length lenses, since changes in focal length affect the actual angle range corresponding to a single pixel, Rpixel2step can be dynamically updated as the focal length changes. Specifically, the system can pre-calibrate the corresponding Rpixel2step for different focal length settings and establish a focal length-pixel step conversion table. When the zoom value changes, the system updates Rpixel2step by looking up the table based on the current focal length value.

[0058] After calculating the number of compensation pulses ΔP, the embedded controller outputs compensation pulses to the stepper motor driver chip, causing the gimbal to perform a slight correction movement, thereby reducing the residual deviation between the actual position of the gimbal and the target preset position.

[0059] Furthermore, to avoid continuing error compensation when the visual reference fails, this embodiment sets up a compensation confidence management mechanism. After each visual verification, the system calculates the image correlation score Scorr between the current frame and the historical reference frame. The historical reference frame is the image frame corresponding to the establishment or updating of the historical reference fingerprint, and the historical reference fingerprint is the feature descriptor extracted from the ROI image patch in the historical reference frame. Preferably, the image correlation score Scorr is calculated using the Normalized Cross-Correlation (NCC) algorithm, and its value range is [-1, 1].

[0060] When Scorr is not lower than the confidence threshold Sth, it indicates that the current frame has a high similarity to the historical reference frame, the visual compensation result is reliable, and the system continues to execute the visual compensation logic. When Scorr is lower than the confidence threshold Sth, it indicates that the current visual environment is unreliable, the system immediately suspends the visual compensation logic, and enters the unreliable visual compensation state. Sth can be set according to the application scenario; preferably, Sth is 0.6.

[0061] In one implementation, after a visual compensation untrusted state persists for C consecutive frames, the system determines that the visual compensation untrusted state meets a preset degradation condition. C can be configured to be 2 to 8 frames, preferably 4. At this time, the system clears the historical reference fingerprint corresponding to the current target preset bit and forces a physical collision zeroing calibration. After the physical collision zeroing calibration is completed, the system re-executes step S1 to re-establish a new historical reference fingerprint.

[0062] By using the above methods, we can avoid the long-term involvement of failed historical reference fingerprints in visual compensation, thereby reducing the risk of miscompensation.

[0063] Step S4: Calibrate urgency management and event-driven collision zeroing calibration.

[0064] During long-term operation of an open-loop stepper motor gimbal, step loss, mechanical return error, changes in frictional resistance, and frequent reversing operations can all lead to a cumulative deviation between the gimbal's pulse count position and its actual mechanical position. Therefore, this embodiment maintains and calibrates the urgency index Ical in real time and dynamically determines whether to perform a physical collision zeroing calibration based on this index.

[0065] In one implementation, the calibration urgency index Ical is dynamically updated based on the time interval Tidle since the last physical calibration and the commutation operation frequency Frev within a preset statistical time window. The update rule is as follows: Ical = α × Tidle + β × Frev Here, α and β are preset weighting coefficients. A larger Tidle indicates a longer time since the last physical collision zeroing calibration; a larger Frev indicates more frequent gimbal reversing operations within the preset statistical time window, leading to a higher risk of accumulated mechanical backtracking and step loss errors. The preset statistical time window is a recent period used to statistically analyze the reversing operation frequency Frev, such as the last 10 minutes, the last 30 minutes, or a sliding time window set by system parameters. Therefore, when Ical exceeds the set urgency threshold Ith, the system triggers event-driven collision zeroing calibration.

[0066] In this embodiment, the event-driven collision zeroing calibration includes a limit-based follow-up zeroing path, an idle silent zeroing path, and a timeout forced zeroing path.

[0067] In the limit-following zeroing path, when the target tracking control causes the gimbal to move near the physical limit, the system continues to move towards the corresponding physical limit direction using the current tracking direction. When the horizontal rotation angle of the gimbal reaches within the preset margin angle θmargin from the physical limit, the controller adds an overshoot pulse based on the current tracking command, causing the gimbal to actively contact the physical limit part or the gimbal's original limit detection structure. This structure is only used for coordinate reset and does not constitute a continuous position closed-loop feedback device such as an encoder, Hall sensor, or IMU. After the collision, the system records the current pulse count and clears the absolute coordinates to zero, thereby completing the gimbal coordinate system reset. The θmargin can be dynamically adjusted according to the horizontal rotation range of the gimbal, the current tracking target's movement trend, or the current gimbal movement speed.

[0068] In the idle silent zeroing path, when Ical exceeds the urgency threshold Ith and the screen is in an idle state, the system performs silent collision zeroing calibration without affecting the main monitoring tasks. Specifically, the system determines whether the screen motion amplitude within K consecutive frames is lower than the stationary determination threshold Mth. If the screen motion amplitude within K consecutive frames is lower than Mth, the current screen is determined to be in an idle state, and silent collision zeroing calibration is triggered. K can be configured to be 3 to 10 frames.

[0069] In the forced zeroing path after timeout, when the cumulative duration of Ical exceeding the urgency threshold Ith exceeds the preset forced zeroing time limit Tforce, the system will forcibly trigger collision zeroing towards the nearest physical limit, regardless of whether the camera image is stationary. After completing collision zeroing, the system will clear the absolute coordinates of the gimbal and clear Ical.

[0070] Through the above three paths, the system does not need to be zeroed frequently according to a fixed cycle. Instead, it selects an appropriate time to zero based on the actual operational risks of the PTZ and the status of the monitoring screen, thereby reducing the impact of the zeroing action on the normal monitoring process.

[0071] Furthermore, in this embodiment, steps S0 to S4 together constitute a three-level progressive virtual closed loop of pulse prediction, visual correction, and physical zeroing. Step S0 is used to reduce the driving layer error before the gimbal reaches the target preset position; steps S1 to S3 are used to correct the remaining positioning deviation after the gimbal reaches the target preset position under visually reliable conditions; and step S4 is used to physically reset the gimbal coordinate system when visual compensation fails or the calibration urgency index Ical meets the trigger condition. In other words, the first level is pulse prediction compensation based on S-curve acceleration / deceleration and hysteresis error compensation; the second level is real-time visual correction based on ROI displacement vector consistency filtering; and the third level is forced coordinate system reset based on event-triggered collision zeroing. When the second-level visual compensation fails due to the image correlation score being lower than the reliability threshold, the system automatically downgrades to the third-level physical collision zeroing calibration. Therefore, this invention can improve the preset position positioning accuracy and long-term operational stability of an open-loop gimbal without the need for an encoder or Hall sensor.

[0072] Experimental data verification To verify the technical effectiveness of this embodiment, preset position switching, semi-static interference filtering, CPU utilization, and calibration state machine tests were conducted on several typical IPC hardware platforms. The following test data are based on the applicant's prototype in a typical IPC PTZ camera scenario. During testing, the test platform adopted an open-loop stepper motor PTZ structure without encoders, Hall sensors, or IMUs for hardware position closed-loop feedback. The test screen resolution was 1920×1080 or 2560×1440, the frame rate was 25fps or 30fps, and the lens focal length was 3.6mm, 4mm, or 6mm. The number of ROIs N was set to 4, the number of consecutive frames L was set to 8, the RANSAC in-point threshold dth was set to 3 pixels, the confidence threshold Sth was set to 0.6, the calibration urgency threshold Ith was set to 70, and the forced zeroing time limit Tforce was set to 30 minutes. The above parameters can be adjusted according to different IPC platforms, lens focal lengths, and PTZ structures. The results in the table are used to illustrate the trend improvement effect of the present invention compared with the uncompensated scheme and the whole image feature matching scheme, and are not intended to limit the scope of protection of the present invention. The test results are as follows: Table 1. Preset Position Offset Accuracy Comparison Test: Table 1

[0073] Table 1 illustrates the improvement in positioning accuracy of the present invention in scenarios with repeated preset position switching. As shown in Table 1, under different IPC hardware platforms and lens focal lengths, the average offset after preset position return is relatively large without compensation. When using the whole-image feature matching comparison scheme, the improvement in average offset is limited due to its susceptibility to local dynamic targets, lighting changes, and the computational load of the whole image. After adopting the scheme of the present invention, through the driving layer pulse pre-compensation in step S0, the static background ROI screening in steps S1 to S3, and RANSAC consistent visual correction, the average preset position offset across platforms can be significantly reduced, and the P95 offset can be controlled at a low level. P95 represents the 95th percentile offset in the test statistics. This result demonstrates that the present invention can improve the target preset position return accuracy in an open-loop gimbal structure without encoders or Hall sensors.

[0074] Table 2. RANSAC Consistency Filtering Performance Test in Semi-Static Interference Scenarios: Table 2

[0075] Table 2 illustrates the effect of this invention on suppressing miscompensation in semi-static interference scenarios. As shown in Table 2, in the interference-free control scenario, the system basically does not experience miscompensation. However, in the presence of stationary vehicles, pedestrians, swaying shrubs, and mixed interference from vehicles and trees, without RANSAC consistency filtering, some interfered ROIs are incorrectly included in the visual compensation calculation, leading to a significant increase in the number of miscompensations. After applying the RANSAC consistency filtering of this invention, abnormal ROIs deviating from the common displacement trend of most static background ROIs are identified as outliers and eliminated, significantly reducing the number of miscompensations. This result demonstrates that this invention can effectively distinguish between static background ROIs and semi-static interference ROIs, improving the reliability of visual compensation in complex monitoring scenarios.

[0076] Table 3 Comparison of CPU utilization and processing latency for each scheme: Table 3

[0077] Table 3 illustrates the computational resource consumption of this invention on an embedded IPC platform. As shown in Table 3, the whole-image feature matching and comparison scheme requires feature extraction and matching of the entire frame, resulting in high CPU utilization and high single-frame processing time, making it generally more suitable for high-end platforms. While the IMU inertial fusion comparison scheme does not rely on a GPU or NPU, it requires additional inertial devices and incurs certain computational overhead. This invention only filters a small number of ROIs based on high-frequency statistics output by the ISP and performs displacement estimation and consistency filtering only on the ROI regions, thus resulting in low CPU utilization and low single-frame processing time. These results indicate that this invention is suitable for deployment on mid-to-low-end IPC SoC platforms and can achieve preset position calibration without relying on GPUs / NPUs or adding hardware feedback devices such as encoders and Hall sensors.

[0078] Table 4. Calibration Urgency Index Ical State Machine and Three-Path Zeroing Trigger Conditions: Table 4

[0079] Table 4 illustrates the state machine division and three-path zeroing scheduling logic of the calibration urgency index Ical. As shown in Table 4, when Ical is in the green normal state, the system mainly relies on visual compensation to maintain the preset position accuracy; when Ical enters the yellow warning state, the system prioritizes the limit-based automatic zeroing path, integrating the zeroing action into subsequent PTZ motion commands to reduce user perception; when Ical enters the orange alert state, the system performs silent zeroing when the screen is idle to minimize the impact on the monitoring process; when Ical enters the red forced state, indicating a high risk of accumulated PTZ coordinate errors, the system triggers a timeout-forced zeroing to ensure timely coordinate system reset. This result demonstrates that the present invention does not employ fixed-period zeroing, but rather dynamically selects the zeroing timing and path based on the Tidle, Frev, and Ical states, thereby achieving a balance between positioning accuracy and monitoring continuity.

[0080] In summary, Tables 1 to 4 verify the technical effectiveness of this invention from four aspects: preset position positioning accuracy, semi-static interference suppression capability, embedded platform computational overhead, and physical zeroing scheduling strategy. The above results demonstrate that this invention can improve the preset position positioning accuracy and long-term operational stability of an open-loop gimbal without adding hardware closed-loop feedback devices such as encoders and Hall sensors. This is achieved through software virtual closed-loop control formed by drive-layer pulse pre-compensation, image-layer visual consistency correction, and event-driven physical zeroing calibration.

[0081] Example 2: This embodiment provides an open-loop gimbal preset position calibration system based on multimodal feedback. The system includes a camera module, a stepper motor-driven gimbal, and an embedded controller. The embedded controller may include a motion pre-compensation module, a dynamic ROI discovery module, a visual consistency verification module, a compensation confidence management module, and a calibration scheduling module.

[0082] The motion pre-compensation module is used to reshape the target velocity sequence into an S-curve acceleration / deceleration pulse frequency sequence before the gimbal performs the target preset position movement. Before each reversal command, it inserts compensation pulses Δp according to the backlash error compensation table to suppress the accumulation of stepper motor step loss error and mechanical backlash error. Δp represents the number of backlash error compensation pulses inserted before the reversal command. The motion pre-compensation module can store a backlash error compensation table indexed by the gimbal reduction ratio, which can be written into the firmware during the equipment factory calibration process.

[0083] The dynamic ROI discovery module is used to read the high-frequency component statistics of the image signal processor (ISP) output of the camera module when the gimbal is stationary, perform texture scoring and dynamic filtering on multiple image sub-regions, and select N stable background sub-regions as regions of interest (ROIs) and extract their historical reference fingerprints.

[0084] The visual consistency verification module calculates the displacement vectors between N ROIs and historical reference fingerprints after the gimbal reaches the target preset position. After filtering out abnormal ROIs using the Random Sample Consensus Algorithm (RANSAC), it calculates the visual compensation amount (Dcorr) using the weighted average of the remaining inlier displacements. Furthermore, the visual consistency verification module can maintain a record of historical ROI displacement vectors within a sliding time window. When the gimbal returns to the same target preset position multiple times consecutively, the average of the historical compensation vectors is used as the initial prior for the RANSAC model estimation.

[0085] The compensation confidence management module converts the visual compensation amount (Dcorr) into the corresponding number of stepper motor compensation pulses (ΔP) and applies it to the gimbal drive. It also calculates the normalized cross-correlation score (Scorr) between the current frame and historical reference frames in real time. When Scorr falls below the confidence threshold (Sth), the compensation confidence management module suspends the visual compensation logic and clears the historical reference fingerprint when preset degradation conditions are met, thereby triggering physical collision zeroing calibration.

[0086] The calibration scheduling module maintains the calibration urgency index Ical and triggers event-driven collision zeroing calibration based on the relationship between Ical and the threshold Ith. The calibration scheduling module may include a timeout-forced zeroing subunit. This subunit forces collision zeroing when the cumulative duration of Ical exceeding the threshold exceeds the forced zeroing timeout Tforce, regardless of whether the screen is idle.

[0087] Furthermore, the calibration scheduling module may also include a timer subunit, a commutation frequency statistics subunit, and a limit sensing subunit. The timer subunit is used to count the time interval Tidle since the last physical collision zeroing calibration; the commutation frequency statistics subunit is used to count the commutation operation frequency Frev within a preset statistical time window; and the limit sensing subunit is used to determine whether the gimbal has moved to the vicinity of a physical limit.

[0088] The modules can communicate with each other with low latency via shared memory, event queues, or message queues. Each module can run on the embedded SoC, microcontroller, or real-time operating system of the PTZ camera. Through modular configuration, the system can achieve preset position calibration, visual compensation, and physical zeroing scheduling on low- to mid-range IPC hardware platforms.

[0089] Example 3: This embodiment uses a semi-static interference scenario at a courtyard entrance as an example to further illustrate the anti-interference compensation process of the present invention. A PTZ camera is installed at the gate of a residential courtyard, and the monitored area includes fixed building structures and vehicles that stop intermittently. The fixed building structures may include walls, gateposts, ground textures, or building edges. Preset position P1 is aligned with the courtyard gate, and preset position P2 is aligned with the parking area. The PTZ camera can cruise back and forth between P1 and P2 according to a preset cruise strategy. When the device uses a fixed-focus lens, Rpixel2step can automatically calibrate through a self-learning process upon initial power-on.

[0090] During a gimbal rotation from P1 to P2, vehicle A leaves the parking area, exposing the ground texture previously obscured by the vehicle. Simultaneously, vehicle B enters and stops at different locations. At this time, some ROIs in the P2 view are located within the vehicle area, and their displacement vectors deviate significantly from the actual background displacement vectors. For example, the system maintains N=4 ROIs, where ROI_1 is located in the gatepost texture area, ROI_4 is located in the wall paint area, and ROI_2 and ROI_3 are located in the vehicle interference area. The displacement vectors V_2 and V_3 corresponding to the vehicle areas may deviate from the background displacement vectors by 15 to 30 pixels.

[0091] After performing RANSAC consistency filtering, ROI_1 and ROI_4 showed high consistency in their displacement vectors (e.g., deviation less than 3 pixels) and were classified as inliers. ROI_2 and ROI_3, affected by vehicle movement or occlusion, were identified as interfering outliers and removed from the compensation calculation. When calculating the weighted mean, ROI_1, with its high AF score and central location, was assigned a higher weight (e.g., a weighting coefficient of 0.65); ROI_4, with its moderate AF score and central location, was assigned a lower weight (e.g., a weighting coefficient of 0.35). The resulting weighted compensation vector exhibited greater stability compared to the equal-weighted mean, reducing the contamination of visual compensation by semi-static interfering targets.

[0092] Furthermore, if there is frequent entry and exit from the courtyard during the day, and pedestrians or vehicles are constantly moving in the frame, causing the idle silent zeroing path to fail to trigger for an extended period, then the timeout forced zeroing path can be intervened as a fallback mechanism. For example, when the cumulative duration of Ical exceeding the urgency threshold Ith reaches Tforce, where Tforce can be 30 minutes, the system forcibly executes collision zeroing during the current cruise interval to ensure the long-term stability of the gimbal coordinate system in extremely active scenarios.

[0093] As can be seen from the above typical scenarios, the present invention can identify and eliminate semi-static interference ROIs by using static background ROI and RANSAC consistency filtering, and at the same time use timeout forced zeroing path to avoid coordinate accumulation deviation caused by long-term inability to calibrate, thereby improving the positioning stability of preset positions in complex monitoring scenarios.

[0094] Example 4: Example of a computer-readable storage medium This embodiment also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the open-loop gimbal preset position calibration method based on multimodal feedback described in Embodiment 1. The computer-readable storage medium can be flash memory, read-only memory, random access memory, memory card, solid-state drive, eMMC, or non-volatile memory in an embedded device. The computer program can be stored as firmware in the embedded SoC or embedded controller of the gimbal camera and loaded and run after the device is powered on.

[0095] Through the above embodiments, the present invention can improve the positioning accuracy of the open-loop gimbal when returning to the preset position and reduce the risk of miscompensation caused by semi-static targets, partial occlusion or reference image failure without adding hardware closed-loop feedback devices such as encoders and Hall sensors.

[0096] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural transformations made using the contents of the present invention's specification and drawings under the inventive concept of the present invention, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.

Claims

1. A method for calibrating an open-loop gimbal preset position based on multimodal feedback, characterized in that, Includes the following steps: Step S0: Before the gimbal performs the target preset position movement, the pulse sequence of the stepper motor is pre-compensated for acceleration and deceleration by S curve, and a compensation pulse Δp is inserted before each commutation command according to the pre-established return error compensation table; Step S1: During the period when the gimbal is stationary, read the high-frequency component statistical data output by the image signal processor, score the texture richness of the sub-regions of the image, and filter out the sub-regions that have changed significantly during the period when the gimbal is stationary by the amount of pixel change in the time domain. Select N static background sub-regions as dynamically selected regions of interest (ROIs) and extract the historical reference fingerprints of the ROIs. Step S2: After the gimbal rotates to the target preset position, calculate the displacement vector Vi between the N ROIs and their respective historical reference fingerprints; perform consistency checks on all displacement vectors based on the random sample consistency algorithm, and remove interfering ROIs that deviate from the common trend of most displacement vectors; calculate the visual compensation amount Decorr by weighted average of the remaining interior point displacement vectors. Step S3: Convert the visual compensation amount Dcorr into the corresponding stepper motor compensation pulse number ΔP and apply it to the gimbal driver; calculate the image correlation score Scorr between the current frame and the historical reference frame corresponding to the extraction of historical reference fingerprint in real time. When Scorr is lower than the set confidence threshold Sth, suspend the visual compensation logic and enter the visual compensation untrusted state. When the visual compensation untrusted state meets the preset degradation condition, clear the historical reference fingerprint and schedule physical collision zeroing calibration. Step S4: Real-time maintenance and calibration urgency index Ical. When Ical exceeds the set urgency threshold Ith, trigger event-driven collision zeroing calibration to complete the gimbal coordinate system reset. Among them, steps S0 to S4 together constitute a three-level progressive virtual closed loop of pulse prediction, visual correction and physical zeroing.

2. The open-loop gimbal preset position calibration method based on multimodal feedback according to claim 1, characterized in that, In step S0, when performing S-curve acceleration and deceleration pre-compensation on the drive pulse sequence of the stepper motor, the acceleration and deceleration phases are controlled to change continuously, and a compensation pulse Δp is inserted according to the return error compensation table before each commutation command. In step S0, the S-curve acceleration / deceleration pre-compensation includes: dividing the entire acceleration phase into three sub-phases: acceleration phase, uniform acceleration phase, and deceleration phase. The duration of each sub-phase is predetermined based on the motor's rated torque and load inertia. The return error compensation table uses the mechanical reduction ratio as an index to store the number of commutation compensation pulses Δp corresponding to each reduction ratio. The Δp is written into the device firmware through the factory calibration process to suppress the accumulation of stepper motor step loss error and mechanical return error at the drive layer.

3. The open-loop gimbal preset position calibration method based on multimodal feedback according to claim 1, characterized in that, In step S4, the calibration urgency index Ical is dynamically updated based on the time interval Tidle from the last physical calibration and the reversing operation frequency Frev within a preset statistical time window; the update rule of the calibration urgency index Ical is: Ical = α × Tidle + β × Frev, where α and β are preset weighting coefficients; When the target tracking control causes the gimbal to move near the physical limit, the limit-based zeroing path is triggered; when Ical exceeds the urgency threshold Ith and the screen is in an idle state, the idle silent zeroing path is triggered; when the cumulative duration of Ical exceeding the urgency threshold Ith exceeds the preset forced zeroing time limit Tforce, the timeout forced zeroing path is triggered. The above limit-based zeroing path, idle silent zeroing path, and timeout forced zeroing path all belong to the event-driven collision zeroing calibration in step S4.

4. The open-loop gimbal preset position calibration method based on multimodal feedback according to claim 3, characterized in that, In the limit-following zeroing path, when the horizontal rotation angle of the gimbal reaches within the preset margin angle θmargin from the physical limit, the system adds an overshoot pulse based on the current tracking command, causing the gimbal to actively contact the physical limit part or the original limit detection structure of the gimbal; after the collision, the system records the current pulse count and clears the absolute coordinates to zero, completing the coordinate system reset; the θmargin is dynamically adjusted according to the horizontal rotation range of the gimbal and the current tracking target's movement trend. In the idle silent zeroing path, when Ical ≥ Ith and the motion amplitude of the image is lower than the static determination threshold Mth within K consecutive frames, silent collision zeroing calibration is triggered, and the value of K ranges from 3 to 10 frames; in the timeout forced zeroing path, regardless of whether the camera image is in a static state, collision zeroing is forcibly triggered in the direction of the nearest physical limit, and Ical is cleared to zero after the coordinate system is reset.

5. The open-loop gimbal preset position calibration method based on multimodal feedback according to claim 1, characterized in that, In step S1, the high-frequency component statistics output by the ISP hardware include at least one of the autofocus evaluation function value or the AC high-frequency component energy value; the sub-region scoring rule is: the average high-frequency component of each sub-region is used as the texture richness score, and the N sub-regions with the highest scores are selected, where N ranges from 3 to 8. The temporal pixel change filtering step includes: if the average inter-frame brightness difference of a certain sub-region exceeds the dynamic filtering threshold Pth within a continuous L frames when the gimbal is stationary, the sub-region is removed from the ROI candidate set, and the value of L ranges from 5 to 15 frames. During the ROI selection process, spatial arrangement constraints are also applied to the selected N ROIs to ensure that the minimum center-to-center distance between the N ROIs in the image is not less than dmin, so as to ensure that the ROIs are distributed in different regions of the image; the dmin is dynamically set according to the image resolution, and by default is not less than 15% of the image width.

6. The open-loop gimbal preset position calibration method based on multimodal feedback according to claim 1, characterized in that, In step S2, the historical reference fingerprint is the feature descriptor of the ROI image block collected when the gimbal reaches the target preset position, including at least one of the normalized grayscale histogram or gradient direction histogram; the inlier determination condition of the consistency test is: the Euclidean distance between the displacement vector Vi and the current RANSAC model estimate is less than the threshold dth, and the value of dth is in the range of 1 to 5 pixels. In step S2, the displacement vectors of the inlier points that pass the consistency test are weighted averages when calculating the visual compensation amount Dicorr. The weight of each inlier ROI is determined based on its texture richness score and spatial location. Among them, ROIs with higher texture scores and closer locations to the center of the image are given higher weights.

7. The open-loop gimbal preset position calibration method based on multimodal feedback according to claim 1, characterized in that, In step S3, the image correlation score Scorr is calculated using a normalized cross-correlation algorithm, and the calculation range is the selected N ROI regions. When Scorr is lower than the confidence threshold Sth, the system marks the current visual compensation confidence as unreliable. After the unreliable state continues for more than C consecutive frames, it is determined that the visual compensation unreliable state meets the preset degradation condition. After the physical collision zeroing calibration is completed, the historical reference fingerprint is re-established. The value range of C is 2 to 8 frames. The conversion formula for converting the visual compensation amount Decorr into the number of stepper motor compensation pulses ΔP is: ΔP=Dcorr×Rpixel2step, where Rpixel2step is the conversion coefficient of the number of motor steps corresponding to a single pixel. Rpixel2step is determined by the current focal length, the image sensor pixel size, and the deceleration ratio.

8. An open-loop gimbal preset position calibration system based on multimodal feedback, characterized in that, The system includes a camera module, a stepper motor-driven gimbal, and an embedded controller. The controller performs gimbal preset position calibration through the following functional modules: The motion pre-compensation module is used to reshape the target velocity sequence into an S-curve acceleration and deceleration pulse frequency sequence before the gimbal performs the target preset position motion, and inserts a compensation pulse Δp according to the return error compensation table before each reversal command to suppress the accumulation of stepper motor step loss error and mechanical return error. The dynamic ROI discovery module is used to read the high-frequency component statistics output by the image signal processor of the camera module when the gimbal is stationary, perform texture scoring and dynamic filtering on multiple sub-regions of the image, and select N stable background sub-regions as regions of interest (ROIs) and extract their historical reference fingerprints. The visual consistency verification module is used to calculate the displacement vector between N ROIs and historical reference fingerprints after the gimbal reaches the target preset position, and after filtering out abnormal ROIs based on the random sample consistency algorithm, calculate the visual compensation amount Drorr based on the weighted average of the remaining in-point displacements. The compensation confidence management module is used to convert the visual compensation amount Dcorr into the corresponding number of stepper motor compensation pulses ΔP and apply it to the gimbal drive, and calculate the normalized cross-correlation score Scorr between the current frame and the historical reference frame in real time. When Scorr is lower than the confidence threshold Sth, the visual compensation logic is suspended, and the historical reference fingerprint is cleared to trigger physical collision zeroing calibration when the preset degradation conditions are met. The calibration scheduling module is used to maintain the calibration urgency index Ical and trigger event-driven collision zeroing calibration based on the relationship between Ical and the threshold Ith, including paths such as limit-based automatic zeroing, idle silent zeroing, and timeout forced zeroing.

9. The system according to claim 8, characterized in that, The calibration scheduling module includes a timeout forced zeroing subunit, which is used to force collision zeroing when the cumulative duration of Ical exceeding the threshold exceeds the forced zeroing time limit Tforce, regardless of whether the screen is in an idle state. The motion pre-compensation module stores a backhaul error compensation table indexed by the gimbal deceleration ratio. The visual consistency verification module is also used to maintain historical ROI displacement vector records within the sliding time window. When the gimbal returns to the same target preset position multiple times in a row, the mean of the historical compensation vectors is used as the initial prior for RANSAC model estimation.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the open-loop gimbal preset position calibration method based on multimodal feedback as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Pan tilt, pan tilt offset compensation correction method, computer storage medium and equipment

    CN111246094A

  • Video preset point calibration method

    CN115941930A

  • Method for stepper motor position referencing

    US7239108B2