A target monitoring method based on wide-angle and long-focus camera cooperative control and related equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-11
AI Technical Summary
[0006]本发明实施例提供的一种基于广角与长焦摄像机协同控制的目标监测方法及相关设备,至少解决相关技术中目标监测方法存在的设备控制精度低、目标监测效率低、目标监测效果差的问题
Smart Images

Figure CN122554725A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of target monitoring, and in particular to a target monitoring method and related equipment based on the coordinated control of wide-angle and telephoto cameras. Background Technology
[0002] In industrial sites such as grain depots, ports, industrial park control centers, chemical plants, and large factories, video surveillance systems typically employ a collaborative approach between wide-angle cameras and PTZ (Pan-Tilt-Zoom, a combination of horizontal and vertical rotation of the pan-tilt unit and lens zoom) telephoto cameras to monitor targets. Wide-angle cameras are used for overall panoramic coverage, encompassing a large area; PTZ cameras are used to magnify and capture detailed images of specific targets. In practice, operators typically locate a target in the wide-angle view and then manually click to activate the PTZ camera or use a simple automatic turning mechanism to steer it towards the target area.
[0003] Related technologies also provide some improvement schemes, such as: a linkage scheme that directly converts the PTZ rotation angle based on the wide-angle click position; a visualization scheme that estimates the theoretical field of view position in the wide-angle base map based on the PTZ attitude return; a control scheme for target detection, automatic tracking or centering fine adjustment in telephoto images; and a synchronization scheme for general timestamp management of video frames, AI metadata and attitude return.
[0004] However, the relevant technologies still suffer from problems such as low initial control accuracy of the camera, the need for repeated fine-tuning, and increased invalid PTZ operation, which in turn lead to low equipment control accuracy, low target monitoring efficiency, and poor target monitoring effect.
[0005] There is currently no effective solution to the aforementioned problems in the relevant technologies. Summary of the Invention
[0006] The present invention provides a target monitoring method and related equipment based on the coordinated control of wide-angle and telephoto cameras, which at least solves the problems of low equipment control accuracy, low target monitoring efficiency and poor target monitoring effect in target monitoring methods in related technologies.
[0007] To address the aforementioned problems, one aspect of this invention provides a target monitoring method based on the coordinated control of wide-angle and telephoto cameras, comprising: Obtain the motion trajectory of the target in the video stream of the wide-angle camera, and obtain the current camera pose of the PTZ telephoto camera; Based on the estimated total delay from the current control decision moment to the effective mechanical action of the PTZ telephoto camera, the target moment for the PTZ telephoto camera control to take effect is determined. Using the target moment as a reference, the predicted state of the target at the target moment is predicted from the motion trajectory. Based on the predicted state and combined with the pre-compensation amount corresponding to the current preset attitude interval into which the current camera attitude falls, the initial control amount of the PTZ telephoto camera is solved, and the PTZ telephoto camera is driven to perform cooperative control so that the telephoto camera can capture and magnify the target. After the PTZ telephoto camera performs cooperative control, the theoretical field of view center is determined based on the attitude feedback information of the PTZ telephoto camera, and the observation deviation is determined based on the actual observation position of the target in the telephoto camera image. Based on the theoretical field of view center and the observation deviation, the control residual under the current preset attitude range is constructed. The control residual is used to update the corresponding pre-compensation amount so that in the subsequent cooperative control process, the initial control amount is corrected based on the updated pre-compensation amount to achieve correction of the observation deviation. The target image captured by the telephoto camera, after being magnified and captured through collaborative control and corrected for the observation deviation, is output as the target monitoring result.
[0008] In some of these embodiments, the total latency estimate includes video decoding latency, AI processing latency, transmission latency, command sending latency, and PTZ telephoto camera mechanical response latency, and each latency component is dynamically updated using a moving average method.
[0009] In some embodiments, the step of predicting the predicted state of the target at the target time from the motion trajectory includes: Calculate the velocity and / or acceleration of the target based on the position information of at least two historical moments in the motion trajectory; Based on the velocity and / or the acceleration, the state of the target is extrapolated from the detection time to the target time to obtain the predicted state.
[0010] In some embodiments, the step of solving the initial control quantity of the PTZ telephoto camera includes: Based on the preset calibration relationship, the predicted state is converted to the control space of the PTZ telephoto camera to obtain the theoretical control quantity; Determine the current preset attitude range into which the current camera attitude falls, obtain the pre-compensation amount corresponding to the current preset attitude range, and superimpose the pre-compensation amount onto the theoretical control amount to obtain the initial control amount.
[0011] In some of these embodiments, it also includes: A mapping table is constructed between preset attitude intervals and pre-compensation values. The mapping table discretizes the range of values for the horizontal angle, pitch angle and zoom magnification of the PTZ telephoto camera into multiple three-dimensional intervals, and each interval stores a pre-compensation value. The current preset attitude interval refers to the corresponding interval in the mapping table into which the current camera attitude falls. During the collaborative control process, when the real-time camera attitude enters a certain interval, the pre-compensation quantity stored in that interval is directly called to participate in the initial control quantity solution.
[0012] In some embodiments, updating the corresponding pre-compensation amount based on the control residual includes: An exponential moving average algorithm is used to weight and fuse the control residual with the original pre-compensation amount corresponding to the current preset attitude interval for updating; wherein... The weighted fusion update includes: adding the product of the first weight and the original pre-compensation amount, and the product of the second weight and the control residual, to obtain the updated pre-compensation amount, wherein the sum of the first weight and the second weight is 1.
[0013] In some embodiments, before updating the corresponding pre-compensation amount based on the control residual, the method further includes: The effectiveness of this collaborative control was verified, and the verification conditions included any one or more of the following: the target identifier in the telephoto camera's image is consistent with the target identifier in the wide-angle camera's video stream; the actual observation position of the target in the telephoto camera's image is within a preset search window based on the theoretical field of view center; the target category observed in the telephoto camera's image is consistent with the target category monitored in the wide-angle camera's video stream; the confidence level of the target observed in the telephoto camera's image is higher than a preset threshold; and the attitude feedback information of the PTZ telephoto camera remains stable over multiple consecutive cycles. If the verification passes, the pre-compensation amount is updated.
[0014] In some of these embodiments, it also includes: The update of the pre-compensation amount is frozen when at least one of the following conditions is met: The target is lost in the telephoto camera footage for more than a preset duration; the attitude feedback information changes abnormally; the pre-compensation amount of the current preset attitude range fluctuates within a preset time window and exceeds a set threshold; scene calibration drift is detected and recalibration has not yet been completed.
[0015] To address the aforementioned problems, one aspect of this invention provides a target monitoring system based on the coordinated control of wide-angle and telephoto cameras, comprising: The acquisition unit is used to acquire the motion trajectory of the target in the video stream of the wide-angle camera, and to acquire the current camera pose of the PTZ telephoto camera; The collaborative control unit is used to determine the target time when the PTZ telephoto camera control takes effect based on the estimated total delay from the current control decision time to the effective mechanical action of the PTZ telephoto camera; using the target time as a reference, it predicts the predicted state of the target at the target time from the motion trajectory; based on the predicted state and combined with the pre-compensation amount corresponding to the current preset attitude interval into which the current camera attitude falls, it solves the initial control amount of the PTZ telephoto camera, drives the PTZ telephoto camera to perform collaborative control, and enables the telephoto camera to capture and magnify the target. A control residual construction unit is used to determine the theoretical field of view center based on the attitude feedback information of the PTZ telephoto camera after the PTZ telephoto camera performs cooperative control, and to determine the observation deviation based on the actual observation position of the target in the telephoto camera image; based on the theoretical field of view center and the observation deviation, construct the control residual under the current preset attitude range; wherein, the control residual is used to update the corresponding pre-compensation amount, so that in the subsequent cooperative control process, the initial control amount is corrected based on the updated pre-compensation amount, so as to achieve the correction of the observation deviation; The output unit is used to output the target image captured by the telephoto camera after being magnified and corrected for the observation deviation, as the target monitoring result.
[0016] To address the aforementioned problems, one aspect of this invention provides a non-transitory machine-readable medium storing computer instructions for causing a computer to execute any of the target monitoring methods based on the coordinated control of wide-angle and telephoto cameras.
[0017] The beneficial effects of this invention are as follows: By acquiring the motion trajectory of the target in the video stream of a wide-angle camera and the current camera attitude of a PTZ telephoto camera; based on the estimated total delay from the current control decision time to the effective mechanical action of the PTZ telephoto camera, the target time for the PTZ telephoto camera control is determined; using the target time as a reference, the predicted state of the target at the target time is predicted from the motion trajectory; based on the predicted state and combined with the pre-compensation amount corresponding to the current preset attitude interval into which the current camera attitude falls, the initial control amount of the PTZ telephoto camera is solved, driving the PTZ telephoto camera to perform cooperative control, enabling the telephoto camera to capture and magnify the target; after the PTZ telephoto camera performs cooperative control, the theoretical field of view center is determined based on the attitude feedback information of the PTZ telephoto camera, and the observation deviation is determined based on the actual observation position of the target in the telephoto camera image; based on the theoretical field of view center and the observation deviation, the control residual under the current preset attitude interval is constructed; wherein, the control residual is used to update the corresponding pre-compensation amount, so that in the subsequent cooperative control process, the updated pre-compensation amount is used for correction. Initial control inputs are used to correct for observation biases. A target image, magnified and corrected for observation biases from a telephoto camera, is used as the target monitoring result output. Target time estimation and forward prediction of target state ensure alignment between the control input and the actual PTZ activation time, reducing initial landing point deviation at the source. Dual-domain feedback (theoretical field of view center versus actual observation deviation) constructs a control residual bound to the attitude interval, quantifying and making the error traceable. Simultaneously, by using the control residual to update the pre-compensation input and directly calling the updated pre-compensation input in subsequent collaborative control, a closed-loop technical chain is formed: "first select the state according to the control activation time, then define the residual according to dual-domain feedback, and finally feed back pre-compensation according to the attitude interval." This significantly reduces the number of repeated fine-tunings, lowers PTZ jitter and invalid searches, and improves the initial control accuracy under the same scene and similar attitudes. It achieves continuous improvement in the initial accuracy of subsequent collaborative control, thereby improving equipment control accuracy, target monitoring efficiency, and enhancing the technical effect of improving poor target monitoring performance.
[0018] Details of one or more embodiments of the present invention are set forth in the following drawings and description, so that other features, objects and advantages of the invention will be more readily understood. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of the main process of a target monitoring method based on the coordinated control of wide-angle and telephoto cameras, according to one embodiment of the present invention. Figure 2 This is a schematic diagram of the main process of a target monitoring method based on the coordinated control of wide-angle and telephoto cameras, which is another embodiment of the present invention. Figure 3 This is a schematic diagram of the main framework of a target monitoring system based on the coordinated control of wide-angle and telephoto cameras, according to one embodiment of the present invention. Figure 4 This is a schematic diagram of the electronic device of the present invention. Detailed Implementation
[0021] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.
[0022] To more clearly compare the improvements of this application with related technologies, the shortcomings of related technologies are further analyzed as follows: First, the solutions in related technologies usually directly use the target state at the detection time to generate control quantities. However, since the PTZ control command needs to go through multiple stages such as video decoding, transmission, command sending, and mechanical response before the actual action is generated, the control takes effect later than the target detection time. Therefore, the PTZ control scheme adopted by related technologies will cause the target to deviate from the expected position when the PTZ first turns, resulting in a systematic first landing point deviation and low initial control accuracy. Second, the PTZ attitude feedback information can only indicate where the PTZ has theoretically turned, and cannot directly reflect whether the target in the telephoto view has actually landed in the ideal center position. The solutions in related technologies usually use attitude feedback only for interface visualization and telephoto observation deviation only for temporary fine-tuning of the current frame. The two lack a mechanism for joint definition and structured utilization. The deviation generated after each linkage is often used as a one-time correction quantity and cannot be accumulated into reusable empirical data. Third, even if the deviation in the telephoto view is used for this fine-tuning, the lack of a mechanism to bind the deviation to a specific PTZ attitude range means that the deviation cannot be automatically reused in subsequent similar scenarios. When the linkage under the same or similar attitude occurs again, the system still needs to start from scratch and undergo the initial landing point deviation and repeated fine-tuning, resulting in a low first-time landing rate, frequent invalid PTZ actions, and obvious image jitter in the same scenario. Fourth, related technical solutions usually treat time synchronization, PTZ pointing solution, attitude feedback, and telephoto observation correction as independent sub-functions, lacking a unified closed-loop technical link built around "how to improve the first-time landing accuracy of the next linkage". The core of the problem is not whether there are several separate modules, but that related technologies have not uniformly solved the following: based on which moment's target state should the control quantity be generated, how should the current linkage error be structurally defined, and how should this error be applied to the control solution under the next similar attitude in a reusable manner.
[0023] Accordingly, in view of the above-mentioned problems existing in related technologies, this application provides a target monitoring method based on the coordinated control of wide-angle and telephoto cameras, aiming to solve at least the following technical problems: (1) How to make the target state on which PTZ control is based correspond to the actual time when the control takes effect, rather than to the time when the detection has expired, so as to reduce the systematic first landing point deviation caused by link delay.
[0024] (2) How to simultaneously utilize the theoretical field of view information formed by PTZ attitude backhaul and the actual target deviation information of telephoto images to construct semantically clear and reusable control residuals, so that the source of residuals is more stable and the physical meaning is clearer.
[0025] (3) How to recursively update the control residual formed by this linkage according to the PTZ attitude range, and use it as a pre-compensation quantity to directly participate in the initial control quantity solution in subsequent linkages, so as to structurally influence the "current linkage result" on the "initial control of the next linkage" and improve the first-time arrival rate under the same scenario and similar attitude.
[0026] (4) How to use the above coupling mechanism to precipitate the results of this linkage into sustainable evolution compensation experience, reduce the number of repeated fine-tuning, reduce PTZ jitter and invalid search, and achieve continuous improvement of the accuracy of the first arrival of subsequent linkages.
[0027] To address the aforementioned problems, embodiments of the present invention provide a target monitoring method based on the coordinated control of wide-angle and telephoto cameras, such as... Figure 1 As shown, this target monitoring method based on the coordinated control of wide-angle and telephoto cameras mainly includes: Step S101: Obtain the motion trajectory of the target in the video stream of the wide-angle camera, and obtain the current camera pose of the PTZ telephoto camera; Step S102: Based on the estimated total delay from the current control decision time to the effective mechanical action of the PTZ telephoto camera, determine the target time when the PTZ telephoto camera control takes effect. Using the target time as a reference, predict the target's state at the target time from the motion trajectory. Based on the predicted state and combined with the pre-compensation amount corresponding to the current preset attitude interval into which the current camera attitude falls, solve the initial control amount of the PTZ telephoto camera, drive the PTZ telephoto camera to perform cooperative control, and enable the telephoto camera to capture and magnify the target. Step S103: After the PTZ telephoto camera performs cooperative control, the theoretical field of view center is determined based on the attitude feedback information of the PTZ telephoto camera, and the observation deviation is determined based on the actual observation position of the target in the telephoto camera image. Based on the theoretical field of view center and the observation deviation, the control residual under the current preset attitude range is constructed. The control residual is used to update the corresponding pre-compensation amount so that in the subsequent cooperative control process, the initial control amount is corrected based on the updated pre-compensation amount to achieve the correction of the observation deviation. Step S104: The target image captured by the telephoto camera after being magnified and photographed through collaborative control and corrected for observation deviations is output as the target monitoring result.
[0028] Based on the above settings, by estimating the target time and predicting the target state forward, the alignment of the control input with the actual PTZ activation time is ensured, reducing the initial landing point deviation from the source. Control residuals bound to the attitude interval are constructed through dual-domain feedback (the deviation between the theoretical field of view center and the actual observation), making the error quantifiable and traceable. Simultaneously, by using the control residuals to update the pre-compensation amount and directly calling the updated pre-compensation amount in subsequent collaborative control, a closed-loop technical link is formed: "first select the state according to the control activation time, then define the residuals according to dual-domain feedback, and finally reinject pre-compensation according to the attitude interval." This reduces the number of repeated fine-tunings, lowers PTZ jitter and invalid searches, and improves the initial control landing rate under the same scenario and similar attitudes, achieving continuous improvement in the initial landing accuracy of subsequent collaborative control, thus achieving an improvement. It can also be understood that the above closed-loop technical link allows the experience generated by each collaborative control (updating the corresponding pre-compensation amount using the control residuals determined after collaborative control) to be accumulated and reused. The initial landing accuracy under the same or similar preset attitude intervals continuously improves with the increase in the number of collaborative control operations, thereby improving the overall monitoring efficiency.
[0029] The wide-angle camera continuously records the target's position changes within the wide-angle frame. Target detection and tracking yield the target's trajectory, which includes information such as position, velocity, and acceleration, reflecting the target's positional change patterns and forming the basis for predicting future states. Simultaneously, the current camera attitude (horizontal angle, pitch angle, zoom level) of the PTZ telephoto camera directly calibrates its real-time pointing in space, a necessary condition for subsequent control calculations and the construction of the theoretical field of view center. Both constitute the necessary inputs for subsequent state prediction and control calculations. Based on step S101, the target's trajectory in the wide-angle camera's video stream and the current camera attitude of the PTZ telephoto camera are obtained, providing spatiotemporal reference data for subsequent prediction and control. This enables the system to describe the target's historical motion patterns and the PTZ telephoto camera's current pointing.
[0030] Related technologies directly use the target state at the detection moment to generate the control quantity. However, due to multiple delays between the control decision and the PTZ mechanical action, such as video decoding, AI processing, transmission, command sending, and mechanical response, the target has already moved by the time the control actually takes effect, leading to initial landing point deviation. This invention, based on step S102, calculates the target moment when the control actually takes effect using the total delay estimate and performs forward prediction of the target state at that moment based on the motion trajectory. This ensures the control quantity targets the target's future true position. By avoiding the use of outdated detection moment states, it fundamentally avoids system deviations caused by control delays. Furthermore, it introduces a pre-compensation quantity corresponding to the preset attitude interval into which the current camera attitude falls, correcting the initial control quantity (this pre-compensation quantity is superimposed on the theoretical control quantity to obtain the initial control quantity). The pre-compensation quantity is error experience accumulated in historical linkages and bound to a specific attitude interval. By correcting the initial control quantity with it, the control result is closer to the ideal position. Therefore, the obtained control quantity considers both the target's future position and incorporates the error experience accumulated in historical linkages, thereby improving the accuracy of the initial control point.
[0031] In some embodiments, based on the above step S103, the stable source of the control residual is clarified, making its semantics clearer. At the same time, the recursive update of the pre-compensation amount is realized, and the error determined in this linkage process is structurally stored for subsequent reuse. Specifically, the control residual is constructed by combining the theoretical field of view center formed by the attitude feedback information (reflecting the direction that PTZ should theoretically see) with the observation deviation determined based on the actual observation position of the target in the telephoto image (reflecting the actual landing point of the target), instead of using one type of data alone. This allows the control residual to reflect the difference between the theoretical pointing and the actual landing point. The theoretical field of view center is calculated by combining the PTZ attitude feedback with the current zoom magnification, and has the characteristics of high frequency and low latency. The actual observation deviation is provided by the target detection result in the telephoto image, reflecting the actual imaging landing point deviation. The control residual constructed by combining the two includes the contribution of the gimbal pointing error, as well as the comprehensive influence of the target prediction deviation and calibration residual error, and has a clear physical semantics. Meanwhile, by using the control residual to update the pre-compensation amount, the pre-compensation amount can be dynamically evolved with each coordinated control, realizing the transformation from "one-time use" to "continuous learning". Furthermore, the error generated by the current coordinated control can be structurally stored in the corresponding preset attitude range, providing reusable empirical data for control correction in subsequent linkages.
[0032] In some examples, based on the above step S104, a clear and accurate target monitoring image can be output, completing the target monitoring task. Specifically, after the prediction state and pre-compensation correction in step S102, the PTZ telephoto camera's first turn has placed the target roughly near the center of the telephoto frame; then, after the observation deviation acquisition and control residual construction in step S103, the system not only obtains the error information of this collaborative control to update the pre-compensation, but also implicitly completes the correction of the current image (through subsequent fine-tuning or directly using the updated pre-compensation), so that the target image in the telephoto frame is in the ideal observation position and is magnified for shooting. The output image at this time is the result that meets the monitoring requirements, effectively demonstrating the application value of this application.
[0033] In some of these embodiments, the total latency estimate includes video decoding latency, AI processing latency, transmission latency, command sending latency, and PTZ telephoto camera mechanical response latency, and each latency component is dynamically updated using a moving average method.
[0034] Based on the above settings, the delay component included in the total delay estimate is provided, and the accurate solution and adaptive adjustment of the total delay estimate are realized, providing reliable parameters for the accurate positioning of the control take-off time, thereby reducing the first landing point deviation caused by delay fluctuations.
[0035] In this embodiment of the invention, the control decision moment refers to the starting moment when the system begins to detect, track, and analyze the motion trajectory of the target in the current frame of the wide-angle camera video stream. Accordingly, this control decision moment is the starting point for the total delay estimation; the target moment when the control takes effect is equal to the sum of the current control decision moment and the total delay estimation value, that is, the moment when the mechanical action of the PTZ telephoto camera actually takes effect.
[0036] Specifically, from the moment of control decision-making to the moment the mechanical action of the PTZ telephoto camera takes effect, multiple sequential processing stages occur. Video decoding latency refers to the time required for a wide-angle video frame to be acquired and processed by AI; AI processing latency refers to the time required for algorithms such as target detection and trajectory extraction; transmission latency refers to the time required for control commands to be transmitted from the decision-making end to the PTZ telephoto camera; command sending latency refers to the time required for protocol encapsulation and queue waiting; and the PTZ telephoto camera mechanical response latency refers to the time required for the gimbal motor to actually complete its rotation after receiving the command. Each of these latency components is accumulated in the total latency estimate. If any latency component is omitted, the estimated target moment will be underestimated, causing the predicted state to occur earlier than the actual control activation moment, resulting in prediction bias. By limiting the total latency estimate, the true delay of control activation can be more accurately reflected, providing a reliable numerical basis for the accurate calculation of the target moment.
[0037] Meanwhile, in actual industrial settings, the various delay components are not fixed constants. For example, video decoding latency may vary due to encoding format and bitrate fluctuations; AI processing latency may increase due to increased computational load caused by a larger number of targets; network transmission latency may fluctuate due to network congestion; and mechanical response latency may drift slowly with equipment aging and temperature changes. If a fixed value is used as the latency estimate, when the actual latency increases, the predicted state will occur earlier than the actual control activation time, resulting in a delayed first landing point; conversely, when the actual latency decreases, the predicted state will occur later than the actual control activation time, resulting in an advanced first landing point. Based on the above settings, a dynamic update mechanism using a moving average is adopted. By weighted averaging the latency observations from historical moments, instantaneous noise can be filtered out while tracking the slow trend of latency changes. This ensures that the estimated value is neither too sensitive nor too lagging, guaranteeing the timeliness of the total latency estimate in the time dimension, enabling it to respond in real time to system operating status and environmental changes.
[0038] It is also understandable that the structural integrity of the total delay estimate and the dynamic adaptability of each delay component work together to ensure that the total delay estimate can maintain high accuracy regardless of the system's load conditions, network conditions, or equipment aging stage. This makes the determination of the control activation target time more reliable and ultimately reduces the first landing point deviation.
[0039] In some of these embodiments, the step of predicting the target's state at the target time from the motion trajectory includes: calculating the target's velocity and / or acceleration based on the position information of at least two historical times in the motion trajectory; and extrapolating the target's state from the detection time to the target time based on the velocity and / or acceleration to obtain the predicted state.
[0040] Based on the above settings, a specific implementation method for predicting the state of a target is provided. On the one hand, velocity and / or acceleration are calculated based on at least two historical moments (preferably several recent moments to improve accuracy), allowing the acquisition of motion parameters to directly originate from the target's actual motion trajectory, without relying on external sensors or prior assumptions, thus exhibiting good scene adaptability. On the other hand, forward extrapolation is performed based on velocity and / or acceleration, using classical kinematic formulas to extend the target state from the detection moment to the target moment. This method has low computational load, high real-time performance, and can flexibly select between velocity (first-order model) or both velocity and acceleration (second-order model) depending on the actual motion mode. This allows the technology to handle both conventional uniform motion scenarios and variable-speed motion scenarios, thereby maintaining high prediction accuracy under various industrial field motion modes and providing a reliable input state for reducing the initial landing point deviation.
[0041] In actual industrial settings, the trajectory of a target is often not a simple straight line; it may involve turning, curvilinear motion, or other changes in direction. Forward extrapolation based solely on the scalar form of velocity and / or acceleration, assuming the target moves in a constant direction, will result in significant prediction errors when the target changes direction.
[0042] To address various motion modes such as constant speed, variable speed, and change of direction, according to a specific embodiment of the present invention, a target state prediction method based on motion vector decomposition and curve fitting is also provided. This method includes: acquiring the target's motion trajectory from a wide-angle camera video stream, the trajectory recording the target's position coordinates across multiple consecutive historical frames. The system selects the position information of several recent historical moments from the trajectory. When two historical moments are selected, the ratio of the target's displacement difference to the time difference between the two frames is calculated to obtain the target's velocity vector, which simultaneously contains information about the speed and direction of movement. When three or more historical moments are selected, the rate of change of velocity can be further calculated to obtain the target's acceleration vector, which reflects the trend of changes in the target's velocity magnitude and direction. After obtaining the velocity vector and / or acceleration vector, the system uses the latest detected target moment as the starting point and calculates the target's position forward according to the time interval between the detection moment and the target moment when control takes effect. If the target is moving at approximately uniform linear speed, only the velocity vector is used; the position at the detection moment is moved along the velocity vector direction by the product of the velocity and the time interval to obtain the predicted state. If the target exhibits acceleration, deceleration, or turning tendencies, both velocity and acceleration vectors are used simultaneously. Second-order corrections are introduced based on velocity changes, making the predicted position more closely match the target's actual motion curve. For cases involving changes in motion direction, since both velocity and acceleration are expressed as vectors, their directions are naturally updated as the historical trajectory changes. Therefore, the forward-calculated predicted state can automatically adapt to scenarios such as target turning or other directional changes.
[0043] Based on the above specific implementation method, the target prediction state can be aligned with the actual target moment controlled by the PTZ telephoto camera, while adapting to complex motion scenarios such as changes in target speed and direction, thereby further reducing the landing point deviation during the first linkage and improving the first landing accuracy of targets with changing speed and direction.
[0044] In some embodiments, the steps of solving the initial control quantity of the PTZ telephoto camera include: converting the predicted state to the control space of the PTZ telephoto camera according to a preset calibration relationship to obtain the theoretical control quantity; determining the current preset attitude interval into which the current camera attitude falls, obtaining the pre-compensation quantity corresponding to the current preset attitude interval, and superimposing the pre-compensation quantity onto the theoretical control quantity to obtain the initial control quantity.
[0045] Based on the above settings, a specific implementation scheme for solving the initial control variables of a PTZ telephoto camera is provided. By constructing a control variable solution mechanism that combines theoretical geometric mapping with attitude-related historical experience, the initial linkage of the PTZ telephoto camera can simultaneously possess accurate geometric pointing and pre-correction capability for system errors, thereby improving the accuracy of the initial landing point.
[0046] Specifically, the preset calibration relationship provides a precise geometric transformation from the predicted state to the theoretical control quantity, ensuring the spatial orientation accuracy of the control quantity. By determining the preset attitude range into which the current camera attitude falls and obtaining the corresponding pre-compensation quantity, error experience highly correlated with the attitude and learned through historical linkage is introduced, enabling targeted correction of the theoretical control quantity. Furthermore, the superposition operation organically integrates the theoretical control quantity and the pre-compensation quantity, so that the final issued initial control quantity simultaneously includes the ideal geometric orientation and the actual error compensation, and the compensation quantity can dynamically change with different attitude ranges. It can be understood that the combination of theoretical mapping and empirical compensation makes the initial control quantity neither a purely theoretical value nor a fixed empirical value, but rather an optimized result of historical error characteristics under the current specific attitude, thereby significantly improving the positioning accuracy of the first turn without increasing the number of subsequent fine-tuning steps.
[0047] According to embodiments of the present invention, a wide-angle camera and a PTZ telephoto camera are mounted on the same device or have a fixed relative positional relationship, and a defined projection geometry model exists between them. The preset calibration relationship includes the intrinsic parameters of the wide-angle camera (such as focal length and distortion parameters), the extrinsic parameters of the wide-angle camera relative to the device coordinate system, and the field-of-view parameters of the PTZ telephoto camera at different zoom magnifications. Through the calibration relationship, the coordinates of the target anchor point in the wide-angle image can be converted into a target direction vector in the device coordinate system. Then, the horizontal and vertical angles required for the PTZ telephoto camera to point in that target direction can be further calculated, i.e., the theoretical control quantities. These theoretical control quantities represent the rotation angle required to align the PTZ telephoto camera with the predicted position under ideal conditions, and are the basis for all subsequent corrections.
[0048] In some embodiments, due to the combined effects of factors such as calibration residual error, mechanical assembly error, servo control error, and target prediction error, the system error under different attitudes exhibits different characteristics. For example, at positions with larger horizontal or pitch angles, the backlash error of mechanical transmission or optical axis misalignment may be more pronounced. Based on the above embodiments, the range of values for the horizontal angle, pitch angle, and zoom magnification of the PTZ telephoto camera is pre-discretized into multiple three-dimensional intervals, each interval corresponding to a pre-compensation amount. This pre-compensation amount is obtained by recursively updating the control residual constructed in historical linkages. By determining the interval in which the current camera attitude falls, the system can quickly find the pre-compensation amount associated with that attitude. This pre-compensation amount reflects the average error experience accumulated in historical linkages within that attitude interval, rather than a fixed global correction value.
[0049] In some cases, the theoretical control values calculated using simple geometric mapping do not consider practical factors such as calibration residual errors and mechanical backlash errors. Directly sending these control values can lead to a systematic deviation between the actual pointing of the PTZ telephoto camera and the true position of the target. By additively superimposing the pre-compensation values onto the theoretical control values—that is, the initial control value's horizontal angle equals the sum of the theoretical horizontal angle and the horizontal component of the pre-compensation value, and the initial control value's pitch angle equals the sum of the theoretical pitch angle and the pitch component of the pre-compensation value—this superposition method allows historical experience to directly correct the current control command without waiting for a deviation to appear in the telephoto image before performing secondary fine-tuning. Since the pre-compensation value is the result of statistical learning of historical residuals within the same attitude range, the initial control value obtained after superposition is statistically closer to the true ideal value that centers the target in the telephoto image, thereby improving the probability of first-time positioning.
[0050] In some embodiments, the method further includes: constructing a mapping table between preset attitude intervals and pre-compensation quantities, wherein the mapping table discretizes the range of values for the horizontal angle, pitch angle, and zoom magnification of the PTZ telephoto camera into multiple three-dimensional intervals, and each interval stores a pre-compensation quantity; wherein, the current preset attitude interval refers to the corresponding interval in the mapping table into which the current camera attitude falls; during the collaborative control process, when the real-time camera attitude enters a certain interval, the pre-compensation quantity stored in that interval is directly called to participate in the initial control quantity solution.
[0051] Based on the above settings, the continuous PTZ telephoto camera attitude space is associated with recursively updated pre-compensation quantities in the form of a discretized mapping table, realizing compact storage, fast indexing and instant retrieval of error experience, so that historical learning results can efficiently serve the current control.
[0052] In the PTZ telephoto camera, the horizontal angle, pitch angle, and zoom magnification are all continuous variables with a theoretically infinite value space, making it impossible to store pre-compensation values individually for every possible attitude value. By dividing the value ranges of the horizontal angle, pitch angle, and zoom magnification into several discrete intervals, these three-dimensional intervals are combined into multiple three-dimensional intervals. For example, the horizontal angle is divided into intervals of 10°, the pitch angle into intervals of 5°, and the zoom magnification into several levels based on magnification. The Cartesian product of these three forms a finite number of three-dimensional intervals. Each interval stores a corresponding pre-compensation value, which represents the statistical characteristics of the historical control residuals of all attitudes within that interval. Through this discretized mapping table (i.e., forming a structured, quickly accessible error experience base), the infinite continuous attitude space is compressed into finite discrete intervals, making the storage and retrieval of pre-compensation values feasible.
[0053] According to an embodiment of the present invention, the PTZ telephoto camera has a current camera attitude (horizontal angle, pitch angle, zoom magnification) before performing cooperative control. By comparing the three dimensions of the current camera attitude with the interval boundaries of each dimension, the three-dimensional interval into which the attitude falls can be uniquely determined. For example, if the current horizontal angle is 35°, it falls into the 30-40° horizontal interval; if the current pitch angle is 12°, it falls into the 10-15° pitch interval; and if the current zoom magnification is 8x, it falls into the 5-10° zoom interval. These combinations result in a unique three-dimensional interval identifier. This mapping process involves only numerical comparison, has extremely low computational complexity, and can complete the fast mapping from attitude to interval within milliseconds without affecting real-time control performance.
[0054] In practical applications, since the pre-compensation value stored in each interval of the mapping table is a stable value obtained through recursive updates of control residuals from previous linkages, it represents the statistical characteristics of the system error within that attitude interval. Therefore, when the real-time camera attitude falls into a certain interval, the system does not need to recalculate the error characteristics under that attitude. It only needs to directly read the corresponding pre-compensation value through the interval index and then superimpose it onto the theoretical control value. The time complexity of this table lookup method is O(1), which is independent of the size of the mapping table and can meet the real-time requirements of industrial sites. At the same time, since the pre-compensation value has already incorporated the average error of historical linkages within that interval, direct calling can enable the initial control value to carry system error correction that matches the attitude interval, thereby improving the accuracy of the first positioning.
[0055] In some embodiments, updating the corresponding pre-compensation amount based on the control residual includes: using an exponential moving average algorithm to perform weighted fusion update of the control residual and the original pre-compensation amount corresponding to the current preset attitude interval; wherein, the weighted fusion update includes: adding the product of the first weight and the original pre-compensation amount and the product of the second weight and the control residual to obtain the updated pre-compensation amount, and the sum of the first weight and the second weight is 1.
[0056] Based on the above settings, a smooth recursive update of the pre-compensation amount is achieved, which integrates the historical learning results with the current observation residuals in a controllable proportion. The pre-compensation amount gradually converges to the optimal estimate of the system error under the preset attitude range as the number of cooperative control cycles increases, thereby continuously improving the accuracy of the first positioning.
[0057] Specifically, the exponential moving average algorithm inherently possesses a memory-decreasing characteristic (i.e., the older the historical residual, the smaller its impact on the current pre-compensation amount; the more recent the residual, the greater its impact). This allows the pre-compensation amount to adaptively track the slow changes in system error without requiring periodic manual calibration. The weighted summation form makes the weight coefficients a single control parameter for adjusting the update rate, simplifying engineering implementation and facilitating the determination of optimal values based on field debugging experience. Simultaneously, the combined characteristic of the first and second weights summing to 1 ensures that the update process of the pre-compensation amount always converges stably, without oscillations or divergence. By performing a weighted fusion update after each linkage between the wide-angle camera and the PTZ telephoto camera, the pre-compensation amount for the current preset attitude range continuously approaches the ideal compensation value that minimizes the control residual within that preset attitude range. When linkage between the same or adjacent attitude ranges occurs again, this pre-compensation amount is invoked to participate in the initial control quantity solution, making the initial control quantity closer to the ideal value, thereby reducing the number of repeated fine-tuning steps and improving the first-time accuracy rate.
[0058] In some embodiments, before updating the corresponding pre-compensation amount based on the control residual, the method further includes: validating the effectiveness of the current cooperative control, wherein the validation conditions include any one or more of the following: the target identifier in the telephoto camera image is consistent with the target identifier in the wide-angle camera video stream; the actual observation position of the target in the telephoto camera image is within a preset search window based on the theoretical field of view center; the target category observed in the telephoto camera image is consistent with the target category monitored in the wide-angle camera video stream; the target confidence level observed in the telephoto camera image is higher than a preset threshold; the attitude feedback information of the PTZ telephoto camera remains stable over multiple consecutive cycles; and when the validation passes, the pre-compensation amount is updated.
[0059] Based on the above settings, a multi-dimensional validity verification mechanism was established to ensure that only reliable linkage results that meet the requirements of target consistency, spatial proximity, category matching, detection reliability, and attitude stability are used to update the pre-compensation quantity. This avoids erroneous or low-quality data from polluting the attitude interval compensation unit and ensures that the learning process of the pre-compensation quantity is healthy, stable, and effective.
[0060] The aforementioned compensation unit refers to a storage entry in the mapping table between preset attitude intervals and pre-compensation quantities that corresponds to a specific preset attitude interval. Each preset attitude interval is uniquely determined by the horizontal angle (pan) interval, pitch angle (tilt) interval, and zoom ratio (zoom) interval of the PTZ telephoto camera, forming a three-dimensional discrete interval. Each compensation unit stores at least the pre-compensation quantity corresponding to that interval, which includes horizontal and vertical compensation components. In addition, the compensation unit may optionally store the following information: the number of samples updated for that interval, the most recent update time, the stability score of the pre-compensation quantity, and the validity label, etc.
[0061] According to an embodiment of the present invention, when multiple targets exist simultaneously in the video stream of a wide-angle camera, the system needs to initiate linkage control monitoring for each target separately. Each target is assigned a unique target identifier on the wide-angle side. When the PTZ telephoto camera turns, multiple candidate targets may appear in the telephoto frame. By comparing whether the target identifier detected in the telephoto frame is consistent with the target identifier initiated by the wide-angle side, it can be confirmed that the currently observed target is the target to be monitored. If the identifiers are inconsistent, it indicates that the PTZ may have mistakenly turned to another target. In this case, the observation deviation and control residual calculated based on the misdetected target do not reflect the error of the actual linkage. If they are used to update the pre-compensation amount, it will contaminate the compensation unit. By ensuring that the target observed in the telephoto camera frame and the target initiated by the wide-angle camera are the same target, it is possible to effectively avoid writing incorrect residuals into the compensation unit due to target identifier mismatch.
[0062] In some embodiments, when the deviation between the theoretical field of view center and the actual observation position is too large (e.g., the target is completely off the edge of the screen), the numerical value of the observation deviation may be distorted, and the linear approximation model based on the small deviation assumption to calculate the angle residual may no longer be applicable. By using a preset search window (e.g., a circular or rectangular area with the theoretical field of view center as the center and a radius equal to a certain proportion of the screen width), the deviation of the current linkage is considered to be within an acceptable correction range only when the actual position of the target falls within this window. The control residual constructed at this time has high reliability. If the target exceeds the window, it indicates that the initial landing point deviation is too large, which may be due to prediction failure, calibration error, or other abnormal reasons. In this case, this abnormal deviation should not be used for learning to avoid the pre-compensation amount being incorrectly and significantly adjusted. By ensuring that the target actually appears in the vicinity of the theoretical field of view center, the observation deviation calculated when the target is severely off the screen is prevented from becoming meaningless.
[0063] In some cases, target detection in the wide-angle camera video stream typically outputs target categories (such as people, vehicles, etc.). When the PTZ telephoto camera turns, multiple objects may be detected in the telephoto view, including non-target category interference (such as birds, leaves, etc.). By comparing the target category observed by the telephoto camera with the target category monitored by the wide-angle camera, false associations caused by category mismatch can be ruled out. If the categories are inconsistent, it means that the object detected in the telephoto view may be another object, rather than the original target to be monitored. In this case, the calculated observation bias cannot represent the error of the actual linkage and should not be used to update the pre-compensation amount. That is, this can avoid misdetecting non-class targets (such as background interference) in the telephoto view as the target to be monitored, thereby preventing the compensation unit from being contaminated by observation bias based on incorrect targets.
[0064] In some cases, target detection algorithms in telephoto camera footage output a confidence score, indicating the reliability of the detection results. Low-confidence detections may be due to image blur, partial occlusion, or false detections by the algorithm, resulting in poor positional accuracy. A pre-set confidence threshold (e.g., 0.7 or 0.8) can filter out these unreliable detections. Only when the confidence score is higher than the threshold is the detected target location considered reliable; only then are the calculated observation bias and control residuals valuable and used to update the pre-compensation, thus ensuring the quality of the learning data source. By ensuring that the target detection results used to calculate the observation bias have sufficiently high reliability, noise introduced by low-confidence detections is avoided.
[0065] In some cases, the attitude feedback information from a PTZ telephoto camera may exhibit transient fluctuations (such as encoder noise or communication interference) or transitional states when the pan-tilt unit (PTZ) has not completely stopped rotating. If attitude feedback data is acquired under such unstable conditions, the calculated theoretical field-of-view center will be biased, leading to inaccurate control residuals. By requiring the attitude feedback information to change less than a set threshold over multiple consecutive periods (e.g., three consecutive frames), it is determined that the PTZ has stably pointed towards the target direction. At this point, the acquired attitude data is more reliable, and the control residuals calculated based on this data accurately reflect system errors, making them suitable for updating pre-compensation values. Ensuring that the attitude feedback data used to calculate the theoretical field-of-view center is in a stable state avoids acquiring inaccurate attitude information while the PTZ is still rotating or when the feedback data abruptly changes.
[0066] In some embodiments, the method further includes freezing the update of the pre-compensation amount when at least one of the following conditions is met: the target loss time in the telephoto camera image exceeds a preset duration; the attitude feedback information changes abnormally; the pre-compensation amount in the current preset attitude range fluctuates within a preset time window beyond a set threshold; or scene calibration drift is detected and recalibration has not yet been completed.
[0067] Based on the above settings, a multi-scenario coverage update freeze protection mechanism was constructed. When the system encounters abnormal states (target loss, attitude change, compensation oscillation, calibration drift), the learning process of the pre-compensation quantity is actively suspended to avoid erroneous or distorted control residuals from contaminating the attitude interval compensation unit. This ensures that the pre-compensation quantity stored in the compensation unit is always based on reliable and stable linkage events, maintaining the data quality of the attitude interval compensation table and the health of the learning process.
[0068] In this scenario, if the target cannot be detected in the PTZ telephoto camera's view after collaborative control (i.e., the target is lost), a valid actual observation position cannot be obtained, thus preventing the construction of reliable control residuals. If the target is only temporarily lost (e.g., temporarily obscured but quickly reappears), updates can continue after recovery. However, if the target is lost for more than a preset duration (e.g., not detected for several consecutive seconds), it indicates that the linkage may have failed (e.g., the PTZ turned in the wrong direction or the target has left the monitoring area). If non-existent observation data or incomplete data from the last moment is used for updates at this time, severely distorted control residuals will be generated, and writing them into the compensation unit will destroy the pre-compensation amount accumulated in history within that attitude range. By freezing the update until the target reappears, it is ensured that only reliable residuals generated by effective linkages can be used for learning.
[0069] In some embodiments, the theoretical field of view center is calculated based on the attitude feedback information from the PTZ telephoto camera. If the attitude feedback information experiences abnormal jumps (e.g., due to communication interference or encoder failure causing a sudden and drastic change in the horizontal or pitch angle that does not conform to physical laws), the calculated theoretical field of view center will be significantly inconsistent with the actual PTZ pointing, and the constructed control residual will not reflect the true deviation. If the pre-compensation amount is updated based on such abnormal data, it will introduce erroneous information into the compensation unit, causing the pre-compensation amount in that attitude range to deviate from the true optimal value, thereby reducing the accuracy of subsequent linkages. By detecting the continuity of attitude feedback (e.g., whether the rate of change of adjacent period feedback values exceeds the physically feasible range), once an abnormal jump is detected, the update is frozen, and learning continues only after the attitude feedback returns to normal, thus protecting the data quality of the pre-compensation amount.
[0070] In some cases, under normal circumstances, the pre-compensation amount for a preset attitude range should gradually approach the optimal compensation value for that range with repeated recursive updates, and its change should be smooth and convergent. If the pre-compensation amount for a certain attitude range fluctuates drastically within a short period (within a preset time window), and the amplitude of the change exceeds a set threshold, it indicates that the range may frequently encounter abnormal events (such as target loss, attitude jumps, external interference, etc.), or that there is an undetected fault in the system. In this case, continuing to update based on the current control residual may further deteriorate the pre-compensation amount. By detecting the fluctuation amplitude and freezing the update, the learning process for that range can be temporarily suspended until the system stabilizes, thereby avoiding continuous contamination of the compensation unit and ensuring the overall reliability of the attitude range compensation table.
[0071] In some embodiments, a preset calibration relationship (spatial mapping parameters between the wide-angle camera and the PTZ telephoto camera) forms the basis for solving the theoretical control quantity. When the equipment experiences mechanical displacement, vibration, or loosening, the calibration relationship may drift. Before recalibration is completed due to calibration drift, the theoretical control quantity calculated based on the old calibration relationship itself contains systematic deviations. The constructed control residual at this time includes not only normal servo errors and prediction errors but also additional deviations caused by calibration changes. If this control residual, mixed with calibration deviations, continues to update the pre-compensation quantity, the compensation unit will incorrectly learn the calibration drift as a systematic error, leading to distortion of the pre-compensation quantity. Once recalibration is completed, the previously learned pre-compensation quantity will become a new source of error. Therefore, when calibration drift is detected but recalibration has not yet occurred, the update is frozen, and learning resumes after recalibration is completed, ensuring the consistency of the pre-compensation quantity with the current calibration reference.
[0072] This invention also provides a target monitoring method based on the coordinated control of wide-angle and telephoto cameras, optionally, such as... Figure 2 As shown, the target monitoring method based on the coordinated control of wide-angle and telephoto cameras provided in this embodiment of the invention mainly includes: Step S201: Obtain wide-angle video frames, target metadata corresponding to the wide-angle video frames, PTZ attitude backhaul data, and telephoto video frames.
[0073] After system startup, the wide-angle camera continuously acquires wide-angle video frames F_w(t) at a fixed frame rate (e.g., 25 frames / second), and processes each frame using a target detection algorithm (such as the YOLOv series), outputting target metadata M(t). The target metadata M(t) includes at least the target identifier (ID), target anchor point (e.g., the pixel coordinates of the target center in the wide-angle frame or the bottom center of the target detection box), target category, and the timestamp of the current frame. Simultaneously, the system receives attitude feedback data P(t) from the PTZ telephoto camera in real time via a communication interface. This data P(t) includes the horizontal angle pan, pitch angle tilt, and current zoom level, as well as the corresponding timestamp. Furthermore, the telephoto camera also continuously acquires telephoto video frames for subsequent target observation and deviation calculation. All data is cached in system memory and associated according to timestamps for subsequent processing.
[0074] Step S202: Establish the calibration relationship between the wide-angle target anchor point and the PTZ control space.
[0075] During the system initialization phase, a spatial mapping model between the wide-angle camera and the PTZ telephoto camera is pre-established. The calibration relationship includes at least the following parameters: the intrinsic parameter matrix K_w of the wide-angle camera (focal length, principal point coordinates, distortion coefficients), the extrinsic parameters of the wide-angle camera relative to the device coordinate system (rotation matrix R_w and translation vector T_w), the null parameters of the PTZ telephoto camera (horizontal null parameter pan_θ, pitch null parameter tilt_θ), and the field of view parameters at different zoom magnifications.
[0076] Specifically, for any target anchor point in a wide-angle image First, the inverse of the wide-angle camera's intrinsic parameter matrix is used to normalize it into a direction vector in the camera coordinate system. Then, the extrinsic rotation matrix is used to transform it into a direction vector in the device coordinate system. * When the target anchor point is taken as the center of the target's bottom edge and the scene conforms to the ground plane constraint, the intersection of the ray and the ground plane can be used to obtain the target's position in three-dimensional space; otherwise, r_b is directly used as the target direction. Based on this, the theoretical horizontal and vertical angles required for the PTZ telephoto camera to point in this direction are calculated according to the target direction vector, serving as the geometric basis for subsequent theoretical control quantities.
[0077] Step S203: Estimate the PTZ control activation time based on the delay of each link.
[0078] The PTZ telephoto camera goes through multiple stages from receiving control commands to the actual mechanical action taking effect. In this embodiment, the delay components corresponding to these stages are collectively referred to as the total delay estimate. Specifically, these include: video decoding delay (the time from acquisition of a wide-angle video frame to when it can be processed by AI), AI processing delay (the computation time of algorithms such as target detection and trajectory extraction), transmission delay (the time for control commands to be transmitted from the decision-making end to the PTZ device), command sending delay (protocol encapsulation and queue waiting time), and PTZ telephoto camera mechanical response delay (the time from receiving the command to completing the rotation of the gimbal motor).
[0079] Each delay component is dynamically updated using a moving average method. After each linkage is completed, the system records the actual time consumed by each step in that linkage and merges the real-time observations with the historical moving averages to obtain updated estimates for each component. Then, the current estimates of all components are summed to obtain the total delay estimate from the current control decision moment to the effective PTZ mechanical action. Finally, the current control decision time t_now is compared with this total delay estimate. Adding _t together yields the target time t_eff when the PTZ telephoto camera control actually takes effect.
[0080] Step S204: Based on the control activation time, select or predict the target state at the corresponding time from the wide-angle target trajectory.
[0081] The system maintains a target trajectory cache, storing the target's position coordinates and corresponding timestamps at several recent historical moments (e.g., within the past second). To obtain the target's predicted state at time t_eff, the system first selects the position information from at least two historical moments in the trajectory cache and calculates the target's velocity and / or acceleration in the horizontal and vertical directions. For example, when three historical moments are selected (t_1, x_1, y_1), (t_2, x_2, y_2), and (t_3, x_3, y_3), the velocity components v_x and v_y in the x and y directions, and the acceleration components acc_x and acc_y in the y direction, can be obtained through quadratic curve fitting.
[0082] Then, a uniformly accelerated motion model is used to forward calculate the target state from the latest detection time t_det to t_eff.
[0083] Step S205: Solve for the initial PTZ control quantity based on the predicted target state and the current attitude interval pre-compensation quantity.
[0084] First, using the calibration relationship established in step S202, the predicted target state is converted to the control space of the PTZ telephoto camera to obtain the theoretical control quantity (pan_theory, tilt_theory). This theoretical control quantity represents the angle required for the PTZ to align with the predicted target under ideal calibration conditions.
[0085] Secondly, the current preset attitude range that the PTZ telephoto camera falls into is determined based on its real-time attitude (obtained from attitude feedback information). The attitude space is discretized according to horizontal angle, pitch angle, and zoom magnification: for example, the horizontal angle is divided into intervals of 10°, the pitch angle into intervals of 5°, and the zoom magnification is divided into several levels according to the magnification range (such as 1-10x, 10-20x, etc.). The Cartesian product of the three forms multiple three-dimensional intervals. Each interval corresponds to a compensation unit, which stores the pre-compensation amount (b_pan, b_tilt) for that attitude interval. This pre-compensation amount is obtained by recursively updating the control residual constructed in the historical linkage.
[0086] Then, the pre-compensation amount corresponding to the current attitude interval is retrieved from the attitude interval compensation table, and this pre-compensation amount is superimposed on the theoretical control amount to obtain the initial control amount. Simultaneously, when the target attitude is near the boundary between two adjacent intervals, linear interpolation of the compensation amounts for adjacent intervals can be performed based on distance weights to obtain a smoother compensation effect.
[0087] Step S206: Send initial control input to the PTZ device.
[0088] The system encapsulates the calculated initial control quantities (pan_init, tilt_init) and the zoom factor (zoom_init) determined based on the target distance estimation, target category, and preset retention boundary ratio into a control command. This command is then sent to the PTZ telephoto camera's pan-tilt controller via a communication interface (e.g., RS485, Ethernet, or HD-SDI) to drive it to perform coordinated control, enabling the telephoto camera to capture and magnify the target.
[0089] Step S207: Calculate the theoretical field of view center based on PTZ attitude backhaul.
[0090] After the PTZ telephoto camera performs cooperative control and stabilizes (e.g., waiting for one mechanical response cycle), the system reads the PTZ attitude feedback data again to obtain the current actual horizontal angle (pan_actual), pitch angle (tilt_actual), and zoom magnification (zoom_actual). Based on the field of view (horizontal field of view (FOV_h) and vertical field of view (FOV_v)) corresponding to the current zoom magnification, combined with the theoretical pointing of the PTZ, the direction vector of the theoretical field of view center can be calculated. This theoretical field of view center represents the true spatial direction that the center of the telephoto camera's image should be aligned with under ideal, unbiased conditions, and serves as the benchmark for subsequent evaluation of observational bias.
[0091] Step S208: Calculate the observation deviation based on the actual position of the target in the telephoto image.
[0092] In the telephoto camera's view, the actual position of the currently bound target is confirmed using a target detection algorithm (which can be the same as the wide-angle side or a lightweight model). Target binding must meet the following conditions: the target's identifier is consistent with the target identifier initiated by the wide-angle side, the target category matches, and the confidence level is higher than a preset threshold (e.g., 0.7). The pixel coordinates (x_obs, y_obs) of the target in the telephoto view are recorded and compared with the ideal center (x_ref, y_ref) of the telephoto view to obtain the pixel deviation (x_ref, y_ref). ):
[0093]
[0094] Meanwhile, to eliminate the influence of different resolutions, the deviation can be further normalized:
[0095]
[0096] Where W_z and H_z are the width and height of the current telephoto image.
[0097] Step S209: Generate control residuals based on the deviation between the theoretical field of view center and the actual observation.
[0098] The key to this specific implementation is that it does not simply use the observation bias for current fine-tuning, but rather binds it to the theoretical field of view center to construct a control residual with physical semantics. Under small bias conditions, the normalized pixel bias ( This can be converted into angular deviation:
[0099]
[0100] Where k_x and k_y are scaling factors (usually taken as 1), and FOV_h(z) and FOV_v(z) are the horizontal and vertical field of view angles at the current zoom level. , This refers to the control residual between the theoretical field of view center and the actual target center within the current attitude range. This control residual comprehensively reflects the combined effects of prediction error, calibration residual error, servo control error, and environmental factors, and can be directly used for subsequent recursive updates.
[0101] Step S210: Update the control residual to the compensation unit in the corresponding attitude range.
[0102] Before updating, a validity verification is performed to ensure the reliability of the linkage results. Verification conditions include, but are not limited to: the target identifier in the telephoto view is consistent with the target identifier in the wide-angle linkage; the actual observed position of the target is within a preset search window based on the theoretical field of view center (e.g., window radius of 50 pixels); the target category is consistent; the target confidence level is higher than a threshold; and the PTZ attitude feedback remains stable over multiple consecutive cycles (e.g., attitude change rate is less than 0.1° / s). If the verification passes, the update is executed.
[0103] The update uses an exponential moving average algorithm to adjust the original pre-compensation amount corresponding to the current attitude range. Control residuals of the new structure Perform weighted fusion:
[0104]
[0105] in, To update the weights, the value range is typically 0.1 to 0.3, and can be set based on on-site debugging experience. Larger values... This makes the pre-compensation amount more sensitive to the latest control residuals, and the smaller the amount, the better. This results in a smoother process. The updated pre-compensation amount is written back to the corresponding compensation unit in the attitude interval compensation table. Simultaneously, auxiliary information such as the update count and last update time of this compensation unit can also be recorded.
[0106] In addition, the system will freeze updates to the pre-compensation amount under the following circumstances: the target loss time in the telephoto view exceeds a preset duration (e.g., 2 seconds); the attitude feedback information undergoes abnormal jumps (e.g., the change between adjacent feedback values exceeds 5°); the pre-compensation amount in the current attitude range fluctuates beyond a set threshold within a preset time window (e.g., the change exceeds 1° in 3 consecutive updates); or scene calibration drift is detected and recalibration has not yet been completed. During the freeze period, the system skips step S210, but does not affect normal linkage control.
[0107] Step S211: When entering the same or adjacent preset attitude ranges in the future, the compensation amount in the compensation unit is superimposed on the PTZ initial control quantity solution process.
[0108] When wide-angle-PTZ linkage is performed again in subsequent steps, the system will query the attitude range into which the current real-time camera attitude falls in step S205 and read the corresponding pre-compensation amount from the attitude range compensation table. Since this pre-compensation amount has been recursively updated through multiple linkages, it reflects the statistically optimal estimate of the system error under the corresponding attitude range. By superimposing this pre-compensation amount onto the theoretical control amount, the corrected initial control amount can be obtained. If the real-time camera attitude falls within an adjacent range of a certain compensation unit in subsequent linkages, and the boundary interpolation function is enabled, the system will interpolate the pre-compensation amounts of the two adjacent ranges according to the distance weight to further smooth the compensation effect and improve the initial positioning accuracy.
[0109] By repeatedly repeating steps S201 to S211, the pre-compensation amount in the attitude interval compensation table will gradually converge to the optimal compensation value under each attitude interval. The first landing point deviation under the same scene and similar attitude will be significantly reduced, the number of repeated fine adjustments will be greatly reduced, the invalid PTZ actions and image jitter will be significantly reduced, and the overall monitoring efficiency and monitoring accuracy will be continuously improved.
[0110] Based on the target monitoring method for coordinated control of wide-angle and telephoto cameras provided in the above-described specific embodiments, this application achieves at least the following technical effects by establishing a complete closed-loop technical link of "control effective time alignment → dual-domain residual construction → attitude interval recursive pre-compensation": (1) Since the target time when the PTZ control actually takes effect is calculated by using the total delay estimate, and the target state is predicted from the target trajectory based on the time as the control input, the first landing point deviation caused by the link delay of multiple links such as video decoding, AI processing, transmission, command sending and mechanical response is effectively overcome, so that the target is closer to the ideal center of the telephoto image after the first turn of PTZ.
[0111] (2) By combining the theoretical field of view center formed by PTZ attitude backhaul with the actual observation deviation of telephoto, the control residual is constructed. The residual has a clear physical semantics (reflecting the deviation between the theoretical orientation and the actual landing point). The data source is stable and traceable, providing high-quality learning samples for subsequent pre-compensation quantity updates.
[0112] (3) By recursively updating the control residuals according to the three-dimensional attitude intervals composed of horizontal angle, pitch angle and zoom magnification (using exponential moving average to fuse historical experience with new residuals), and storing the updated pre-compensation amount in the attitude interval compensation table, the structured accumulation of historical linkage errors is realized. In subsequent identical or adjacent attitude intervals, the pre-compensation amount is directly superimposed on the initial control quantity solution process, so that each linkage can "stand on the shoulders of historical experience" and continuously improve the accuracy of the first positioning.
[0113] (4) As the pre-compensation amount is continuously optimized with the number of linkages, the number of repeated fine-tunings required by the system in the same scenario is significantly reduced, the invalid search and mechanical jitter of the PTZ camera are reduced, the equipment life is extended, and the smoothness of the picture is improved.
[0114] (5) By introducing validity verification (consistent target identification, observation position within the window, category matching, high confidence, stable attitude) and update freezing mechanism (target loss, attitude jump, compensation oscillation, calibration drift), the purity of data in the compensation unit is guaranteed, preventing errors or abnormal events from polluting the learning process, and ensuring that the pre-compensation amount always converges in the correct direction.
[0115] It is also understandable that the above-described specific implementation reduces the first landing point deviation caused by delay from the source, and continuously improves the first landing rate under similar scenes and similar postures through closed-loop learning, effectively improving the overall efficiency, accuracy and robustness of collaborative target monitoring between wide-angle and telephoto cameras.
[0116] Based on the target monitoring method based on the collaborative control of wide-angle and telephoto cameras provided in the embodiments of the present invention, the embodiments of the present invention also provide a target monitoring system based on the collaborative control of wide-angle and telephoto cameras, applicable to moving target monitoring scenarios in industrial settings, such as... Figure 3 As shown, the target monitoring system 300 based on the coordinated control of wide-angle and telephoto cameras includes: The acquisition unit 301 is used to acquire the motion trajectory of the target in the video stream of the wide-angle camera and to acquire the current camera pose of the PTZ telephoto camera; The cooperative control unit 302 is used to determine the target time when the PTZ telephoto camera control takes effect based on the estimated total delay from the current control decision time to the effective mechanical action of the PTZ telephoto camera. Using the target time as a reference, it predicts the target's state at the target time from the motion trajectory. Based on the predicted state and combined with the pre-compensation amount corresponding to the current preset attitude interval into which the current camera attitude falls, it solves the initial control amount of the PTZ telephoto camera and drives the PTZ telephoto camera to perform cooperative control, so that the telephoto camera can capture and magnify the target. The control residual construction unit 303 is used to determine the theoretical field of view center based on the attitude feedback information of the PTZ telephoto camera after the PTZ telephoto camera performs cooperative control, and to determine the observation deviation based on the actual observation position of the target in the telephoto camera image; based on the theoretical field of view center and the observation deviation, it constructs the control residual under the current preset attitude range; wherein, the control residual is used to update the corresponding pre-compensation amount, so that in the subsequent cooperative control process, the initial control amount is corrected based on the updated pre-compensation amount, so as to achieve the correction of the observation deviation; The output unit 304 is used to output the target image from the telephoto camera, which has been magnified and captured under coordinated control and has been corrected for observation deviations, as the target monitoring result.
[0117] Based on the above settings, the target monitoring system based on the coordinated control of wide-angle and telephoto cameras provided in this application ensures that the control input is aligned with the actual PTZ activation time through target time estimation and forward prediction of target state, reducing the initial landing point deviation from the source. By constructing a control residual bound to the attitude interval through dual-domain feedback (the theoretical field of view center versus actual observation deviation), the error is quantified and traceable. Simultaneously, by using the control residual to update the pre-compensation amount and directly calling the updated pre-compensation amount in subsequent coordinated control, a closed-loop technical link is formed: "first select the state according to the control activation time, then define the residual according to dual-domain feedback, and finally reinject pre-compensation according to the attitude interval." This reduces the number of repeated fine-tunings, lowers PTZ jitter and invalid searches, and improves the initial control landing rate under the same scene and similar attitudes, achieving continuous improvement in the initial landing accuracy of subsequent coordinated control, thereby improving performance. It is also understandable that the aforementioned closed-loop technology link enables the experience generated by each collaborative control (updating the corresponding pre-compensation amount using the control residual determined after collaborative control) to be accumulated and reused. The first arrival accuracy under the same or similar preset attitude range continues to improve with the increase of the number of collaborative control, thereby improving the overall monitoring efficiency.
[0118] Meanwhile, the target monitoring system 300 based on the coordinated control of wide-angle and telephoto cameras is configured to execute any of the aforementioned target monitoring methods based on the coordinated control of wide-angle and telephoto cameras. Therefore, the relevant units in the target monitoring system based on the coordinated control of wide-angle and telephoto cameras are also used to execute the corresponding operations in any of the aforementioned target monitoring methods based on the coordinated control of wide-angle and telephoto cameras. Accordingly, it also has all the beneficial effects of any of the aforementioned target monitoring methods based on the coordinated control of wide-angle and telephoto cameras, which will not be elaborated here.
[0119] It should also be noted that the specific units / modules within the target monitoring system based on the coordinated control of wide-angle and telephoto cameras are defined primarily based on the corresponding operations performed, and are not intended to limit the specific units / modules. For example, according to a specific embodiment of the present invention, a target monitoring system based on the coordinated control of wide-angle and telephoto cameras is also provided, mainly comprising: The multi-source data acquisition module is used to acquire wide-angle video frames, target metadata corresponding to the wide-angle video frames, PTZ attitude feedback data, and telephoto video frames; The calibration relationship establishment module is used to establish the calibration relationship from the wide-angle target anchor point to the PTZ control space; The control activation time estimation module is used to estimate the PTZ control activation time based on the delay of each link; The target state prediction module is used to select or predict the target state at the corresponding moment from the wide-angle target trajectory, based on the control effective moment. The initial control solution module is used to solve the PTZ initial control quantity based on the predicted target state and the current attitude interval pre-compensation quantity; The PTZ control transmission module is used to send initial control signals to the PTZ device. The theoretical field of view calculation module is used to calculate the center of the theoretical field of view based on the PTZ attitude feedback. The telephoto observation deviation calculation module is used to calculate the observation deviation based on the actual position of the target in the telephoto image; The control residual generation module is used to generate control residuals based on the deviation between the theoretical field of view center and the actual observation. The attitude interval compensation update module is used to update the control residual to the compensation unit of the corresponding attitude interval; The pre-compensation call module is used to add the compensation amount in the compensation unit to the PTZ initial control quantity solution process when entering the same or adjacent preset attitude ranges in the future.
[0120] Among them, the target state prediction module is controlled by the t_eff output of the control effective time estimation module; the control residual generation module depends on the outputs of both the theoretical field of view calculation module and the telephoto observation deviation calculation module; the pre-compensation call module and the attitude interval compensation update module together form a closed-loop compensation link.
[0121] This invention also provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of this invention.
[0122] This invention also provides a computer program product, including a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform the methods of embodiments of this invention. The computer program product should be understood as a software product that primarily implements the methods described above through a computer program.
[0123] This invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, which, when executed by the at least one processor, causes the electronic device to perform the method of this invention.
[0124] refer to Figure 4 The present invention will now be described in the form of a structural block diagram of an electronic device that can serve as an embodiment of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0125] like Figure 4 As shown, the electronic device includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. The RAM 403 may also store various programs and data required for the operation of the electronic device. The computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0126] Multiple components in the electronic device are connected to I / O interface 405, including: input unit 406, output unit 407, storage unit 408, and communication unit 409. Input unit 406 can be any type of device capable of inputting information into the electronic device. Input unit 406 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of the electronic device. Output unit 407 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 408 may include, but is not limited to, disks and optical discs. Communication unit 409 allows the electronic device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, and / or wireless communication transceivers, such as Bluetooth devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0127] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, CPUs, graphics processing units (GPUs), various special-purpose artificial intelligence (AI) computing units, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above. For example, in some embodiments, the method embodiments of the present invention can be implemented as a computer program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on an electronic device via ROM 402 and / or communication unit 409. In some embodiments, the computing unit 401 can be configured to perform the methods described above by any other suitable means (e.g., by means of firmware).
[0128] Computer programs for implementing the methods of embodiments of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0129] In the context of embodiments of the present invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, or infrared systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0130] It should be noted that the term "comprising" and its variations used in the embodiments of the present invention are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "multiple" mentioned in the embodiments of the present invention are illustrative and not restrictive. Those skilled in the art should understand that, unless explicitly indicated otherwise in the context, they should be understood as "one or more".
[0131] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0132] The steps described in the method embodiments provided by this invention can be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of protection of this invention is not limited in this respect.
[0133] The term "embodiment" in this specification refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of the invention. The appearance of this phrase in various places in the specification does not necessarily imply the same embodiment, nor does it imply independence or alternativeity from other embodiments. The various embodiments in this specification are described in a related manner, with reference to each other for similar or identical parts. In particular, for apparatus, device, and system embodiments, since they are substantially similar to method embodiments, the description is relatively simple, and relevant details are referred to in the description of the method embodiments.
[0134] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.
Claims
1. A target monitoring method based on the coordinated control of wide-angle and telephoto cameras, characterized in that, include: Obtain the motion trajectory of the target in the video stream of the wide-angle camera, and obtain the current camera pose of the PTZ telephoto camera; Based on the estimated total delay from the current control decision moment to the effective mechanical action of the PTZ telephoto camera, the target moment for the PTZ telephoto camera control to take effect is determined. Using the target moment as a reference, the predicted state of the target at the target moment is predicted from the motion trajectory. Based on the predicted state and combined with the pre-compensation amount corresponding to the current preset attitude interval into which the current camera attitude falls, the initial control amount of the PTZ telephoto camera is solved, and the PTZ telephoto camera is driven to perform cooperative control so that the telephoto camera can capture and magnify the target. After the PTZ telephoto camera performs cooperative control, the theoretical field of view center is determined based on the attitude feedback information of the PTZ telephoto camera, and the observation deviation is determined based on the actual observation position of the target in the telephoto camera image. Based on the theoretical field of view center and the observation deviation, the control residual under the current preset attitude range is constructed. The control residual is used to update the corresponding pre-compensation amount so that in the subsequent cooperative control process, the initial control amount is corrected based on the updated pre-compensation amount to achieve correction of the observation deviation. The target image captured by the telephoto camera, after being magnified and captured through collaborative control and corrected for the observation deviation, is output as the target monitoring result.
2. The method according to claim 1, characterized in that, The total latency estimate includes video decoding latency, AI processing latency, transmission latency, command sending latency, and PTZ telephoto camera mechanical response latency, and each latency component is dynamically updated using a moving average method.
3. The method according to claim 1, characterized in that, The step of predicting the target's predicted state at the target time from the motion trajectory includes: Calculate the velocity and / or acceleration of the target based on the position information of at least two historical moments in the motion trajectory; Based on the velocity and / or the acceleration, the state of the target is extrapolated from the detection time to the target time to obtain the predicted state.
4. The method according to claim 1, characterized in that, The steps for solving the initial control variables of the PTZ telephoto camera include: Based on the preset calibration relationship, the predicted state is converted to the control space of the PTZ telephoto camera to obtain the theoretical control quantity; Determine the current preset attitude range into which the current camera attitude falls, obtain the pre-compensation amount corresponding to the current preset attitude range, and superimpose the pre-compensation amount onto the theoretical control amount to obtain the initial control amount.
5. The method according to claim 1, characterized in that, Also includes: A mapping table is constructed between preset attitude intervals and pre-compensation values. The mapping table discretizes the range of values for the horizontal angle, pitch angle and zoom magnification of the PTZ telephoto camera into multiple three-dimensional intervals, and each interval stores a pre-compensation value. The current preset attitude interval refers to the corresponding interval in the mapping table into which the current camera attitude falls. During the collaborative control process, when the real-time camera attitude enters a certain interval, the pre-compensation quantity stored in that interval is directly called to participate in the initial control quantity solution.
6. The method according to claim 1, characterized in that, The corresponding pre-compensation amount is updated based on the control residual, including: An exponential moving average algorithm is used to weight and fuse the control residual with the original pre-compensation amount corresponding to the current preset attitude interval for updating; wherein... The weighted fusion update includes: adding the product of the first weight and the original pre-compensation amount, and the product of the second weight and the control residual, to obtain the updated pre-compensation amount, wherein the sum of the first weight and the second weight is 1.
7. The method according to claim 1, characterized in that, Before updating the corresponding pre-compensation amount based on the control residual, the following steps are also included: The effectiveness of this collaborative control was verified, and the verification conditions included any one or more of the following: the target identifier in the telephoto camera's image is consistent with the target identifier in the wide-angle camera's video stream; the actual observation position of the target in the telephoto camera's image is within a preset search window based on the theoretical field of view center; the target category observed in the telephoto camera's image is consistent with the target category monitored in the wide-angle camera's video stream; the confidence level of the target observed in the telephoto camera's image is higher than a preset threshold; and the attitude feedback information of the PTZ telephoto camera remains stable over multiple consecutive cycles. If the verification passes, the pre-compensation amount is updated.
8. The method according to claim 7, characterized in that, Also includes: The update of the pre-compensation amount is frozen when at least one of the following conditions is met: The target in the telephoto camera footage was lost for longer than the preset duration. The attitude feedback information changes abnormally; the pre-compensation amount of the current preset attitude range fluctuates beyond a set threshold within a preset time window; Scene calibration drift was detected and recalibration has not yet been completed.
9. A target monitoring system based on the coordinated control of wide-angle and telephoto cameras, characterized in that, include: The acquisition unit is used to acquire the motion trajectory of the target in the video stream of the wide-angle camera, and to acquire the current camera pose of the PTZ telephoto camera; The collaborative control unit is used to determine the target time when the PTZ telephoto camera control takes effect based on the estimated total delay from the current control decision time to the effective mechanical action of the PTZ telephoto camera; using the target time as a reference, it predicts the predicted state of the target at the target time from the motion trajectory; based on the predicted state and combined with the pre-compensation amount corresponding to the current preset attitude interval into which the current camera attitude falls, it solves the initial control amount of the PTZ telephoto camera, drives the PTZ telephoto camera to perform collaborative control, and enables the telephoto camera to capture and magnify the target. A control residual construction unit is used to determine the theoretical field of view center based on the attitude feedback information of the PTZ telephoto camera after the PTZ telephoto camera performs cooperative control, and to determine the observation deviation based on the actual observation position of the target in the telephoto camera image; based on the theoretical field of view center and the observation deviation, construct the control residual under the current preset attitude range; wherein, the control residual is used to update the corresponding pre-compensation amount, so that in the subsequent cooperative control process, the initial control amount is corrected based on the updated pre-compensation amount, so as to achieve the correction of the observation deviation; The output unit is used to output the target image captured by the telephoto camera after being magnified and corrected for the observation deviation, as the target monitoring result.
10. A non-transitory machine-readable medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 8.