A method and device for automatically locking a face in cruise

By using a local lightweight algorithm and a direct-plug camera architecture, combined with PTZ control and the UVC/VISCA protocol, the stability and privacy security issues of existing cameras in face recognition and tracking are solved, enabling accurate locking and continuous tracking of specific target faces.

CN122227064APending Publication Date: 2026-06-16XIAMEN RGBLINK SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAMEN RGBLINK SCI & TECH CO LTD
Filing Date
2026-05-19
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing cameras cannot achieve accurate positioning and exclusive locking of specific target faces, resulting in problems such as high network latency, data transmission risks, privacy leaks, misidentification, tracking jumps, target loss, and unstable tracking.

Method used

Employing a local lightweight algorithm and a direct-plug camera architecture, it constructs a face loss detection mechanism through real-time image acquisition, face comparison and matching, and exclusive detection, combined with PTZ control algorithm and UVC/VISCA protocol, to implement a hierarchical recapture strategy and support extended functions such as multi-target priority binding, dynamic adaptive tracking sensitivity, and hot-plug tracking without loss.

Benefits of technology

It achieves stable, anti-interference, and adaptive automatic cruise lock tracking of designated faces, features low power consumption and high real-time performance, supports automatic compatibility with multiple protocols, improves device compatibility and ease of use, and ensures tracking stability and privacy security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122227064A_ABST
    Figure CN122227064A_ABST
Patent Text Reader

Abstract

The application discloses a processing method and device for automatically cruising and locking a face, applied to the technical field of data processing, and realizes target face sample registration, real-time picture collection and exclusive accurate detection by means of a local lightweight algorithm matched with a camera direct-insert architecture, and effectively filters irrelevant face interference. The system uses a PTZ control algorithm based on a face position, relies on a UVC / VISCA protocol to drive a pan-tilt head adjustment, keeps the target face in the center of the picture, constructs a loss detection and hierarchical recapture mechanism, and combines space-time memory to form a complete tracking closed loop. Meanwhile, the system is integrated with multiple target priority, adaptive sensitivity, hot plug follow-up tracking and other expansion capabilities, and finally outputs a high-robustness and high-stability specific face automatic cruising and locking tracking result through layered control engine and parameter iteration optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for automatically cruise and lock onto a face. Background Technology

[0002] With the rapid popularization of video surveillance, intelligent shooting, and human-computer interaction, higher demands are being placed on the intelligent tracking capabilities of cameras. Currently, ordinary cameras only have basic video capture functions and cannot perform local AI face analysis and precise positioning, making it difficult to achieve exclusive locking and continuous tracking of specific target faces. Specifically, they suffer from the following technical deficiencies: Ordinary cameras do not have lightweight AI algorithms built in. Face detection, feature extraction and identity comparison need to be completed by uploading the video stream to the server, which has problems such as high network latency, data transmission risks and privacy leaks. They cannot meet the needs of offline, low power consumption and real-time tracking.

[0003] Ordinary cameras can only detect the presence of facial areas in the image, but cannot distinguish between target faces and irrelevant faces. In scenarios such as multiple people on the same screen, others obscuring the view, or the camera crossing the screen, it is prone to misidentification, tracking jumps, and target loss, and cannot achieve accurate positioning and exclusive locking of a specific face.

[0004] Ordinary cameras lack real-time gimbal control logic based on face position, do not integrate PTZ closed-loop control algorithm, and are not compatible with mainstream gimbal protocols such as UVC / VISCA. They cannot automatically adjust the gimbal attitude according to face offset, making it difficult to keep the target face stably centered in the image.

[0005] When the target face briefly moves out of the frame or is obscured, ordinary cameras lack loss detection, spatiotemporal memory, and hierarchical recapture strategies. They can only passively wait for the target to reappear, which can easily lead to tracking interruption and prevent the formation of a complete closed loop of detection-tracking-loss-recapture.

[0006] Ordinary camera drivers and protocols are complex to adapt to, do not support hot-swapping or automatic compatibility with multiple protocols, and do not have extended functions such as adaptive sensitivity, safe framing, anti-snap-in, and static sleep. In complex scenarios, tracking stability, anti-interference and ease of use are difficult to guarantee. Summary of the Invention

[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution: An automatic cruise face locking method includes: acquiring core data for the entire process of automatic cruise locking and continuous tracking of a specified face; based on a local lightweight algorithm and a direct-plug camera architecture, uploading and registering the target face sample, and accurately locating the target face and filtering out irrelevant face interference through real-time image acquisition, face comparison and matching, and exclusive detection; based on the target face position, driving gimbal attitude adjustment through a PTZ control algorithm and relying on the UVC / VISCA protocol to ensure that the target face is stably kept in the center of the image; and constructing a face loss detection mechanism to record the spatiotemporal information of the target disappearance, initiate a hierarchical recapture strategy, and perform short-term loss detection. It performs localized, small-scale recapture and enters full-domain intelligent cruise after prolonged loss, forming a complete closed loop of detection-tracking-loss-recapture. Through extended mechanisms such as multi-target priority binding, dynamic adaptive tracking sensitivity, hot-swappable tracking without loss, multi-protocol automatic compatibility, secure mapping tracking, anti-sniping lock protection, target static hibernation, and high-speed focus tracking, it optimizes tracking stability, anti-interference, and device compatibility. Real-time face position, gimbal attitude, loss status, and motion feature data are imported into the hierarchical control engine, which uses multi-parameter adaptive adjustment and iterative optimization, combined with mapping calibration and image stabilization compensation modules, to generate highly robust automatic cruise lock tracking results for specific faces.

[0008] An automatic cruise face-locking processing device is configured to execute the automatic cruise face-locking processing method by executing executable instructions.

[0009] Its beneficial effects are as follows: Through a local lightweight algorithm and direct camera connection architecture, it achieves designated face registration, real-time acquisition, exclusive recognition, and PTZ closed-loop tracking. The system focuses on feature extraction, identity comparison, and exclusive locking, combined with adaptive adjustment of gimbal parameters based on offset and motion speed. It also constructs a hierarchical loss and recapture mechanism, supporting short-term on-site recapture and long-term full-domain cruise. Furthermore, it integrates eight extended functions: multi-target priority, spatiotemporal memory recapture, automatic multi-protocol compatibility, hot-swappable continuation tracking, secure image composition, anti-camera snatching, static sleep mode, and high-speed focus tracking. Through a hierarchical control engine and a multi-parameter adaptive adjustment model, the system iteratively optimizes tracking parameters under four constraints, ultimately forming a complete closed loop of detection-tracking-loss-recapture. It can run offline on embedded devices, achieving stable, interference-resistant, and adaptive automatic cruise locking tracking for specific faces.

[0010] This application aims to achieve exclusive face locking for a designated person, free from interference from others, without skipping or false locking, operating locally offline, requiring no cloud access, low power consumption, high real-time performance, and ensuring privacy and security. The gimbal parameters adaptively adjust, ensuring smooth movement at both fast and slow speeds, preventing image jitter and target drift. It features spatiotemporal memory and graded recapture for faster recovery of lost targets and stronger tracking continuity. It supports automatic multi-protocol recognition, hot-swappable devices without losing tracking, significantly improving device compatibility and ease of use. Golden composition and anti-camera-stealing mechanisms ensure aesthetically pleasing images, stable tracking, and strong scene adaptability. Attached Figure Description

[0011] Figure 1 A flowchart illustrating an automatic cruise face-locking processing method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a processing device for automatically cruise and lock onto a face, provided in an embodiment of the present invention. Detailed Implementation

[0012] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention. Figure 1 This application describes an automatic cruise face-locking processing method according to exemplary embodiments thereof. In one embodiment, this application also proposes an automatic cruise face-locking processing method.

[0013] In this application embodiment, an automatic cruise face locking processing method is provided, such as... Figure 1 As shown: S101, acquires core data for the entire process of automatic cruise locking and continuous tracking of a specified face.

[0014] In one implementation, the system first performs feature extraction processing on the local operating environment, camera access link, and target face registration process to generate three types of basic features: local lightweight algorithm features, camera direct connection features, and specified face sample registration features. The local lightweight algorithm features characterize the algorithm's attributes of running independently on the local device, not relying on the cloud, and operating with low computational power consumption and low computational cost. The system adopts an embedded processing architecture, deploying face detection, feature comparison, and gimbal control logic all on local hardware, eliminating the need for server connections and image data uploads; all computational tasks are completed locally. The system does not rely on an external network environment and requires no complex background services, enabling stable and continuous operation on low-power, low-computing-power embedded devices, ensuring smooth, uninterrupted, and data-free operation.

[0015] Taking a UVC standard camera as an example, the direct-connect feature of the camera indicates that the device conforms to the UVC protocol, enabling plug-and-play functionality without the need for additional driver installation or manual parameter configuration. The UVC protocol is a universal video device standard protocol; cameras compliant with this protocol do not require separate driver installation and are automatically recognized after connection. Users insert a UVC-compliant camera into the device interface, and the system automatically completes device enumeration, communication link establishment, and video stream format matching. There is no need for manual protocol selection, resolution setting, or parameter configuration; the system directly reads the video output data from the camera, ensuring rapid device connection and deployment. Furthermore, cameras include, but are not limited to, those with HDMI, SDI, and CVBS interfaces, each supporting its corresponding protocol, which will not be discussed further here.

[0016] The designated facial sample registration features characterize the process of a user uploading a facial photo, the system extracting unique facial features, and binding the target's identity. After the user uploads a photo of the target face to be locked, the system extracts the facial region from the photo, analyzes the texture distribution, contour shape, and key point structure of the face, converts this information into unique feature data that can be used for comparison, and saves it as a dedicated matching template. This template is only used for subsequent facial comparison and recognition, does not participate in other calculations, and will not be confused with other targets, providing a stable and unique basis for identity determination for subsequent locking and tracking.

[0017] Specifically, based on specified face locking and tracking constraints, the system constructs a multi-feature fusion decision algorithm. This algorithm integrates local lightweight algorithm features, direct camera connection features, specified face sample registration features, real-time video frame acquisition features, face identity comparison features, and specific face exclusive recognition features for unified input and joint analysis. The algorithm employs serial verification logic, first determining device and environment compatibility, then confirming face identity, and finally performing exclusive filtering to ensure the legality and validity of data throughout the entire process.

[0018] During the tracking and positioning process, the system uses a dual threshold verification algorithm for real-time determination. The first layer is identity consistency verification. The algorithm calculates the Euclidean distance between the real-time collected facial feature vector and the registered template feature vector, maps the distance value to a similarity value, and compares it with a preset similarity threshold. If the value is greater than or equal to the threshold, it is determined to be a registered target; if the value is less than the threshold, it is determined to be a non-target. The second layer is exclusive locking verification. The algorithm traverses all detected facial targets in the current frame, retains only the target data that passes the identity consistency verification, and forcibly masks and discards the remaining facial data without generating any control output, ensuring that the specified target is uniquely locked.

[0019] After all dual verifications pass, the multi-feature fusion decision algorithm outputs the precise localization result of the target face. This result includes the horizontal and vertical center coordinates of the target face in the video frame, a successful identity matching marker, and a tracking effective enable marker. The localization result is output via a standardized data interface, serving as the sole input for the subsequent PTZ control module. This ensures that PTZ attitude adjustment is performed only on valid targets, improving tracking stability and anti-interference capabilities.

[0020] The system performs feature extraction processing on the real-time image acquisition, face comparison and matching, and exclusive detection processes, generating three types of process features: real-time video frame acquisition features, face identity comparison features, and exclusive face recognition features. Real-time video frame acquisition features include streaming data from continuous camera image acquisition, frame synchronization processing, and real-time position resolution. The system continuously reads the video stream through standard interfaces such as UVC, HDMI, SDI, and CVBS, acquiring image information frame by frame at a fixed frame rate to ensure that each frame remains continuous and unlost in time. The system performs frame synchronization calibration on each frame to ensure that image acquisition, data processing, and PTZ control maintain the same temporal rhythm, avoiding image delays or asynchrony. The system performs a full-area scan on each frame, identifying all facial regions appearing in the image and calculating the horizontal and vertical coordinates of each face. This coordinate information, along with the frame number and timestamp, forms real-time position data, providing precise positional basis for subsequent PTZ attitude adjustments.

[0021] The facial identity comparison features include matching data for calculating the similarity between the features of the face to be identified and the registered face, and determining identity consistency. The system extracts the current facial features from the real-time captured facial regions using the same feature extraction method as in the registration phase. The system compares the current features with the stored registration template features item by item, calculates the feature difference between the two, and converts it into a similarity value. The system pre-sets an identity determination threshold. When the calculated similarity is higher than this threshold, the current face is determined to be the registration target, and a target matching tag is generated. When the similarity is lower than this threshold, it is determined to be a non-target face, and no valid matching tag is generated. For example, after the system registers a user's facial features, it compares the captured faces in real-time detection. If the similarity meets the requirements, it can confirm that they are the same target, ensuring accurate identification.

[0022] The specific face exclusive recognition feature includes filtered data that locks onto only the target face, ignores irrelevant faces, and resists interference from multiple people. After completing face identity comparison, the system filters all faces in the frame, retaining only the face data marked as the registered target, and excluding the location and feature information of all other non-target faces. When multiple people appear in the frame simultaneously, or when others approach and obstruct or cross the frame, the system maintains a unique lock on the registered target, without switching the tracking object, responding to the location of irrelevant faces, or triggering gimbal jump actions. Through exclusive recognition processing, the system can stably focus on the specified target in complex scenes, unaffected by surrounding people, providing reliable single-target data for continuous tracking.

[0023] Based on the overall constraints of specified face locking and tracking, the system performs joint analysis and processing on multiple features extracted in the early stages. These features include local lightweight algorithm features, direct camera connection features, registered features of specified face samples, real-time video frame acquisition features, face identity comparison features, and exclusive face recognition features. The system integrates and logically verifies these features, creating a complete data loop encompassing device operating status, video acquisition process, face registration information, and real-time detection results, providing a reliable basis for accurate positioning. During the tracking and positioning process, the system initiates a dual real-time verification mechanism. The first layer of verification determines whether the currently detected face is the target face pre-uploaded and registered by the user. The system confirms identity by comparing real-time features with registered template features, ensuring accurate identification. The second layer of verification confirms whether the current process strictly adheres to the exclusive locking rule, i.e., responding only to registered faces and ignoring other unregistered faces in the frame.

[0024] Once both verifications pass, the system generates and outputs the precise location result of the target face. This location result includes three key pieces of information: the stable horizontal and vertical coordinates of the target face in the video frame, a target identity verification marker to indicate successful identity verification, and a valid tracking marker to indicate the available tracking process. For example, in a real-world scenario, after a user completes face registration, the system detects the face in the real-time frame and passes identity comparison and exclusion verification, then outputs location information including the face center coordinates, the verification marker, and the valid marker. This result serves as the sole input data to the subsequent gimbal control module, enabling the gimbal to perform attitude adjustments based on the accurate location information, thereby achieving stable and reliable tracking of the specified face.

[0025] S102, based on a local lightweight algorithm and a direct-plug camera architecture, completes the uploading of target face samples and feature registration. Through real-time image acquisition, face comparison and matching, and exclusive detection, it accurately locates the target face and filters out irrelevant face interference.

[0026] In one implementation, the system uses a local lightweight facial feature extraction model as its core to complete target face sample registration and feature generation. This model consists of a feature extraction layer, a feature normalization layer, and a template storage layer. Data is transferred between these three layers via a serial forward connection, without backpropagation or training. The user uploads a photo of the face to be locked as model input. The feature extraction layer extracts the texture distribution, contour shape, and facial feature key point information of the face through facial landmark detection and local feature descriptor operators. The feature normalization layer performs numerical scaling and uniform quantization on the extracted multi-dimensional features to generate a fixed-dimensional feature vector. Key parameters include the feature vector length and the global normalization coefficient. The template storage layer stores the normalized feature vector in local storage units, forming a unique registration template as the benchmark for subsequent identity comparison. The entire registration process only performs feature extraction and storage, without model training, ensuring rapid completion on low-computing-power embedded devices.

[0027] The system establishes a direct data channel with the camera via standard protocols such as UVC, HDMI, SDI, and CVBS, continuously reading video stream data at a fixed frame rate. A real-time video frame acquisition algorithm performs frame synchronization calibration on consecutive video frames, ensuring consistency in timing between image acquisition, data processing, and pan-tilt control. The algorithm performs brightness equalization and spatial filtering denoising on each frame, improving image clarity and the stability of subsequent face detection. After processing, standardized video frame data is output, containing complete pixel information, frame number, and timestamp, providing a timely and stable input source for subsequent face detection.

[0028] The system employs a sliding window detection method to perform face region localization within standardized video frames, obtaining the bounding rectangle coordinates of all faces in the frame. For each detected face region, a real-time feature vector is generated using the same lightweight feature extraction model as in the registration phase. The system uses the Euclidean distance algorithm to calculate the difference between the real-time feature vector and the registered template feature vector dimension by dimension, mapping the distance value to a feature similarity value. The system pre-sets a similarity threshold; when the calculated similarity is higher than the threshold, the current face is determined to be the registration target, and a target matching label and corresponding coordinates are output. When the similarity is lower than the threshold, it is determined to be a non-target face, and no valid label is output. This process uses real-time features as input and matching labels and coordinates as output, combining the feature distance calculation algorithm with identity verification business logic to achieve accurate identity confirmation in offline mode.

[0029] The exclusive filtering algorithm takes face detection and identity determination results as input and iterates through all face entries in the frame. The algorithm only retains data with target matching markers and discards face entries without matching markers. When multiple people are on the same screen, others obscure the view, or they cross the screen, the algorithm maintains its logic of only responding to the registered target, without switching tracking objects, responding to irrelevant face positions, or triggering gimbal actions. The algorithm ultimately outputs a unique target face coordinate and a valid positioning marker, completely filtering out irrelevant face interference and providing a single, stable positional basis for gimbal control.

[0030] The system integrates registration features, real-time data acquisition, identity comparison results, and exclusion screening results to form the final accurate target face localization result. The output includes the target face's horizontal and vertical coordinates within the image, a target identity confirmation marker, and a tracking validity marker. This localization result serves as the sole input to the subsequent gimbal control module, driving PTZ attitude adjustment and achieving a complete closed loop from face registration, real-time detection, identity comparison to accurate localization.

[0031] The S103, based on the target face position, uses the PTZ control algorithm and relies on the UVC / VISCA protocol to drive the gimbal attitude adjustment, ensuring that the target face is stably kept in the center of the image.

[0032] In one implementation, the system takes four data points as input: real-time positioning coordinates of the target face, PTZ gimbal motion characteristics, UVC and VISCA protocol adaptation specifications, and tracking response timeliness. These are then fused using a closed-loop control strategy generation algorithm to form a unified gimbal control logic. This algorithm employs multi-dimensional matching rules. First, it matches the rate of change of the face position in the image with the maximum rotation speed and minimum rotation step size of the gimbal in the horizontal and vertical directions, ensuring that the gimbal's motion amplitude and the face's movement speed are synchronized. Second, it matches the corresponding instruction format and data encapsulation method based on the communication protocol type supported by the camera, ensuring that the control commands can be correctly parsed by the gimbal. Finally, based on the tracking response timeliness requirements set by the system, it determines the duration of a single control cycle, keeping the gimbal adjustment frequency synchronized with the face position update frequency.

[0033] After fusion calculation, the algorithm outputs the final gimbal attitude control and closed-loop tracking strategy. This strategy defines three core elements: first, the execution priority of horizontal and vertical rotation, typically prioritizing the axis with the greater face offset. Second, the applicable protocol call method for the current scenario, automatically selecting the VISCA instruction set for PTZ control under the UVC direct connection architecture. Third, a real-time feedback mechanism, immediately reading the new face position after each gimbal adjustment to form a closed-loop verification.

[0034] When the target face moves rapidly to the left in the frame, with a lateral offset significantly greater than a vertical offset, the algorithm prioritizes horizontal rotation and outputs a horizontal rotation control command. The system automatically encapsulates the leftward rotation command using the VISCA protocol and sends the control signal according to the set response time, causing the gimbal to quickly follow to the left. This ensures that the tracking response delay does not exceed the system's set threshold, achieving fast and stable tracking.

[0035] The system uses four input criteria: face image offset dimension, gimbal rotation angle range, real-time tracking accuracy requirements, and stable composition constraints. Through a parameter configuration algorithm, it quantifies and calculates the gimbal control parameters, generating three core control parameters: gimbal adjustment step size, dynamic response speed, and position deviation compensation threshold. The parameter configuration algorithm employs a graded offset mapping rule, dividing the face's offset distance in the image into multiple levels. Combining this with the maximum and minimum rotation angles supported by the gimbal hardware, the offset levels are linearly mapped to the gimbal adjustment step size and dynamic response speed. Simultaneously, the position deviation compensation threshold is calculated based on the real-time tracking accuracy requirements.

[0036] The gimbal adjustment step size controls the angle range of a single rotation movement of the gimbal. The dynamic response speed controls the rotation speed per unit time. The position deviation compensation threshold defines the acceptable range of face position deviation and determines whether compensation adjustment needs to be initiated. When the face deviates slightly from the center of the image, the parameter configuration algorithm classifies it as a small deviation, outputting a small adjustment step size and a low dynamic response speed to allow the gimbal to correct its position smoothly and prevent image jitter. When the face deviates significantly from the center of the image, the algorithm classifies it as a large deviation, outputting a larger adjustment step size and a higher dynamic response speed to allow the gimbal to quickly return to its original position, ensuring the target does not fly out of the frame. The position deviation compensation threshold serves as the criterion for initiating compensation adjustment. When the face offset is less than this threshold, the system considers the deviation acceptable and does not perform gimbal adjustment; when the offset is greater than this threshold, the system immediately initiates compensation adjustment, driving the gimbal to rotate and correct the position.

[0037] In stable image tracking mode, the system sets a small range of face offset center as acceptable, with a compensation threshold set to one-tenth of the screen width. When the face offset is less than this value, the gimbal remains stationary; when the face offset exceeds this value, the algorithm allocates step size and speed according to the offset magnitude. Small offsets are corrected smoothly with small step sizes and slow speeds, while large offsets are corrected quickly with large step sizes and fast speeds to center the face, ensuring the target remains within the safe image composition area, balancing tracking accuracy and image stability.

[0038] The system comprehensively matches the determined gimbal adjustment step size, dynamic response speed, and position deviation compensation threshold with the camera's acquisition frame rate, data transmission bandwidth, and processing latency. A frequency calibration algorithm precisely calculates and sets three hierarchical control frequencies: face position resolution frequency, gimbal command transmission frequency, and posture feedback verification frequency. The frequency calibration algorithm employs timing synchronization constraints, using the camera's acquisition frame rate as the reference clock. The face position resolution frequency is set to be completely consistent with the acquisition frame rate, ensuring timely resolution of each frame. The gimbal command transmission frequency is set to a reasonable division value of the acquisition frame rate to avoid gimbal congestion caused by excessively rapid command transmission. The posture feedback verification frequency is set to be the same as the resolution frequency, enabling closed-loop verification for each frame. The algorithm uses frame count alignment and delay compensation mechanisms to maintain strict timing synchronization between the control flow and the video acquisition flow, preventing command backlog, response lag, or feedback misalignment.

[0039] In typical operating mode, when the camera outputs video at 25 frames per second, the frequency calibration algorithm sets the face position resolution frequency to 25 times per second to ensure face position calculation for each frame; it sets the gimbal command transmission frequency to 10 times per second to ensure smooth tracking while avoiding gimbal jitter caused by frequent commands; and it sets the attitude feedback verification frequency to 25 times per second to verify the deviation of each resolution result. Through this frequency configuration, the system achieves coordinated operation of the entire process of image acquisition, position resolution, gimbal control, and effect verification, ensuring both real-time tracking and maintaining system stability without lag.

[0040] The system operates continuously according to the pre-calibrated frequency of face position resolution, gimbal command transmission, and posture feedback verification. Using real-time face position data as the sole input, it sequentially executes three consecutive operations—position resolution, command encapsulation, and posture verification—through a closed-loop tracking algorithm, forming a frame-level synchronized control flow. First, the system resolves the face position in each video frame at the set frequency. Through face detection and coordinate calculation, it obtains the horizontal and vertical center coordinates of the target face within the frame, achieving accurate real-time position extraction. Second, based on the deviation direction and magnitude between the current face coordinates and the frame center, the system automatically encapsulates the corresponding gimbal horizontal or vertical rotation control commands according to the UVC or VISCA protocol instruction format and sends the rotation signal to the gimbal via the communication interface. Finally, the system reads the current gimbal posture angle and simultaneously acquires the face position information from the latest frame, jointly verifying both to determine if the face has returned to the center of the frame and confirming that the tracking effect meets the standards.

[0041] The closed-loop tracking algorithm employs a frame-level timing synchronization mechanism, combining the parsing, sending, and verification steps into an indivisible closed-loop execution cycle. Each completed set of operations immediately triggers the next loop, ensuring continuous and uninterrupted tracking. Example: When the system detects a face deviating to the right of the image and the deviation exceeds the compensation threshold, the algorithm immediately calculates the required correction angle, encapsulates a leftward rotation control command according to the VISCA protocol, and sends it to the gimbal. After the gimbal executes the rotation, the system immediately reads the new gimbal angle and the face position, comparing whether the deviation has decreased, until the face returns to the center of the image, completing one full closed-loop adjustment.

[0042] During closed-loop tracking, the system continuously monitors four operational states using a real-time anomaly detection algorithm: excessive face offset, PTZ response lag, protocol adaptation anomaly, and excessive image deviation. The anomaly detection algorithm takes real-time collected face position data, PTZ feedback angle, protocol communication status, and image position data as input, and compares them item by item with preset deviation thresholds, response time thresholds, communication verification thresholds, and image security thresholds. When any data exceeds the corresponding threshold range, the algorithm immediately determines it as an anomaly, sends an anomaly flag and anomaly type information to the upper-layer module, and triggers a step size adjustment. speed Threshold triad adaptive readjustment mechanism.

[0043] The readjustment mechanism employs a parameter self-optimization algorithm. Taking the anomaly type, current deviation level, and historical tracking results as inputs, it recalculates four core parameters: gimbal adjustment step size, dynamic response speed, deviation compensation threshold, and graded control frequency. The algorithm dynamically adjusts the parameter amplitude based on the anomaly level, rapidly reducing the face position deviation while ensuring stable gimbal operation, until the face returns to the center of the image and remains stable. Afterward, it returns to the normal closed-loop control process.

[0044] When the gimbal experiences tracking lag due to hardware response delay, causing the face to continuously deviate from the center of the image and fail to return to its original position, the anomaly detection algorithm identifies the gimbal's response lag anomaly. The readjustment algorithm automatically increases the gimbal's dynamic response speed level, increases the single adjustment step size, and tightens the deviation compensation threshold based on the lag duration and deviation magnitude. Simultaneously, it slightly adjusts the control frequency, enabling the gimbal to quickly correct its position with stronger responsiveness. After several iterative adjustments, the face returns to the center of the image, the deviation falls within the threshold, and the system resumes normal and stable tracking. Through this anomaly detection and adaptive readjustment logic, the system can autonomously correct tracking deviations under complex scenarios and equipment fluctuations, maintaining stable locking throughout the entire process and significantly improving tracking robustness.

[0045] S104 constructs a face loss detection mechanism, records the spatiotemporal information of the target disappearance, and initiates a hierarchical recapture strategy. For short-term loss, it performs on-site small-range recapture; for long-term loss, it enters full-domain intelligent navigation, forming a complete closed loop of detection-tracking-loss-recapture.

[0046] In one implementation, the target face tracking state dataset is classified and split, including real-time face location data, gimbal pose data, loss trigger time data, disappearance direction location data, and tracking state marker data, generating state category classification results, a spatiotemporal feature association table, and a loss-recapture label mapping relationship. During continuous tracking, the system uses a state data dimensionality splitting algorithm to structurally decompose the real-time acquired target face tracking state dataset, dividing the originally continuous and mixed data stream into five independent and logically complementary data categories, providing a standardized data foundation for subsequent hierarchical recapture. The system first extracts real-time face location data, which is used to accurately represent the horizontal and vertical coordinates of the target face in the current video frame, reflecting the target's specific distribution position in the frame, and serving as the direct basis for gimbal adjustment and deviation calculation.

[0047] The system synchronously extracts gimbal attitude data to record the current horizontal and vertical rotation angles of the gimbal, reflecting the physical position of the device and providing a baseline for the initial attitude during recapture. The system extracts loss trigger time data to record the precise timestamp when the target face switches from normal tracking to an undetectable state, serving as a core criterion for distinguishing between short-term and long-term loss. The system extracts disappearance direction and location data to record the last image area, movement direction, and exit position of the target face before loss, providing directional guidance for subsequent spatiotemporal memory recapture. The system extracts tracking status marker data to mark the current process as being in normal tracking, temporary loss, or cruise search modes, ensuring stable switching between different states.

[0048] The state data dimensionality splitting algorithm organizes and correlates the five types of data mentioned above, generating standardized state category classification results, and establishes a spatiotemporal feature association table corresponding to both time and space dimensions, binding the five pieces of information—position, attitude, time, direction, and state—at the same moment into a complete record. Simultaneously, the algorithm establishes a loss-recapture label mapping relationship based on different data combination characteristics, assigning a corresponding recapture strategy label to each loss scenario.

[0049] When the target face moves from the right side of the screen outwards and eventually disappears, the system records in real time that the face is located in the right-side area of ​​the screen, the gimbal attitude is horizontally facing right, the loss trigger time is the current system time, the disappearance direction is the right side of the screen, and the tracking status flag switches from normal tracking to loss status. These five data points are integrated by the algorithm to form a complete spatiotemporal correlation record and generate a corresponding recapture label for disappearance on the right side, providing complete data support for subsequent priority right-side reverse search.

[0050] Based on a pre-defined target face loss and recapture decision strategy, the system uses a data normalization algorithm to unify and logically calibrate the state category classification results, spatiotemporal feature association tables, and loss / recapture label mapping relationships. This ensures consistency in field definitions, numerical ranges, and temporal references across all data, guaranteeing stable execution of subsequent grouping, calculation, and control processes. The normalization process comprises three core operations. First, data format unification converts different types of data, such as face location, gimbal angle, timestamps, orientation markers, and status identifiers, into standard data of the same length and encoding format. Second, temporal alignment synchronizes and calibrates the loss trigger time, location update time, and gimbal response time using the system clock as a reference, eliminating temporal offset errors. Third, label mapping normalization maps different recapture methods for different loss scenarios to standard instruction labels, facilitating direct calls from subsequent modules. After normalization, the system divides the search area into different levels of precision based on a pre-set loss duration threshold using a regional granularity partitioning algorithm. Short-term loss corresponds to a smaller local area, while long-term loss corresponds to a larger global area. Meanwhile, based on the hierarchical recapture rules, the system determines two types of working parameters through a cycle and interval configuration algorithm: the in-situ recapture cycle under short-term loss conditions, and the global cruise interval under long-term loss conditions.

[0051] The system sets a loss duration threshold of three seconds. Losses of three seconds or less are classified as short-term losses, while losses exceeding three seconds are classified as long-term losses. For short-term losses, the system sets a shorter in-situ recapture cycle, allowing the gimbal to perform small-area, high-frequency reciprocating searches near the last disappearance location. For long-term losses, the system sets an appropriate global patrol interval, ensuring the gimbal smoothly scans the entire field along a fixed path, avoiding target omissions due to excessively fast or slow speeds.

[0052] If the target face disappears briefly within two seconds, the system determines it as a short-term loss and initiates a small-scale recapture in place. The gimbal moves back and forth slightly left and right in the central area, with a cycle of 0.5 seconds. If the target does not reappear for more than three seconds, the system automatically switches to a long-term loss mode. The gimbal begins a full-area patrol in a vertical, horizontal, and vertical sequence, with a patrol interval of one second, ensuring complete coverage of the field of view and a stable search rhythm. Through standardized processing and hierarchical parameter settings, the system provides clear, stable, and quantifiable execution standards for different loss scenarios, providing a reliable data foundation and control basis for the recapture process.

[0053] Based on the granularity standard of the recapture area and the data update cycle, the system classifies, collects and restructures five types of status data after normalization: real-time face location, gimbal attitude, loss trigger time, disappearance direction location, and tracking status marker. These data are then combined into four types of dedicated subsets and combined into recapture control batches that can be directly scheduled and executed. This enables precise matching of different recapture logics with dedicated data, improving search efficiency and positioning accuracy.

[0054] The system constructs four datasets: The first is a short-term loss recapture dataset, consisting of small-scale face offset data, gimbal near-field pose data, and short timestamp sequences. This dataset supports in-situ, small-area searches in short-term loss scenarios. The data focuses on the central area where the target last stopped, with high sampling density and a small area, suitable for rapid, iterative recapture. The second dataset is a long-term loss cruise dataset, consisting of full-field-of-view coordinates, gimbal full-angle range, and complete time series. This dataset supports full-field scanning in long-term loss scenarios. The data covers all rotatable areas of the gimbal and uses uniform sampling to ensure complete coverage without blind spots. The third dataset is a spatiotemporal memory recapture dataset, consisting of the target's last disappearance location, direction of movement, and corresponding gimbal pose. It has clear directional orientation and is used for high-priority direction-based searches. The fourth dataset is a full-field search dataset, containing panoramic image partition data and full-path cruise trajectory data. This dataset is used to initiate a full-coverage search when spatiotemporal memory recapture yields no results, serving as a final fallback search mechanism. The recapture data grouping and integration algorithm binds the above four types of sub-data sets in terms of time sequence and logic to form a unified recapture control batch, so that each recapture strategy has its own dedicated data support, avoiding misjudgment and inefficiency caused by data mixing.

[0055] In actual operation, when the duration of target face loss is less than a set threshold, the system determines it as a short-term loss and automatically calls the short-term loss recapture subset, controlling the gimbal to perform a small-range, high-frequency reciprocating search near the target's disappearance point. When the loss duration exceeds the threshold, the system determines it as a long-term loss and prioritizes calling the spatiotemporal memory recapture subset, controlling the gimbal to perform a reverse priority search along the direction of target disappearance; if the directional search is unsuccessful, then the long-term loss cruise subset is called, and a full-domain scan is performed according to the planned path.

[0056] If the target disappears from the right side of the screen and does not reappear within two seconds, the system determines it as a short-term loss and uses the short-term loss recapture subset to perform a small, back-and-forth search in the center-right area of ​​the screen. If the target is not found after more than three seconds, the system switches to long-term loss mode. It first searches outwards to the right based on the spatiotemporal memory recapture subset. If the search is unsuccessful, it activates the long-term loss cruise subset to perform a full-area scan in all directions until the target is found.

[0057] The system employs a recapture priority ranking algorithm to intelligently sort the generated recapture control batches. The ranking process comprehensively considers three core indicators: loss duration, disappearance direction, and target reappearance probability, assigning an execution order to different recapture strategies to ensure the highest efficiency in recovering target faces. The recapture priority ranking algorithm uses a weighted scoring mechanism, assigning weight coefficients to each of the three indicators. Shorter loss durations, more clearly defined disappearance directions, and higher target reappearance probabilities all receive higher weights. The algorithm weightedly sums the scores from the three dimensions to obtain a comprehensive priority score for each recapture batch, arranging the batches by score from highest to lowest to ensure that the recapture method with the highest probability and fastest speed is initiated first.

[0058] The system combines the tracking closed-loop synchronization rhythm with the hierarchical recapture timing requirements, and uses a cruise path planning algorithm to determine a unique execution order and search path for each recapture batch, generating a complete and orderly cruise plan. This plan includes five key pieces of information: the recapture batch number, detailed composition of the corresponding search area, the cruise timing plan for gimbal rotation, data synchronization verification rules, and hierarchical compensation rules corresponding to the duration of data loss. All rules are consistent with the previously established criteria for judging short-term and long-term data loss, ensuring logical consistency throughout the entire process.

[0059] The system strictly follows an ordered cruise protocol to sequentially initiate recapture actions, adhering to the principles of prioritizing high-probability targets over larger areas, and directional targets over the entire field of view, thus avoiding redundant searches and wasting system resources. The system prioritizes recapture based on spatiotemporal memory direction, using the target's last disappearance location and direction for a reverse search; next, it performs in-situ recapture, repeatedly searching a small area at the target's last appearance location; finally, it initiates a full-field cruise, systematically scanning the entire field of view covered by the gimbal. For each batch of recapture actions, the system controls the gimbal rotation according to a pre-set sequence, simultaneously performing face detection and identity verification to check if the registered target appears in the current frame.

[0060] If the target face disappears from the left side of the screen for two seconds without reappearing, the system determines it as a short-term loss. The recapture priority ranking algorithm prioritizes the spatiotemporal memory recapture batch based on three criteria: clear disappearance direction, short loss duration, and high target reappearance probability. The system first controls the gimbal to search directionally in the left-side outer area. If the target is not found, it switches to the second priority, in-situ recapture, performing small, iterative searches near the center of the screen. If still no results are found, and the loss duration exceeds three seconds, the system enters long-term loss mode, initiating the third priority, full-area navigation. It scans the entire field in a left-right-up-down sequence, performing a face matching verification after each rotation, until the registered target is detected again and normal tracking resumes.

[0061] The face tracking closed-loop control engine is the core decision-making unit for achieving face loss and recapture. This engine consists of four functional layers: a feature input layer, a weighted fusion layer, a policy adaptation layer, and a signal output layer. Data is transferred between layers via sequential forward connections, with the output of the upper layer directly serving as the input of the lower layer. It does not include backfeedback or training stages and operates entirely in a lightweight inference manner. The feature input layer serves as the engine's data entry point, uniformly receiving three external input data streams: real-time face location data, gimbal pose data, and target disappearance spatiotemporal data. This layer performs format standardization, temporal alignment, and numerical normalization on the input data, providing a standard input source for subsequent calculations.

[0062] The weighted fusion layer receives the regularized data output from the feature input layer and employs a multi-feature weighted fusion algorithm to assign independent weight coefficients to three key information components: short-term loss state, long-term loss state, and spatiotemporal memory feature. The algorithm dynamically adjusts the weight allocation based on the loss duration: increasing the weight for in-situ recapture in short-term loss states and increasing the weight for global navigation in long-term loss states. Regardless of the loss scenario, a high fixed weight is assigned to the spatiotemporal memory feature to strengthen the direction-first search logic. The weighted fusion layer outputs a comprehensive decision feature value through weighted summation, serving as the basis for strategy selection.

[0063] The strategy adaptation layer matches and switches the corresponding recapture mode based on comprehensive decision feature values. Dynamically adaptable modes include in-situ recapture mode, full-domain cruise mode, and direction-priority recapture mode. The strategy adaptation layer combines the target's historical movement trajectory, the gimbal's horizontal and vertical movement range, and actual tracking scenario constraints to finalize execution parameters such as recapture path, rotation step size, response speed, and search interval. The signal output layer encapsulates the execution parameters generated by the strategy adaptation layer into a directly executable face recapture control signal according to the UVC or VISCA protocol format. This signal carries spatiotemporal memory information and is used to precisely drive the gimbal to complete search actions within a specified direction and range.

[0064] During operation, the engine continuously inputs face location, gimbal posture, and spatiotemporal data of target disappearance. It learns the mapping between data and recapture strategies through forward computation and calculates optimal recapture execution parameters in real time according to a weighted fusion logic based on short-term loss, long-term loss, and spatiotemporal memory. Throughout the process, the engine prioritizes recapture based on the target's disappearance direction, assigning higher weight to spatiotemporal memory recapture and controlling the gimbal to rotate in the opposite direction of the target's disappearance. If no target is detected within the directional search cycle, the engine automatically switches to full-field cruise mode, controlling the gimbal to complete a full-field scan along a preset path. Once a face matching the registered template is detected, the engine immediately terminates the recapture process, outputs a tracking recovery command, and switches the gimbal and algorithm back to normal centered tracking, completing the full closed loop of detection-tracking-loss-recapture.

[0065] The target face disappears from the right side of the screen for two seconds. The feature input layer obtains data such as the disappearance location being on the right, the gimbal's orientation being rightward, and the disappearance time being the current moment. The weighted fusion layer determines it as a short-term loss and assigns a higher weight to the spatiotemporal memory feature. The strategy adaptation layer selects a direction-priority recapture mode and generates execution parameters for a leftward reverse search. The signal output layer outputs a leftward rotation control signal according to the VISCA protocol, driving the gimbal to search to the left. If the target is not found after three seconds of continuous searching, the engine automatically determines the loss state as a long-term loss, reduces the spatiotemporal memory weight, increases the global cruise weight, switches to a global scanning mode, and controls the gimbal to search in an orderly manner left, right, up, and down. When the registered face is detected again, the engine immediately resumes the normal centering tracking process, completing a full closed-loop recapture.

[0066] The S105 optimizes tracking stability, anti-interference capabilities, and device compatibility through extended mechanisms such as multi-target priority binding, dynamic adaptive tracking sensitivity, hot-swappable tracking without loss of tracking, automatic multi-protocol compatibility, secure mapping tracking, anti-sniping lock protection, target static sleep mode, and high-speed focus tracking.

[0067] In one implementation, the system supports the registration and priority binding of multiple specified face samples. A multi-target priority management algorithm is used to achieve ordered target tracking. The algorithm structure includes a target registration unit, a priority encoding unit, a target detection and filtering unit, and a tracking switching unit, which are connected sequentially. Users upload multiple face photos sequentially. The system assigns an independent number and priority score to each face, with the primary target having the highest score, and the scores of secondary targets decreasing sequentially. The system detects all faces in the frame in real time and compares them with the registered template, retaining only the valid target with the highest score. When the highest priority target disappears from the frame, the system automatically switches the tracking target to the visible target with the highest current score. When all registered targets are invisible, the system performs a full-domain cruise search in descending priority order. Key parameters include the maximum number of registered targets, the priority score range, the number of frames for determining target disappearance, and the switching delay coefficient. Users register a primary target (number one) and a secondary target (number two). The system sets the primary target's priority to level ten and the secondary target's priority to level five. While the primary target is in the frame, the system continuously outputs its coordinates. After the primary target disappears for more than three frames, the system automatically switches to tracking the secondary target. After all targets disappear, the system prioritizes scanning the areas where the primary target frequently appears, and then scans the areas of secondary targets to ensure that the tracking is orderly and not chaotic.

[0068] The system employs a motion rate estimation and sensitivity adaptive algorithm to dynamically adjust gimbal control parameters based on the target's actual movement. The algorithm comprises a continuous frame displacement calculation unit, a velocity grading unit, a parameter mapping table unit, and a gimbal command generation unit. The algorithm performs interpolation calculations on the face center coordinates of three consecutive frames to obtain lateral and longitudinal displacements, and calculates the movement speed by combining this with the frame interval. The system categorizes speed into four levels: stationary, low-speed, medium-speed, and high-speed. At low speed, the gimbal rotation step size is reduced, decreasing the response speed; at medium speed, standard parameters are used; at high speed, the step size and rotation speed are increased. When the image jitter exceeds a threshold, anti-shake constraints are activated, stopping large-angle gimbal movements. Key parameters include the displacement calculation frame count, velocity grading threshold, anti-shake trigger threshold, and gimbal rotation speed coefficient. When the target walks slowly (less than five pixels per second), the gimbal makes smooth, fine-tuning adjustments in increments of 0.5 degrees per step. When the target walks quickly (more than fifteen pixels per second), the gimbal follows rapidly in increments of 2 degrees per step. When significant image jitter occurs, the gimbal locks its current angle and does not follow the shaking to prevent the target from flying out of the frame.

[0069] The system employs a link status monitoring and context saving algorithm to ensure uninterrupted tracking even after camera disconnection and reconnection. The algorithm includes a hardware link monitoring unit, a runtime context storage unit, a reconnection recovery verification unit, and a process continuation unit. The system continuously monitors the camera's USB enumeration status and video stream reception status. When the link is lost, the system immediately saves the currently registered face template, the pan-tilt horizontal and vertical angles, the tracking status, the time lost, and the direction of disappearance. When the camera is reconnected, the system does not restart, re-register, or reset the process; it directly reads the stored context data and restores the tracking state to its state before the disconnection. Key parameters include the link detection cycle, context storage refresh frequency, and reconnection verification timeout. If the user unplugs the camera during tracking, the system saves the target template and pan-tilt angle as 10 degrees to the right of the horizontal. After replugging the camera, the system directly continues tracking the same target from that angle without interruption, jumps, or re-enrollment of the face.

[0070] The system employs a protocol detection and command adaptation algorithm to automatically identify and switch between multiple protocols, including UVC, VISCA, and PELCO. The algorithm includes a protocol detection frame sending unit, a response parsing unit, a protocol matching unit, and a command library switching unit. The system sequentially sends standard protocol detection commands to the camera, determining the supported protocol type based on the returned response format, checksum, and control word. Upon identification, the system automatically loads the corresponding command encapsulation format, verification rules, and communication timing, requiring no manual configuration. PTZ control commands are automatically packetized and sent according to the matched protocol, ensuring correct parsing and execution. Key parameters include the detection command sequence, response timeout, protocol matching feature code, and command header format. When connected to a VISCA protocol PTZ camera, the system encapsulates rotation commands in the VISCA format after identification. When connected to a PELCO-D protocol device, the system automatically switches the command structure to maintain normal control; it is plug-and-play and requires no setup.

[0071] The system employs a golden ratio composition area constraint algorithm to stably keep the target face within the safe area of ​​the image. The algorithm includes a safe area division unit, a real-time deviation calculation unit, a compensation and correction decision unit, and a gimbal fine-tuning unit. The system divides the image horizontally and vertically into a central safe area, an edge transition area, and a boundary-crossing danger area. The safe area is 60% of the center of the image, representing an ideal composition position. The algorithm calculates the deviation between the face center and the safe area center in real time. If the deviation is small, no action is taken; if the deviation enters the transition area, a small compensation is initiated; if the deviation approaches the danger area, a rapid pullback is performed to bring the target back to the safe area. Key parameters include the safe area percentage, the transition area threshold, the compensation step size, and the pullback speed coefficient. When the face approaches the right edge of the image, entering the transition area, the gimbal slowly fine-tunes to the left to bring the face back to the central safe area, avoiding edge-grabbing, out-of-frame, and image jitter.

[0072] The system employs an exclusive target masking algorithm to achieve strong locking, no screen-stealing, and no jumps. The algorithm includes a multi-face detection unit, a feature comparison unit, a non-target masking unit, and a tracking-and-hold unit. When the system detects multiple faces, it performs feature comparison on each face, retaining only those matching the registered template. The coordinates, size, and movement information of other faces are discarded and not included in the gimbal's calculations. When others approach, obstruct, or cross the screen, the system does not respond, switch, or jitter. Key parameters include the feature comparison threshold, the number of effective masking frames, the interference judgment area ratio, and the jump suppression coefficient. If someone crosses in front of the target or approaches from the side, the system only outputs the original target's coordinates, and the gimbal maintains its original trajectory without turning, jumping, or switching the tracked object.

[0073] The system employs a static state determination and low-power scheduling algorithm to enable the gimbal to sleep when the target is stationary and wake up when it moves. The algorithm includes a multi-frame position variance calculation unit, a static threshold determination unit, a gimbal sleep control unit, and a motion wake-up unit. The algorithm continuously calculates the variance of the face coordinates for five frames; if the variance is less than a set threshold, the target is considered stationary. After a set duration of static state, the system stops all gimbal rotation, shuts down unnecessary computational tasks, and enters a low-power sleep mode. When the variance exceeds the wake-up threshold, the target is determined to have resumed movement, and the system immediately exits sleep mode and restarts the tracking process. Key parameters include the number of decision frames, static variance threshold, sleep trigger time, and wake-up sensitivity. If the target remains stationary for more than three seconds, the gimbal stops rotating. When the target begins to move, the coordinate variance increases, and the system instantly wakes up the gimbal to resume centered tracking, achieving a balance between power saving and sensitivity.

[0074] The system employs a rapid motion detection and focus-tracking mode switching algorithm to achieve strong tracking of high-speed moving targets. The algorithm includes an instantaneous velocity calculation unit, a focus-tracking trigger decision unit, a high-frequency control unit, and a gimbal speed-up unit. The algorithm calculates instantaneous velocity through continuous frame displacement. When the velocity exceeds a conventional motion threshold, it is determined to be fast running, and the system automatically enters high-speed focus-tracking mode. In this mode, the face resolution frequency, gimbal command transmission frequency, rotation speed, and response priority are increased to ensure the target does not leave the frame. Key parameters include the focus-tracking trigger velocity threshold, high-frequency mode frame rate, gimbal speed-up ratio, and focus-tracking duration. When the target switches from walking to fast running, and the speed exceeds the threshold, the system immediately increases the gimbal rotation speed by 50% and increases the position detection frequency to achieve high-speed tracking, ensuring the target does not leave the frame.

[0075] The system integrates eight extended capabilities into the main tracking process through a comprehensive scheduling algorithm based on an extended mechanism. The algorithm takes priority rules, sensitivity parameters, hot-swappable context, protocol type, mapping constraints, anti-interference rules, sleep strategies, and focus conditions as input parameters, and performs real-time corrections across the entire process, including face localization, gimbal control, anomaly handling, and loss reacquisition. The algorithm follows constraint priority rules, ensuring that core constraints such as safe mapping, anti-interference, and image stabilization are executed first, while also being compatible with adaptive adjustment and extended functions. Ultimately, it outputs stable, robust, and widely adaptable tracking results, achieving overall performance optimization.

[0076] The S106 imports real-time face position, gimbal attitude, loss status, and motion feature data into the hierarchical control engine. It adopts multi-parameter adaptive adjustment and iterative optimization, combined with the mapping calibration and anti-shake compensation modules, to generate highly robust specific face automatic cruise lock tracking results.

[0077] In one implementation, the system uses four types of data—real-time face location data, gimbal pose data, loss status data, and target motion feature data—as unified inputs, and connects them to a hierarchical control engine for integrated decision-making. This engine is a local lightweight inference engine that does not rely on cloud computing. It employs forward weighted computation and parameter iteration logic throughout the process, forming a complete data connection and execution loop with the face detection, exclusive identification, gimbal closed-loop control, and hierarchical recapture processes described above.

[0078] The hierarchical control engine employs a multi-dimensional feature weighting algorithm. This algorithm uses four types of input data as its computational basis, assigning independent weight coefficients to each type of data based on its actual contribution to tracking stability, response speed, and anti-interference effect. Weight allocation follows a priority principle of tracking performance: face offset data weights are used to adjust the gimbal response intensity; gimbal attitude data weights are used to ensure rotational smoothness; lost state data weights are used to trigger tiered recapture strategies; and motion feature data weights are used to match sensitivity and focus tracking modes. Within each frame's computation cycle, the engine performs weighted fusion calculations on the real-time values ​​of the four feature types and their corresponding weights to obtain comprehensive control parameters. Based on these parameters, it dynamically updates the gimbal adjustment step size, response speed, compensation threshold, recapture timing, and other execution quantities, ensuring that the output control commands align with the actual tracking needs of the current scene.

[0079] The specific execution rules of the multi-dimensional feature weighting algorithm are as follows: the greater the face offset, the more dynamically the position offset weight is increased, and the engine outputs a larger gimbal adjustment step size and a higher response speed to achieve rapid repositioning correction. When the target's movement speed is low or it is stationary, the motion feature weight is increased, and the engine simultaneously reduces the gimbal rotation frequency and movement amplitude, entering a stable tracking or sleep standby state. When the target is lost, the weight of the lost state increases, and the engine automatically switches to the hierarchical recapture logic, starting the corresponding search process in short-term or long-term mode. When the gimbal attitude approaches the limit angle, the attitude constraint weight increases, and the engine automatically limits the rotation amplitude to avoid operating beyond the range.

[0080] In actual operation, when a face deviates significantly from the center of the image, exceeding a preset safety range, the engine increases the positional offset weight, rapidly increasing the gimbal correction step size and rotation speed to quickly bring the face back to the center of the image. When the target stops moving and remains stationary for more than a set duration, the engine increases the motion / static weight, reduces the gimbal control frequency, and decreases the amplitude of movements until it enters a static sleep state, achieving a balance between stability and energy saving. When the target is briefly lost but the duration is within the short-term loss range, the engine increases the weight of the lost state and initiates a small-scale in-situ recapture process. The entire algorithm requires no training or sample learning, operating solely through numerical weighting and threshold decision-making, meeting the low power consumption and high real-time performance requirements of embedded devices.

[0081] Through the aforementioned multi-dimensional feature weighting and parameter adaptive adjustment, the hierarchical control engine can continuously output stable, smooth, and precise control commands in various scenarios such as face offset, target stillness, short-term loss, and gimbal limitation. Working in conjunction with the aforementioned extension mechanism, closed-loop control, and hierarchical recapture process, it ultimately achieves highly robust automatic cruise locking and tracking of designated faces.

[0082] In the hierarchical control engine's computational process, the system uses four tracking indicators as its foundation: face position offset, gimbal rotation angle, duration of data loss, and target motion speed. A multi-dimensional tracking control weight statistical algorithm performs real-time statistical and quantifiable calculations for each indicator, generating corresponding tracking control weights. Weight calculation employs standardized numerical mapping rules. Face position offset is normalized based on the horizontal and vertical distances from the center of the image; the greater the offset distance, the higher the weight. Gimbal rotation angle is calculated based on the difference between the current angle and the target angle; the larger the difference, the higher the weight. Duration of data loss is graded according to real-time timing values; the longer the duration, the higher the weight. Target motion speed is calculated using the difference in displacement between consecutive frames; the faster the motion speed, the higher the weight. The weights of the four indicators are independent and do not interfere with each other, serving as input parameters for subsequent model calculations.

[0083] The system synchronously inputs the above four weight data to construct a multi-parameter adaptive adjustment model. This model is a lightweight forward inference model, with an overall structure consisting of three functional layers: an input layer, a weight fusion layer, and a parameter output layer. Data is transferred between each layer using a unidirectional forward connection. The output of the upper layer is directly used as the input of the lower layer. It does not include backpropagation, feedback adjustment, or model training processes. All calculations are completed locally on the embedded device, meeting the requirements for low power consumption and real-time operation.

[0084] The input layer receives four tracking control weights, unifies the data format, normalizes the values, and aligns the timing, providing standard input for subsequent calculations. The weight fusion layer uses a weighted superposition fusion algorithm to linearly sum the four input weights according to preset coefficients, obtaining the comprehensive control parameters. The parameter output layer performs threshold judgment and interval mapping on the comprehensive control parameters, converting them into directly executable gimbal control parameters, recapture mode parameters, sensitivity parameters, and mapping compensation parameters, completing the final output. Key parameters for model operation include the weight coefficients of each of the four indicators, the parameter iteration update step size, and the output result judgment threshold. The weight coefficients are used to balance the influence of each indicator on the final result, the update step size is used to control the smoothness of parameter adjustment, and the output threshold is used to distinguish different working modes and execution actions. By reasonably configuring key parameters, the model can respond quickly and output stably in different scenarios.

[0085] Without a training process, the model outputs optimal control parameters solely through linear weight superposition and threshold mapping, achieving a deep integration of tracking logic and business scenarios. When the target moves rapidly, the weight of the motion speed indicator increases significantly. After model fusion calculation, high-speed tracking parameters are output, improving gimbal rotation speed and response frequency. When the target loss duration exceeds a set threshold, the weight of loss duration increases, and the model outputs global cruise parameters, initiating a large-scale search mode. When the face position deviates significantly, the position deviation weight increases, and the model outputs large-angle gimbal correction parameters. When the gimbal approaches its maximum rotation angle, the weight of the gimbal rotation angle increases, and the model outputs limit protection parameters.

[0086] When the target switches from normal walking to rapid running, the weight of the motion speed indicator increases from low to high. The model increases the weight of motion speed in the weight fusion layer, and the parameter output layer generates a high-speed tracking mode command. The gimbal immediately increases its rotation speed and tracking frequency to keep the target in the frame. If the target is lost for more than three seconds, the weight of the duration of loss increases, the model switches to output full-field cruise parameters, and the gimbal begins a full-field scan in the order of up, down, left, and right. Through this model, the system achieves adaptive parameter adjustment for different tracking scenarios, making tracking control more precise and stable.

[0087] After the system completes the parameter output of the multi-parameter adaptive adjustment model, it uses four core constraints as a unified optimization benchmark: face composition centering constraint, gimbal motion stability constraint, loss recapture timeliness constraint, and anti-interference locking constraint. Through constraint-guided iterative optimization algorithm, the real-time tracking control parameters are cyclically corrected and iteratively optimized so that the final output parameters simultaneously meet all the requirements of tracking accuracy, operational stability, recapture efficiency, and anti-interference stability.

[0088] The iterative optimization algorithm employs a combination of serial verification and parallel correction, with the iteration terminating when all four constraints are satisfied. Within each control frame cycle, the algorithm sequentially checks the constraint compliance of the current tracking control parameters. For parameters that do not meet the constraints, directional adjustments are performed, and the parameters are re-substituted for verification until all constraints are within the compliant range. Finally, the optimized control parameters are output. The entire iterative process is completed in real-time on local embedded hardware, requiring no cloud intervention or model training. Optimal control is achieved solely through numerical adjustment and threshold determination.

[0089] The specific judgment rules and parameter optimization methods for the four constraints are as follows: The system uses the center area of ​​the video frame as the target range and calculates the deviation between the coordinates of the face center and the coordinates of the frame center in real time. When the deviation exceeds the preset allowable range, it is determined that the composition centering constraint is not met. The iterative optimization algorithm increases the gimbal compensation step size and response amplitude according to the direction and magnitude of the deviation, quickly pulling the face back to the center area; the smaller the deviation, the smaller the adjustment amplitude, achieving smooth return to the center. The system judges whether the gimbal has jittered, jumped, or overshooted by the gimbal angle change in consecutive frames. When the angle change exceeds the smooth operation threshold, it is determined that the smoothness constraint is not met. The iterative optimization algorithm automatically reduces the gimbal rotation speed, reduces the adjustment step size, and extends the control cycle to reduce motion impact, eliminate jitter, and ensure smooth and vibration-free gimbal operation.

[0090] The system uses the recapture time and search response interval after target loss as the criteria for judgment. When the recapture response is too slow or the search cycle is too long, it is determined that the recapture timeliness constraint is not met. The iterative optimization algorithm increases the face detection frequency, shortens the cruise search interval, and speeds up the gimbal turning speed, thereby reducing the recapture time while ensuring coverage. The system performs identity comparison on all faces in the frame. When a non-target face is detected, causing tracking drift, gimbal erratic movement, or target switching, it is determined that the anti-interference locking constraint is not met. The iterative optimization algorithm strengthens the exclusion filtering rules, increases the feature matching threshold, retains only registered target data, and masks the position and feature information of all irrelevant faces to ensure that the tracked object is not interfered with. In each iteration, the iterative optimization algorithm simultaneously detects and corrects the above four constraints, prioritizing the anti-interference constraint and stability constraint, while taking into account image centering and recapture timeliness, so that the system can maintain stable tracking even in complex scenes.

[0091] When the gimbal causes image jitter due to excessive adjustment, the algorithm detects that the gimbal motion stability constraint is not met and immediately reduces the rotation speed and single-step adjustment angle to smooth the gimbal movement. When a face deviates from the center of the image to one side, the algorithm increases the horizontal compensation amplitude based on the composition centering constraint, quickly bringing the face back to the center of the image. When someone approaches, causing tracking fluctuations, the algorithm strengthens the exclusion selection based on the anti-interference locking constraint, keeping the target from switching. When the search is slow after the target is lost, the algorithm increases the scanning frequency based on the loss recapture timeliness constraint, accelerating the recapture speed. After multiple rounds of iterative corrections, all constraints are satisfied, and the system enters the optimal tracking operation state.

[0092] After completing multi-parameter adaptive adjustment and iterative optimization of four constraints, the system enters the stage of tracking control data integration and final result output. In this stage, tracking parameter matching and data encapsulation algorithms are used to match and logically verify the iteratively optimized tracking parameters, such as gimbal adjustment step size, dynamic response speed, deviation compensation threshold, and control frequency, with specified exclusive face tracking features and full-domain cruise recapture requirements, generating stable tracking control data in a unified format that can directly drive hardware execution.

[0093] The tracking control data is standardized structured data, containing four core fields, as follows: Gimbal control commands: including horizontal / vertical rotation angle, rotation speed, step size, and start / stop signals, conforming to the UVC / VISCA / PELCO protocol format; Tracking status markers: including status codes such as normal tracking, temporary loss, short-term re-capture, full-domain cruise, target lock, and sleep / standby; Recapture mode commands: including mode codes such as in-situ re-capture, spatiotemporal memory direction search, full-domain cruise, and priority traversal search; and Composition calibration parameters: including face center coordinates, safe area deviation value, compensation correction amount, and golden composition area constraint value.

[0094] The system uses a data integration and scheduling algorithm to package the above four types of fields into a fixed frame structure, adding frame sequence number, timestamp, and checksum to form a complete output frame structure. The integration process strictly follows local offline computing rules, does not transmit images externally, does not rely on the cloud, and all data processing is completed within the embedded chip. The final output of highly robust specific face automatic cruise locking and tracking results has four core capabilities: output commands can directly drive PTZ gimbal movements without secondary parsing; maintains stable image without jumps in scenarios such as target movement, jitter, slow speed, and fast speed; does not switch targets when multiple people are on the same screen, occluded, crossing, or approaching; supports short-term loss recapture, long-term loss cruise, hot-swappable recovery, and automatic switching of multiple protocols. This result is continuously output at a fixed frame rate, forming a complete closed loop from "face registration - real-time detection - exclusive recognition - gimbal closed loop - loss recapture - optimized output".

[0095] When multiple people interfere with the image, the target moves slightly, and a short-term loss and recovery has just occurred, the system outputs the following information: Gimbal control command: Horizontal +0.5° / step, low speed, anti-shake enabled; Tracking status marker: Normal tracking, target locked; Recapture mode command: Recapture off, enter normal tracking; Composition calibration parameters: Face center offset <5%, no compensation, the gimbal will make smooth fine adjustments accordingly to keep the target in the golden composition area and completely ignore interference from others.

[0096] When the target is lost for more than 3 seconds, the system outputs the following: PTZ control command: full-domain cruise, 120° left and right scan; tracking status marker: long-term loss; recapture mode command: spatiotemporal memory priority + full-domain cruise; composition calibration parameters: full-screen search, no composition constraints, until the registered face is re-identified, then automatically switch back to center tracking. Through the above data matching, encapsulation, integration, and output, the system achieves offline, highly robust, anti-interference, and recaptureable full-process locking and tracking of a specified face, completing the entire process from data input to final execution.

[0097] like Figure 2 As shown, an automatic cruise face-locking processing device includes: The acquisition module 201 is used to acquire core data of the entire process of automatic cruise locking and continuous tracking of a specified face, providing basic data support for subsequent face registration, real-time detection, gimbal control and lost recapture; The face registration and detection module 202 is used to process based on local lightweight algorithms and camera plug-in architecture, complete the upload of target face samples and feature registration, and accurately locate the target face and filter out irrelevant face interference through real-time image acquisition, face comparison and matching and exclusive detection. The gimbal control module 203 is used to process the target face position information, drive the gimbal attitude adjustment through the PTZ control algorithm and rely on the UVC / VISCA protocol to ensure that the target face is stably kept in the center of the image. The lost and recapture module 204 is used to build a face loss detection mechanism, record the spatiotemporal information of the target disappearance, and start a hierarchical recapture strategy. For short-term loss, it performs on-site small-range recapture, and for long-term loss, it enters full-domain intelligent cruise, forming a complete closed loop of detection-tracking-loss-recapture. The extended optimization module 205 is used to optimize tracking stability, anti-interference and device compatibility through extended mechanisms such as multi-target priority binding, dynamic adaptive tracking sensitivity, hot-swappable tracking without loss, multi-protocol automatic compatibility, safe mapping tracking, anti-snatching lock protection, target static sleep, and high-speed focus tracking. The tracking generation module 206 is used to import real-time face position, gimbal attitude, loss status, and motion feature data into the hierarchical control engine. It adopts multi-parameter adaptive adjustment and iterative optimization, combined with the composition calibration and anti-shake compensation modules, to generate highly robust specific face automatic cruise lock tracking results.

[0098] A computing device includes a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute any automatic cruise face locking processing method.

[0099] The methods and / or embodiments in this application can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by a processing unit, it performs the functions defined in the methods of this application.

[0100] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0101] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application.

Claims

1. A method for automatically cruise and lock onto a face, characterized in that, include: Acquire core data for the entire process of automatic cruise control and continuous tracking of a specified face; Based on a local lightweight algorithm and a direct-plug camera architecture, the target face sample is uploaded and its features are registered. Through real-time image acquisition, face comparison and matching and exclusive detection, the target face is accurately located and irrelevant face interference is filtered out. Based on the target face position, the PTZ control algorithm is used to drive the gimbal attitude adjustment through the UVC / VISCA protocol to ensure that the target face is stably kept in the center of the image; A face loss detection mechanism is constructed to record the spatiotemporal information of the target disappearance and to initiate a hierarchical recapture strategy. For short-term loss, a small-scale on-site recapture is performed, while for long-term loss, full-domain intelligent navigation is initiated, forming a complete closed loop of detection, tracking, loss and recapture. Through extended mechanisms such as multi-target priority binding, dynamic adaptive tracking sensitivity, hot-swappable tracking without loss, multi-protocol automatic compatibility, secure mapping tracking, anti-snatching lock protection, target static sleep mode, and high-speed focus tracking, tracking stability, anti-interference and device compatibility are optimized. Real-time face location, gimbal attitude, loss status, and motion feature data are imported into the hierarchical control engine. Multi-parameter adaptive adjustment and iterative optimization are adopted, combined with the mapping calibration and anti-shake compensation modules, to generate highly robust automatic cruise locking and tracking results for specific faces.

2. The automatic cruise face locking processing method according to claim 1, characterized in that, Based on a local lightweight algorithm and a direct-plug-in camera architecture, the system completes the uploading and feature registration of target face samples. Through real-time image acquisition, face comparison and matching, and exclusive detection, it accurately locates the target face and filters out irrelevant face interference, including: Feature extraction processing is performed on the local operating environment, camera access link, and target face registration process to generate local lightweight algorithm features, camera direct connection features, and specified face sample registration features; Feature extraction processing is performed on the real-time image acquisition, face comparison and matching, and exclusive detection processes to generate real-time video frame acquisition features, face identity comparison features, and specific face exclusive recognition features. Among them, the real-time video frame acquisition features include streaming data of continuous camera image acquisition, frame synchronous processing, and real-time position parsing; the face identity comparison features include matching data for calculating the similarity of features between the face to be identified and the registered face and determining identity consistency; and the specific face exclusive recognition features include filtering data that only locks the target face, ignores irrelevant faces, and resists interference from multiple people. Based on specified face locking and tracking constraints, the algorithm combines local lightweight algorithm features, camera direct connection features, specified face sample registration features, real-time video frame acquisition features, face identity comparison features, and specific face exclusive recognition features for analysis and processing. When performing tracking and positioning, it verifies whether the target is a registered face and meets the exclusive locking requirements, and generates accurate target face positioning results.

3. The automatic cruise face locking processing method according to claim 1, characterized in that, Based on the target face position, the PTZ control algorithm, relying on the UVC / VISCA protocol, drives the gimbal attitude adjustment to ensure that the target face is stably kept in the center of the image, including: Based on the real-time positioning coordinates of the target face, the motion characteristics of the PTZ gimbal, the UVC / VISCA protocol adaptation specifications, and the tracking response time, the gimbal attitude control and closed-loop tracking strategies are determined collaboratively. The system sets the gimbal adjustment step size, dynamic response speed and position deviation compensation threshold based on the face image offset dimension, gimbal rotation angle range, real-time tracking accuracy requirements and stable composition constraints. The system will adjust the step size, response speed, compensation threshold to match the camera's acquisition capabilities, and calibrate and determine the hierarchical control frequency for face position resolution, gimbal command transmission, and posture feedback verification. The system analyzes face location data in real time at a determined frequency, and completes protocol instruction encapsulation, PTZ control transmission and position deviation verification according to closed-loop tracking timing rules. If an error occurs such as excessive face offset, PTZ response lag, protocol compatibility anomaly, or excessive image deviation, the system reports the anomaly to the upper-layer module and triggers a step-up mechanism. speed The threshold three-in-one adaptive readjustment mechanism re-optimizes the gimbal adjustment parameters, control frequency and deviation compensation strategy until the face position meets the requirements for stable tracking with the image centered.

4. The automatic cruise face locking processing method according to claim 1, characterized in that, A face loss detection mechanism is constructed, recording the spatiotemporal information of the target's disappearance, and initiating a tiered recapture strategy. For short-term loss, a small-scale on-site recapture is performed; for long-term loss, full-domain intelligent navigation is initiated, forming a complete closed loop of detection-tracking-loss-recapture, including: The target face tracking state dataset is classified and split, including real-time face location data, gimbal pose data, loss trigger time data, disappearance direction location data, and tracking state label data, generating state category classification results, spatiotemporal feature association table, and loss-recapture label mapping relationship; Based on the target face loss and recapture decision strategy, the state category classification results, spatiotemporal feature association table, and loss and recapture label mapping relationship are standardized. The recapture area division granularity is set in combination with the loss duration threshold, and the short-term recapture cycle and long-term cruise interval are determined according to the hierarchical recapture rules. Based on the recapture region granularity standard and update cycle, each state category is grouped and integrated to form a recapture control batch that includes short-term loss recapture subset, long-term loss cruise subset, spatiotemporal memory recapture subset, and global search subset. Prioritize the patrols for each batch of recapture control, and formulate an orderly patrol plan based on the tracking closed-loop synchronization rhythm and the hierarchical recapture timing requirements. Generate recapture process control information that includes recapture batch number, regional composition details, patrol timing plan, synchronization verification rules, and hierarchical compensation rules for lost time. The face tracking closed-loop control engine learns the mapping relationship between face position, pose information, disappearance spatiotemporal data and loss recapture strategy. It optimizes the recapture execution parameters according to the weighted fusion logic of short-term loss, long-term loss and spatiotemporal memory. Combined with the target movement trajectory, gimbal movement range and tracking scenario requirements, it dynamically adapts to in-situ recapture, full-domain cruise and direction-priority recapture modes, and generates face recapture control signals with spatiotemporal memory basis, forming a complete closed loop of detection-tracking-loss-recapture.

5. The automatic cruise face locking processing method according to claim 4, characterized in that, Real-time face location, gimbal attitude, loss status, and motion feature data are imported into the hierarchical control engine. Multi-parameter adaptive adjustment and iterative optimization, combined with mapping calibration and image stabilization compensation modules, generate highly robust specific face-based automatic cruise tracking results, including: Combining the specific requirements for automatic face cruise locking and tracking with the hierarchical control engine's computational logic, a multi-dimensional tracking feature weight-driven approach is used to complete real-time iterative adjustment of tracking parameters and generation of stable tracking output. Based on the face position offset, gimbal rotation angle, loss duration and target movement speed mapped to each tracking stage, the corresponding tracking control weights are statistically calculated and used as input parameters to construct a multi-parameter adaptive adjustment model. Based on four constraints—face composition centering constraint, gimbal motion stability constraint, loss and recapture time constraint, and anti-interference locking constraint—the real-time tracking control parameters are iteratively optimized. Generate stable tracking and control data that matches the exclusive tracking features of the specified face and meets the requirements of full-domain cruise recapture, and complete the output processing of highly robust automatic cruise locking and tracking results for specific faces.

6. A processing device for automatically cruise and lock onto a face, characterized in that, The device is configured to perform the automatic cruise face locking processing method according to any one of claims 1 to 5 by executing the executable instructions.

7. An electronic device, characterized in that, include: First processor; and memory for storing executable instructions of the first processor; The first processor is configured to execute the automatic cruise face locking processing method according to any one of claims 1 to 5 by executing the executable instructions.

8. A computing device, the device comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein, When the computer program instructions are executed by the processor, the device is triggered to execute the automatic cruise face locking processing method according to any one of claims 1 to 5.