Motion decoupling control method and system based on motion mirror big data
By aligning multi-source data and estimating global image motion and detecting the subject on camera movement data, a parameter template is generated to separate the intended camera movement from shaky disturbances. This solves the problems of misaligned multi-source data and parameter dependence on human experience in existing technologies, and achieves stable and interpretable camera movement control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING KANGHUI INTELLIGENT INNOVATION TECHNOLOGY CO LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-04-17
AI Technical Summary
Existing video stabilization and camera movement control methods suffer from several problems, including a lack of unified time reference and frame-by-frame alignment for multi-source data, difficulty in decoupling image motion leading to overlapping of intended camera movement and shake disturbances, discontinuous subject tracking and composition constraints, and parameters relying on human experience.
By collecting video, gimbal and camera control records and inertial attitude data, the data is aligned frame by frame according to a unified time reference and segmented into the database to output camera movement data. The global motion of each frame is estimated based on the camera movement data, and the subject is detected and tracked to generate composition errors. The global motion, subject and composition summary are used as search conditions to generate parameter templates. The global motion is separated into intentional camera movement components and jitter disturbance components. The final control quantity is calculated and the data is limited and archived. At the same time, the results are summarized and fed back to update the sample database index.
It achieves consistency, traceability, and long-term adaptability in camera movement control under complex scenarios, explicitly decouples the main trend of lens language from high-frequency disturbances, unifies stability suppression and compositional following in the same control output, reduces the sensitivity of manual parameter tuning, and forms a data closed loop for sustainable iteration.
Smart Images

Figure CN121888097A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing and video camera movement control technology, specifically to a motion decoupling control method and system based on big data of camera movement. Background Technology
[0002] In recent years, with the development of computational photography, video understanding, and intelligent shooting equipment, camera motion estimation, target detection and tracking, and video stabilization technologies based on image sequences have matured. Gimbals, PTZ cameras, and virtual cameras have been widely used in film and television production, live broadcasting, drone aerial photography, and mobile terminal shooting. The industry has gradually evolved from early electronic image stabilization and mechanical image stabilization to joint image stabilization schemes that integrate optical flow, feature matching, homography estimation, and inertial measurement. Furthermore, it has introduced large-scale sample training for image composition and shot language modeling, extending camera movement from experience-driven to data-driven, and from offline editing to online control.
[0003] However, existing technologies still have significant shortcomings in the engineering closed loop for camera movement control: First, multi-source data often lack a unified time reference and frame-by-frame alignment mechanism. There are sampling rate differences, time drift, and missing segments among video frames, gimbal control records, and inertial attitude data, making it difficult to stably reproduce subsequent motion estimation and control derivation. Second, existing image stabilization or camera movement control often treats image motion as a single whole, mainly using low-pass smoothing or global suppression, lacking a decomposable representation of intentional camera movement and shake disturbances, which can easily lead to over-stabilization and weaken the visual language. The problems include: 1) insufficient rhythm or stability leading to high-frequency oscillations; 2) insufficient decoupling between subject tracking and composition control and image stabilization, making it difficult to continuously constrain composition errors in scenarios with occlusion, missed detection, or rapid movement, resulting in mutual constraints between tracking and stabilization; 3) parameter configuration still relies heavily on manual experience and scene-specific adjustments, lacking a templated parameter generation mechanism based on global motion, subject motion, and composition summary as retrieval conditions, and also lacking a continuously updated sample database index system formed by feedback of operational results, thus making it difficult to achieve a consistent and recalculated camera movement decoupling control process in different scenarios. Summary of the Invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by this invention is that existing video stabilization and camera movement control methods have problems such as lack of unified time reference and frame-by-frame alignment for multi-source data, difficulty in decoupling image motion leading to overlapping of intended camera movement and shaking disturbance, discontinuity of subject tracking and composition constraints, and parameter dependence on human experience. The invention also addresses how to achieve decoupling of motion components based on camera movement big data and output control quantities that can be limited and updated by backfeeding the sample library index.
[0006] To address the aforementioned technical problems, this invention provides the following technical solution: a motion decoupling control method based on big data of camera movement, comprising acquiring video, gimbal and camera control records and inertial attitude data, aligning and segmenting them frame by frame according to a unified time reference and storing them in a database, and outputting camera movement data; estimating the global frame-by-frame motion based on the camera movement data, and simultaneously detecting and tracking the subject to generate composition errors; using the global motion, subject and composition summary as search conditions to generate parameter templates; separating the global motion into intentional camera movement components and jitter disturbance components; calculating the final control quantity according to the fusion weights of the jitter disturbance components, composition errors and intentional camera movement components, limiting and archiving them, and simultaneously summarizing and feeding back the results to update the sample database index.
[0007] As a preferred embodiment of the motion decoupling control method based on big data of camera movement described in this invention, the step of aligning and segmenting the data frame by frame according to a unified time reference includes establishing a video frame time axis and recording the frame rate and frame-by-frame timestamp; retaining the original sampling timestamps for the gimbal and camera control records and inertial attitude data; resampling the control records and inertial attitude data to a frame-by-frame time grid consistent with the video frames, with the resampling using nearest neighbor sampling and linear interpolation and recording the methods used; writing an explicit missing flag for missing samples, and truncating and removing outliers that exceed the device range and physical upper limit according to a threshold and recording the threshold.
[0008] As a preferred embodiment of the motion decoupling control method based on big data of camera movement described in this invention, the estimation of frame-by-frame global motion includes extracting and matching local feature points between adjacent frames, solving the global motion model based on robust estimation and outputting the number of inliers, the proportion of inliers, and residual statistics; when the proportion of inliers is lower than a preset threshold and the residual is abnormal, switching to dense optical flow calculation and aggregating the pixel motion vector field to obtain the global motion increment; archiving the global motion in the form of a frame-by-frame incremental sequence, and writing unreliable markers and corresponding frame ranges for unreliable frames.
[0009] As a preferred embodiment of the motion decoupling control method based on big data of camera movement described in this invention, the synchronous detection and tracking of the subject to generate composition error includes outputting the subject bounding box, center point coordinates and confidence level for each frame, and constructing an association cost based on position distance, scale change and appearance consistency to achieve cross-frame tracking; when occlusion, missed detection or confidence level is below a threshold, the system enters a prediction-preservation mode with a limited number of frames and writes a preservation flag; the system generates the coordinates of the composition target point according to preset composition rules, the composition error is the difference between the coordinates of the subject's center point and the coordinates of the composition target point, and writes a missing flag for the frame composition error when the subject result is missing.
[0010] As a preferred embodiment of the motion decoupling control method based on big data of camera movement described in this invention, the step of generating a parameter template using global motion, subject, and composition summary as retrieval conditions includes calculating amplitude quantiles, mean square values, and peak values of the global motion sequence and composition error sequence to form continuous features, and forming discrete features of scene description information, subject category, and motion type; firstly filtering candidate segments according to discrete features, and then sorting them by nearest neighbor according to the distance of continuous features to obtain a reference set; and generating a parameter template containing a separation parameter set and a fusion parameter set from the reference set. The separation parameter set includes the smoothing window length and cutoff frequency or equivalent bandwidth parameters, and the fusion parameter set includes the initial value range of fusion weights and a reference to the device limiting configuration.
[0011] As a preferred embodiment of the motion decoupling control method based on big data of camera movement described in this invention, the step of separating global motion into intentional camera movement components and jitter disturbance components includes performing robust smoothing on the global motion amount frame by frame according to the parameter template to obtain the intentional camera movement component, and using the difference between the global motion amount frame by frame and the intentional camera movement component as the jitter disturbance component; for frames with unreliable labels or missing labels, the labels are maintained and the labels are propagated to the intentional camera movement component and the jitter disturbance component; the jitter disturbance component, composition error and intentional camera movement component are written into the frame-by-frame result sequence according to the same frame time index.
[0012] As a preferred embodiment of the motion decoupling control method based on big data of camera movement described in this invention, the calculation of the final control quantity and the limiting archiving includes reading jitter disturbance components, composition error and intended camera movement components on the same frame-by-frame time grid, calculating the unlimited control quantity according to the fusion weight given by the parameter template, and performing limiting processing on the amplitude and rate of change of the control quantity according to the device limiting configuration; during archiving, the limiting trigger flag and the rollback flag are written simultaneously, and the statistical summary of the segment, the parameter group selection results and the version information are summarized and fed back to update the sample library index field.
[0013] Another objective of this invention is to provide a motion decoupling control system based on big data of camera movement. This system can estimate the global motion of each frame based on camera movement data and simultaneously detect and track the subject to generate composition errors. It uses the global motion, the subject, and the composition summary as retrieval conditions to generate parameter templates, and separates the global motion into the intended camera movement component and the jitter disturbance component. This solves the problem in current video stabilization and camera movement control methods where the inability to decouple the image motion leads to the overlap of intended camera movement and jitter disturbance.
[0014] As a preferred embodiment of the motion decoupling control system based on camera movement big data described in this invention, it includes: a data alignment and storage module, a motion estimation and decoupling module, and a fusion control backfeed module; the data alignment and storage module is used to collect video, gimbal and camera control records and inertial attitude data, and align and segment them frame by frame according to a unified time reference to generate camera movement data; the motion estimation and decoupling module is used to estimate the global frame-by-frame motion on the camera movement data, detect and track the subject to generate composition errors, and retrieve and generate parameter templates to separate the global motion into intentional camera movement components and jitter disturbance components; the fusion control backfeed module is used to fuse jitter disturbance components, composition errors and intentional camera movement components to calculate the final control quantity and limit and archive it, while summarizing the results and backfeeding back to update the sample library index.
[0015] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement a motion decoupling control method based on big data of camera movements.
[0016] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of a motion decoupling control method based on big data of camera movements.
[0017] The beneficial effects of this invention are as follows: The motion decoupling control method based on big data of camera movement provided by this invention collects video, gimbal and camera control records and inertial attitude data, aligns them frame by frame according to a unified time reference, segments and stores them in a database, forming a consistent binding of frame sequences and control / attitude sequences on the same time grid. This makes the source of input fields for subsequent calculations clear and the units and coordinate semantics consistent, thus providing a recalcible data foundation for global motion estimation, subject trajectory extraction and parameter retrieval. It estimates frame-by-frame global image motion on camera movement data and simultaneously detects and tracks subject-generated composition errors, allowing the overall image motion and subject composition constraints to be directly associated under the same frame index, used to stably express the lens motion state and composition deviation. Furthermore, it uses global motion, subject, and composition summaries as detection... The system generates parameter templates based on conditions, enabling the separation and fusion parameters to match the scene and motion intensity, reducing the sensitivity of manual parameter tuning to scene transitions. Based on this, the global motion is separated into intentional camera movement components and shake disturbance components, explicitly decoupling the main trend of the shot language from high-frequency disturbances, providing interpretable and independently adjustable inputs for subsequent control. Finally, the shake disturbance components, composition error, and intentional camera movement components are calculated into control quantities according to fusion weights and archived with amplitude limiting. This ensures that stability suppression, composition following, and style retention are uniformly constrained in the same control output and meet the device capability boundaries. At the same time, the running results are summarized and fed back to update the sample library index, forming a continuously iterative data closed loop, ensuring that camera movement control has consistency, traceability, and long-term adaptability in complex scenes. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 The first embodiment of the present invention provides an overall flowchart of a motion decoupling control method based on big data of camera movement. Detailed Implementation
[0020] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0021] Example 1, referring to Figure 1 As an embodiment of the present invention, a motion decoupling control method based on big data of camera movement is provided, comprising: S1: Collects video, gimbal and camera control records and inertial attitude data, aligns them frame by frame according to a unified time reference and segments them into the database, and outputs camera movement data.
[0022] Furthermore, in the data construction phase, the big data on camera movement is limited to a reproducible set of data composed of video frame sequences and their corresponding shooting end control data, lens parameter data, and inertial and attitude data. The acquired data includes at least: video files and their frame sequence information (resolution, frame rate, encoding timestamp or frame display timestamp); attitude and control records output by the gimbal or camera (yaw, pitch, roll angles or angular velocities, as well as obtainable fields such as zoom magnification, focus distance, aperture, shutter speed, and exposure compensation); inertial measurement data (gyroscope angular velocity, accelerometer acceleration, attitude calculation results, etc.); and if the device has an encoder or odometer, encoder angles and angular velocities are acquired simultaneously. For each acquisition device and each acquisition task, the device model, lens model, firmware version, sensor range, sampling rate, output unit, and coordinate system definition are recorded, and the above metadata is bound and stored with the acquired data to ensure that the physical meaning and unit of the same field can be distinguished on different devices.
[0023] To ensure consistent referencing of multi-source data, a unified time base is used for alignment. On the video side, a frame-level timeline is generated using timestamp information within the video container or frame-by-frame sequence numbers, recording the start time, frame rate, time base, and corresponding timestamps for each frame. When a reliable timestamp is missing from the video, an equally spaced timeline is generated using frame sequence numbers and frame rates, and this generation rule is embedded in the metadata. Control data and inertial data retain their original sampling timestamps and record the timestamp source and clock stability information. When the device supports external time synchronization or hardware synchronization triggering, the time synchronization status, synchronization period, and estimated clock deviation are recorded. When no time synchronization conditions exist, significant motion segments are selected within a preset alignment window. By comparing the peak positions of the global motion intensity curve of the video frame with those of the control angular velocity and inertial angular velocity curves, the initial time offset is estimated, and the offset and estimated interval are written into the alignment parameters of that data segment.
[0024] It should be noted that after establishing a unified timeline, control and inertial data are resampled to a frame-by-frame time grid consistent with the video frames. Specifically, each frame time is used as a sampling point. For control data, nearest neighbor sampling or linear interpolation is used to generate frame-by-frame control sequences. For inertial data, a fixed window aggregation method is used to generate frame-by-frame inertial sequences based on the sampling rate. The interpolation and aggregation methods used are recorded. Missing samples are explicitly marked and their missing locations are preserved; no implicit padding is performed. Abnormal jump values are removed or truncated based on the device's physical upper limit and range threshold, and the abnormal location, threshold, and processing method are saved in the data. For fields with unit differences, a unified unit conversion is performed. For example, angles are unified to radians or degrees, angular velocities are unified to angles per second or radians per second, and accelerations are unified to meters per second squared. The unit, dimension, and value range of each field are clearly defined in the field dictionary.
[0025] After frame-by-frame alignment, the entire footage is segmented to ensure consistent sample granularity for subsequent retrieval and training. Segmentation is based on two combined conditions: one is shot boundary conditions, which can be obtained by detecting sudden changes in image content or keyframe structural changes; the other is motion continuity conditions, where segmentation points are formed when curves for global motion, gimbal angular velocity, and subject motion intensity consistently exceed a threshold. A parameter table is fixed for the segmentation process, including at least the shortest segment duration, maximum segment duration, number of buffer frames before and after the segmentation point, whether overlap is allowed, and rules for merging and secondary segmentation. When a segment is too short or too long, adjacent segments are merged or secondary segmented according to the parameter table, and the location of the adjustment and the threshold used are recorded. For each segment, the start and end frame numbers, start and end times, frame rate, resolution, and encoding information are saved. If the video has been cropped or scaled, the cropped area and scaling ratio are saved, along with their correspondence with frame-by-frame control and inertial sequences.
[0026] Furthermore, for each segment, structured index fields are created and written into an index table for subsequent retrieval of similar camera movement samples based on scene and motion conditions. The index fields include at least scene description information (quantifiable or enumerable fields such as indoor / outdoor, lighting intensity, and background texture complexity), subject category and quantity, subject occlusion ratio range, motion type labels (follow shot, pan, push-pull, circle, rapid horizontal movement, etc., allowing multiple labels and recording their primary and secondary order), and a motion intensity summary for retrieval. The motion intensity summary can be calculated from a frame-by-frame aligned sequence, such as quantile statistics of global motion, quantile statistics of gimbal angular velocity, quantile statistics of subject center displacement velocity, and zoom change rate statistics, etc. The statistical window and quantile values are fixed in the parameter table. If a quality score field is required, the score value and scoring rule version are saved; when the score is generated by rule calculation, the rule input fields, threshold, and calculation range must be saved simultaneously to ensure that the score for the same segment can be recalculated.
[0027] In terms of storage organization, a layered archiving system is adopted, consisting of raw data, aligned data, and indexed data. The raw data layer stores video files and raw log files without irreversible processing. The aligned data layer stores the control sequence and inertial sequence after frame-by-frame alignment, and preserves the correspondence with the frame sequence, missing data flags, and interpolation methods. The indexed data layer stores the segment index table, field dictionary, segmentation parameter table, and alignment parameter table. A checksum is generated for each archived file and written to a checksum list, which includes the file path, size, generation time, checksum algorithm, and checksum result for consistency verification. For materials involving identifiable information, an anonymized version can be generated according to preset rules, and the anonymization range and parameters can be recorded, while maintaining the association and traceability between the raw data and the anonymized data.
[0028] It should be noted that, ultimately, the sample database uses fragments as the smallest retrieval unit, providing an interface capability to query and return data sets based on index fields. The returned data set includes at least the frame sequence reference for that fragment, the control and inertial sequences for frame-by-frame alignment, fragment metadata, index fields, alignment parameters, and segmentation parameters, and is constrained by a field dictionary. The field dictionary clearly defines the meaning, unit, coordinate system, missing field flag rules, and value range for each field. The field dictionary, segmentation parameter table, and alignment parameter table are all versioned, and the version number used is recorded each time the data is added to the database, ensuring that the data addition process and field meanings for any fragment can be fully reproduced.
[0029] S2: Estimate the frame-by-frame global motion based on the camera movement data, and simultaneously detect and track the subject to generate composition errors. Use the global motion, subject, and composition summary as retrieval conditions to generate parameter templates, and separate the global motion into the intended camera movement component and the shaky perturbation component.
[0030] Furthermore, in the camera motion estimation stage, using the formed sample library as the data entry point, the frame sequence references and frame-by-frame alignment control and inertial sequences of each segment are read as the smallest processing unit. Segment metadata, alignment parameters, and segmentation parameters are read simultaneously, and all fields are constrained by a field dictionary to ensure consistency in the time position, unit, coordinate system, and missing data rules for each frame. Before computation, the resolution, frame rate, encoding information, and the presence of cropped regions and scaling ratios of the segment are determined based on the segment metadata. When cropping or scaling occurs, the frame sequence is first restored to the unified image coordinate system used for alignment, or the subsequent calculation results are converted according to the cropped region and scaling ratio, and the adopted coordinate conversion rules are written into the calculation parameters of the segment to ensure that the results for the same segment can be recalculated.
[0031] Image domain preprocessing is performed on the frame sequence to ensure numerical stability of motion estimation. Preprocessing includes: color space conversion or grayscale conversion as determined by the field dictionary and parameter table; construction of a fixed-size multi-scale pyramid to adapt to different motion amplitudes; effective region cropping or windowing of image edge regions according to preset rules to avoid interference from black edges and occlusion bands; if lens distortion parameters are recorded in the segment metadata or can be mapped to distortion configuration from device and lens models, radial and tangential distortion corrections are performed during preprocessing, and the parameter source, applicable resolution, and interpolation method used for distortion correction are recorded; when distortion parameters are unavailable, the uncorrected state is directly recorded.
[0032] It should be noted that observational relationships for global motion estimation are established between adjacent frames. For each pair of adjacent frames, repeatable local features are first extracted and descriptors are calculated on the reference frame. Then, feature matching is performed on the target frame. Matching can employ consistency rules such as nearest neighbor and ratio constraints. Simultaneously, a geometric consistency screening is performed on the matched pairs, for example, removing matches whose displacement exceeds the image diagonal ratio threshold and matches that clearly violate scale consistency. Robust estimation is used to solve the global motion model for the matched pairs that pass the initial screening. The robust estimation process fixes the upper limit of the number of iterations, the inlier threshold, and the confidence threshold, and saves statistics such as the number of inliers, the inlier ratio, and the residual distribution. When the inlier ratio is lower than the threshold or the residual distribution shows a multi-peak structure, the frame pair is marked as an unreliable frame pair and enters the backup calculation path. At the same time, this mark is written into the quality field of the frame-by-frame result sequence. The quality field is also constrained by the field dictionary.
[0033] The selection of the global motion model is determined based on the segment metadata and lens parameter status. For segments that do not involve significant zoom or can be considered as slow changes in zoom, affine or similarity transformations are used to describe the global motion. For segments with significant viewpoint changes or plane-dominated scenes, homography models are used to describe the global motion. After solving, the model matrix is normalized and numerical conditions are checked to eliminate solutions that are irreversible or have abnormal condition numbers. When zoom magnification is recorded frame by frame in a segment, the zoom change rate of adjacent frames is first used as a reference for model selection. When the zoom change rate exceeds a threshold, models that allow scale changes are selected first. When there are missing markers in the zoom field, the explicit missing marker rule for missing samples is followed. No implicit filling is performed on the missing segments. Instead, the model is adaptively selected for that segment based on image observations, and the result is recorded as the model selection source being image observations.
[0034] Furthermore, when feature matching cannot provide a sufficiently stable set of inliers, dense optical flow is used as a backup observation to estimate global motion. Dense optical flow calculation uses fixed pyramid layers and iteration parameters. After outputting the motion vector field for each pixel, the motion vector field is aggregated according to a preset background consistency rule to obtain the global motion. The aggregation process includes at least the removal of outlier vectors, masking of occluded areas, and weight reduction processing of large-area main motion regions.
[0035] The global motion results estimated from each pair of adjacent frames are uniformly encoded into a frame-by-frame motion sequence and bound to a frame-by-frame time grid. Specifically, for each frame, a global motion increment record relative to the previous frame is generated, along with fields such as the model type used, robust estimation interior point statistics, residual statistics, and whether dense optical flow backup paths are enabled. Simultaneously, based on segment alignment parameters, it is ensured that this motion increment record corresponds one-to-one with the frame-by-frame aligned control sequence and inertial sequence on the same frame order. If a segment has missing or duplicate frames, it is processed according to the frame timestamp and frame sequence number relationship recorded in the segment metadata: for duplicate frames, only one frame is retained, and the range of duplicate frames is recorded in the result; for missing frames, a missing frame marker is inserted into the result sequence, and the location of the missing frame is recorded. Image motion observations are not falsified by interpolation, ensuring the traceability of the motion sequence's origin.
[0036] Furthermore, in the subject detection and tracking stage, the frame-by-frame global motion sequence is used as one of the basic inputs. Simultaneously, according to the formed sample library data entry, the frame sequence reference of each segment, as well as the frame-by-frame alignment control and inertial sequences, are read as the smallest processing unit. Segment metadata, alignment parameters, and segmentation parameters are read synchronously, and all fields are constrained by the field dictionary to ensure consistency in the time position, unit, coordinate system, and missing flag rules for each frame. To maintain image coordinate consistency, if distortion correction, cropping region conversion, or scaling ratio conversion is performed on the frame sequence, the same coordinate conversion rules are used to ensure that the subject position output is directly in the unified image coordinate system used for global motion estimation. When the detection end uses multi-scale input or a fixed network input size, the output result is reverse-calculated back to the unified image coordinate system according to the coordinate conversion rules before being written into the frame-by-frame result sequence, and the scaling ratio and offset used for the conversion are recorded.
[0037] Subject detection outputs candidate subject bounding boxes or keypoint sets using a frame-by-frame processing approach. Before detection, each frame undergoes an image domain preprocessing procedure, including color space conversion or grayscale conversion, multi-scale pyramid construction, and effective region cropping or windowing. Parameters such as input size, confidence threshold, and non-maximum suppression threshold are fixed in the detection parameters. When multiple candidate subjects exist within the same frame, the primary and secondary subjects are determined according to preset subject selection rules. These rules can be based on factors such as subject area, detection confidence, consistency with the subject position in the previous frame, and compatibility of the subject's motion direction with the image's motion direction. These rules are fixed in the parameter table and do not change during runtime. For the selected subject, its bounding box coordinates, center point coordinates, width, height, and area in a unified image coordinate system are output, along with the detection confidence. For secondary subjects, necessary fields are retained, and priority markers are set to avoid a lack of fallback candidates in case of occlusion or short-term false detections during subsequent tracking.
[0038] It should be noted that, in order to form a continuous subject trajectory frame by frame, the detection results are temporally correlated and tracked. Temporal correlation uses a frame-by-frame time grid as an index, the subject position, scale, and motion trend of the previous frame as prediction quantities, and the detection candidates of the current frame as observations. Correlation is completed through cost matrix minimization or threshold matching. The correlation cost includes at least positional distance cost, scale change cost, and appearance consistency cost; among which, appearance consistency can be calculated from local features or shallow appearance descriptors, and the calculation window size and update frequency are fixed in the parameter table. For frames with occlusion, missed detection, or detection confidence below the threshold, the tracking end enters a short-term prediction hold mode: in hold mode, it is allowed to predict the subject position and scale based only on the state of the previous moment within a limited number of frames, and write an explicit tag to the prediction result of this segment. The tag field is constrained by the field dictionary, and the prediction result is not written back to the detection module as a detection observation. If the hold mode continues to exceed the preset upper limit, the subject trajectory is terminated, and the number of subject tracking interruptions and the range of interruption locations for this segment are recorded in the index data layer.
[0039] After the main trajectory is generated, the motion-related quantities of the main body are calculated on a frame-by-frame time grid and aligned and bound to the global motion sequence. For each frame, the discrete approximate values of the main body center point coordinates, the displacement of the center point of adjacent frames, the displacement velocity, and the displacement acceleration are calculated. For changes in the main body scale, the relative rate of change of the bounding box area or width and height is calculated, and the sign and magnitude of the rate of change are recorded. All of the above motion-related quantities are calculated in a unified image coordinate system, and the pixel-level displacement is converted into displacement per second by combining the frame rate in the segment metadata. When the segment metadata records valid camera intrinsic parameters or can be mapped to intrinsic parameter configuration by device model and lens model, it is allowed to further convert the pixel displacement into angular displacement in the view domain. However, when the intrinsic parameters are unavailable or there is a missing flag, the explicit missing flag rule is followed, and no forced conversion is performed. Only the pixel-domain motion quantity is retained, and the unit is clearly stated as pixels or pixels per second in the field. For any velocity or acceleration metric obtained by differential calculation, the boundary frame processing rule and the missing propagation rule are adopted: when there is a missing flag or unreliable flag in an adjacent frame, the differential result of that frame is written into the missing flag and not filled by interpolation, so as to ensure that each derived quantity can be traced back to a valid observation frame.
[0040] Based on the availability of the subject's position and motion, a method for determining the composition target is provided for subsequent control fusion. The composition target is limited to a target point or target area in the image coordinate system, without introducing control commands in the device coordinate system. The composition target is generated by preset composition rules, which are fixed in the parameter table and include at least: the normalized coordinates of the target point relative to the screen width and height, the allowable deviation tolerance range, and the offset rule for reserving space in the subject's motion direction. The subject's motion direction is determined by the displacement direction of the subject's center point. When the subject's displacement direction is unstable within a short window or has missing markers, the composition target reverts to a fixed target point and a reversion marker is recorded. For each frame, the coordinates of the composition target point and the coordinate difference between the subject's center point and the composition target point are output. The coordinate difference serves as the error input in subsequent control fusion, and is expressed in pixels or normalized coordinates, with the unit and value range clearly specified in the field dictionary.
[0041] Ultimately, two types of frame-by-frame results are generated and archived into the aligned data layer: the first is a sequence of subject detection and tracking results, including the bounding box, center point, confidence score, prediction preservation marker, and necessary secondary subject candidate information for each frame; the second is a sequence of subject motion derivatives, including displacement, velocity, acceleration, scale change rate, and mapping target point and mapping error. Simultaneously, the index data layer supplements the segment with subject-related summary statistical fields, such as the subject center displacement velocity quantile, occlusion or missed detection ratio, number of preservation mode triggers, and mapping error quantiles. The statistical window and quantile values are fixed in the parameter table and a version number is recorded, making the calculation process of subject trajectory and mapping error reproducible.
[0042] Furthermore, in the style retrieval and parameter template generation stage driven by big data in camera movement, the aforementioned sample library is still used as the data entry point, and calculations and retrieval are carried out by segment as the smallest processing unit. For the current segment to be processed, the frame sequence reference of the segment and the frame-by-frame alignment control sequence and inertial sequence are read, and the segment metadata, alignment parameters and segmentation parameters are read simultaneously. All fields are constrained by the field dictionary to ensure consistency in time position, unit, coordinate system and missing label rules. At the same time, the output frame-by-frame global motion sequence, as well as the output subject detection and tracking result sequence, subject motion derivation sequence and composition error sequence are read as the main observation source for the retrieval conditions. If unreliable labels or missing labels are found in the results, the corresponding frame segments are excluded from the statistical and summary calculations, and the proportion and range of valid frames are recorded in the retrieval conditions to avoid treating missing segments as real observations in similarity calculations.
[0043] To construct a searchable fragment description vector, retrieval features are extracted from the current fragment in a unified image coordinate system and structured encoding is performed. The retrieval features include at least three categories: The first category is global motion summary features, calculated from frame-by-frame global motion sequences, including quantile statistics, mean square statistics, and short-window peak statistics for global motion amplitude, rotation components, and scale change rates. The statistical window length and quantile values are fixed in the parameter table. The second category is subject motion and composition summary features, consisting of quantile statistics, mean square statistics, and short-window peak statistics for subject center displacement velocity, displacement acceleration, subject scale change rate, and composition error. These features are only included in encoding when the subject detection and tracking result sequence has not triggered termination and the proportion of effective frames exceeds a threshold. The third category is discrete attribute features in the fragment metadata and index fields, such as scene description information, subject category and quantity, subject occlusion ratio range, and motion type label. Discrete attributes are encoded using enumerated values or multi-hot vectors, and the encoding rules are fixed in the field dictionary and parameter table. The three types of features mentioned above are normalized before entering the similarity calculation. The scale parameter used for normalization comes from the global scale table pre-statistically calculated in the sample database index data layer. It includes the minimum, maximum, quantile range or mean variance of each continuous feature dimension. The scale table is also versioned and the version number is recorded to ensure that the feature encoding of the same segment is consistent in different batches.
[0044] On the sample database side, to improve retrieval efficiency, a retrieval feature vector with the same structure as the current segment is pre-calculated and stored for each already included segment, and a retrieval index structure is established at the index data layer. The retrieval index structure includes at least: an inverted index bucketed by discrete attributes, used to quickly filter candidate segments based on conditions such as subject category, motion type label, and scene description information; and a nearest neighbor retrieval index built by continuous feature vectors, used to sort the filtered candidate set by similarity. When building the index, the continuous feature vectors adopt the same normalization rules and scale table version as the current segment, and for each already included segment, the version numbers of the field dictionary, segmentation parameter table, and alignment parameter table used to generate its feature vector are saved. This ensures that the segment range requiring recalculation can be located when the index is updated or the parameter version is upgraded, and feature vectors generated by different versions are not mixed.
[0045] The retrieval process is divided into two stages: candidate filtering and similarity ranking. In the candidate filtering stage, a candidate set is determined in the inverted index based on the discrete attribute features of the current segment. Filtering conditions must include at least consistent or compatible subject categories, consistent or compatible motion type tags, and scene description information belonging to the same enumeration set or satisfying a preset mapping table relationship. When the filtering conditions are too strict, causing the number of candidates to fall below a threshold, the conditions are relaxed level by level according to the relaxation order specified in the parameter table. For example, first relax the matching of scene description information, then relax the compatible mapping of motion type tags, and finally relax the compatible set of subject categories. The level of relaxation and the corresponding number of candidates are recorded for each relaxation. In the similarity ranking stage, continuous feature similarity is calculated within the candidate set. Continuous feature similarity is implemented using a weighted distance method, with the distance weights fixed in the parameter table. Weights can be set separately for two main categories: global motion summary and subject motion and composition summary. Weights can also be reduced based on the proportion of effective frames. When there are large areas of missing markers or mode-preserving markers in the subject detection and tracking result sequence, making the subject-related summary unusable, the similarity ranking automatically removes the subject-related feature dimensions and only uses the available fields in the global motion summary and fragment metadata to participate in the ranking. The feature dimension set marker used in this retrieval is written into the retrieval results.
[0046] After obtaining the similarity ranking results, the top few nearest neighbor segments are selected from the candidate set as a reference set, and the camera movement parameter template for the current segment is generated from the reference set. The camera movement parameter template is expressed in the form of a structured parameter table, which does not contain the control commands themselves, but rather the required separation and fusion parameters. This parameter template contains at least two types of parameters: the first type is motion component separation parameters, used to separate the frame-by-frame global motion sequence. These parameters include the smoothing window length, cutoff frequency or equivalent bandwidth parameters, robust smoothing residual threshold and upper limit of iterations, as well as backoff parameter groups when the proportion of unreliable frames is high; the second type is fusion control weights and clipping configuration parameters, which include the initial value range of the weight coefficients and the upper limit of the update step size, as well as the configuration reference corresponding to the clipping boundary given by the device capabilities. When generating the template, the corresponding parameters of each segment in the reference set are first checked for consistency to ensure that their field meanings, units, coordinate systems, and missing flag rules are consistent with the current segment; the numerical fusion of parameters with the same name adopts the median or truncated mean method, and the fusion weights are weighted according to the similarity score. The fusion rules are fixed in the parameter table. If there are missing parameters or inconsistent versions in the reference set that prevent direct fusion, then the reference fragment subset with the same version number as the current fragment field dictionary is selected for fusion first, and the number and version distribution of the reference fragments that ultimately participate in the fusion are recorded.
[0047] To ensure the parameter template matches the observation intensity of the current segment, an intra-segment self-calibration is performed after template generation. This self-calibration does not introduce new scoring or subjective judgment; it only scales or recalibrates the template parameters based on the statistical amplitudes of the frame-by-frame global motion sequence and the main motion derivative sequence of the current segment. For example, it limits the smoothing window length to a range derived from the frame rate and motion peak interval, and limits the cutoff frequency to a candidate interval derived from the spectral energy distribution of the global motion sequence. The scaling factor and the limitation boundaries are written into the template parameter record for that segment. If the segment's metadata contains frame-by-frame records of zoom magnification and the zoom change rate exceeds a threshold, a separation parameter group allowing scale variation is enabled in the template, and this enabling condition is recorded along with the threshold. When a missing marker exists in the zoom field, the explicit missing marker rule is followed; no interpolation is performed on the missing segment, and the zoom change rate is calculated only for the valid frame segments, with the percentage of valid frames recorded in the template.
[0048] It should be noted that during the motion decoupling and separation phase, the frame sequence references and frame-by-frame aligned control and inertial sequences of each segment are read as the smallest processing unit. Segment metadata, alignment parameters, and segmentation parameters are read simultaneously, and all fields are constrained by a field dictionary to ensure consistency in time position, unit, coordinate system, and missing frame rules. Simultaneously, the output frame-by-frame global motion sequence is read, along with the generated and archived camera movement parameter template. The camera movement parameter template contains the required separation parameter sets, which at least include the smoothing window length, cutoff frequency or equivalent bandwidth parameter, robust smoothing residual threshold, upper limit of iterations, and a backoff parameter set for high unreliable frame ratios. When multiple separation parameter sets exist in the template, the actual parameter set used for this segment is determined based on the frame rate, reliable flag ratio, and global motion amplitude summary statistics in the segment metadata, according to the selection rules fixed in the parameter table. The selected parameter set identifier, the values of each parameter, and their version numbers are written into the separation record field of this segment.
[0049] To ensure that the separation results and the quantities used in subsequent control fusion are in the same coordinate and unit system, a step-by-step global motion sequence is used as the motion observation input. The global motion model is not re-estimated, and no new irreversible processing is introduced into the frame sequence. A consistency check is first performed on the frame-by-frame global motion sequence. This check includes whether the motion increment amplitude exceeds a reasonable range derived from the resolution and frame rate, whether there are abnormal abrupt changes between consecutive frames, and whether the distribution of unreliable and missing markers exceeds a threshold. Frame segments that trigger the check are only explicitly marked, and their frame range and triggering conditions are recorded; no interpolation, patching, or replacement is performed on these data segments. Subsequently, the effective frame range is determined based on alignment parameters, and separation calculations are performed only within the effective frame range. For the frame times where unreliable or missing markers are located, their markers are retained, and the corresponding markers are synchronously written into the output field.
[0050] Within the effective frame range, component separation is performed on the frame-by-frame global motion sequence. The separation method adopts a process of first obtaining the smoothed main trend and then obtaining the high-frequency perturbation from the residual: robust smoothing is performed on the frame-by-frame global motion quantity according to the smoothing window and cutoff frequency given by the camera movement parameter template. The residual threshold and the upper limit of the number of iterations for robust smoothing are derived from the template parameters; the smoothed output is used as the main trend sequence, and the difference between the main trend and the original sequence is used as the residual sequence. If the zoom magnification is recorded frame-by-frame in the segment metadata and the zoom change rate exceeds the template threshold, the separation parameter group that allows scale changes in the template is activated; when there is a missing flag in the zoom field, the explicit missing flag rule is followed, and the zoom change rate is judged and the parameter group is selected only in the effective frame segment, without inferring from the missing segment. Consistent window and frequency band parameters are used for each dimension of motion component to avoid time alignment deviation caused by using different parameters for different dimensions; when the output global motion quantity is in vector form, the same smoothing and residual calculations are performed on each component according to the component order specified in the field dictionary, and the vector dimension and component definition consistent with the input are retained in the output.
[0051] The global motion quantity is decomposed using a motion decoupling expression, as follows:
[0052] in, This represents the frame time index on the frame-by-frame time grid; Indicates the output at frame time The global motion quantities are in vector form, and their components are explicitly mapped and assigned units by a field dictionary. For example, they may include one or more of the following: translation, rotation, and scale variation components. When using an incremental representation of the global motion model... When this increment vector is used in dense optical flow aggregation representation, Take the global motion increment vector obtained from the aggregation; Representing frame time The intentional camera movement component is defined as the... The output vector is robustly smoothed according to the smoothing window length and cutoff frequency or equivalent bandwidth parameters given by the camera movement parameter template; the residual threshold and the upper limit of the number of iterations used in its calculation process are derived from the camera movement parameter template and are saved in the separate record field. Representing frame time The jitter disturbance component is defined as and The difference residual vector is obtained by calculating element-wise for each frame. At a certain frame time When unreliable markers or missing markers exist, the frame's and Write the same type of tag.
[0053] When archiving and Explicitly write the frame-by-frame result file to the aligned data layer, and supplement the segment with a separation result summary field in the index data layer. The frame-by-frame result file should contain at least the following fields: frame time index and timestamp, vector, vector, Vectors, unreliable markers, and missing markers, as well as the identifiers for the separated parameter groups used in this segment. The summary field of the index data layer must contain at least the following: and The amplitude quantile statistics, energy percentage statistics, and unreliable frame ratio are fixed in the parameter table, and the statistical window and quantile values are recorded with the version number.
[0054] S3: Calculate the final control quantity based on the fusion weights of the jitter disturbance component, composition error and intentional camera movement component, and archive it with a limit. At the same time, summarize the results and feed them back to update the sample library index.
[0055] Furthermore, in the decoupling and fusion control output stage, the frame sequence reference and frame-by-frame alignment control and inertial sequences of each segment are read as the smallest processing unit. Segment metadata, alignment parameters, and segmentation parameters are read simultaneously, with all fields constrained by a field dictionary to ensure consistency in time position, unit, coordinate system, and missing flag rules. Simultaneously, the output composition error sequence, intended camera movement components, and jitter perturbation components are read, along with the generated and archived camera movement parameter template. The camera movement parameter template contains the required fusion parameter sets, which at least include the initial value range of each fusion weight, the upper limit of the weight update step size, and the configuration reference corresponding to the limiting boundary given by the device capability. When multiple fusion parameter sets exist in the template, the actual fusion parameter set used for this segment is determined based on the frame rate, unreliable marker ratio, hold-mode trigger ratio in the segment metadata, and the selection rules fixed in the parameter table for the separation result summary. The selected parameter set identifier, the values of each parameter, and their version numbers are written into the fusion record field of this segment.
[0056] After forming the fused input record, the consistency of the scale and unit of each input quantity is confirmed. The composition error sequence is the coordinate difference under a unified image coordinate system, which can be in pixel or normalized coordinate form, and the unit is specified in the field dictionary; the intentional camera movement component and the shake perturbation component are the decomposition results of the global motion quantity, inheriting the vector dimension and unit definition. To avoid numerical instability caused by the direct addition of different units, the input quantities are mapped to a unified control quantity representation space according to the scale mapping parameters given by the camera movement parameter template. The scale mapping parameters come from the self-calibration record or the global scale table version during the template generation stage; the mapping method can adopt linear proportional mapping or component independent scaling, and the scaling factor is written into the fused record field and bound to the template version number. If there are available camera intrinsic parameters in the fragment metadata or if the intrinsic parameter configuration can be mapped from the device model and lens model, the pixel domain composition error can be converted into the viewpoint domain error to match the angular velocity control quantity; when the intrinsic parameter is unavailable or there is a missing flag, the explicit missing flag rule is followed, and no forced conversion is performed. Only the pixel domain error is used and the corresponding pixel domain scaling factor in the template is adopted.
[0057] The control values are calculated on a frame-by-frame time grid, and written as a frame-by-frame result sequence into the aligned data layer, while retaining intermediate input fields and parameter fields for recalculation. The frame-by-frame control value calculation uses the following expression:
[0058] in, Representing frame time The final control output is in vector form, and its components are determined by the controlled object and specified in the field dictionary. For example, it may include one or more of the following: yaw rate component, pitch rate component, roll compensation component, and zoom rate component. When the controlled object is only the gimbal rotation... It must include at least yaw and pitch control components; This indicates a limiting operator, used to limit the unlimited control quantity within the brackets to the equipment's capacity on a component-by-component basis; the limiting process includes amplitude limiting and rate of change limiting as specified in the parameter table. When using rate of change limiting, the upper limit of the rate of change is also derived from the equipment capacity configuration reference. Indicates the output frame time. The composition error is defined as the difference between the coordinates of the subject's center point and the coordinates of the target point in the composition. It is defined in a unified image coordinate system, and the specific unit is specified in the field dictionary. This error occurs when subject detection and tracking enters preserve mode or when a missing marker exists in the frame. Retain missing flags in the input records and trigger rollback rules; , , These represent the fusion weight coefficients for the stability, follow, and style items, respectively. The initial value range and update step size upper limit of the weight coefficients are derived from the fusion parameter group in the camera movement parameter template and are saved in the fusion record field of this segment. When online fine-tuning is enabled, the fine-tuning only slowly changes the weights within the step size upper limit specified in the parameter table, and writes the weight value of each frame into the frame-by-frame result field for recalculation. , These represent the lower and upper limit vectors of the control quantity, respectively, used for... Each component's limiting boundary is constrained, and the boundary values are derived from the device capability configuration reference and bound to the device model, firmware version, and unit definition for storage.
[0059] When calculating the unlimited control value, it is read from the fused input record according to the frame-by-frame time grid. , , The unlimited control vector is obtained by performing a linear combination of the weight coefficients and scale mapping coefficients given by the fusion parameter set. The linear combination calculation is performed according to the component alignment rule: when and When the dimension includes translation, rotation, and scale components, it is first converted into a representation consistent with the control components according to the mapping table specified in the field dictionary. For example, the rotation component is mapped to the yaw or pitch control component, and the translation component is mapped to the virtual camera translation control component. The mapping table is then fixed as part of the fusion parameter set. When a certain mapping relationship is not applicable under the current control object, the corresponding component is not included in the combination and a "not included" mark is written in the frame-by-frame record. Subsequently, amplitude limiting processing is performed on the combination result. The amplitude limiting processing is performed component-by-component according to the amplitude boundary and rate of change boundary in the device capability configuration reference, and the final control quantity is output. .
[0060] Archive the frame-by-frame control quantity sequence to the aligned data layer. The archived fields must at least include the frame time index and timestamp. The data includes vectors, pre-limiting control vectors, missing and unreliable flags for each input, weight coefficient values, scale mapping coefficient values, and limit trigger flags and trigger reasons. Simultaneously, the index data layer supplements the fusion output summary field for this segment, including quantiles of control amplitude, quantiles of rate of change, limit trigger ratio, and backoff trigger ratio. The statistical window and quantile values are fixed in the parameter table and their version numbers are recorded, ensuring that the fusion calculation process, the limit constraint process, and the meaning of the output fields can be fully reproduced under the field dictionary constraints.
[0061] It should be noted that during the online feedback recording and sample re-feeding phase, the frame sequence references and frame-by-frame alignment control and inertial sequences of each segment are read as the smallest processing unit. Segment metadata, alignment parameters, and segmentation parameters are read simultaneously, and all fields are constrained by the field dictionary to ensure consistency in time position, unit, coordinate system, and missing flag rules. Simultaneously, the output frame-by-frame global motion sequence, the output subject detection and tracking result sequence and composition error sequence, the output intentional camera movement components and jitter perturbation components, and the output final control quantity sequence along with its limiting trigger flag and backtracking trigger flag are read. To ensure that the re-feeding data can still serve as a valid basis for subsequent retrieval and parameter template generation in the sample library, the generation rules for frame-by-frame results are not changed; only structured summarization, quality field calculation, re-feeding index updates, and version recording are performed.
[0062] First, a completeness check is performed on each frame-by-frame result. The check includes at least: whether the frame sequence references are consistent with the frame-by-frame time grid; whether there are consecutive missing segments in the frame-by-frame aligned control and inertial sequences; whether each frame-by-frame result has a one-to-one correspondence under the same frame time index; and whether the missing and unreliable flags meet the constraints of the flag values and propagation rules in the field dictionary. For inconsistencies in timeline, frame count, or units found during the check, only explicit error marking is performed, and the scope and cause are recorded; no resampling or repair of historical records is performed. When the inconsistency affects subsequent statistical calculations, the affected frame segments are excluded from the statistical range, and the effective frame percentage and effective frame range fields are written into the statistical results.
[0063] After verification or exclusion, a summary record for the segment is generated for re-implantation. The summary record includes two parts: a summary of frame-by-frame results and a process event statistics. The summary statistics include at least: the amplitude quantiles of global motion and the proportion of unreliable frames, the amplitude quantiles and energy proportions of jitter perturbation components, the amplitude quantiles and main trend change rate statistics of the intended camera movement component, the amplitude quantiles of compositional errors and the proportion of frames that continuously exceed the threshold, the amplitude quantiles of the final control quantity, the change rate quantiles and the amplitude limiting trigger ratio; the window length, quantile values, and energy calculation methods of the above statistics are all fixed in the form of parameter tables and written with version numbers to ensure that the statistical caliber is consistent when the same segment is re-implanted in different batches. The process event statistics include at least: the number of triggers and the distribution of the number of frames for main body tracking to enter hold mode, the number of interruptions and the range of interruption frames for main body tracking, the trigger ratio of global motion estimation to enable dense optical flow backup path, the selection results of separation parameter group and fusion parameter group, the component category and trigger frame range of amplitude limiting trigger, and the type and trigger frame range of fallback rule trigger; the event statistics are stored in structured fields, without using free text, and the meaning and value range of the fields are constrained by the field dictionary.
[0064] After the summary record is generated, the fragment is written as a new sample into the index data layer of the sample library, and the index fields related to retrieval are updated. The updates include at least: writing the calculated summary statistics field into the index table of the fragment so that it can be used as one of the continuous feature dimensions for similarity ranking; writing the process event statistics field into the index table of the fragment so that it can be used as a constraint or exclusion condition during candidate filtering; and writing the dictionary version number, alignment parameter table version number, segmentation parameter table version number, and statistics parameter table version number used by this fragment into the index record to ensure that candidate fragments can be filtered according to version consistency during subsequent retrieval. If the sample library adopts a nearest neighbor retrieval index structure, a retrieval feature vector with a consistent structure is generated at the same time as writing into the index table, and written into the nearest neighbor index according to a unified normalization rule and scale table version; when the scale table or normalization rule version changes, the historical fragments are not recalculated immediately, but only the version difference is marked in the index record so that subsequent batch processing updates can locate the set of fragments that need to be recalculated.
[0065] To ensure that the reprocessing chain can be reproduced in the re-injected samples, the parameter templates related to the segment and the actual operating parameters are archived together. The archived content includes at least: the generated camera movement parameter template and its search conditions, the number of candidate sets and similarity score statistics; the actual separation parameter sets used and their values; the actual fusion parameter sets used, the weight coefficient value sequence or its summary statistics, scale mapping coefficients, and references to the limiting boundary configuration. All of the above parameters are written into the index data layer in the form of a structured parameter table and bound to the segment's metadata. For fields involving equipment capability configuration references, the equipment model, firmware version, unit definition, and version identifiers of the upper and lower limit vectors are also saved to avoid inconsistencies in units or capability boundaries for the same control variable value on different devices, which could lead to unreproducibility.
[0066] Finally, the checksums and checksums generated after the fragment is re-implanted are updated in the sample library management structure. Checksums are generated for all new files in the alignment data layer and index data layer and written to the checksums. The checksums include the file path, size, generation time, checksum algorithm, and checksum result. At the same time, an association identifier is generated for the re-implantation record of this fragment, which is used to associate and locate the original data layer files, the frame-by-frame result files of the alignment data layer, the summary of the index data layer, and the parameter files.
[0067] Example 2, one embodiment of the present invention, provides a motion decoupling control method based on big data of camera movement. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiment.
[0068] First, a camera system with a gimbal control interface was selected as the shooting end. The data collected included: video frame sequences, gimbal or camera control records (yaw / pitch / roll angles or angular velocities, zoom magnification, and other obtainable fields), and inertial attitude data (gyroscope angular velocity, acceleration, attitude calculation results, etc.). To reduce the non-reproducibility caused by differences in multi-source sampling rates, the original timestamps of each data source were retained, and the frame rate, resolution, sampling rate, unit, and coordinate system definition were recorded. Then, a unified time reference was established: a frame-level time axis was constructed using the video frame-by-frame timestamps. The control records and inertial attitude data were resampled to a frame-by-frame time grid. During resampling, missing flags were retained, and abnormal jump values were removed or truncated according to the device's physical upper limit threshold, with the processing range recorded. After frame-by-frame alignment, the continuously shot footage was segmented according to lens boundaries and motion continuity conditions, ensuring that each segment had frame sequence references and frame-by-frame aligned control and inertial sequences. Segment metadata and alignment parameters were stored synchronously, forming searchable camera movement data.
[0069] Based on this, frame-by-frame global motion estimation is performed for each segment. Global motion increments are calculated between adjacent frames using feature point matching and robust estimation. When the proportion of matched intrapoints is insufficient or residuals are abnormal, dense optical flow is used as a backup observation and aggregated to obtain the global motion increment. Unreliable frames are marked as unreliable. In parallel, subject detection and cross-frame tracking are performed for each frame, outputting the subject's bounding box and center point coordinates. In case of occlusion or short-term loss of detection, a prediction-preserving mode with a limited number of frames is entered, and a preservation flag is written. Based on preset composition rules, composition target points are generated. The coordinate difference between the subject's center point and the composition target points is calculated as a composition error sequence. When the subject result is missing, a missing flag is retained without interpolation filling, enabling subsequent links to clearly identify the effective frame range.
[0070] Subsequently, based on the global motion summary, main motion, and composition error summary of the segment as search criteria, similar segments are retrieved from the camera movement sample library, and a parameter template is generated. The template includes at least a separation parameter group, a fusion parameter group, and a reference to the device limiting configuration. The global motion sequence is then component-separated according to this template to obtain the intended camera movement component and the jitter disturbance component, which are written to the frame-by-frame result file according to the same frame index. Finally, the jitter disturbance component, composition error, and intended camera movement component are fused on the same frame-by-frame time grid to obtain the final control quantity sequence. The amplitude and rate of change of the control quantity are then limited according to the device capability boundary. Simultaneously, the limiting trigger flag, backtracking flag, parameter selection results, and summary statistics are archived together, and the segment results are fed back to update the sample library index, ensuring that subsequent searches and template generation can utilize the statistical features of the latest samples.
[0071] The comparison scheme selected from existing technical approaches: it adopts a fixed-parameter image stabilization and tracking control process, lacks a unified frame-by-frame alignment and database entry mechanism, lacks a parameter template generation mechanism based on sample database retrieval, and does not explicitly separate the intent and jitter of global motion, achieving stabilization only through overall smoothing. Table 1 Experimental Data
[0072] As shown in Table 1, existing technologies exhibit significant cross-source time inconsistencies (18.4–25.2 ms) across all three scenarios, while this invention compresses these inconsistencies to 2.6–3.8 ms. This demonstrates that by aligning each frame into the database and explicitly handling missing / abnormal data, a stable intra-frame index binding is established between video frames and control / attitude data. This difference directly impacts the reproducibility of subsequent global motion estimation and control derivation: in high-dynamic scenarios such as sports tracking and vehicle lateral movement, the proportion of unreliable frames in existing technologies reaches 10.8% and 12.1%, respectively, while this invention reduces them to 2.4% and 2.8%. This indicates that with the support of a unified time reference, motion estimation more easily forms continuous and stable frame-by-frame motion sequences, and unreliable frame segments can be identified and isolated, reducing error propagation at the source.
[0073] The number of subject tracking interruptions and the composition error P95 together reflect the continuity of composition constraints. Existing technologies experience 5 interruptions within 30 seconds and a composition error P95 of 91 pixels in a vehicle-mounted lateral movement scene, indicating that tracking control is prone to instability and subject deviation from the composition target under occlusion, rapid movement, or background changes. This invention reduces the number of interruptions to 2 and the composition error P95 to 45 pixels, demonstrating that the continuous constraint of subject detection and tracking and composition error on the same frame-by-frame grid can effectively reduce composition drift. Furthermore, the retention of prediction-preserving markers and missing markers provides a clear basis for subsequent fusion calculations to back off, thereby preventing erroneous observations from being forcibly incorporated into the control.
[0074] The shake disturbance energy and the intent camera fidelity score demonstrate the ability of the motion decoupling mechanism to protect the camera language. Existing technologies mainly suppress shake through overall smoothing, but the shake disturbance energy still reaches 0.58 in the vehicular lateral movement scene, while the intent camera fidelity score is only 62, indicating that a single smoothing strategy cannot simultaneously ensure stability and the rhythm of the camera language. This invention reduces the shake disturbance energy to 0.24 in the same scene, while improving the fidelity score to 83. This shows that after separating the global motion into the intent camera component and the shake disturbance component, stability suppression can be more focused on high-frequency disturbances, while the intent component is preserved and participates in the fusion, thus avoiding the contradiction of becoming more dead if the motion is too stable and more shaky if the motion is too close.
[0075] The control limiting trigger rate further demonstrates the executability of the control output. Existing technologies have trigger rates ranging from 9.4% to 14.6%, indicating that under high dynamic input, the control quantity is more likely to exceed the device's capability limits and be frequently pruned, leading to abrupt changes or discontinuities in the control chain. This invention reduces the trigger rate to 4.8%–6.8%, demonstrating that through parameter template matching and decoupling fusion, the amplitude and rate of change of the control quantity better conform to the device's capability constraints, and limiting becomes more of a safety boundary rather than a high-frequency, routine event. The overall experience score improves in all three scenarios (e.g., sports tracking 71→89, stage slow push 74→90, vehicle panning 68→86), and this improvement shows a consistent trend with objective indicators such as alignment error, unreliable frame ratio, composition error, and jitter energy. This indicates that this invention, through a chain process of alignment and database entry, parallel modeling of motion estimation and subject composition, parameter template retrieval, decoupling of motion components, and fusion limiting archiving and refeeding updates, achieves a comprehensive improvement in stability, shot language fidelity, and executable control output that is difficult to achieve simultaneously in existing technologies.
[0076] Example 3, one embodiment of the present invention, provides a motion decoupling control system based on big data of camera movement, including a data alignment and storage module, a motion estimation decoupling module, and a fusion control feedback module.
[0077] The data alignment and storage module is used to collect video, gimbal and camera control records and inertial attitude data, and align and segment them frame by frame according to a unified time reference to generate camera movement data; the motion estimation and decoupling module is used to estimate the global motion of each frame on the camera movement data, detect and track the subject to generate composition error, and retrieve the generated parameter template to separate the global motion into the intended camera movement component and the jitter disturbance component; the fusion control backfeed module is used to fuse the jitter disturbance component, composition error and intended camera movement component to calculate the final control quantity and limit the archive, while summarizing the results and backfeeding to update the sample library index.
[0078] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0079] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0080] More specific examples (a non-exhaustive list) of computer-readable media include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0081] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc. It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
[0082] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A motion decoupling control method based on lens movement big data, characterized in that, include: Collect video, gimbal and camera control records and inertial attitude data, align them frame by frame according to a unified time base and segment them into the database, and output camera movement data; Estimating frame-by-frame global motion based on camera movement data, and simultaneously detecting and tracking the subject to generate composition errors, using global motion, subject, and composition summary as retrieval conditions to generate parameter templates, and separating global motion into intentional camera movement components and jitter perturbation components; The jitter, composition error, and intentional camera movement components are weighted according to the final control values and then archived with a limit. At the same time, the results are summarized and fed back to update the sample library index. 2.The motion decoupling control method based on the mirror-based big data according to claim 1, wherein: The process of aligning and segmenting video frames according to a unified time reference includes establishing a video frame timeline and recording the frame rate and frame timestamp. The original sampling timestamps are retained for the gimbal and camera control records and inertial attitude data; The control recording and inertial attitude data are resampled to a frame-by-frame time grid consistent with the video frames. The resampling uses nearest neighbor sampling and linear interpolation, and the method used is recorded. Write an explicit missing flag for missing samples, and truncate and remove outliers that exceed the equipment range and physical limits according to a threshold, and record the threshold.
3. The motion decoupling control method based on big data of camera movement as described in claim 2, characterized in that: The estimation of frame-by-frame global motion includes extracting and matching local feature points between adjacent frames, solving the global motion model based on robust estimation, and outputting the number of inliers, the proportion of inliers, and residual statistics. When the proportion of inliers is lower than a preset threshold and the residual is abnormal, switch to dense optical flow calculation and aggregate the pixel motion vector field to obtain the global motion increment; Global motion is archived as a frame-by-frame incremental sequence, and unreliable frames are marked with an unreliable tag and the corresponding frame range.
4. The motion decoupling control method based on big data of camera movement as described in claim 3, characterized in that: The synchronous detection and tracking of the subject's generated composition error includes outputting the subject's bounding box, center point coordinates, and confidence level for each frame, and constructing an association cost based on positional distance, scale change, and appearance consistency to achieve cross-frame tracking. When occlusion, missed detection, or confidence level below a threshold occurs, the prediction hold mode with a limited number of frames is entered and a hold flag is written. The coordinates of the target point are generated according to the preset composition rules. The composition error is the difference between the coordinates of the center point of the subject and the coordinates of the target point. When the subject result is missing, a missing flag is written to the frame composition error.
5. The motion decoupling control method based on big data of camera movement as described in claim 4, characterized in that: The parameter template generated by using global motion, subject, and composition summary as retrieval conditions includes calculating amplitude quantiles, mean square values, and peak values of global motion sequences and composition error sequences to form continuous features, and forming discrete features of scene description information, subject category, and motion type. First, filter candidate segments according to discrete features, and then sort them by nearest neighbor according to continuous feature distance to obtain a reference set; The reference set is fused to generate a parameter template containing a separation parameter group and a fusion parameter group. The separation parameter group contains the smoothing window length and cutoff frequency or equivalent bandwidth parameters, while the fusion parameter group contains the initial value range of the fusion weight and a reference to the device limiting configuration.
6. The motion decoupling control method based on big data of camera movement as described in claim 5, characterized in that: The step of separating global motion into intended camera movement component and jitter perturbation component includes performing robust smoothing on the frame-by-frame global motion amount according to the parameter template to obtain the intended camera movement component, and using the difference between the frame-by-frame global motion amount and the intended camera movement component as the jitter perturbation component. For frames with unreliable or missing markers, retain the markers and propagate them to the intended camera movement component and the jitter perturbation component; The jitter perturbation component, composition error, and intended camera movement component are written into the frame-by-frame result sequence according to the same frame time index.
7. The motion decoupling control method based on big data of camera movement as described in claim 6, characterized in that: The calculation of the final control quantity and the limiting archive includes reading the jitter disturbance component, composition error and intended camera movement component on the same frame-by-frame time grid, calculating the unlimited control quantity according to the fusion weight given by the parameter template, and performing limiting processing on the amplitude and rate of change of the control quantity according to the device limiting configuration. During archiving, both the amplitude limit trigger flag and the rollback flag are written simultaneously, and the statistical summary of the fragment, the parameter group selection results, and the version information are summarized and fed back to update the sample library index fields.
8. A system employing the motion decoupling control method based on big data of camera movement as described in any one of claims 1 to 7, characterized in that: This includes a data alignment and storage module, a motion estimation decoupling module, and a fusion control feedback module; The data alignment and storage module is used to collect video, gimbal and camera control records and inertial attitude data, and align and segment them frame by frame according to a unified time reference to generate camera movement data. The motion estimation decoupling module is used to estimate the frame-by-frame global motion on the camera movement data, detect and track the subject to generate composition errors, and retrieve the generation parameter template to separate the global motion into the intended camera movement component and the jitter disturbance component. The fusion control backfeed module is used to fuse jitter disturbance components, composition error and intentional camera movement components to calculate the final control quantity and limit and archive it, while summarizing the results and backfeeding to update the sample library index.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the motion decoupling control method based on big data of camera movement as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the motion decoupling control method based on big data of camera movement as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Image pick-up apparatus and method executed in same
CN109391767A
Intelligent mobile shooting method and device and controller
CN114173060A
Automatic following shooting system and method based on artificial intelligence
CN119485000A
Shooting method, electronic equipment and storage medium
CN120343395A
Video generation method and device with controllable mirror operation mode, and electronic equipment
CN121099153A