IMU-based camera motion compensation method, device and storage medium
By filtering and modeling the IMU sensor data and combining image feature information, high-quality image stability in complex motion scenarios is achieved, solving the problem that traditional methods are difficult to accurately compensate for camera movement, and improving image stability effect and visual quality.
Patent Information
- Application Number
- CN202411500548.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2044-10-25
AI Technical Summary
Traditional optical and digital anti-shake technologies are difficult to meet the needs of high-quality image stability in complex motion scenarios. There is noise and drift in IMU data, and existing compensation methods are difficult to accurately capture the camera's six-degree of freedom motion and ignore the semantic information of the image content.
By performing low-pass filtering and adaptive Kalman filtering on multiple IMU sensor data, combining linear convolutional aliasing model and BP neural network, high-dimensional motion feature vectors are extracted, advanced network compensation model is constructed, combined with dense optical flow field optimization and time-weighted background pixel pool modeling, content-aware filling and time-domain filtering are performed, and multiple compensation modes are provided.
It improves the accuracy and reliability of motion data, can accurately compensate camera motion in complex motion scenarios, maintain image stability and retain foreground information, and improves the visual quality of the image sequence.
Smart Images

Figure CN119342346B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of camera motion compensation, and in particular to an IMU-based camera motion compensation method, device, and storage medium. Background Art
[0002] Camera stability is a key factor affecting image and video quality. Traditional optical and digital image stabilization technologies often struggle to meet the requirements for high-quality image stabilization in complex motion scenes. In recent years, camera motion compensation methods based on inertial measurement units (IMUs) have attracted widespread attention due to their fast response and high precision.
[0003] However, IMU data suffers from noise and drift, making it difficult to achieve ideal results when directly using it for motion compensation. Furthermore, traditional motion compensation algorithms struggle to accurately capture the camera's six degrees of freedom (6DOF) when handling complex motion scenes, resulting in poor compensation results. Furthermore, existing compensation methods often neglect the semantic information of image content, making it difficult to preserve important foreground information while maintaining image stability. Summary of the Invention
[0004] The present invention provides an IMU-based camera motion compensation method, device, and storage medium. The present invention can effectively fuse IMU data and image information, accurately estimate camera motion parameters, and dynamically adjust the compensation strategy according to scene characteristics to achieve camera motion compensation.
[0005] In a first aspect, the present invention provides an IMU-based camera motion compensation method, the IMU-based camera motion compensation method comprising:
[0006] The raw data collected by multiple IMU sensors are processed by low-pass filtering and adaptive Kalman filtering to obtain six-degree-of-freedom IMU fusion data;
[0007] Inputting the six-degree-of-freedom IMU fusion data into a linear convolution aliasing model for processing to obtain simulated motion data and extracting the corresponding high-dimensional motion feature vector;
[0008] According to the high-dimensional motion feature vector and the original image sequence, a six-degree-of-freedom motion model based on a sine wave combination and a BP neural network prediction model are fused to obtain a camera motion parameter prediction value;
[0009] Building a look-ahead network compensation model based on the camera motion parameters, and dynamically adjusting the compensation model parameter matrix based on the jitter intensity evaluation results of the IMU data variance and peak analysis;
[0010] Performing geometric transformation on the original image data based on the compensation model parameter matrix, and combining dense optical flow field optimization and time-weighted background pixel pool modeling to obtain an initial stable image sequence and a foreground mask;
[0011] Content-aware filling and temporal filtering are performed on the initial stabilized image sequence and the foreground mask, and target stabilized image sequences in multiple compensation modes are output.
[0012] In a second aspect, the present invention provides an IMU-based camera motion compensation device, the IMU-based camera motion compensation device comprising:
[0013] The acquisition module is used to perform low-pass filtering and adaptive Kalman filtering on the raw data collected by multiple IMU sensors to obtain six-degree-of-freedom IMU fusion data;
[0014] An extraction module is used to input the six-degree-of-freedom IMU fusion data into a linear convolution aliasing model for processing to obtain simulated motion data and extract the corresponding high-dimensional motion feature vector;
[0015] A fusion module is used to fuse the high-dimensional motion feature vector and the original image sequence using a six-degree-of-freedom motion model based on a sine wave combination and a BP neural network prediction model to obtain a camera motion parameter prediction value;
[0016] An evaluation module is used to build a look-ahead network compensation model based on the camera motion parameters and dynamically adjust the compensation model parameter matrix according to the jitter intensity evaluation results of the IMU data variance and peak analysis;
[0017] a transformation module, configured to perform geometric transformation on the original image data based on the compensation model parameter matrix, and obtain an initial stable image sequence and a foreground mask by combining dense optical flow field optimization and time-weighted background pixel pool modeling;
[0018] An output module is configured to perform content-aware filling and temporal filtering on the initial stabilized image sequence and the foreground mask, and output a target stabilized image sequence in a plurality of compensation modes.
[0019] A third aspect of the present invention provides a computer device comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the computer device executes the above-mentioned IMU-based camera motion compensation method.
[0020] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, enable the computer to execute the above-mentioned IMU-based camera motion compensation method.
[0021] The technical solution provided by the present invention effectively reduces IMU data noise and drift by performing low-pass filtering and adaptive Kalman filtering on multiple IMU sensor data, thereby improving the accuracy and reliability of motion data. A linear convolution aliasing model is used to process IMU fusion data, enabling better capture of complex motion patterns and extraction of high-dimensional motion feature vectors, providing richer information for subsequent motion parameter prediction. Combining a six-degree-of-freedom motion model based on a sine wave combination with a BP neural network prediction model, accurate prediction of camera motion parameters is achieved, effectively addressing the difficulty of traditional methods in accurately describing complex motion. An advance network compensation model and dynamic adjustment mechanism are introduced to adjust the compensation strategy in real time based on the jitter intensity assessment results of the IMU data, improving the flexibility and adaptability of the compensation. Combining dense optical flow field optimization with time-weighted background pixel pool modeling, foreground information is effectively preserved while image stabilization is performed, enhancing the visual quality of the stabilization effect. Content-aware padding and temporal filtering address edge gaps and temporal discontinuities that may occur during image stabilization, further improving the quality of the stabilized image sequence. It provides multiple compensation modes (global compensation, local compensation, and adaptive compensation) and dynamically selects the optimal mode through jitter evaluation to meet the stability requirements in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0023] Figure 1 Schematic diagram of the steps of the IMU-based camera motion compensation method in an embodiment of the present invention;
[0024] Figure 2 Schematic diagram of the structure of an IMU-based camera motion compensation device in an embodiment of the present invention;
[0025] Figure 3 It is a schematic block diagram of the structure of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] Embodiments of the present invention provide an IMU-based camera motion compensation method, apparatus, and storage medium. The terms "first," "second," "third," "fourth," and so forth (if any) in the specification and claims of the present invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. Furthermore, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatuses.
[0027] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 An embodiment of the camera motion compensation method based on IMU in the embodiment of the present invention includes:
[0028] Step S1: low-pass filtering and adaptive Kalman filtering are performed on the raw data collected by multiple IMU sensors to obtain six-degree-of-freedom IMU fusion data;
[0029] It is understood that the execution subject of the present invention can be an IMU-based camera motion compensation device, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking the server as the execution subject as an example.
[0030] Specifically, the raw acceleration and angular velocity data collected by each IMU sensor is digitally low-pass filtered to remove high-frequency noise components and extract smoother motion data, resulting in first IMU data. The first IMU data is then subjected to zero-bias correction and temperature compensation. Zero-bias correction eliminates errors caused by internal sensor bias, while temperature compensation corrects for measurement errors caused by sensor drift in different temperature environments, resulting in more accurate second IMU data. The second IMU data is then input into an adaptive Kalman filter. The adaptive Kalman filter is a dynamic filter that adjusts the process noise covariance matrix in real time based on the current motion state to adaptively optimize the filtering results. In this way, the filter can effectively cope with noise variations under different motion modes, improving the accuracy of the filtered data and obtaining third IMU data. Allan variance analysis is performed on the third IMU data to calculate the noise characteristic parameters of each IMU sensor, including indicators such as random walk and offset instability. A noise parameter matrix is then constructed using these indicators. Based on the noise parameter matrix, the third IMU data from multiple IMU sensors is weightedly fused. The weighted fusion process assigns weights based on the noise characteristics of each IMU sensor, prioritizing sensor data with lower noise levels to improve the overall data reliability and accuracy. This fused fourth IMU data is then used to perform attitude calculations, using a quaternion algorithm to compute the camera's rotation matrix. Quaternions avoid singularities (such as gimbal lock) that occur in traditional Euler angle representations, and their numerical stability helps maintain attitude continuity, providing a more accurate description of camera pose changes. The camera pose data is then combined with the fourth IMU data and fused using a complementary filtering algorithm to enhance the stability and accuracy of the pose calculation. Complementary filtering is a lightweight fusion method that combines high-frequency acceleration data with low-frequency gyroscope data to effectively filter out long-term drift errors during pose changes while ensuring dynamic response, resulting in corrected 6-DOF motion data. This corrected 6-DOF motion data is time-synchronized and interpolated to align with the camera's frame rate. The data acquisition frequency of IMU sensors is typically much higher than the camera's frame rate, necessitating time synchronization to ensure that the motion data corresponds to the image sequence captured by the camera. During this time synchronization process, interpolation is used to compensate for the discrepancies between sampling frequencies, enabling accurate motion state information to be obtained at the camera sampling moment. This processing ultimately results in 6DOF IMU fusion data aligned with the camera's frame rate.
[0031] Step S2: Input the six-degree-of-freedom IMU fusion data into the linear convolution aliasing model for processing to obtain simulated motion data and extract the corresponding high-dimensional motion feature vector;
[0032] Specifically, the linear convolutional model consists of three one-dimensional convolutional layers, two max pooling layers, and two fully connected layers. The first convolutional layer uses 32 one-dimensional convolutional kernels of size 5 with a stride of 1 to capture short-term motion features in the input data. The second convolutional layer uses 64 one-dimensional convolutional kernels of size 3 with a stride of 1 to extract deeper motion patterns and detailed features. The third convolutional layer uses 128 one-dimensional convolutional kernels of size 3 with a stride of 1 to enrich the data representation with higher dimensions. To adapt the 6DOF IMU fusion data to the convolutional network, the data is split into time windows to produce multiple sub-signal sequences. These sub-signal sequences are input to the first one-dimensional convolutional layer of the linear convolutional model for a one-dimensional convolution operation. The one-dimensional convolution captures local dependencies and characteristic patterns of the input signal in the temporal dimension by sliding the convolution kernel. Using the ReLU activation function, 32 first feature maps are generated, each with a size of 96. Max pooling is performed on the 32 first feature maps with a pooling window size of 2 and a stride of 2. Max pooling can reduce the size of feature maps, lowering the data dimensionality and computational complexity, while also providing some noise immunity. This pooling operation yields 32 downsampled second feature maps, each with a size of 48. These 32 downsampled second feature maps are fed into the second one-dimensional convolutional layer for convolution and activated with the ReLU function to yield 64 third feature maps, each with a size of 46. Max pooling is again performed on these 64 third feature maps with a pooling window size of 2 and a stride of 2. This step reduces the resolution and data size of the feature maps, yielding 64 downsampled fourth feature maps, each with a size of 23. These 64 downsampled fourth feature maps are fed into the third one-dimensional convolutional layer for convolution and activated with the ReLU function to yield 128 fifth feature maps, each with a size of 21. After convolution and pooling, the 128 fifth feature maps are flattened to facilitate input into the fully connected layer. The resulting flattened feature vector has a dimension of 2688. This feature vector is then fed into the first fully connected layer, which contains 1024 neurons and uses a Reluctant Unit (ReLU) activation function to perform nonlinear mapping of the input features, extracting more expressive features and generating a 1024-dimensional feature vector. Batch normalization is performed on this 1024-dimensional feature vector to ensure a more uniform distribution of features, preventing vanishing or exploding gradients during model training due to a large data set, thereby improving model convergence speed and stability. To extract the core features of camera motion, the normalized feature vector is fed into an autoencoder network. The autoencoder network consists of two parts: an encoder that reduces the dimensionality of the features and a decoder that remaps the compressed features back to a higher-dimensional space.The encoder consists of three fully connected layers with 512, 256, and 128 neurons, respectively. This gradually reduces the feature dimensionality, capturing key information and removing redundancy. The encoder generates 128-dimensional compressed features, which contain the most important motion information from the original high-dimensional features. The 128-dimensional compressed features are used as input for reconstruction by the decoder, which also consists of three fully connected layers with 256, 512, and 1024 neurons, respectively. These layers restore the compressed features to their original high-dimensional space, resulting in a high-dimensional motion feature vector.
[0033] Step S3: Based on the high-dimensional motion feature vector and the original image sequence, a six-degree-of-freedom motion model based on a sine wave combination and a BP neural network prediction model are fused to obtain a camera motion parameter prediction value;
[0034] Specifically, feature point extraction and matching are performed on adjacent frames in the original image sequence. Using feature point detection methods such as SIFT and SURF, significant feature points in the image are identified and matched, resulting in a feature point correspondence matrix. Based on this feature point correspondence matrix, outlier matching points are removed, eliminating incorrectly matched feature points and improving the accuracy of subsequent motion estimation. By calculating the fundamental matrix, the camera's relative rotation matrix R and translation vector t are solved to obtain an initial estimate of image motion. The high-dimensional motion feature vector is input into a BP neural network prediction model. The output layer of the BP neural network consists of six neurons, one for each of the six degrees of freedom (DOF) motion parameters, including the camera's 3D translation and 3D rotation. During the forward propagation process, the BP neural network uses a series of nonlinear transformations to map the learned motion features to the motion state to calculate initial motion parameter predictions based on the IMU data. These predictions reflect the motion information provided by the IMU sensor and, combined with the image motion characteristics, provide a reliable basis for motion compensation. The initial image motion estimate and the initial motion parameter prediction based on IMU data are fused. A weighted average method is used to combine the motion estimates from the two sources, image and IMU, to obtain fused camera motion parameters. The weights are dynamically adjusted based on the uncertainty of the IMU and image motion estimates, adapting to changes in motion state in different motion scenarios. The fused camera motion parameters are then fed into a six-degree-of-freedom motion model based on a combination of sine waves. This model represents the camera's 3D translation and rotation as a superposition of six sine wave functions, providing a smoother and more periodic description of motion. It is particularly effective in modeling small-scale, periodic jitter. The sine wave combination calculations yield the camera's current position and attitude, reflecting its spatial location and including specific information about its orientation. A Kalman filter is then applied to the camera's current position and attitude to improve the smoothness and accuracy of the predictions. The Kalman filter's state vector contains position, velocity, and acceleration, while the measurement vector represents the fused camera motion parameters. Through the Kalman filter's prediction-update loop, the state estimate is continuously updated using the prediction model combined with actual measurements to obtain smoothed camera motion parameters. The Kalman filter can optimally estimate the state of a dynamic system and is suitable for processing linear dynamic systems with Gaussian noise. Based on the smoothed camera motion parameters, the current frame in the original image sequence is back-projected, and the Euclidean distance between the projected point and the actual feature point is calculated to obtain the reprojection error. The reprojection error is used to evaluate the accuracy of the predicted camera motion parameters. If the difference between the projection result and the actual feature point position is small, it indicates that the predicted motion parameters have high accuracy.When the reprojection error is less than the preset threshold, the current smoothed camera motion parameters are output as the predicted values of the camera motion parameters. If the reprojection error exceeds the preset threshold, the data fusion weights are adjusted and the data fusion and processing are repeated. This process will continue to iterate until the accuracy requirements are met or the maximum number of iterations is reached.
[0035] Step S4: constructing an advance network compensation model based on the camera motion parameters, and dynamically adjusting the compensation model parameter matrix according to the jitter intensity evaluation results of the IMU data variance and peak analysis;
[0036] Specifically, a time series analysis of camera motion parameters is performed, and the motion parameters are predicted to obtain a predicted sequence of camera motion parameters. This predicted sequence of camera motion parameters is then input into a lookahead network compensation model. This lookahead network compensation model consists of three fully connected layers: the first layer contains 64 neurons, the second layer contains 32 neurons, and the third layer contains 16 neurons. Reinforced linear unit (ReLU) activation functions are used between each layer to introduce nonlinearity. Through forward propagation, the lookahead network compensation model generates a 3×3 initial compensation model parameter matrix, which describes the geometric transformation of the camera during the compensation process. Due to its multi-layer, fully connected structure, the model can effectively capture the complex dynamic characteristics of camera motion. To dynamically adjust the compensation model parameter matrix, the jitter characteristics of the IMU data are analyzed. The variance of the 6-DOF IMU fusion data is calculated within a sliding time window to obtain the IMU data variance sequence. The variance calculation reflects the degree of change in the IMU data over different time periods. The IMU data variance sequence is spectrally analyzed using a fast Fourier transform to obtain the spectral characteristics of the IMU data. Spectral features contain amplitude information at different frequencies in motion data, helping to identify the frequency components and intensity of motion jitter. Based on the spectral features of the IMU data, the peak amplitudes and corresponding frequencies are calculated to generate a jitter intensity feature vector that describes the motion jitter characteristics. This vector contains the primary jitter frequencies and their corresponding intensities during motion. This information effectively characterizes the degree of camera instability during motion. This jitter intensity feature vector is then input into the jitter intensity assessment model, which uses the Support Vector Regression (SVR) algorithm. Through regression analysis, the SVR model calculates a jitter intensity assessment score. SVR has excellent generalization capabilities and can accurately assess jitter intensity even with limited sample sizes. Based on the jitter intensity assessment score, a pre-set adaptive weighting function is used to calculate a compensation intensity coefficient. This adaptive weighting function flexibly adjusts the compensation strength based on the jitter intensity, ensuring stronger compensation for stronger jitter and less compensation for weaker jitter to avoid artifacts caused by over-compensation. The calculated compensation intensity coefficient is multiplied by each element in the initial compensation model parameter matrix to dynamically adjust the compensation model parameter matrix. This adjustment method can optimize the compensation parameters in real time according to the actual motion state, so that the compensation effect is more in line with the needs of the camera during actual motion, ensuring the accuracy and stability of the compensation.
[0037] Step S5: performing geometric transformation on the original image data based on the compensation model parameter matrix, and combining dense optical flow field optimization and time-weighted background pixel pool modeling to obtain an initial stable image sequence and foreground mask;
[0038] Specifically, the compensation model parameter matrix is applied to the original image data to perform an affine transformation, thereby adjusting the geometric shape of the image to obtain a first image sequence. The affine transformation performs operations such as rotation, scaling, and translation on the image based on the compensation model parameter matrix to eliminate irregular changes caused by camera motion and preliminarily align the image sequence. A dense optical flow field is calculated between adjacent frames in the first image sequence to obtain a motion vector field for each pixel position. The dense optical flow field calculation can capture the displacement changes of all pixels in the image. Based on the motion vector field obtained from the dense optical flow field, local motion compensation and pixel resampling are performed on the first image sequence to obtain a second image sequence. Local motion compensation uses the information of the optical flow field to eliminate local subtle movements by accurately repositioning pixels, significantly improving the stability of the image sequence. On this basis, a time-weighted background pixel pool is constructed for the second image sequence to capture a stable background in a dynamic environment. A fixed-size queue of pixel values is maintained for each pixel position. Newly acquired pixel values are queued with a time-decaying weighting scheme. New pixels are assigned higher weights, giving them a greater influence on the current background state, while older pixel values gradually decay. When the queue reaches capacity, the oldest pixel values are removed. A dynamic update mechanism enables the background model to adapt to dynamic conditions such as lighting changes and slowly moving objects in the scene, improving the reliability of background modeling. Based on the dynamically updated background model, foreground detection is performed on each frame in the second image sequence to separate moving objects from the background. By comparing the current frame with the background model, pixels belonging to the foreground are determined and a foreground mask is generated based on this information. The foreground mask is a binary image, with foreground portions marked as 1 and background portions marked as 0, representing moving objects in the image. A pixel-by-pixel AND operation is performed on the foreground mask with the second image sequence to extract the foreground region, resulting in a foreground image sequence. Furthermore, to enhance image stabilization, inter-frame difference analysis is performed between the foreground and second image sequences. By calculating the absolute difference between adjacent frames, regions of the image that have changed due to camera motion are identified. An adaptive thresholding method is applied to the calculated absolute difference to segment the image, distinguishing between static and dynamic components. This method dynamically adjusts the threshold based on the grayscale characteristics of the image in different scenes, adapting to varying lighting conditions and motion intensity, and achieving foreground-background separation. This process ultimately yields an initial, stable image sequence.
[0039] Step S6: performing content-aware filling and temporal filtering on the initial stabilized image sequence and the foreground mask, and outputting target stabilized image sequences in multiple compensation modes.
[0040] Specifically, the edge regions of each frame in the initial stabilized image sequence are detected and extended. This edge extension effectively increases the image fill area in subsequent operations, resulting in an extended image sequence. Content-aware filling is then performed on blank areas within the image based on the extended image sequence and a foreground mask. This content-aware filling process utilizes the values of surrounding pixels to appropriately fill the blank areas, ensuring a natural visual transition in the filled image sequence and avoiding noticeable artifacts caused by compensation and transformation operations. The filled image sequence is then processed using a temporal Kalman filter. The Kalman filter's state vector represents the pixel value at each pixel position, and the observation vector represents the pixel value at the corresponding position in the adjacent frame. The Kalman filter's prediction and update process smoothes temporal variations in the image to reduce image quality degradation caused by random jitter, resulting in a temporally filtered image sequence. Kalman filtering effectively smooths noise in the image sequence, thereby improving the temporal continuity of the image and making it appear smoother and more fluid in motion scenes. Multi-scale image processing is then performed on the temporally filtered image sequence, decomposing each frame into multiple scales to obtain the initial multi-scale image sequence. The goal of multi-scale processing is to separate details from global information in an image, enabling more precise manipulation of features at different scales. Bilateral filtering is applied to each scale of the initial multi-scale image sequence. Bilateral filtering is a nonlinear filtering method that preserves edges while removing noise, ensuring sharp edges and low noise at all scales. This results in a filtered multi-scale image sequence. The filtered multi-scale image sequence is then reconstructed to restore the original resolution, yielding a reconstructed image sequence. The reconstruction process fuses image information from each scale, ensuring that the final image retains detail clarity while maintaining overall smoothness, improving visual quality. Three different compensation modes are applied to the reconstructed image sequence to improve its stability and adaptability. The first compensation mode, global compensation, uses a global affine transformation to process the image sequence. This mode applies global rotation, scaling, and translation operations to the entire image, making it suitable for relatively simple global motion compensation. The second compensation mode, local compensation, uses a grid deformation algorithm to divide the image into multiple grids and independently deform each grid to accommodate local motion variations. The local compensation method can handle non-uniform motion in images, resulting in a more refined compensation effect. The third compensation mode is the adaptive compensation mode. This mode dynamically selects a compensation strategy based on image content, combining global and local motion characteristics to compensate the image in a manner most appropriate for the current scene. The adaptive compensation mode is suitable for complex scenes and diverse motion. After applying the three different compensation modes, image sequences using the global compensation mode, the local compensation mode, and the adaptive compensation mode are obtained.To select the optimal compensation result, image sequences using the three compensation modes were evaluated for jitter. The root mean square error (RMS error) of the inter-frame differences was calculated. This error measures the stability between adjacent frames; smaller errors indicate a more stable image sequence. The compensation mode with the smallest error was selected as the target stabilized image sequence for output, ensuring the resulting image sequence has the best stabilization effect and visual quality in terms of motion compensation.
[0041] In embodiments of the present invention, low-pass filtering and adaptive Kalman filtering are performed on multiple IMU sensor data, effectively reducing IMU data noise and drift, and improving the accuracy and reliability of motion data. A linear convolution aliasing model is used to process IMU fusion data, enabling better capture of complex motion patterns and extracting high-dimensional motion feature vectors, providing richer information for subsequent motion parameter prediction. Combining a six-degree-of-freedom motion model based on a sine wave combination with a BP neural network prediction model enables accurate prediction of camera motion parameters, effectively addressing the difficulty of traditional methods in accurately describing complex motion. A look-ahead network compensation model and dynamic adjustment mechanism are introduced to adjust the compensation strategy in real time based on jitter intensity assessment results from the IMU data, enhancing compensation flexibility and adaptability. Combining dense optical flow optimization with temporally weighted background pixel pooling effectively preserves foreground information during image stabilization, improving the visual quality of the stabilization effect. Content-aware padding and temporal filtering address edge gaps and temporal discontinuities that may occur during image stabilization, further improving the quality of the stabilized image sequence. It provides multiple compensation modes (global compensation, local compensation, and adaptive compensation) and dynamically selects the optimal mode through jitter evaluation to meet the stability requirements in different scenarios.
[0042] In a specific embodiment, the process of executing step S1 may specifically include the following steps:
[0043] (1) Perform digital low-pass filtering on the acceleration and angular velocity raw data collected by each IMU sensor to obtain the first IMU data, and perform zero bias correction and temperature compensation on the first IMU data to obtain the second IMU data;
[0044] (2) Input the second IMU data into the adaptive Kalman filter, dynamically adjust the process noise covariance matrix according to the current motion state, obtain the third IMU data, and perform Allan variance analysis on the third IMU data to calculate the noise characteristic parameters of each IMU sensor and obtain the noise parameter matrix;
[0045] (3) Based on the noise parameter matrix, the third IMU data of multiple IMU sensors are weighted fused to obtain the fourth IMU data, and the attitude of the fourth IMU data is solved. The rotation matrix of the camera is calculated using the quaternion algorithm to obtain the camera attitude data;
[0046] (4) The camera attitude data is combined with the fourth IMU data and fused through a complementary filtering algorithm to obtain the corrected six-degree-of-freedom motion data. The corrected six-degree-of-freedom motion data is then time-synchronized and interpolated to obtain the six-degree-of-freedom IMU fusion data aligned with the camera frame rate.
[0047] Specifically, the acceleration and angular velocity raw data collected by each IMU sensor are digitally low-pass filtered to eliminate high-frequency noise components in the data, making the sensor data smoother and more suitable for subsequent processing. Assume that the raw acceleration data is , through a digital low-pass filter , the filtered data is expressed as:
[0048] ;
[0049] in, is the frequency response function of the filter, which can suppress the signal portion above a certain cutoff frequency to obtain the first IMU data. The first IMU data is subjected to zero bias correction and temperature compensation. The IMU sensor is affected by bias and temperature changes, which can cause the measured value to deviate from the true value. In order to correct these effects, zero bias correction is performed, and the correction process is expressed as:
[0050] ;
[0051] in, Is the bias value of the IMU in a static state, which is measured experimentally and used for subsequent compensation. Temperature compensation is achieved by modeling the change of sensor data with temperature. The temperature model is expressed as ,in The temperature is the temperature, and the compensated IMU data is the second IMU data. The second IMU data is input into the adaptive Kalman filter for processing. The Kalman filter is an optimal estimator that estimates the state of the system by combining the prediction model and measurement data. In order to adapt to the dynamically changing motion state, the adaptive Kalman filter dynamically adjusts the process noise covariance matrix according to the current motion state. Assume that the system state is , the process model is:
[0052] ;
[0053] in, is the state transition matrix, is the process noise, and its covariance is Adaptive Kalman filter adjusts according to the motion state , to improve the accuracy of filtering. After Kalman filtering, the third IMU data is obtained. In order to analyze the noise characteristics of the IMU sensor, AIlan variance analysis is performed on the third IMU data. Allan variance analysis is a tool for evaluating noise characteristics. By calculating the variance of the data at different time intervals, different noise sources in the sensor can be identified. Through Allan variance analysis of the third IMU data, the noise characteristic parameters of each IMU sensor are obtained, such as random walk, bias instability, etc., and a noise parameter matrix is constructed based on these parameters for weighted processing of subsequent data fusion. Based on the noise parameter matrix, the third IMU data of multiple IMU sensors are weightedly fused. According to the noise characteristics of each IMU sensor, the data of each sensor are weightedly summed to improve the accuracy of the overall data. Assume For the The data of IMU sensors, weighted by , then the fourth IMU data after weighted fusion Expressed as
[0054] ;
[0055] in, is the number of IMU sensors, weight The noise parameter matrix is used to allocate the characteristic parameters, and the sensor data with less noise is given a greater weight. The attitude of the fourth IMU data is solved, and the camera rotation matrix is calculated using the quaternion algorithm. The quaternion representation has an important application in attitude solution, especially in avoiding the singularity problem in the Euler angle representation (i.e., gimbal lock). Assume that the current quaternion is , the camera's rotation matrix Calculated by quaternion:
[0056] ;
[0057] Quaternion calculations are used to obtain camera attitude data, which describes the camera's rotation. This camera attitude data is combined with the fourth IMU data and fused using a complementary filtering algorithm to obtain corrected 6DOF motion data. Complementary filtering combines low-frequency and high-frequency data. The IMU data provides short-term, high-frequency motion information, while the attitude data provides long-term, low-frequency stability information. This fusion yields more accurate 6DOF motion data. The corrected 6DOF motion data is time-synchronized and interpolated to align with the camera's frame rate. The IMU sensor's data acquisition frequency is typically higher than the camera's frame rate. Therefore, the IMU data is downsampled or interpolated to ensure temporal alignment between the motion data and the image sequence. Interpolation is performed using methods such as linear interpolation and spline interpolation, ensuring accurate 6DOF motion data at every point in time during camera acquisition. Ultimately, 6DOF IMU-fused data is aligned with the camera's frame rate and used for camera motion compensation and image stabilization.
[0058] In a specific embodiment, the process of executing step S2 may specifically include the following steps:
[0059] (1) The linear convolutional aliasing model consists of three one-dimensional convolutional layers, two maximum pooling layers, and two fully connected layers. The first convolutional layer uses 32 one-dimensional convolution kernels of size 5 with a stride of 1, the second convolutional layer uses 64 one-dimensional convolution kernels of size 3 with a stride of 1, and the third convolutional layer uses 128 one-dimensional convolution kernels of size 3 with a stride of 1.
[0060] (2) The six-degree-of-freedom IMU fusion data is divided into time windows to obtain multiple sub-signal sequences, and the multiple sub-signal sequences are input into the first one-dimensional convolution layer of the linear convolution aliasing model to perform a one-dimensional convolution operation. The ReLU activation function is used to obtain 32 first feature maps, each of which has a size of 96;
[0061] (3) Perform a maximum pooling operation on the 32 first feature maps, with a pooling window size of 2 and a step size of 2, to obtain 32 downsampled second feature maps with a size of 48. The 32 downsampled second feature maps are input into the second one-dimensional convolution layer for a one-dimensional convolution operation. Using the ReLU activation function, 64 third feature maps are obtained, each with a size of 46.
[0062] (4) Perform a maximum pooling operation on the 64 third feature maps, with a pooling window size of 2 and a step size of 2, to obtain 64 downsampled fourth feature maps with a size of 23. The 64 downsampled fourth feature maps are input into the third one-dimensional convolution layer for a one-dimensional convolution operation. Using the ReLU activation function, 128 fifth feature maps are obtained, each with a size of 21.
[0063] (5) Flatten the 128 fifth feature maps to obtain a 2688-dimensional feature vector, and input it into the first fully connected layer. The number of neurons in the fully connected layer is 1024, and the activation function uses ReLU to obtain a 1024-dimensional feature vector. The 1024-dimensional feature vector is batch normalized to obtain a normalized feature vector.
[0064] (6) The normalized feature vector is input into the autoencoder network. The autoencoder network includes an encoder and a decoder. The encoder consists of three fully connected layers with 512, 256 and 128 neurons respectively. The decoder consists of three fully connected layers with 256, 512 and 1024 neurons respectively. The encoder of the autoencoder network performs dimensionality reduction on the normalized feature vector to obtain a 128-dimensional compressed feature. The 128-dimensional compressed feature is used as input and reconstructed by the decoder to obtain a high-dimensional motion feature vector.
[0065] Specifically, the linear convolutional model consists of three one-dimensional convolutional layers, two maximum pooling layers, and two fully connected layers. The first convolutional layer uses 32 one-dimensional convolution kernels of size 5 and a stride of 1 to capture the short-term time-related features in the original input data. Assume that the input six-degree-of-freedom IMU fusion data is , its feature representation after convolution operation is:
[0066] ;
[0067] in, represents the weight of the first layer of convolution kernel, Represents the bias term, and 32 first feature maps are obtained through convolution operation, and the size of each feature map is 96. Apply the ReLU activation function to the feature map , introducing nonlinearity to help the network gain stronger expressive power when learning complex motion features. A maximum pooling operation is performed on the 32 first feature maps, with a pooling window size of 2 and a stride of 2. The purpose of the maximum pooling operation is to downsample the feature maps, reducing the feature dimensionality and computational complexity, while also increasing feature invariance. Through maximum pooling, the feature map size is reduced from 96 to 48, resulting in 32 downsampled second feature maps. These downsampled feature maps are then fed into the second convolutional layer, which uses 64 one-dimensional convolution kernels of size 3 with a stride of 1 to further extract and refine features. After convolution, the ReLU activation function is applied to produce 64 third feature maps, each with a size of 46. The 64 third feature maps are again max-pooled with a pooling window size of 2 and a stride of 2, reducing the feature map size to 23, resulting in 64 downsampled fourth feature maps. The 64 downsampled fourth feature maps are input to the third convolutional layer, which uses 128 one-dimensional convolution kernels of size 3 with a stride of 1. After convolution and pooling, the 128 fifth feature maps are flattened into a vector with a total dimension of 128 × 21 = 2688. This flattened feature vector contains rich motion feature information extracted by the convolutional neural network. This flattened feature vector is input to the first fully connected layer, which contains 1024 neurons and uses the ReLU activation function to perform nonlinear mapping on the input features, resulting in a 1024-dimensional feature vector. To enhance feature stability and accelerate network convergence during training, the 1024-dimensional feature vector is batch normalized. Batch normalization adjusts the mean and variance of the features to conform to a standard normal distribution, improving model training efficiency and performance. In order to reduce the dimension of the feature and extract the most representative motion information, the normalized feature vector is input into the autoencoder network. The autoencoder network consists of an encoder and a decoder. The encoder part includes three fully connected layers with 512, 256 and 128 neurons respectively. The main function of the encoder is to gradually compress the input high-dimensional features to extract a more compact feature representation. Assume that the feature representation after the encoder is , then the calculation process is:
[0068] ;
[0069] in, Represent the weight matrices of the three fully connected layers, represents the bias term, The activation function (such as ReLU) is used to generate 128-dimensional compressed features. These compressed features contain the core motion information in the original high-dimensional features while eliminating redundancy and noise. The 128-dimensional compressed features are input into the decoder, which consists of three fully connected layers with 256, 512, and 1024 neurons, respectively. Through layer-by-layer reverse mapping, the decoder restores the compressed features to the high-dimensional space, obtaining a motion feature vector with the same dimensionality as the original high-dimensional features. Through autoencoder training, the network learns how to effectively represent and reconstruct motion features, ensuring that key information is not lost during the feature compression process.
[0070] In a specific embodiment, the process of executing step S3 may specifically include the following steps:
[0071] A1: Extract and match feature points from adjacent frames in the original image sequence to obtain a feature point correspondence matrix. Based on this matrix, remove abnormal matching points and calculate the basic matrix. This solves the camera's relative rotation matrix R and translation vector t to obtain an initial estimate of the image motion.
[0072] A2: Input the high-dimensional motion feature vector into the BP neural network prediction model. The output layer of the BP neural network prediction model corresponds to the motion parameters of the six degrees of freedom. The initial motion parameter prediction values based on the IMU data are obtained through forward propagation calculation.
[0073] A3: Data fusion is performed on the initial estimated image motion value and the initial motion parameter prediction value based on IMU data to obtain the fused camera motion parameters;
[0074] A4: Substitute the fused camera motion parameters into a six-degree-of-freedom motion model based on a sine wave combination. This six-degree-of-freedom motion model represents the camera's three-dimensional translation and three-dimensional rotation as the superposition of six sine wave functions. The camera's current position and attitude are calculated.
[0075] A5: Kalman filtering is performed on the camera's current position and attitude. The filter's state vector contains position, velocity, and acceleration, and the measurement vector is the fused camera motion parameter. Smoothed camera motion parameters are obtained through a prediction-update loop.
[0076] A6: Based on the smoothed camera motion parameters, the current frame in the original image sequence is back-projected, and the Euclidean distance between the projected point and the actual feature point is calculated to obtain the reprojection error.
[0077] A7: When the reprojection error is less than the preset threshold, the current smoothed camera motion parameters are output as the camera motion parameter prediction values. If it is greater than the preset threshold, return to step A3, adjust the data fusion weights, and re-fuse and process the data until the accuracy requirements are met or the maximum number of iterations is reached.
[0078] Specifically, feature point extraction and feature point matching are performed on adjacent frames in the original image sequence. The SIFT or ORB feature extraction algorithm is used to extract unique feature points from the image. These feature points are consistent between different frames and are used to estimate camera motion. After extracting the feature points, a feature point correspondence matrix between adjacent frames is established through a matching algorithm (such as nearest neighbor matching). The accuracy of matching is improved by removing abnormal matching points. The RANSAC (random sampling consensus) algorithm is used to remove abnormal matching points to ensure that the final matched feature point pairs are reliable. Based on the feature point matching relationship, the basic matrix F is calculated. The basic matrix defines the geometric relationship between the two images. The camera's relative rotation matrix R and translation vector t are solved through the basic matrix. These parameters describe the relative motion of the camera between two adjacent frames, that is, the initial estimate of the image motion. The high-dimensional motion feature vector is input into the BP neural network prediction model. The BP neural network is a multi-layer perceptron with nonlinear mapping capabilities. By learning from training data, it can predict the motion state of the camera. The output layer of the BP neural network corresponds to the motion parameters of six degrees of freedom, which are three-dimensional translation, and three-dimensional rotation In the process of forward propagation, the BP neural network uses the high-dimensional features of the input and outputs the predicted six-degree-of-freedom motion parameters through step-by-step calculations in the hidden layer, which can be expressed as:
[0079] ;
[0080] in, is the input high-dimensional motion feature vector, is the weight matrix of each layer, is the bias, Represents the activation function (such as ReLU). Through forward propagation, the initial motion parameter prediction value based on IMU data is obtained. The initial image motion estimation value and the initial motion parameter prediction value based on IMU data are fused to obtain more accurate camera motion parameters. Data fusion is achieved by weighted averaging, where the weights of the image motion estimation and the IMU prediction value are assigned according to their respective uncertainties. Assume that the image estimation value is , the IMU prediction value is , then the fused camera motion parameters Expressed as:
[0081] ;
[0082] in, and Represent the weights of image and IMU data respectively, and satisfy By properly allocating weights, we balance the advantages and disadvantages of image and IMU data, thus improving the accuracy of motion estimation. The fused camera motion parameters are substituted into a six-degree-of-freedom motion model based on a sine wave combination, which is used to model the camera motion more smoothly. Assuming that the three-dimensional translation and three-dimensional rotation of the camera are represented as the superposition of six sine wave functions, they can be expressed as:
[0083] ;
[0084] in, are the amplitudes of the sine waves, is the angular frequency, is the initial phase. By superimposing the sine function, the periodic motion characteristics of the camera are better described, so that the position and attitude of the camera at the current moment are calculated based on the fused motion parameters. The Kalman filter is used to process the camera position and attitude at the current moment. The Kalman filter is an optimal state estimation method suitable for state estimation of linear dynamic systems. The state vector of the filter contains position, velocity and acceleration, which is recorded as , the measurement vector is the fused camera motion parameter, and the smoothed camera motion parameter is obtained through the prediction-update cycle of the Kalman filter. The prediction equation of the Kalman filter is expressed as:
[0085] ;
[0086] in, is the state transition matrix, is the control input matrix, To control the input vector, the motion trajectory is smoothed through prediction and update steps to reduce the influence of noise. Based on the smoothed camera motion parameters, the current frame in the original image sequence is back-projected and the Euclidean distance between the projected point and the actual feature point is calculated. Assume that the coordinates of the projected point are , the actual feature point coordinates are , then the Euclidean distance Expressed as:
[0087] ;
[0088] By calculating this distance, we obtain the reprojection error, which is used to evaluate the accuracy of the motion parameters. When the reprojection error is less than a preset threshold, the current smoothed camera motion parameters are output as the predicted values of the camera motion parameters. If the reprojection error is greater than the preset threshold, we return to step A3, adjust the data fusion weights, and repeat the data fusion and processing until the accuracy requirements are met or the maximum number of iterations is reached.
[0089] In a specific embodiment, the process of executing step S4 may specifically include the following steps:
[0090] (1) Perform time series analysis on the camera motion parameters to obtain the predicted camera motion parameter sequence, and input the predicted camera motion parameter sequence into the advance network compensation model. The initial compensation model parameter matrix is obtained by forward propagation calculation. The advance network compensation model contains three fully connected layers, with 64, 32, and 16 neurons in each layer, respectively. The ReLU activation function is used, and the output layer corresponds to a 3×3 compensation model parameter matrix.
[0091] (2) Calculate the variance of the six-degree-of-freedom IMU fusion data within the sliding time window to obtain the IMU data variance sequence, and perform fast Fourier transform on the IMU data variance sequence to obtain the spectral characteristics of the IMU data;
[0092] (3) Based on the spectral characteristics of the IMU data, the peak amplitude and frequency are calculated to obtain the jitter intensity feature vector, and the jitter intensity feature vector is input into the jitter intensity assessment model. The jitter intensity assessment score is calculated by the jitter intensity assessment model, and the jitter intensity assessment model adopts the support vector regression algorithm;
[0093] (4) According to the jitter intensity evaluation score, the compensation intensity coefficient is calculated using the preset adaptive weight function, and the compensation intensity coefficient is multiplied by each element of the initial compensation model parameter matrix to dynamically adjust the compensation model parameter matrix.
[0094] Specifically, the camera motion parameters are analyzed in time series to predict the camera motion state at the future moment. Assuming the camera motion parameters are , predict the future motion parameter sequence through time series analysis model (such as autoregressive model or long short-term memory network LSTM), expressed as Time series analysis can capture the trend and periodicity of camera motion. The predicted motion parameter sequence is input into the advance network compensation model, which consists of three fully connected layers, each with 64, 32 and 16 neurons respectively. The ReLU activation function is used to introduce nonlinearity, thereby enhancing the network's expressive power. Through forward propagation calculation, the 3×3 compensation model parameter matrix of the output layer is obtained. . Assume that the input vector is , after each layer of transformation, it is expressed as:
[0095] ;
[0096] in, is the weight matrix, is the bias, ReLU function ReLU Used to introduce nonlinearity. The output of the forward propagation process is a 3×3 initial compensation model parameter matrix, which is used to describe the preliminary geometric transformation relationship of camera motion compensation. In order to optimize the parameters of the compensation model, the jitter characteristics in the IMU data are analyzed. The variance of the six-degree-of-freedom IMU fusion data is calculated in the sliding time window to obtain the variance sequence of the IMU data. The variance sequence can reflect the degree of fluctuation of the IMU data at different times. Assume that the IMU data is , then in the sliding time window The variance within is calculated as:
[0097] ;
[0098] in, is the mean of the data in the window, For in time The variance at each moment. Perform a fast Fourier transform on this variance sequence to obtain the spectral characteristics of the IMU data. The spectral characteristics describe the energy distribution of the data at different frequency components and can help identify periodic jitter in camera motion. Based on the spectrum analysis results, calculate the peak amplitude and corresponding frequency in the spectrum to obtain the jitter intensity feature vector ,in Indicates the The peak amplitude, Represents the corresponding frequency. These features reflect the impact of different frequency components on the camera during movement. The input is fed into the jitter intensity assessment model, which uses the support vector regression (SVR) algorithm to calculate the jitter intensity assessment score by learning the relationship between spectrum features and jitter intensity. The goal of support vector regression is to find a function that can accurately describe the relationship between input and output, so as to minimize the error. The jitter intensity evaluation score is used to represent the degree of jitter in the current camera motion. , calculate the compensation intensity coefficient using the preset adaptive weight function The adaptive weight function is obtained based on experience or through training to ensure that appropriate compensation can be provided under different jitter intensities. For example, the compensation intensity coefficient Expressed as:
[0099] ;
[0100] in, To adjust the parameters, is the threshold value, and its function form is similar to the Sigmoid function, which is used to map the jitter assessment score to a reasonable compensation intensity range. When the jitter intensity is large, the compensation intensity coefficient is large, and vice versa. Multiply each element of the initial compensation model parameter matrix to achieve dynamic adjustment of the compensation model parameter matrix. The adjusted compensation model parameter matrix Expressed as:
[0101] ;
[0102] The dynamic adjustment mechanism enables the compensation model to adapt to different motion environments. When the jitter is large, the compensation strength is increased to reduce the unstable effects caused by motion, while when the jitter is small, the compensation is reduced to avoid unnatural images caused by over-correction.
[0103] In a specific embodiment, the process of executing step S5 may specifically include the following steps:
[0104] (1) Applying the compensation model parameter matrix to the original image data, performing affine transformation to obtain a first image sequence, and performing dense optical flow field calculation between adjacent frames in the first image sequence to obtain a motion vector field;
[0105] (2) Based on the motion vector field, local motion compensation and pixel resampling are performed on the first image sequence to obtain a second image sequence. A time-weighted background pixel pool is constructed for the second image sequence. A fixed-size pixel value queue is maintained for each pixel position. New pixel values are queued with a time-attenuated weight. When the queue is full, the oldest pixel value is dequeued to obtain a dynamically updated background model.
[0106] (3) Based on the dynamically updated background model, foreground detection is performed on each frame of the second image sequence to obtain a foreground mask;
[0107] (4) Perform pixel-by-pixel AND operation on the foreground mask and the second image sequence to extract the foreground area and obtain the foreground image sequence. Perform inter-frame difference between the foreground image sequence and the second image sequence, calculate the absolute difference, and perform adaptive threshold segmentation to obtain the initial stable image sequence.
[0108] Specifically, the compensation model parameter matrix is applied to the original image data for affine transformation. Affine transformation is a two-dimensional spatial transformation that performs operations such as rotation, scaling, and translation on the image to correct the irregular transformation caused by camera motion. Assume that the compensation model parameter matrix is , the coordinates of the original image are , the new coordinates after affine transformation are , the formula of affine transformation is expressed as:
[0109] ;
[0110] in, are the scaling and rotation parameters, is the translation parameter. Through affine transformation, a first image sequence is obtained, in which the original image data is compensated so that the effect of camera motion on the image is preliminarily corrected. Dense optical flow field is calculated between adjacent frames in the first image sequence to obtain the motion vector field of each pixel position. Dense optical flow field is used to describe the pixel motion between adjacent frames. Assume that the pixel point in the first frame The corresponding position in the next frame is ,in Respectively expressed in and The purpose of optical flow calculation is to find the corresponding position of each pixel in the next frame and obtain the motion vector field. The calculation of optical flow field is achieved through the optical flow constraint equation:
[0111] ;
[0112] in, Respectively represent the images in and The gradient in direction, Represents the gradient of the image in time. By solving this equation, we can get the motion vector of each pixel. . Based on the obtained motion vector field, local motion compensation and pixel resampling are performed on the first image sequence to obtain the second image sequence. The purpose of local motion compensation is to eliminate subtle motions in local areas of the image so that objects in the image remain aligned between adjacent frames and improve image stability. According to the motion vector of each pixel position, the pixel is repositioned and the pixel value is interpolated to ensure that the compensated image is visually continuous and consistent. A time-weighted background pixel pool is constructed for the second image sequence. For each pixel position, a fixed-size pixel value queue is maintained, and new pixel values are queued with time-attenuated weights. When the queue is full, the oldest pixel value is dequeued to achieve dynamic updating of the background model. Assume that at time The pixel value of a certain pixel position at the moment is , its weight when joining the queue is ,in If is the temporal decay coefficient, the influence of pixel values in the background pixel pool gradually decreases over time, and new pixel values gain a greater weight. This approach ensures that the background model adapts to changes in the scene, such as gradual changes in lighting and slowly moving objects in the background. Based on the dynamically updated background model, foreground detection is performed on each frame in the second image sequence to generate a foreground mask. The goal of foreground detection is to separate moving objects from the static background. This is typically achieved by comparing the current frame with the background model. If the difference between the position of a pixel and the background model exceeds a preset threshold, the pixel is considered to be foreground; otherwise, it is considered background. This generates a foreground mask, a binary image used to identify moving objects in the image. The foreground mask is then ANDed pixel by pixel with the second image sequence to extract the foreground region, resulting in a foreground image sequence. Pixels marked as foreground in the foreground mask are multiplied with the corresponding pixels in the original image, and the pixel values at all other locations are set to 0, resulting in a foreground image containing only moving objects. Inter-frame difference is performed on the foreground image sequence and the second image sequence. By calculating the absolute difference between adjacent frames, changing regions in the scene can be identified. Assume that the two frames in the second image sequence are and , then the inter-frame difference is expressed as:
[0113] ;
[0114] in, Represents the absolute difference image between frames, reflecting the pixel difference between the two frames. , an adaptive threshold segmentation method is applied to determine which differences belong to valid foreground motion and which belong to noise.,Adaptive threshold segmentation dynamically selects the threshold according to the grayscale histogram of the image, thereby adapting to the lighting and motion conditions of different scenes, and finally obtaining an initial stable image sequence.
[0115] In a specific embodiment, the process of executing step S6 may specifically include the following steps:
[0116] (1) Detect and extend the edge region of each frame in the initial stabilized image sequence to obtain an extended image sequence, and fill the blank region based on the extended image sequence and the foreground mask to obtain a filled image sequence;
[0117] (2) Perform time-domain Kalman filtering on the padded image sequence, with the state vector being the pixel value and the observation vector being the pixel value at the corresponding position of the adjacent frame, to obtain the image sequence after time-domain filtering;
[0118] (3) Performing multi-scale image processing based on the image sequence after time domain filtering to obtain an initial multi-scale image sequence, and performing bilateral filtering on each scale of the initial multi-scale image sequence to obtain a filtered multi-scale image sequence;
[0119] (4) Reconstructing the filtered multi-scale image sequence to the original resolution to obtain a reconstructed image sequence;
[0120] (5) Apply three different compensation modes to the reconstructed image sequence: global compensation mode, local compensation mode, and adaptive compensation mode. The global compensation mode uses global affine transformation, the local compensation mode uses a grid deformation algorithm, and the adaptive compensation mode dynamically selects a compensation strategy based on the image content. The image sequences of the global compensation mode, the local compensation mode, and the adaptive compensation mode are obtained.
[0121] (6) Perform jitter evaluation on the image sequence of the global compensation mode, the image sequence of the local compensation mode, and the image sequence of the adaptive compensation mode, calculate the root mean square error of the inter-frame difference, and select the compensation mode with the smallest error as the target stable image sequence output.
[0122] Specifically, the edge region of each frame in the initial stable image sequence is detected and edge extended. The edge detection adopts the classic Canny edge detection method, which can identify the significant edges in the image. Assume that the image is , after edge detection, we get the edge image , then expand the edge area to obtain the extended image sequence Fill the blank area based on the expanded image and the foreground mask. The blank area is filled using a content-aware filling method based on the expanded area, assuming that the pixel value of the blank area is The filling process interpolates or copies the pixel values from the adjacent known areas to obtain the filled pixel values. To ensure that the blank part of the image is naturally integrated into the overall content, forming a filled image sequence The padded image sequence is subjected to time-domain Kalman filtering to smooth the temporal noise in the image sequence. Kalman filtering is an optimal state estimation method for linear dynamic systems. The state vector is defined as the pixel value, and the observation vector is the pixel value at the corresponding position of the adjacent frame. Assume that the state vector is , the observation vector is , the prediction equation of Kalman filter is:
[0123] ;
[0124] in, is the state transition matrix, is the control matrix, is the control input. Through the prediction and update steps, Kalman filtering is performed on each pixel value to obtain the image sequence after time domain filtering. This process effectively reduces the short-term jitter in the image sequence, making the image more stable. Multi-scale image processing is performed based on the image sequence after time domain filtering, and the image is decomposed into multiple scales to process different levels of detail respectively. Each frame of the image is decomposed into multiple scales through the Gaussian pyramid to obtain the initial multi-scale image sequence , each scale contains images of different resolutions. The lower scale is used to capture the global features of the image, and the higher scale is used to process the details. At each scale, bilateral filtering is applied to the image. Bilateral filtering is an edge-preserving filtering method that can remove noise while keeping edges clear. The formula for bilateral filtering is:
[0125] ;
[0126] in, and are Gaussian functions of spatial distance and pixel intensity difference, respectively, is the normalization coefficient. The filtered multi-scale image sequence is obtained by bilateral filtering The filtered multi-scale image sequence is reconstructed to the original resolution to form a reconstructed image sequence The reconstruction process is to perform step-by-step upsampling and additive fusion of the multi-scale pyramid to restore the original resolution, ensuring that the image maintains global consistency while retaining details, thereby improving visual quality. Three different compensation modes are applied to the reconstructed image sequence, namely global compensation mode, local compensation mode, and adaptive compensation mode. In the global compensation mode, the image is processed using the overall affine transformation, and the affine transformation formula is as follows:
[0127] ;
[0128] in is the affine transformation matrix, including rotation, scaling and translation parameters. Global compensation is applicable to situations where the overall motion of the camera is relatively consistent. In the local compensation mode, the image is processed using a grid deformation algorithm, and the image is divided into multiple grids. Each grid is deformed independently to adapt to the local motion. This method can compensate for local complex motion more finely. In the adaptive compensation mode, the compensation strategy is dynamically selected according to the image content. If global motion is detected, affine transformation is used for compensation. If strong local motion is detected, the grid deformation method is used. The compensation strategy is automatically adjusted according to the image content to adapt to complex scenes. After applying the three compensation modes, the image sequence of the global compensation mode is obtained. , image sequence of local compensation mode and image sequences in adaptive compensation mode In order to select the best compensation result, the jitter evaluation of the three image sequences is performed and the root mean square error (RMS error) of the inter-frame difference is calculated to measure the smoothness of the compensated image sequence. Assume that the two frames of image are and , then the root mean square error of the inter-frame difference is calculated as:
[0129] ;
[0130] in, is the total number of pixels in the image, and Respectively represent pixels in time and By calculating the RMS error of the image sequences under the three compensation modes, the compensation mode with the smallest error is selected as the target stable image sequence output, which means that this compensation mode performs best in smoothing the image sequence and reducing jitter.
[0131] The above describes the camera motion compensation method based on IMU in the embodiment of the present invention. The following describes the camera motion compensation device based on IMU in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, an IMU-based camera motion compensation device includes:
[0132] The acquisition module is used to perform low-pass filtering and adaptive Kalman filtering on the raw data collected by multiple IMU sensors to obtain six-degree-of-freedom IMU fusion data;
[0133] The extraction module is used to input the six-degree-of-freedom IMU fusion data into the linear convolution aliasing model for processing, obtain simulated motion data, and extract the corresponding high-dimensional motion feature vector;
[0134] The fusion module is used to fuse the high-dimensional motion feature vector and the original image sequence using a six-degree-of-freedom motion model based on a sine wave combination and a BP neural network prediction model to obtain the predicted value of the camera motion parameter;
[0135] The evaluation module is used to build a look-ahead network compensation model based on camera motion parameters and dynamically adjust the compensation model parameter matrix based on the jitter intensity evaluation results of IMU data variance and peak analysis;
[0136] The transformation module is used to perform geometric transformation on the original image data based on the compensation model parameter matrix, and combines dense optical flow field optimization and time-weighted background pixel pool modeling to obtain the initial stable image sequence and foreground mask;
[0137] The output module is used to perform content-aware filling and temporal filtering on the initial stabilized image sequence and foreground mask, and output target stabilized image sequences with multiple compensation modes.
[0138] Through the collaborative efforts of the aforementioned components, low-pass filtering and adaptive Kalman filtering of multiple IMU sensor data effectively reduce IMU data noise and drift, improving the accuracy and reliability of motion data. A linear convolution aliasing model is used to process IMU fusion data, enabling better capture of complex motion patterns and extracting high-dimensional motion feature vectors, providing richer information for subsequent motion parameter prediction. Combining a six-degree-of-freedom motion model based on a sine wave combination with a BP neural network prediction model enables accurate prediction of camera motion parameters, effectively addressing the difficulty of traditional methods in accurately describing complex motion. A look-ahead network compensation model and dynamic adjustment mechanism are introduced to adjust the compensation strategy in real time based on jitter intensity assessments of IMU data, enhancing compensation flexibility and adaptability. Combining dense optical flow optimization with temporally weighted background pixel pooling modeling, this approach effectively preserves foreground information during image stabilization, improving the visual quality of the stabilization effect. Content-aware padding and temporal filtering address edge gaps and temporal discontinuities that may occur during image stabilization, further improving the quality of the stabilized image sequences. It provides multiple compensation modes (global compensation, local compensation, and adaptive compensation) and dynamically selects the optimal mode through jitter evaluation to meet the stability requirements in different scenarios.
[0139] Reference Figure 3 In an embodiment of the present invention, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3 As shown. The computer device includes a processor, memory, display screen, input device, network interface and database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above method is implemented.
[0140] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0141] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the above-described method when executed by a processor. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0142] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media provided herein and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM.
[0143] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0144] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0145] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A camera motion compensation method based on IMU, characterized in that: The method comprises: The acceleration and angular velocity raw data collected by multiple IMU sensors are processed by low-pass filtering and adaptive Kalman filtering to obtain six-degree-of-freedom IMU fusion data; Inputting the six-degree-of-freedom IMU fusion data into a linear convolution aliasing model for processing to obtain simulated motion data and extracting the corresponding high-dimensional motion feature vector; According to the high-dimensional motion feature vector and the original image sequence, a six-degree-of-freedom motion model based on a sine wave combination and a BP neural network prediction model are fused to obtain a camera motion parameter prediction value; Building a look-ahead network compensation model based on the camera motion parameters, and dynamically adjusting the compensation model parameter matrix based on the jitter intensity evaluation results of the IMU data variance and peak analysis; Performing geometric transformation on the original image sequence based on the compensation model parameter matrix, and combining dense optical flow field optimization and time-weighted background pixel pool modeling to obtain an initial stable image sequence and a foreground mask; Content-aware filling and temporal filtering are performed on the initial stabilized image sequence and the foreground mask, and target stabilized image sequences in multiple compensation modes are output.
2. The camera motion compensation method based on IMU according to claim 1, characterized in that: The acceleration and angular velocity raw data collected by multiple IMU sensors are processed by low-pass filtering and adaptive Kalman filtering to obtain six-degree-of-freedom IMU fusion data, including: Performing digital low-pass filtering on the raw acceleration and angular velocity data collected by each IMU sensor to obtain first IMU data, and performing zero bias correction and temperature compensation on the first IMU data to obtain second IMU data; Input the second IMU data into an adaptive Kalman filter, dynamically adjust the process noise covariance matrix according to the current motion state, obtain third IMU data, and perform Allan variance analysis on the third IMU data to calculate the noise characteristic parameters of each IMU sensor to obtain a noise parameter matrix; Based on the noise parameter matrix, weightedly fuse the third IMU data of the multiple IMU sensors to obtain fourth IMU data, perform attitude calculation on the fourth IMU data, calculate the rotation matrix of the camera using a quaternion algorithm, and obtain camera attitude data; The camera posture data is combined with the fourth IMU data and fused through a complementary filtering algorithm to obtain corrected six-degree-of-freedom motion data, and the corrected six-degree-of-freedom motion data is time synchronized and interpolated to obtain six-degree-of-freedom IMU fusion data aligned with the camera frame rate.
3. The camera motion compensation method based on IMU according to claim 2, characterized in that: The six-degree-of-freedom IMU fusion data is input into the linear convolution aliasing model for processing to obtain simulated motion data and extract the corresponding high-dimensional motion feature vector, including: The linear convolutional aliasing model includes three one-dimensional convolutional layers, two maximum pooling layers and two fully connected layers, wherein the first convolutional layer uses 32 one-dimensional convolution kernels of size 5 with a step size of 1, the second convolutional layer uses 64 one-dimensional convolution kernels of size 3 with a step size of 1, and the third convolutional layer uses 128 one-dimensional convolution kernels of size 3 with a step size of 1; Performing time window segmentation on the six-degree-of-freedom IMU fusion data to obtain multiple sub-signal sequences, and inputting the multiple sub-signal sequences into the first one-dimensional convolution layer of the linear convolution aliasing model to perform a one-dimensional convolution operation, using a ReLU activation function to obtain 32 first feature maps, each of which has a size of 96; Perform a maximum pooling operation on the 32 first feature maps with a pooling window size of 2 and a step size of 2 to obtain 32 downsampled second feature maps with a size of 48. Input the 32 downsampled second feature maps into the second one-dimensional convolution layer for a one-dimensional convolution operation. Use the ReLU activation function to obtain 64 third feature maps with a size of 46. Perform a maximum pooling operation on the 64 third feature maps, with a pooling window size of 2 and a step size of 2, to obtain 64 downsampled fourth feature maps, the size of the fourth feature map is 23, and the 64 downsampled fourth feature maps are input into the third one-dimensional convolution layer for a one-dimensional convolution operation. A ReLU activation function is used to obtain 128 fifth feature maps, each of which has a size of 21; Flatten the 128 fifth feature maps to obtain a 2688-dimensional feature vector, and input it into the first fully connected layer. The number of neurons in the fully connected layer is 1024, and the activation function adopts ReLU to obtain a 1024-dimensional feature vector. The 1024-dimensional feature vector is batch normalized to obtain a normalized feature vector. The normalized feature vector is input into the autoencoder network, which includes an encoder and a decoder. The encoder consists of three fully connected layers with 512, 256 and 128 neurons respectively, and the decoder consists of three fully connected layers with 256, 512 and 1024 neurons respectively. The normalized feature vector is subjected to dimensionality reduction processing by the encoder of the autoencoder network to obtain 128-dimensional compressed features. The 128-dimensional compressed features are used as input and reconstructed by the decoder to obtain a high-dimensional motion feature vector.
4. The IMU-based camera motion compensation method according to claim 3, wherein: The method of obtaining a camera motion parameter prediction value by fusing the high-dimensional motion feature vector and the original image sequence with a six-degree-of-freedom motion model based on a sine wave combination and a BP neural network prediction model includes: A1: Extract and match feature points from adjacent frames in the original image sequence to obtain a feature point correspondence matrix. Based on this feature point correspondence matrix, remove abnormal matching points and calculate the basic matrix. This solves the camera's relative rotation matrix R and translation vector t to obtain an initial estimate of the image motion. A2: Input the high-dimensional motion feature vector into a BP neural network prediction model. The output layer of the BP neural network prediction model corresponds to the motion parameters of six degrees of freedom. The initial motion parameter prediction value based on the IMU data is obtained through forward propagation calculation; A3: performing data fusion on the initial image motion estimation value and the initial motion parameter prediction value based on the IMU data to obtain fused camera motion parameters; A4: Substituting the fused camera motion parameters into a six-degree-of-freedom motion model based on a sine wave combination, the six-degree-of-freedom motion model based on a sine wave combination represents the camera's three-dimensional translation and three-dimensional rotation as a superposition of six sine wave functions, and calculating the camera's position and posture at the current moment; A5: Perform Kalman filtering on the camera's current position and attitude. The filter's state vector contains position, velocity, and acceleration. The measurement vector is the fused camera motion parameter. Smoothed camera motion parameters are obtained through a prediction-update loop. A6: Based on the smoothed camera motion parameters, reversely project the current frame in the original image sequence, calculate the Euclidean distance between the projected point and the actual feature point, and obtain a reprojection error; A7: When the reprojection error is less than a preset threshold, the current smoothed camera motion parameters are output as camera motion parameter prediction values; if it is greater than the preset threshold, return to step A3, adjust the data fusion weights, and re-fuse and process the data until the accuracy requirements are met or the maximum number of iterations is reached.
5. The camera motion compensation method based on IMU according to claim 4, characterized in that: The advanced network compensation model is constructed based on the camera motion parameters, and the compensation model parameter matrix is dynamically adjusted according to the jitter intensity evaluation results of the IMU data variance and peak analysis, including: Performing time series analysis on the camera motion parameters to obtain a predicted camera motion parameter sequence, and inputting the predicted camera motion parameter sequence into the advance network compensation model. An initial compensation model parameter matrix is obtained by forward propagation calculation. The advance network compensation model includes three fully connected layers, with 64, 32, and 16 neurons in each layer, respectively. A ReLU activation function is used, and the output layer corresponds to a 3×3 compensation model parameter matrix. Calculating the variance of the six-degree-of-freedom IMU fusion data within a sliding time window to obtain an IMU data variance sequence, and performing a fast Fourier transform on the IMU data variance sequence to obtain a spectral feature of the IMU data; Based on the spectral characteristics of the IMU data, the peak amplitude and frequency are calculated to obtain a shake intensity feature vector, and the shake intensity feature vector is input into a shake intensity assessment model. The shake intensity assessment model calculates a shake intensity assessment score using a support vector regression algorithm. According to the jitter intensity evaluation score, a compensation intensity coefficient is calculated using a preset adaptive weight function, and the compensation intensity coefficient is multiplied by each element of the initial compensation model parameter matrix to dynamically adjust the compensation model parameter matrix.
6. The IMU-based camera motion compensation method according to claim 5, characterized in that: The geometric transformation of the original image sequence based on the compensation model parameter matrix is performed, and dense optical flow field optimization and time-weighted background pixel pool modeling are combined to obtain an initial stable image sequence and a foreground mask, including: Applying the compensation model parameter matrix to the original image sequence to perform affine transformation to obtain a first image sequence, and performing dense optical flow field calculation between adjacent frames in the first image sequence to obtain a motion vector field; Based on the motion vector field, local motion compensation and pixel resampling are performed on the first image sequence to obtain a second image sequence, and a time-weighted background pixel pool is constructed for the second image sequence. A fixed-size pixel value queue is maintained for each pixel position, new pixel values are queued with a time-decaying weight, and the oldest pixel value is dequeued when the queue is full, thereby obtaining a dynamically updated background model; performing foreground detection on each frame of the second image sequence based on the dynamically updated background model to obtain a foreground mask; The foreground mask and the second image sequence are subjected to a pixel-by-pixel AND operation to extract the foreground area to obtain a foreground image sequence, and inter-frame difference is performed on the foreground image sequence and the second image sequence to calculate the absolute difference and perform adaptive threshold segmentation to obtain an initial stable image sequence.
7. The IMU-based camera motion compensation method according to claim 6, characterized in that: The performing content-aware filling and temporal filtering on the initial stabilized image sequence and the foreground mask to output target stabilized image sequences in multiple compensation modes includes: Detecting and extending the edge region of each frame image in the initial stabilized image sequence to obtain an extended image sequence, and filling blank regions according to the extended image sequence and the foreground mask to obtain a filled image sequence; Performing a time-domain Kalman filter on the padded image sequence, where the state vector is a pixel value and the observation vector is a pixel value at a corresponding position of an adjacent frame, to obtain a time-domain filtered image sequence; Performing multi-scale image processing on the image sequence after the time domain filtering to obtain an initial multi-scale image sequence, and performing bilateral filtering on each scale of the initial multi-scale image sequence to obtain a filtered multi-scale image sequence; Reconstructing the filtered multi-scale image sequence to the original resolution to obtain a reconstructed image sequence; Applying three different compensation modes to the reconstructed image sequence: a global compensation mode, a local compensation mode, and an adaptive compensation mode, wherein the global compensation mode uses a global affine transformation, the local compensation mode uses a grid deformation algorithm, and the adaptive compensation mode dynamically selects a compensation strategy based on image content, thereby obtaining an image sequence in the global compensation mode, an image sequence in the local compensation mode, and an image sequence in the adaptive compensation mode; The image sequence of the global compensation mode, the image sequence of the local compensation mode, and the image sequence of the adaptive compensation mode are subjected to jitter evaluation, the root mean square error of the inter-frame differences is calculated, and the compensation mode with the smallest error is selected as the target stable image sequence output.
8. A camera motion compensation device based on IMU, characterized in that: The device is configured to perform the IMU-based camera motion compensation method according to any one of claims 1 to 7, the device comprising: The acquisition module is used to perform low-pass filtering and adaptive Kalman filtering on the acceleration and angular velocity raw data collected by multiple IMU sensors to obtain six-degree-of-freedom IMU fusion data; An extraction module is used to input the six-degree-of-freedom IMU fusion data into a linear convolution aliasing model for processing to obtain simulated motion data and extract the corresponding high-dimensional motion feature vector; A fusion module is used to fuse the high-dimensional motion feature vector and the original image sequence using a six-degree-of-freedom motion model based on a sine wave combination and a BP neural network prediction model to obtain a camera motion parameter prediction value; An evaluation module is used to build a look-ahead network compensation model based on the camera motion parameters and dynamically adjust the compensation model parameter matrix according to the jitter intensity evaluation results of the IMU data variance and peak analysis; A transformation module, configured to perform a geometric transformation on the original image sequence based on the compensation model parameter matrix, and obtain an initial stable image sequence and a foreground mask by combining dense optical flow field optimization and time-weighted background pixel pool modeling; An output module is configured to perform content-aware filling and temporal filtering on the initial stabilized image sequence and the foreground mask, and output a target stabilized image sequence in a plurality of compensation modes.
9. A computer device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and is characterized in that when the processor executes the computer program, the IMU-based camera motion compensation method described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor is caused to perform the IMU-based camera motion compensation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Video processing method and device, electronic equipment and readable storage medium
CN113286194A
Multi-target visual tracking method based on camera motion trend estimation
CN117495900A