Video stabilization method that deeply integrates optical flow and IMU data
Through a video image stabilization method that deeply integrates optical flow and IMU data, combined with RAFT network and IMU sensors, the problem of insufficient video stability in complex motion scenarios is solved, and high stability and efficient video processing is achieved.
Patent Information
- Application Number
- PCT/CN2023/136324
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-12
AI Technical Summary
Traditional video image stabilization technology is not ideal when dealing with complex motion and vibration scenarios, making it difficult to provide high-stability video output.
Using a video image stabilization method that deeply integrates optical flow and IMU data, efficiently estimates the video optical flow through the RAFT network, and combines the camera posture and motion information provided by the IMU sensor to achieve a more comprehensive and accurate video image stabilization.
This method can adapt to complex motion scenarios, provide high-stability video output, and improves the efficiency of video processing and analysis.
Smart Images

Figure CN2023136324_12062025_PF_FP_ABST
Abstract
Description
A video stabilization method based on deep fusion of optical flow and IMU data Technical Field
[0001] The present invention relates to the field of video processing technology, and in particular to a video stabilization method that deeply fuses optical flow and IMU data. Background Art
[0002] Video stabilization is a technique that reduces jitter and blur by estimating and compensating for motion in video sequences. Real-time video stabilization is a key research area in image processing, primarily used to obtain stable video images on mobile platforms. Its applications range from drones, action cameras, and mobile phones.
[0003] Video stabilization involves stabilizing raw video sequences using relevant equipment or algorithms to improve viewing comfort. Traditional video stabilization techniques are primarily based on computer vision or traditional signal processing techniques. The algorithm consists of three steps: camera motion estimation, camera motion smoothing, and mapping from a shaky perspective to a stable perspective.
[0004] Traditional video stabilization methods are not ideal when dealing with complex motion and vibration scenes. In recent years, deep learning technology has made significant progress in the field of computer vision. Among them, the RAFT (Recurrent All-Pairs Field Transforms) network, as a deep learning architecture, is widely used for video optical flow estimation. The core idea of the RAFT network is to iteratively update the optical flow field by combining the architecture of convolutional neural networks and recurrent neural networks to estimate the motion in the scene frame by frame. The network's method of fusing multi-scale correlation volumes of information at different scales can better handle motion at different scales and effectively capture the spatiotemporal information in image sequences, thereby improving the performance of optical flow estimation.
[0005] At the same time, inertial measurement unit (IMU) sensors are widely used in modern mobile devices and can provide camera pose and motion information. This data is crucial for understanding the camera's motion in three-dimensional space and more accurately estimating the relative motion between image frames.
[0006] Summary of the Invention
[0007] To overcome the shortcomings of existing technologies and provide a video stabilization method that can adapt to complex motion scenes and provide highly stable video output, the present invention provides a video stabilization method that deeply integrates optical flow and IMU data. By combining the efficient estimation capability of the RAFT network for video optical flow and the camera posture and motion information provided by the IMU sensor, a more comprehensive and accurate video stabilization is achieved. The main steps are as follows:
[0008] Step 1: Acquire and preprocess data. Obtain the data required for video stabilization, including video frames, IMU acceleration, and IMU angular velocity data. The camera's rotation angle is obtained by processing the angular velocity, and the camera's translation is obtained by processing the acceleration.
[0009] This step includes:
[0010] Step 1.1, obtain IMU acceleration and angular velocity data, acceleration and angular velocity are expressed as: a=(a x ,a y ,a z ,t) (1) ω=(ω x ,ω y ,ω z ,t) (2)
[0011] Among them, a x ,a y ,a z is the acceleration value in the three axes, ω x ,ω y ,ω z are the angular velocity values along the three axes, and t is the sensor timestamp corresponding to the acceleration and angular velocity.
[0012] Step 1.2: For the acceleration signal a, obtain its frequency domain value through Fourier transform:
[0013] One point earns:
[0014] The second integral is:
[0015] A(k) is the frequency domain conversion of acceleration a; V(r) is the frequency domain conversion of velocity v; S(r) is the frequency domain conversion of displacement s; j is the imaginary unit.
[0016] In step 1.3, the spectrum is filtered based on a preset filter. The useless part of the integral spectrum is the trend item, which is removed to obtain the effective value set in each spectrum.
[0017] Where H(k) is the filter, Δf is the frequency resolution; f m is the lower cutoff frequency; f n is the upper cutoff frequency.
[0018] Step 1.4: After completing the spectrum screening, the frequency domain integration results (spectra) corresponding to acceleration a, velocity v, and displacement s are subjected to inverse Fourier transform to obtain the time domain integration results, i.e., displacement.
[0019] In step 1.5, the camera translation T is obtained by calculating the difference in the camera displacement between two adjacent frames.
[0020] Step 1.6, angular velocity data is expressed as (ω x ,ω y ,ω z ,t), ω is the angular velocity, and the camera rotation angle R(t) is obtained by integrating R(t) = Sω(t)*R(tS) (7), where S refers to the sampling interval.
[0021] In step 1.7, the camera rotation angle R(t) is converted into a quaternion representation and saved in a queue.
[0022] Step 2: Calculate the optical flow using deep learning methods. Use the RAFT network to calculate the optical flow of two adjacent video frames.
[0023] This step includes:
[0024] In step 2.1, two adjacent video frames are input into the feature extraction network, which is a convolutional neural network that uses a CNN (Convolutional Neural Networks) architecture with 6 residual layers to extract pixel-by-pixel features.
[0025] Step 2.2: Initialize the optical flow field. A small neural network is used as the optical flow initialization module to initialize an optical flow field for each pixel. This initial optical flow field will serve as the starting point for subsequent steps.
[0026] In step 2.3, the initial optical flow field is iteratively updated using a recurrent neural network structure. The RAFT model gradually improves the accuracy of the estimated optical flow field by iteratively updating the optical flow between all pixel pairs.
[0027] In step 2.4, a pyramid structure is used to process feature pyramids at different scales, so that the RAFT model can capture motion information at different scales.
[0028] In step 2.5, after multiple iterations, the optical flow field generated by the model is used as the final output, which includes the displacement vector of each pixel representing the movement between adjacent frames, and the optical flow between two adjacent frames is obtained.
[0029] Step 3: Calculate the original pose and stable pose history data. The camera pose is represented and combined with the previous and next frame camera poses to form the original pose history data and stable pose history data.
[0030] This step includes:
[0031] Step 3.1, define the two-dimensional offset S(t) of the camera principal point at time t as: S(t) = αΔT + βΔI (8)
[0032] Where T is the translation calculated from the acceleration data in step 1, I is the optical flow extracted from step 2, and α and β are related weight coefficients.
[0033] In step 3.2, the camera posture is described by rotation and translation to describe the position and orientation of the camera in three-dimensional space. The original posture of the camera is expressed as: P o =(R o ,S o ) (9)
[0034] where R o represents the rotation of the camera, S o is the 2D offset of the camera's principal point.
[0035] Step 3.3: For the current frame, the original camera pose at the current frame, the original camera pose at the previous M frames, and the original camera pose at the next M frames constitute the original pose history data, which is expressed as: H o =(P o (t-MΔt),…,P o (t),…,P o (t+MΔt)) (10)
[0036] Wherein, Δt represents the time interval between two frames.
[0037] Step 3.4, the stable attitude is a quaternion used for image stabilization transformation, that is, the rotation R used for image stabilization transformation s , expressed as: P s =R s (11)
[0038] Step 3.5, the stable posture of the camera at the previous M frames is used to form the stable posture history data, which is expressed as: H s =(P s(t-MΔt),…,P o (t-Δt)) (12)
[0039] Step 4: Predict the stabilized camera pose. Based on the original pose, the historical stabilized pose data, and the optical flow data, the stabilized quaternion of the current frame can be predicted through unsupervised training.
[0040] This step includes:
[0041] Step 4.1: Use the ResNet network to further extract the optical flow features. The video frame optical flow information extracted in step 2 is input into the ResNet network for convolution processing to obtain a high-level spatiotemporal feature representation.
[0042] In step 4.2, the extracted optical flow feature information is input into a network composed of a series of convolutional networks, and the optical flow is encoded into the latent space using an encoder with 2D convolution to obtain the optical flow mapping to a low-dimensional representation z.
[0043] Step 4.3: The original posture history H calculated in step 3 is o and the stable attitude history H s and z calculated in step 4.1, the combined representation is the joint motion representation [z,H o ,H s ].
[0044] In step 4.4, the joint motion representation is input into a deep learning network consisting of an LSTM (Long Short Term Memory) unit and a fully connected layer.
[0045] In step 4.5, the LSTM unit in the deep learning network is used to fuse the hidden layer representation and maintain the time information to predict the stable camera pose. The fully connected layer is used to decode the hidden layer representation into the stable camera pose.
[0046] In step 4.6, the stabilized camera pose output by the network is placed into the historical stabilized pose queue.
[0047] Step 5: Use the stabilized pose to stabilize the current video frame to form a stabilized video. By dividing the video frame into grids, transforming each grid vertex, and performing secondary interpolation on the pixels in the grid, the camera's true pose can be transformed into the stabilized pose, thereby achieving the transformation of the stabilized video frame.
[0048] This step includes:
[0049] Step 5.1: For camera imaging, the projection point x of a point X in the real 3D space in the 2D image is: x = KRX (13)
[0050] Among them, R is the rotation angle of the camera, K is the intrinsic parameter matrix of the camera
[0051] f is the focal length of the camera, and (u,v) is the principal point of the camera.
[0052] Step 5.2, divide the video frame image into 64x64 rectangular boxes (called Grid), and process the vertices of each Grid. For each vertex, the original pose P given by the camera o =(R o ,S o ) and stable posture P s =R s , from the original point x o To the stable point x s The image stabilization transformation is x s =K s R s R o -1 K o -1 x o (15)
[0053] In step 5.3, the pixel values within the rectangular frame are obtained by performing quadratic interpolation on the four vertices of the grid. After processing all the pixel values of the image, a stabilized video frame is obtained.
[0054] Step 6: Train the deep learning model.
[0055] This step includes:
[0056] In step 6.1, the smoothness loss L1 is defined to evaluate the stability and smoothness of the camera's stable posture.
[0057] In step 6.2, the edge loss L2 is defined to evaluate the extent to which the stabilized video image exceeds the true edge, so that the camera stabilization posture remains stable.
[0058] In step 6.3, the deformation loss L3 is defined to evaluate the degree of deformation of the stabilized video image.
[0059] In step 6.4, a translation loss L4 is defined to minimize the translation motion between adjacent frames.
[0060] In step 6.5, the final loss is calculated by weighting the four losses defined in steps 6.1, 6.2, 6.3, and 6.4: L = w1L1 + w2L2 + w3L3 + w4L4 (16)
[0061] Among them, w1, w2, w3, and w4 are the weight systems of smoothness loss, edge loss, deformation loss, and translation loss, respectively.
[0062] In step 6.6, the network is trained using multiple loss calculation methods (loss functions). In the first training phase, only L1 and L3 are trained to minimize them. In the second training phase, L2 is added for training. In the final phase, L4 is added for training to minimize the final loss. This achieves the goal of training a deep learning model.
[0063] The beneficial effects of the present invention are: a video stabilization method that deeply integrates optical flow and IMU data is proposed, which can adapt to complex motion scenes and provide highly stable video output. Compared with traditional video stabilization algorithms based on image content, the video stabilization method of the present invention integrates video content and sensor motion data, and uses RAFT technology to extract optical flow. The optical flow data can effectively reduce the shaking and jitter of the video sequence, but if there is obvious global motion in the video, the optical flow will fail. Therefore, the present invention also uses IMU data to estimate the camera posture, and adopts an unsupervised learning method of multi-stage training to predict the stable camera posture, and achieve stable processing of the video sequence, thereby improving the efficiency of video processing and analysis. This method can adapt to complex motion scenes such as parallax changes, foreground occlusion, and low-quality video, and produce highly stable video. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] FIG1 shows a flow chart of a video stabilization method for deep fusion of optical flow and IMU data according to the present invention;
[0065] FIG2 shows the specific structure of the joint motion representation input deep neural network of an example of the present invention; DETAILED DESCRIPTION
[0066] The preferred embodiments of the present invention are further described below with reference to the accompanying drawings and examples.
[0067] The flowchart shown in FIG1 shows the specific process of the entire implementation of the present invention:
[0068] Step 1: Acquire and preprocess data. Obtain the data required for video stabilization, including video frames, IMU acceleration, and IMU angular velocity data. The camera's rotation angle is obtained by processing the angular velocity, and the camera's translation is obtained by processing the acceleration.
[0069] This step includes:
[0070] Step 1.1, obtain IMU acceleration and angular velocity data, acceleration and angular velocity are expressed as: a=(a x ,a y ,a z ,t) (1) ω=(ω x ,ω y ,ω z ,t) (2)
[0071] Among them, a x ,a y ,a z is the acceleration value in the three axes, ω x ,ω y ,ω z are the angular velocity values along the three axes, and t is the sensor timestamp corresponding to the acceleration and angular velocity.
[0072] Step 1.2: For the acceleration signal a, obtain its frequency domain value through Fourier transform:
[0073] One point earns:
[0074] The second integral is:
[0075] A(k) is the frequency domain conversion of acceleration a; V(r) is the frequency domain conversion of velocity v; S(r) is the frequency domain conversion of displacement s; j is the imaginary unit.
[0076] In step 1.3, the spectrum is filtered based on a preset filter. The useless part of the integral spectrum is the trend item, which is removed to obtain the effective value set in each spectrum.
[0077] Where H(k) is the filter, Δf is the frequency resolution; f m is the lower cutoff frequency; f n is the upper cutoff frequency.
[0078] Step 1.4: After completing the spectrum screening, the frequency domain integration results (spectra) corresponding to acceleration a, velocity v, and displacement s are subjected to inverse Fourier transform to obtain the time domain integration results, i.e., displacement.
[0079] In step 1.5, the camera translation T is obtained by calculating the difference in the camera displacement between two adjacent frames.
[0080] Step 1.6, angular velocity data is expressed as (ω x ,ω y ,ω z,t), ω is the angular velocity, and the camera rotation angle R(t) is given by R(t)=Sω(t)*R(tS) (7)
[0081] Integrate to get . Where S refers to the sampling interval.
[0082] In step 1.7, the camera rotation angle R(t) is expressed as a quaternion and saved in a queue.
[0083] Step 2: Calculate the optical flow using deep learning methods. Use the RAFT network to calculate the optical flow of two adjacent video frames.
[0084] This step includes:
[0085] In step 2.1, two adjacent video frames are input into the feature extraction network, which is a convolutional neural network that uses a CNN architecture with 6 residual layers to extract pixel-by-pixel features.
[0086] Step 2.2: Initialize the optical flow field. A small neural network is used as the optical flow initialization module to initialize an optical flow field for each pixel. This initial optical flow field will serve as the starting point for subsequent steps.
[0087] In step 2.3, the initial optical flow field is iteratively updated using a recurrent neural network structure. The RAFT model gradually improves the accuracy of the estimated optical flow field by iteratively updating the optical flow between all pixel pairs.
[0088] In step 2.4, by calculating visual similarity, we perform the inner product calculation on all pixel pairs and construct a multi-scale correlation volume for each pixel pair. Using a pyramid structure, we process feature pyramids at different scales, enabling the RAFT model to capture motion information at different scales. By searching the correlation volume, the iterative unit iteratively updates the optical flow field.
[0089] In step 2.5, after multiple iterations, the optical flow field generated by the model is used as the final output, which includes the displacement vector of each pixel representing the movement between adjacent frames, and the optical flow between two adjacent frames is obtained.
[0090] Step 3: Calculate the original pose and stable pose history data. The camera pose is represented and combined with the previous and next frame camera poses to form the original pose history data and stable pose history data.
[0091] This step includes:
[0092] Step 3.1, define the two-dimensional offset S(t) of the camera principal point at time t as: S(t) = αΔT + βΔI (8)
[0093] Where T is the translation calculated from the acceleration data in step 1, I is the optical flow extracted from step 2, and α and β are related weight coefficients.
[0094] In step 3.2, the camera posture is described by rotation and translation to describe the position and orientation of the camera in three-dimensional space. The original posture of the camera is expressed as: P o =(R o ,S o ) (9)
[0095] where R o represents the rotation of the camera, S o is the 2D offset of the camera's principal point.
[0096] Step 3.3: For the current frame, the original camera pose at the current frame, the original camera pose at the previous M frames, and the original camera pose at the next M frames constitute the original pose history data, which is expressed as: H o =(P o (t-MΔt),…,P o (t),…,P o (t+MΔt)) (10)
[0097] Wherein, Δt represents the time interval between two frames.
[0098] Step 3.4, the stable attitude is a quaternion used for image stabilization transformation, that is, the rotation R used for image stabilization transformation s , expressed as: P s =R s (11)
[0099] Step 3.5, the stable posture of the camera at the previous M frames is used to form the stable posture history data, which is expressed as: H s =(P s (t-MΔt),…,P o (t-Δt)) (12)
[0100] Step 4: Predict the stabilized camera pose. Based on the original pose, the historical stabilized pose data, and the optical flow data, the stabilized quaternion for the current frame can be predicted through unsupervised training.
[0101] This step includes:
[0102] Step 4.1: Use the ResNet network to further extract the optical flow features. The video frame optical flow information extracted in step 2 is input into the ResNet network for convolution processing to obtain a high-level spatiotemporal feature representation.
[0103] In step 4.2, the extracted optical flow feature information is input into a network composed of a series of convolutional networks, and the optical flow is encoded into the latent space using an encoder with 2D convolution to obtain the optical flow mapping to a low-dimensional representation z.
[0104] Step 4.3: The original posture history H calculated in step 3 is o and the stable attitude history H s and z calculated in step 4.1, the combined representation is the joint motion representation [z,H o ,H s ].
[0105] In step 4.4, the joint motion representation is input into a deep learning network consisting of an LSTM unit and a fully connected layer.
[0106] In step 4.5, the LSTM unit in the deep learning network is used to fuse the hidden layer representation and maintain the time information to predict the stable camera pose. The fully connected layer is used to decode the hidden layer representation into the stable camera pose.
[0107] In step 4.6, the stabilized camera pose output by the network is placed into the historical stabilized pose queue.
[0108] This example uses stable posture data from ten frames before and after to form historical data for image stabilization, and the joint motion representation is used as the input of the deep learning network. The specific structure is shown in Figure 2.
[0109] Step 5: Use the stabilized pose to stabilize the current video frame to form a stabilized video. By dividing the video frame into grids, transforming each grid vertex, and performing secondary interpolation on the pixels in the grid, the camera's true pose can be transformed into the stabilized pose, thereby achieving the transformation of the stabilized video frame.
[0110] Step 5.1: For camera imaging, the projection point x of a point X in the real 3D space in the 2D image is: x = KRX (13)
[0111] Among them, R is the rotation angle of the camera, K is the intrinsic parameter matrix of the camera
[0112] f is the focal length of the camera, and (u,v) is the principal point of the camera.
[0113] Step 5.2, divide the video frame image into 64x64 rectangular boxes (called Grid), and process the vertices of each Grid. For each vertex, the original pose P given by the camera o =(R o ,S o ) and stable posture Ps =R s , from the original point x o To the stable point x s The image stabilization transformation is x s =K s R s R o -1 K o -1 x o (15)
[0114] In step 5.3, the pixel values within the rectangular frame are obtained by performing quadratic interpolation on the four vertices of the grid. After processing all the pixel values of the image, a stabilized video frame is obtained.
[0115] Step 6: Train the deep learning model.
[0116] Step 6.1, define the smoothness loss L1 to evaluate the stability and smoothness of the camera's stable posture: L1 = ‖R s (t)-R s (t-Δt)‖ 2 (16)
[0117] In step 6.2, the edge loss L2 is defined to evaluate the extent to which the stabilized video image exceeds the true edge, so that the camera stabilization posture remains stable:
[0118] Among them, w p,i is the Gaussian normal distribution weight coefficient centered on the current frame, α is the tolerable edge protrusion value, and the prot function is used to evaluate the coefficient of the video frame after stabilization transformation that exceeds the edge of the original frame after cropping.
[0119] Step 6.3, define the deformation loss L3 to evaluate the degree of deformation of the stabilized video image:
[0120] Among them, Ω(R s ,R o ) is the angle between the original camera pose and the stabilized pose, β0 is a threshold, and β1 is a parameter that controls the slope of the function.
[0121] Step 6.4, define the translation loss L4 to minimize the translation motion between adjacent frames:
[0122] Among them, x o,n and y o,n+1 It is the corresponding pixel point of the translation in two adjacent frames of the original camera space. The pixel point of the video frame at time t (x t ,y tThe two-dimensional offset S(t) of the image is obtained by fusion of the translation calculated by optical flow and acceleration data, and is expressed as S(t) = αΔT + βΔI (20)
[0123] T is the translation calculated from the acceleration data in step 1, I is the optical flow extracted from step 2, α, β are the relevant weight coefficients. Tr is the transformation from the original camera space to the stabilized camera space, and x s,n =Tr n (x o,n ),y s,n+1 =Tr n+1 (y o,n+1 ), is the forward translation, For backward translation, the pixel point after stabilization can be expressed as:
[0124] In step 6.5, the final loss is calculated by weighting the four losses defined in steps 6.1, 6.2, 6.3, and 6.4: L = w1L1 + w2L2 + w3L3 + w4L4 (23)
[0125] Among them, w1, w2, w3, and w4 are the weight systems of smoothness loss, edge loss, deformation loss, and translation loss, respectively.
[0126] In step 6.6, the network is trained using multiple loss calculation methods (loss functions). In the first training phase, only L1 and L3 are trained to minimize them. In the second training phase, L2 is added for training. In the final phase, L4 is added for training to minimize the final loss. This achieves the goal of training a deep learning model.
Claims
1. A video stabilization method that deeply fuses optical flow and IMU data, characterized in that, it includes the following steps: Step 1, obtain data and perform preprocessing; obtain the data required for video stabilization, including video frames, IMU acceleration, and IMU angular velocity data, and obtain the rotation angle of the camera by processing the angular velocity and the translation of the camera by processing the acceleration; Step 2, calculate the optical flow using the deep learning method; use the RAFT network to calculate the optical flow between two adjacent video frames; Step 3, calculate the original pose and the stable pose history data; represent the camera pose and form the original pose history data and the stable pose history data with the camera poses of the previous and next frames; Step 4, predict the stable camera pose; according to the original pose, the stable pose history data, and the optical flow data, the stabilized quaternion of the current frame can be predicted after unsupervised training; Step 5, perform a stabilization transformation on the current video frame with the stable pose to form a stabilized video; by dividing the video frame into Grids, perform transformation processing on the vertices of each Grid, and perform bilinear interpolation on the pixels in the Grid to achieve the transformation of the true pose of the camera to the stable camera pose, thereby realizing the transformation of the stabilized video frame; Step 6, train the deep learning model.
2. The video stabilization method that deeply fuses optical flow and IMU data according to claim 1, characterized in that, in step 1 of obtaining data and performing preprocessing, the sampling frequency domain integration algorithm is used to process the IMU acceleration and calculate the displacement of the camera; step 1 further includes: Step 1.1, obtain the IMU acceleration and angular velocity data, and the acceleration and angular velocity are expressed as: a = (a x , a y , a z , t) (1) ω = (ω x , ω y , ω z , t) (2) where a x , a y , a z are acceleration values in three axial directions, ω x , ω y , ω z are angular velocity values in three axial directions, and t is the sensor timestamp corresponding to the acceleration and angular velocity; Step 1.2, for the acceleration signal a, its frequency domain value is obtained through Fourier transform: The first integration gives: The double integral gives: A(k) is the frequency domain conversion of the acceleration a; V(r) is the frequency domain conversion of the velocity v; S(r) is the frequency domain conversion of the displacement s; j is the imaginary unit; Step 1.3, screen the spectrum based on a preset filter. The useless part in the integrated spectrogram is the trend term, which is removed to obtain the valid value set in each spectrum; Among them, H(k) is a filter, and Δf is the frequency resolution; f m is the lower cut-off frequency; f n is the upper cut-off frequency; Step 1.4, after completing the screening of the spectrum, perform the inverse Fourier transform on the frequency domain integration results (spectrums) corresponding to the acceleration a, the velocity v, and the displacement s to obtain the integration result in the time domain, that is, the displacement; Step 1.5, obtain the translation T of the camera by calculating the difference in displacement between two adjacent frames of the camera; Step 1.6, the angular velocity data is expressed as (ω x , ω y , ω z , t), where ω is the angular velocity, and the rotation angle R(t) of the camera is determined by R(t) = Sω(t) * R(t - S) (7) is obtained by integration; where S refers to the sampling interval time; Step 1.7, represent the rotation angle R(t) of the camera with quaternions and save it in a queue.
3. The video stabilization method that deeply fuses optical flow and IMU data according to claim 1, characterized in that, in step 2 of calculating the optical flow using the deep learning method, the RAFT network is used to calculate the optical flow between two adjacent video frames; step 2 further includes: Step 2.1, input two adjacent video frames into the feature extraction network, which is a convolutional neural network, and use the CNN architecture with 6 residual layers to extract pixel-by-pixel features; Step 2.2, initialize the optical flow field; use a small neural network as the optical flow initialization module to initialize an optical flow field for each pixel, and this initial optical flow field will be used as the starting point for subsequent steps; Step 2.3, use a recurrent neural network structure to iteratively update the initial optical flow field; the RAFT model gradually improves the accuracy of the estimated optical flow field by iteratively updating the optical flow between all pixel pairs; Step 2.4, adopt a pyramid structure, and by performing feature pyramid processing at different scales, enable the RAFT model to capture motion information at different scales; Step 2.5, after multiple iterations, the optical flow field generated by the model is used as the final output, including the displacement vector of each pixel representing the motion between adjacent frames, and the optical flow between two adjacent frames is obtained.
4. A video stabilization method for deeply fusing optical flow and IMU data according to claim 1, characterized in that, In step 3, calculate the original pose and the historical data of the stable pose, and the camera pose is represented as P t =(R t , S t ) for representation; in step 3, the camera pose is defined as: a.R t represents the rotation of the camera. The rotation angle R(t) of the camera at time t is given by R(t) = Sω(t) * R(t - S) (8) is obtained by integration; where S refers to the sampling interval time and ω is the angular velocity; b.S t is the two-dimensional offset of the camera principal point, which is obtained by fusing the optical flow extracted in step 2 and the translation calculated from the acceleration data in step 1; S(t) = αΔT + βΔI (9).
5. A video stabilization method for deeply fusing optical flow and IMU data according to claim 1, characterized in that, In step 4 for predicting the stable camera pose, the input of the deep neural network is defined as a joint motion representation composed of the camera pose historical data and the optical flow data of the video frames; in step 4, the joint motion representation [z, H o , H s is defined as: a. Input the optical flow extracted in step 2 into a network composed of a series of convolutional networks, and use an encoder with 2D convolution to encode the optical flow into the latent space to obtain the low-dimensional representation z of the optical flow mapping; b. For the current frame, here the original pose history data is composed of the original pose of the camera at the current frame time, the original poses of the camera at the previous M frame times, and the original poses of the camera at the next M frame times, which is expressed as: H o = (P o (t - MΔt), …, P o (t), …, P o (t + MΔt)) (10) where Δt represents the time interval between two frames; c. Use the stable pose of the camera at the previous M frame times to form the stable pose history data, which is expressed as: H s = (P s (t - MΔt), …, P o (t - Δt)) (11) d. The original pose history H o and the stable pose history H s along with the optical flow extracted from the input frame mapped to a low-dimensional representation z, combined as a joint motion representation [z, H o , H s .
6. A video stabilization method for deeply fusing optical flow and IMU data according to claim 1, characterized in that, the method of stabilizing the current video frame with the stable pose in step 5; by dividing the video frame into grids, performing transformation processing on the vertices of each grid, and performing bilinear interpolation on the pixels in the grid, the transformation from the true pose of the camera to the stable camera pose can be realized, thereby realizing the transformation of the stabilized video frame; Using this method is more efficient than performing stabilization transformation on each pixel value; The step 5 further includes: Step 5.1, for camera imaging, for a point X in the real 3D space, its projection point x in the 2D image is: x = KRX (13) Wherein, R is the rotation angle of the camera, and K is the internal parameter matrix of the camera f is the camera focal length, and (u, v) is the camera principal point; Step 5.2, divide the video frame image into 64x64 rectangular frames (referred to as Grid), and process the vertices of each Grid; for each vertex, the original pose P given by the camera o =(R o , S o ) and the stable pose P s =R s , the video stabilization transformation from the original point x o to the stable point x s is x s = K s R s R o -1 K o -1 x o (15) Step 5.3, the pixel values within the rectangle are obtained by performing bilinear interpolation on the four vertices of the grid. After processing all the pixel values of the image, the stabilized video frame is obtained.
7. A video stabilization method for deeply fusing optical flow and IMU data according to claim 1, characterized in that, In step 6 as described above, the deep learning model is trained; a variety of loss calculation methods (loss functions) and a multi-stage method are used to train the network, including smooth loss, edge loss, deformation loss, translation loss, etc.; Step 6 as described above further includes: Step 6.1, define the smooth loss L 1 It is used to evaluate the stability and smoothness of the stable pose of the camera; Step 6.2, define the edge loss L 2 which is used to evaluate the degree to which the stabilized video image exceeds the real edge, so as to keep the stable attitude of the camera stable; Step 6.3, define the distortion loss L 3 which is used to evaluate the distortion degree of the stabilized video image; Step 6.4, define the translation loss L 4 Minimize the translational motion of adjacent frames; Step 6.5, the final loss is obtained by weighted calculation of the 4 losses defined in steps 6.1, 6.2, 6.3, and 6.4: L = w 1 L 1 + w 2 L 2 + w 3 L 3 + w 4 L 4 (12) where, w 1 , w 2 , w 3 , w 4 are the weight systems of the smooth loss, edge loss, deformation loss, and translation loss, respectively; Step 6.6, use multiple loss calculation methods (loss functions) to train the network. In the first training stage, only minimize L 1 and L 3 for training; in the second training stage, add L 2 for training; in the final stage, add L 4 for training to minimize the final loss; implement the training of the deep learning model.
Citation Information
Patent Citations
Satellite-borne push-scan optical load high-frequency error compensation method based on angular displacement sensor
CN109668579A
Camera pose estimation method and device based on depth visual odometer and IMU
CN112648994A
Video data processing method and device, equipment and storage medium
CN113556582A
Hybrid anti-shake method and system based on deep learning
CN115174817A
Data processing method and device, electronic equipment and computer readable storage medium
CN117812467A
Cited By
Video image stabilization method based on online path smoothing network
CN120583314A
Camera motion estimation method and system under foreground mask based on optical flow guidance
CN120602799A
Multi-camera cooperative anti-shake control system and method
CN121000969A
Video anomaly detection method based on Mask grid stability
CN121259736A