Lightweight 3D video image stabilization method and device based on CNN-Transform hybrid framework, and electronic equipment
By image reconstruction network based on CNN-Transformer hybrid framework, depth map and pose prediction results are generated, and smoothed, the problem of poor stability in traditional video stabilization methods when processing complex scenes is solved, and effective reduction of video jitter and stability improvement is achieved.
Patent Information
- Application Number
- CN202510314354.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-20
AI Technical Summary
Traditional video stabilization methods have limitations in handling scenes with significant parallax and large depth changes caused by translational motion, and hardware limitations make it difficult to capture stable video from exactly the same viewing angle as jitter video, resulting in poor video stability.
The lightweight 3D video image stabilization method based on the CNN-Transformer hybrid framework is adopted. By inputting the to-process jitter video frame sequence to be processed, the network is constructed by inputting the pre-trained image sequence to be processed, the depth map prediction results and pose prediction results are generated, and the smoothing process is performed, and the video frame sequence is finally reconstructed to generate a stable video.
It effectively reduces jitter in the video, improves the stability and viewing experience of the video, and does not require manual annotation of image data, and has strong generalization ability.
Smart Images

Figure CN120182497A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to a lightweight 3D video stabilization method, apparatus, and electronic device based on a CNN-Transformer hybrid framework. Background Art
[0002] Video stabilization technology is crucial for improving the quality of visual data in various Internet of Things (IoT) applications, such as smartphones, robotics, and drones. By enhancing the stability and clarity of videos, these technologies support more accurate post-processing and analysis, which are essential for effective IoT deployment and advanced data-driven insights. Video stabilization methods typically begin with estimating the global motion trajectory in a shaky video, which includes both the expected motion and unwanted jitter. Subsequently, filters are applied to smooth this trajectory, ultimately generating a stabilized video that retains the expected motion while minimizing distortion and cropping.
[0003] Traditional video stabilization methods have limitations in scenarios involving significant parallax and large depth variations caused by translational motion, and hardware limitations make it almost impossible to capture a stabilized video from exactly the same perspective as the shaky video, resulting in poor stability of the final video. Summary of the Invention
[0004] In view of this, embodiments of this application provide a lightweight 3D video stabilization method, apparatus, and electronic device based on a CNN-Transformer hybrid framework to overcome or at least partially solve the above problems.
[0005] A first aspect of the embodiments of this application provides a lightweight 3D video stabilization method based on a CNN-Transformer hybrid framework, the method including: Inputting a sequence of frames of a to-be-processed shaky video into a pre-trained image reconstruction network to obtain a depth map prediction result and a pose prediction result of the sequence of frames of the to-be-processed shaky video, where the training samples of the pre-trained image reconstruction network are sample image pairs in a sequence of sample shaky video frames, and the sample image pairs include sample source images and corresponding sample target images; the pre-trained image reconstruction network is obtained through unsupervised training with the goal of reconstructing the sample target images; Determining the predicted poses of each shaky video frame in the sequence of frames of the to-be-processed shaky video by using the pose prediction result of the sequence of frames of the to-be-processed shaky video; Performing smoothing processing on the predicted poses of each shaky video frame in the sequence of frames of the to-be-processed shaky video to obtain smoothed predicted poses; Based on the depth map prediction result of the to-be-processed jitter video frame sequence and the smoothed predicted pose, each jitter video frame in the to-be-processed jitter video frame sequence is reconstructed to obtain a stabilized video frame sequence.
[0006] Optionally, the image reconstruction network includes: a depth estimation network and a pose estimation network; the training process of the image reconstruction network is as follows: According to the sample jitter video frame sequence, a plurality of sample image pairs are generated; For each sample image pair in the plurality of sample image pairs, the sample source image in the sample image pair is input into the depth estimation network to obtain the sample depth map prediction result of the sample source image; The sample source image and the sample target image in the sample image pair are input into the pose estimation network to obtain a sample pose prediction result, and the sample pose prediction result is the predicted relative pose between the sample source image and the sample target image; According to the sample depth map prediction result, the sample pose prediction result, and the sample source image, the sample target image is reconstructed to obtain a sample reconstructed image; Based on the loss function value between the sample reconstructed image and the sample target image, the parameters of the image reconstruction network are updated.
[0007] Optionally, the loss function value of the image reconstruction network includes: a graph reconstruction loss and an edge smoothing loss of the depth map, and the loss function value of the image reconstruction network is determined according to the following steps: Based on the sample source image and the sample target image, a first threshold is determined; Based on the magnitude relationship between the graph reconstruction loss and the first threshold, a first weight value corresponding to the graph reconstruction loss is determined; According to the first weight value corresponding to the graph reconstruction loss, the graph reconstruction loss, the edge smoothing loss of the depth map, and the second weight value corresponding to the edge smoothing loss of the depth map, the loss function value of the image reconstruction network is determined.
[0008] Optionally, the determining of the first threshold based on the sample source image and the sample target image includes: Taking the previous frame of the sample target image as the sample source image to form the first sample image pair corresponding to the sample target image, and taking the next frame of the sample target image as the sample source image to form the second sample image pair corresponding to the sample target image; Based on the first pair of sample images, perform photometric reconstruction calculation on the previous frame of the sample target image and the sample target image to obtain the first photometric reconstruction loss between the previous frame of the sample target image and the sample target image; Based on the second pair of sample images, determine the second photometric reconstruction loss between the next frame of the sample target image and the sample target image; Determine the minimum value of the first photometric reconstruction loss and the second photometric reconstruction loss as the first threshold; Wherein, when the graphic reconstruction loss is less than the first threshold, the first weight value is 1, and when the graphic reconstruction loss is not less than the first threshold, the first weight value is 0.
[0009] Optionally, the graphic reconstruction loss is determined according to the following steps: Based on the sample reconstruction image and the sample target image, determine the photometric reconstruction loss between the sample reconstruction image and the sample target image; Determine the graphic reconstruction loss according to the photometric reconstruction loss between the sample reconstruction image and the sample target image.
[0010] Optionally, the determining the graphic reconstruction loss according to the photometric reconstruction loss between the sample reconstruction image and the sample target image includes: Use the previous frame of the sample target image as the sample source image, as the first pair of sample images corresponding to the sample target image, and use the next frame of the sample target image as the sample source image, as the second pair of sample images corresponding to the sample target image; The sample reconstruction image includes: the first sample reconstruction image based on the first pair of sample images, and the second sample reconstruction image based on the second pair of sample images; Based on the first sample reconstruction image and the sample target image, determine the photometric reconstruction loss between the first sample reconstruction image and the sample target image as the third photometric reconstruction loss; Based on the second sample reconstruction image and the sample target image, determine the photometric reconstruction loss between the second sample reconstruction image and the sample target image as the fourth photometric reconstruction loss; Determine the minimum value of the third photometric reconstruction loss and the fourth photometric reconstruction loss as the graphic reconstruction loss.
[0011] Optionally, the reconstructing each jitter video frame in the to-be-processed jitter video frame sequence according to the depth map prediction result and the smoothed predicted pose of the to-be-processed jitter video frame sequence to obtain a stabilized video frame sequence includes: Taking each jittery video frame in the sequence of jittery video frames to be processed as the current jittery video frame to be processed, reconstruct the next frame of the current jittery video frame according to the depth map prediction result of the current jittery video frame to be processed, the smoothed predicted pose of the current jittery video frame to be processed, the smoothed predicted pose of the next frame of the current jittery video frame to be processed, and the current jittery video frame to be processed, to obtain the stabilized next frame of the current jittery video frame; The stabilized next frames of the jittery video frames in the sequence of jittery video frames to be processed form a sequence of stabilized video frames.
[0012] Optionally, the process of performing photometric reconstruction calculation includes: For a first image and a second image, determine the structural similarity index SSIM between the first image and the second image; For the first image and the second image, determine the L1 loss between the first image and the second image; According to a preset weight, perform weighted summation on the structural similarity index SSIM and the L1 loss to obtain the photometric reconstruction loss between the first image and the second image.
[0013] A second aspect of the embodiments of the present application provides a lightweight 3D video stabilization device based on a CNN-Transformer hybrid framework. The device includes: An input module, configured to input a sequence of jittery video frames to be processed into a pre-trained image reconstruction network, to obtain a depth map prediction result and a pose prediction result of the sequence of jittery video frames to be processed. The training samples of the pre-trained image reconstruction network are sample image pairs in a sample sequence of jittery video frames, and the sample image pairs include sample source images and corresponding sample target images; the pre-trained image reconstruction network is obtained through unsupervised training with the goal of reconstructing the sample target images; A determination module, configured to use the pose prediction result of the sequence of jittery video frames to be processed to determine the predicted poses of the jittery video frames in the sequence of jittery video frames to be processed; A smoothing processing module, configured to perform smoothing processing on the predicted poses of the jittery video frames in the sequence of jittery video frames to be processed to obtain smoothed predicted poses; A reconstruction module, configured to reconstruct each jittery video frame in the sequence of jittery video frames to be processed according to the depth map prediction result of the sequence of jittery video frames to be processed and the smoothed predicted poses, to obtain a sequence of stabilized video frames.
[0014] In a third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the method described in the first aspect.
[0015] Advantages of the present application: The present application provides a lightweight 3D video stabilization method, device, and electronic device based on a CNN-Transformer hybrid framework, relating to the technical field of image processing. The method includes: inputting a sequence of jittery video frames to be processed into a pre-trained image reconstruction network to obtain a depth map prediction result and a pose prediction result. The training samples of the pre-trained image reconstruction network are sample image pairs in a sample sequence of jittery video frames, and the sample image pair includes a sample source image and a corresponding sample target image; the pre-trained image reconstruction network is obtained through unsupervised training with the goal of reconstructing the sample target image; using the pose prediction result, determining the predicted poses of each jittery video frame in the sequence of jittery video frames to be processed; performing smoothing processing on the predicted poses of each jittery video frame to obtain the smoothed predicted poses; and reconstructing each jittery video frame according to the depth map prediction result and the smoothed predicted poses to obtain a stabilized sequence of video frames.
[0016] A lightweight 3D video stabilization method based on a CNN-Transformer hybrid framework proposed by the present application can effectively reduce jitter in the video and improve the stability and viewing experience of the video. Specifically, through the image reconstruction network, a corresponding depth map prediction result and pose prediction result are generated for the sequence of jittery video frames to be processed. Combining with the smoothed predicted poses, each jittery video frame is reconstructed, making the generated sequence of video frames more stable, thereby reducing the jitter of the original video and improving the viewing experience of the video. At the same time, the image reconstruction network uses an unsupervised training method, without using manually labeled image data, and is trained with the goal of reconstructing stable images, which can accurately estimate the depth map prediction result and the pose prediction result, enabling the image reconstruction network to adapt to different video scenarios and having strong generalization ability. Description of the Drawings
[0017] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application.
[0018] To more clearly illustrate the technical solutions of the present application, the drawings required for the description of the present application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a flowchart of a lightweight 3D video stabilization method based on a CNN-Transformer hybrid framework provided by an embodiment of the present application; Figure 2 It is a schematic diagram of the training process and inference process of an image reconstruction network provided by an embodiment of the present application; Figure 3 It is a schematic framework diagram of a lightweight 3D video stabilization device based on a CNN-Transformer hybrid framework provided by an embodiment of the present application. Detailed implementation manners
[0020] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0022] With the progress of novel view synthesis for monocular depth estimation, it has become possible to reconstruct stable videos through self-supervised learning by leveraging the estimated depth maps and smooth camera poses. Self-supervised monocular depth estimation treats this problem as a novel view synthesis problem, where a neural network makes predictions based on the appearance of the target image.
[0023] On the source image. Inspired by this, Deep3D is the first self-supervised 3D video stabilization method that incorporates monocular depth estimation into the stabilization process.
[0024] Good stability has been achieved in multiple categories. In addition, a three-dimensional multi-frame perspective is also adopted to generate stable images while retaining the structure. Although these methods effectively handle parallax and achieve good results, they are still subject to residual jitter due to geometric estimation errors. In addition, the introduction of training during testing and the large model architecture limit the application of these methods in the Internet of Things.
[0025] Reference Figure 1 , Figure 1 It is a flowchart of a lightweight 3D video stabilization method based on a CNN-Transformer hybrid framework provided by an embodiment of the present application. As Figure 1 shown, the lightweight 3D video stabilization method based on the CNN-Transformer hybrid framework provided in this embodiment at least includes the following steps: Step S11: Input the sequence of jittery video frames to be processed into a pre-trained image reconstruction network to obtain the depth map prediction result and pose prediction result of the sequence of jittery video frames to be processed. The training samples of the pre-trained image reconstruction network are sample image pairs in the sample sequence of jittery video frames. Each sample image pair includes a sample source image and a corresponding sample target image. The pre-trained image reconstruction network is obtained through unsupervised training with the goal of reconstructing the sample target image.
[0026] In this application, we propose a lightweight 3D video stabilization framework that combines a hybrid convolutional neural network and a Transformer architecture, namely the image reconstruction network mentioned in this article. Refer to Figure 2 , Figure 2 which is a schematic diagram of the training process and inference process of the image reconstruction network provided in an embodiment of this application.
[0027] As Figure 2 shown, in this embodiment, the sequence of jittery video frames to be processed contains multiple consecutive jittery video frames. For example, a video captured by a camera. The image reconstruction network is a pre-trained network for reconstructing the sequence of multiple consecutive jittery video frames to be processed. During the inference process, first input the sequence of jittery video frames to be processed into the pre-trained image reconstruction network. By extracting local features and global features from multiple consecutive jittery video frames, obtain the depth map prediction result of the sequence of jittery video frames to be processed; and estimate the camera pose between multiple consecutive jittery video frames to obtain the pose prediction result of the sequence of jittery video frames to be processed.
[0028] Exemplarily, for the sequence of jittery video frames to be processed , the corresponding depth map prediction result is , and the corresponding pose prediction result is , where N represents the number of jittery video frames, represents the jittery video frame at the t-th frame.
[0029] Among them, the training samples of the image reconstruction network are sample image pairs in the sample sequence of jittery video frames. Each sample image pair includes a sample source image and a corresponding sample target image. The image reconstruction network is optimized through unsupervised training with the goal of reconstructing the sample target image. During the training process, the parameters of the image reconstruction network are optimized by minimizing the loss between the synthesized frame (i.e., the reconstructed sample target image) and the target frame (i.e., the sample target image), obtaining the pre-trained image reconstruction network.
[0030] Step S12: Determine the predicted poses of each jitter video frame in the to-be-processed jitter video frame sequence by using the pose prediction result of the to-be-processed jitter video frame sequence.
[0031] In this embodiment, use the pose prediction result obtained in step S11 to determine the pose prediction result of each jitter video frame included in the to-be-processed jitter video frame sequence. The pose prediction result represents the camera pose of each jitter video frame, including the rotation and translation information of the camera. By applying the pose prediction result to each jitter video frame, the predicted pose of each jitter video frame in the three-dimensional space can be obtained.
[0032] Specifically, the steps to obtain the predicted pose of each jitter video frame are as follows: First, set the predicted pose of the first frame of each jitter video frame in the to-be-processed jitter video frame sequence as the unit pose, that is , starting from the second frame, the predicted pose corresponding to each jitter video frame is calculated according to the formula . Finally, a set of predicted poses corresponding to each jitter video frame in the to-be-processed jitter video frame sequence is obtained , that is , which contains the predicted pose of each jitter video frame in the to-be-processed jitter video frame sequence in the three-dimensional space.
[0033] Step S13: Smooth the predicted poses of each jitter video frame in the to-be-processed jitter video frame sequence to obtain the smoothed predicted poses.
[0034] In this embodiment, since the to-be-processed jitter video frame sequence is the original video obtained by shooting, the predicted trajectory of each jitter video frame in the to-be-processed jitter video frame sequence in the three-dimensional space is unstable and contains unnecessary motion components. Therefore, it is necessary to smooth the predicted poses of each jitter video frame in the to-be-processed jitter video frame sequence to obtain the smoothed predicted poses, so as to stabilize the predicted poses.
[0035] Specifically, in an optional embodiment, a Gaussian filter can be used to smooth the predicted pose of each jitter video frame in the three-dimensional space included in , that is . Thus, a set of predicted poses corresponding to each jitter video frame in the to-be-processed jitter video frame sequence after smoothing is obtained , , which contains the smoothed predicted pose of each jitter video frame. Among them, is the implicit function of the Gaussian filter.
[0036] Step S14: Reconstruct each jittery video frame in the to-be-processed jittery video frame sequence according to the depth map prediction result of the to-be-processed jittery video frame sequence and the smoothed predicted pose, to obtain a stabilized video frame sequence.
[0037] In this embodiment, according to the depth map prediction result obtained in step S11 and the smoothed predicted pose obtained in step S13, each jittery video frame in the to-be-processed jittery video frame sequence is reconstructed to generate a stabilized video frame sequence.
[0038] Specifically, during the reconstruction process, using the depth map prediction result and the smoothed predicted pose, each pixel in each jittery video frame in the to-be-processed jittery video frame sequence is reconstructed to obtain a new pixel. After each pixel is reconstructed, a stabilized jittery video frame corresponding to each jittery video frame is obtained. In this way, each jittery video frame in the to-be-processed jittery video frame sequence is reconstructed into a stabilized jittery video frame, thereby generating the entire stabilized video frame sequence, reducing the jitter in the original video (i.e., the to-be-processed jittery video frame sequence), and ensuring the stability and video quality of the video.
[0039] Combining the above embodiments, in an optional implementation manner, the present application further provides a lightweight 3D video stabilization method based on a CNN-Transformer hybrid framework. In this method, the image reconstruction network includes: a depth estimation network and a pose estimation network; and, refer Figure 2 , the training process of the image reconstruction network includes steps S21 to S25: Step S21: Generate a plurality of sample image pairs according to the sample jittery video frame sequence.
[0040] In this embodiment, the image reconstruction network includes a depth estimation network DepthNet and a pose estimation network PoseNet. The depth estimation network is responsible for generating the depth map prediction result of the jittery video frame, and the pose estimation network is responsible for generating the pose prediction result between the jittery video frames.
[0041] Specifically, during the process of training the image reconstruction network, first, a sample video needs to be selected. According to the sample video, a sample jittery video frame sequence is obtained, and a plurality of sample image pairs are extracted from the sample jittery video frame sequence. Each sample image pair includes a sample source image and a sample target image corresponding to the sample source image. In the sample image pair, the sample source image and the sample target image are consecutive video frames. For example, the sample source image can be the t-th frame, and the sample target image can be the (t + 1)-th frame or the (t - 1)-th frame.
[0042] Step S22: For each sample image pair among the multiple sample image pairs, input the sample source image in the sample image pair into the depth estimation network to obtain the sample depth map prediction result of the sample source image.
[0043] In this embodiment, for each sample image pair, the sample source image is input into the depth estimation network. The depth estimation network estimates the depth map prediction result of the sample source image, that is, the distance from each pixel point of the sample source image to the camera. The depth estimation network internally combines the CNN and Transformer architectures. The encoder of the depth estimation network uses dilated convolutions to extract local features of the sample source image, effectively expanding the receptive field without increasing additional computational costs. At the same time, in order to capture the global features of the sample source image, a Transformer architecture integrating the cross-covariance attention mechanism is also used to extract global features.
[0044] Step S23: Input the sample source image and the sample target image in the sample image pair into the pose estimation network to obtain the sample pose prediction result, where the sample pose prediction result is the predicted relative pose between the sample source image and the sample target image.
[0045] In this embodiment, next, the sample source image and the sample target image in the sample image pair are simultaneously input into the pose estimation network. According to the sample source image and the sample target image, the relative pose between frames is estimated , and the sample pose prediction result between the sample source image and the sample target image is obtained, that is, the motion trajectory of the camera from the sample source image to the sample target image.
[0046] Exemplarily, in this process, the pose estimation network extracts the features of the sample source image and the sample target image through a pre-trained ResNet18 network and outputs a 6D vector through 4 convolutional layers. This 6D vector represents the rotation and translation process of the camera from the sample source image to the sample target image. Subsequently, this 6D vector is converted into a 4×4 homogeneous transformation matrix, and this homogeneous transformation matrix is used to represent the sample pose prediction result between the sample source image and the sample target image.
[0047] Step S24: Reconstruct the sample target image according to the sample depth map prediction result, the sample pose prediction result, and the sample source image to obtain the sample reconstructed image.
[0048] In this embodiment, further using the sample depth map prediction result obtained in step S22 and the sample pose prediction result obtained in step S23, combined with the sample source image, the sample target image is reconstructed to obtain a sample reconstructed image. Specifically, in this process, according to the depth map prediction result and the pose prediction result, each pixel in the sample source image coordinate system is transformed to the sample target image coordinate system, so that the pixels in the sample source image are mapped to the perspective of the sample target image. After all the pixels of the sample source image are mapped, a sample reconstructed image is finally generated.
[0049] In this process, for the i-th pixel of the sample reconstructed image the homogeneous pixel coordinate can be calculated as follows: , where is the homogeneous pixel coordinate of the i-th pixel of the sample source image , K is the camera intrinsic matrix used for the conversion between the image coordinate system and the camera coordinate system, refers to the corresponding depth value in the sample depth map prediction result for the i-th pixel of the sample source image, refers to the value representing the relative pose between frames in the sample pose prediction result from the sample source image to the sample target image.
[0050] Step S25, based on the loss function value between the sample reconstructed image and the sample target image, update the parameters of the image reconstruction network.
[0051] In this embodiment, finally, calculate the loss function value between the sample reconstructed image and the sample target image. By minimizing the loss function value, use backpropagation to update the parameters of the depth estimation network and the pose estimation network in the image reconstruction network, and finally optimize the performance of the image reconstruction network, so that the network can accurately reconstruct images.
[0052] Through the above steps, the lightweight 3D video stabilization method based on the CNN-Transformer hybrid framework provided by this application can effectively train the image reconstruction network, enabling it to accurately estimate the depth map and pose and generate high-quality reconstructed images. This training method combines optimizing network parameters through unsupervised learning, improving the generalization ability and stability of the network. Finally, the trained network can be used to process actual jittery video frame sequences, generate stabilized video outputs, reduce jitter, and improve the quality of the video.
[0053] Combined with the above embodiments, in an alternative embodiment, the present application further provides a lightweight 3D video stabilization method based on a CNN-Transformer hybrid framework. In this method, the loss function value of the image synthesis network and the image reconstruction network includes: a graph reconstruction loss and an edge smoothing loss of the depth map. The loss function value of the image synthesis network and the image reconstruction network is determined according to the following steps S31 to S33: Step S31: Based on the sample source image and the sample target image, determine a first threshold.
[0054] In this embodiment, during the process of training the depth estimation network and the pose estimation network of the image reconstruction network, it is necessary to determine a threshold for adjusting the weight of the graph reconstruction loss, that is, the first threshold. The first threshold is set according to the difference between the sample source image and the sample target image.
[0055] Specifically, in the process of setting the first threshold, it can be set based on experience, or it can be dynamically determined by analyzing the photometric reconstruction loss between sample image pairs, which is the sum of the Structural Similarity Index (SSIM) and the L1 norm loss (L1 loss). The Structural Similarity Index (SSIM) is an index for measuring the similarity between two images, which takes into account the brightness, contrast, and structural information of the images. The value range of SSIM is between -1 and 1, and the closer the value is to 1, the more similar the two images are. The L1 loss is another commonly used index for measuring the difference between the predicted value and the true value. It calculates the average of the absolute differences between the predicted value and the true value.
[0056] In an alternative embodiment, the process of performing photometric reconstruction calculation includes: When calculating the photometric reconstruction loss between any two images, such as the first image and the second image, first, for the first image and the second image, determine the Structural Similarity Index (SSIM) between the first image and the second image, and then for the first image and the second image, determine the L1 loss between the first image and the second image; according to a preset weight, perform a weighted sum of the Structural Similarity Index (SSIM) and the L1 loss to obtain the photometric reconstruction loss between the first image and the second image.
[0057] For example, calculate according to the formula where, is the first image, is the second image, and α is the preset weight (the first weight value in the following text) .
[0058] It should be noted that the first image and the second image refer to any image. For example, the first image can be a sample target image, and the second image is a sample reconstructed image obtained after reconstruction, or the first image is a sample source image, and the second image is a sample target image. This part will be described in detail later.
[0059] Step S32: Determine a first weight value corresponding to the graph reconstruction loss based on the magnitude relationship between the graph reconstruction loss and the first threshold.
[0060] In this embodiment, after determining the first threshold, it is then necessary to determine the first weight value corresponding to the graph reconstruction loss according to the magnitude relationship between the graph reconstruction loss and the first threshold. The graph reconstruction loss is calculated by comparing the photometric reconstruction loss values between the target frames reconstructed from different original frames and the original target frame.
[0061] Specifically, this self-supervised training is usually carried out under the assumption of a moving camera in a static scene. When these assumptions fail, for example, when the camera is stationary or there is object motion in the scene, the performance will be greatly affected, manifested as holes with infinite depth in the predicted depth map. Introducing μ (the first threshold) is to filter out the pixels that do not change in appearance from one frame to the next in the sequence. The effect of this is to make the network ignore the objects moving at the same speed as the camera, and even ignore the entire frame in a monocular video when the camera stops moving.
[0062] We observe that the pixels that remain consistent between adjacent frames in the sequence usually indicate that the camera is stationary, or the object is translating at the same speed as the camera, or a low-texture area. Therefore, we set μ (the first threshold) to only include the pixels where the graph reconstruction loss is lower than the loss between the original unprocessed source image and the target frame.
[0063] Step S33: Determine the loss function value of the image synthesis network image reconstruction network according to the first weight value corresponding to the graph reconstruction loss, the graph reconstruction loss, the second weight value, the edge smoothing loss of the depth map, and the second weight value corresponding to the edge smoothing loss of the depth map.
[0064] In this embodiment, finally, by comprehensively considering the graph reconstruction loss and the edge smoothing loss of the depth map , and their corresponding weight values, to determine the loss function value of the image reconstruction network .
[0065] Specifically, the calculation formula for the loss function value of the image reconstruction network is: , where is the first weight value, is the second weight value, which is a preset value and can be set to 0.001, and is used to balance the loss of graph reconstruction and the edge smoothing loss of the depth map.
[0066] The edge smoothing loss of the depth map is used to optimize the parameters of the depth estimation network. Specifically, the calculation formula of the edge smoothing loss of the depth map is: 。
[0067] Among them, in the above calculation formula, refers to the inverse depth map after normalization, specifically expressed as , refers to the prediction result of the sample depth map, refers to the normalization factor. and respectively refer to the gradients of the depth map in the sample depth map prediction result in the horizontal direction (x-axis) and vertical direction (y-axis). refers to the sample target image, and respectively refer to the absolute values of the gradients of the depth map in the sample depth map prediction result in the horizontal direction and vertical direction, which are used to measure the change rate of the depth map in different directions. and are respectively the edge perception weights of the depth map in the sample depth map prediction result in the horizontal direction and vertical direction. Since the edge regions of the sample source image usually contain more noise and discontinuities, they can be used to suppress the gradient changes of the depth map in the edge regions of the sample source image in the depth map prediction result, and reduce the influence of these edge regions on the edge smoothing loss of the depth map.
[0068] Combining the above embodiments, in an optional implementation manner, the present application further provides a lightweight 3D video stabilization method based on a CNN-Transformer hybrid framework. In this method, the step of "determining the first threshold based on the sample source image and the sample target image" in step S31 specifically includes steps S31-1 to S31-4: Step S31-1, using the previous frame of the sample target image as the sample source image to form the first sample image pair corresponding to the sample target image, and using the next frame of the sample target image as the sample source image to form the second sample image pair corresponding to the sample target image.
[0069] In this embodiment, first, in one of the methods, the first threshold is determined based on the photometric reconstruction loss between the sample target image and the adjacent frame (sample source image). The stability of the reconstructed sample target image (i.e., the sample reconstruction image) is evaluated through the first threshold. Selecting different frames as the sample source image and calculating the first threshold will result in different photometric reconstruction losses, making the value of the first threshold different. Therefore, in order to more accurately reflect the reconstruction effect of the sample reconstruction image, during this process, for the same sample target image, two sample image pairs need to be constructed. In the first sample image pair, the sample source image is the previous frame of the sample target image, and in the second sample image pair, the sample source image is the next frame of the sample target image, so as to calculate the first threshold using the two sample image pairs.
[0070] Step S31-2: Based on the first sample image pair, perform photometric reconstruction calculation on the previous frame of the sample target image and the sample target image to obtain the first photometric reconstruction loss between the previous frame of the sample target image and the sample target image.
[0071] Step S31-3: Based on the second sample image pair, determine the second photometric reconstruction loss between the next frame of the sample target image and the sample target image.
[0072] Among them, the calculation process of the photometric reconstruction loss in Step S31-2 and Step S31-3 refers to the calculation formula of the photometric reconstruction loss between the first image and the second image shown above to calculate the photometric reconstruction loss between the sample source image and the sample target image. That is, at this time, the first image is the sample source image, the second image is the sample target image, the first sample image pair calculates the photometric reconstruction loss between the sample target image and the sample source image (the previous frame of the sample target image), and the second sample image pair calculates the photometric reconstruction loss between the sample target image and the sample source image (the next frame of the sample target image).
[0073] Step S31-4: Determine the minimum value of the first photometric reconstruction loss and the second photometric reconstruction loss as the first threshold.
[0074] In this embodiment, by minimizing the reconstruction loss, problems such as pixels exceeding the boundary and occlusion caused by object movement are avoided, thereby improving the accuracy and stability of image reconstruction. This method can effectively reduce the influence caused by object movement, occlusion, or image edges, and improve the quality of target image reconstruction.
[0075] Among them, when the graphic reconstruction loss is less than the first threshold, the first weight value is 1, and when the graphic reconstruction loss is not less than the first threshold, the first weight value is 0.
[0076] Through the above embodiments, by obtaining the respective photometric reconstruction losses corresponding to the sample target image and its adjacent frames, and selecting the smaller photometric reconstruction loss from the first photometric reconstruction loss and the second photometric reconstruction loss as the first threshold to evaluate the graphic reconstruction loss, the stability and consistency of the sample reconstructed image can be evaluated more accurately.
[0077] Combined with the above embodiments, in an alternative embodiment, the present application further provides a lightweight 3D video stabilization method based on a CNN-Transformer hybrid framework. In this method, the graphic reconstruction loss is determined according to the following steps: Step S41: Based on the sample reconstructed image and the sample target image, determine the photometric reconstruction loss between the sample reconstructed image and the sample target image.
[0078] In this embodiment, in order to evaluate the similarity between the sample reconstructed image and the sample target image, we choose the photometric reconstruction loss to measure the similarity between the two images. The photometric reconstruction loss needs to be calculated by combining the structural similarity index (SSIM) and the L1 loss. The specific steps are as follows: First, based on the sample reconstructed image and the sample target image, according to the formula , determine the photometric reconstruction loss between the sample reconstructed image and the sample target image ; where is the sample target image, is the sample reconstructed image, is the structural similarity index SSIM between the sample reconstructed image and the sample target image, is the L1 loss between the sample reconstructed image and the sample target image, that is = ; α is a weight parameter, usually taking a value of 0.85, which is used to balance the SSIM and the L1 loss.
[0079] Step S42: Determine the graphic reconstruction loss according to the photometric reconstruction loss between the sample reconstructed image and the sample target image.
[0080] In this embodiment, after determining the photometric reconstruction loss between the sample reconstructed image and the sample target image, according to the photometric reconstruction loss between the sample reconstructed image and the sample target image, determine the corresponding graphic reconstruction loss.
[0081] Combined with the above embodiments, in an alternative embodiment, the present application further provides a lightweight 3D video stabilization method based on a CNN-Transformer hybrid framework. In this method, "determining the graphic reconstruction loss according to the photometric reconstruction loss between the sample reconstructed image and the sample target image" in step S42 specifically includes steps S42-1 to S42-4: Step S42-1: Use the previous frame of the sample target image as the sample source image to form the first sample image pair corresponding to the sample target image, and use the next frame of the sample target image as the sample source image to form the second sample image pair corresponding to the sample target image; the sample reconstructed image includes: the first sample reconstructed image based on the first sample image pair and the second sample reconstructed image based on the second sample image pair.
[0082] In this embodiment, to solve the viewpoint pixel problem and occlusion problem caused by moving objects near the image boundary, we use two frames adjacent to the sample target image as the sample source images respectively, that is, Is ∈ {I t−1 , I t+1}. Selecting different frames as the sample source images will result in different photometric reconstruction losses. Specifically, we obtain two sets of sample image pairs, namely the first sample image pair and the second sample image pair. The first sample image pair is: the sample source image (the previous frame of the sample target image) and the sample target image; the second sample image pair is: the sample source image (the next frame of the sample target image) and the sample target image; since different frames are selected as the sample source images, two sample reconstructed images can be obtained, namely the first sample reconstructed image corresponding to the first sample image pair and the second sample reconstructed image corresponding to the second sample image pair.
[0083] Step S42-2: Based on the first sample reconstructed image and the sample target image, determine the photometric reconstruction loss between the first sample reconstructed image and the sample target image as the third photometric reconstruction loss.
[0084] In this embodiment, referring to the formula calculate the photometric reconstruction loss (i.e., the third photometric reconstruction loss) between the first sample reconstructed image and the sample target image, where represents the third photometric reconstruction loss, represents the first sample reconstructed image, represents the structural similarity index between the first sample reconstructed image and the sample target image, represents the L1 loss between the first sample reconstructed image and the sample target image.
[0085] Step S42-3: Based on the second sample reconstructed image and the sample target image, determine that the photometric reconstruction loss between the second sample reconstructed image and the sample target image is the fourth photometric reconstruction loss.
[0086] In this embodiment, referring to the formula calculate the photometric reconstruction loss (i.e., the fourth photometric reconstruction loss) between the second sample reconstructed image and the sample target image, where represents the fourth photometric reconstruction loss, represents the second sample reconstructed image, represents the structural similarity index between the second sample reconstructed image and the sample target image, represents the L1 loss between the second sample reconstructed image and the sample target image.
[0087] Step S42-4: Determine the minimum value between the third photometric reconstruction loss and the fourth photometric reconstruction loss as the graphic reconstruction loss.
[0088] In this embodiment, to avoid the problem that the pixels of the moving objects at the image edge exceed the frame and occlusion, we select the minimum loss as the graphic reconstruction loss. Specifically, referring to the formula determine the minimum value between the third photometric reconstruction loss and the fourth photometric reconstruction loss as the graphic reconstruction loss. When the third photometric reconstruction loss is less than the fourth photometric reconstruction loss, is equal to ; conversely, is equal to .
[0089] It should be noted that to avoid the influence of moving objects, we introduce a binary mask to remove the loss caused by static pixels. As shown in the following formula, when the minimum value is the third photometric reconstruction loss, it means that the sample source image is the previous frame of the sample target image. Therefore, when determining the first threshold, the first photometric reconstruction loss is compared with the graphic reconstruction loss to determine the first weight value. Conversely, when the minimum value is the fourth photometric reconstruction loss, it means that the sample source image is the next frame of the sample target image. Therefore, when constructing the first threshold, the second photometric reconstruction loss is compared with the graphic reconstruction loss to determine the first weight value.
[0090] .
[0091] Through the above embodiments, different photometric reconstruction losses generated between the sample target image and its adjacent frames are combined, and the smaller loss value is selected as the graphic reconstruction loss, so as to dynamically evaluate the accuracy of pose prediction, optimize the training process of the network, and improve the optimization ability of the pose estimation network for pose prediction.
[0092] Combined with the above embodiments, in an alternative embodiment, the present application further provides a lightweight 3D video stabilization method based on a CNN-Transformer hybrid framework. In this method, the step of "reconstructing each jitter video frame in the to-be-processed jitter video frame sequence according to the depth map prediction result of the to-be-processed jitter video frame sequence and the smoothed predicted pose to obtain a stabilized video frame sequence" in step S14 includes steps S14-1 to S14-2: Step S14-1: Taking each jitter video frame in the to-be-processed jitter video frame sequence as the current to-be-processed jitter video frame, and reconstructing the next frame of the current to-be-processed jitter video frame according to the depth map prediction result of the current to-be-processed jitter video frame, the smoothed predicted pose of the current to-be-processed jitter video frame, the smoothed predicted pose of the next frame of the current to-be-processed jitter video frame, and the current to-be-processed jitter video frame, to obtain the stabilized next frame of the current to-be-processed jitter video frame.
[0093] In this embodiment, during the process of reconstructing the to-be-processed jitter video frames, the first jitter video frame is used as the current to-be-processed jitter video frame, the second jitter video frame is used as the next frame of the current to-be-processed jitter video frame. The depth map prediction result of the first jitter video frame is obtained using a depth estimation network, and the smoothed predicted pose of the first jitter video frame and the smoothed predicted pose of the second jitter video frame, as well as the first jitter video frame, are used to reconstruct the second jitter video frame. After the second jitter video frame is reconstructed, the reconstructed second jitter video frame is used as the current to-be-processed jitter video frame, the third jitter video frame is used as the next frame of the current to-be-processed jitter video frame, and the third jitter video frame is reconstructed, and so on, to reconstruct each jitter video frame in the entire to-be-processed jitter video frame sequence.
[0094] For example, during the process of reconstructing each jitter video frame, taking the t-th jitter video frame as the current to-be-processed jitter video frame, the process of reconstructing the (t + 1)-th jitter video frame is taken as an example for illustration. Please refer to the formula 'Determine the homogeneous coordinates corresponding to each pixel of the (t + 1)-th jitter video frame to obtain the reconstructed (t + 1)-th jitter video frame. After reconstruction, the (t + 1)-th jitter video frame becomes a stabilized video frame, reducing the unstable situation caused by camera jitter.
[0095] Among them, represents the homogeneous coordinate of the i-th pixel of the t-th jitter video frame, Denote the homogeneous coordinates of the i-th pixel in the (t + 1)-th reconstructed jitter video frame at the (t + 1)-th time, Denote the depth map prediction result of the i-th pixel in the t-th jitter video frame in the sequence of jitter video frames to be processed, Denote the predicted pose of the t-th jitter video frame after smoothing processing, Denote the predicted pose of the (t + 1)-th jitter video frame after smoothing processing, where K is the camera intrinsic matrix.
[0096] Step S14-2, the next frames of each jitter video frame in the sequence of jitter video frames to be processed that are stabilized form a sequence of stabilized video frames.
[0097] In this embodiment, all the reconstructed and stabilized 3D videos are arranged in sequence to form a final sequence of stabilized jitter video frames. Each frame in this sequence has undergone depth map prediction, pose prediction, and smoothing processing, thereby reducing the jitter of the original video and making the overall video obtained finally more stable.
[0098] Through the technical solution of this application, in this article, we propose an image reconstruction network, which shows a video stabilization algorithm based on lightweight 3D deep learning foundation, and it integrates a hybrid CNN and Transformer architecture.
[0099] By using the estimated depth map prediction result and pose prediction result, the sample target image is reconstructed to obtain a sample reconstructed image, and the image reconstruction network is optimized by minimizing the loss between the sample target image and the sample reconstructed image. Moreover, in order to further improve the stability of the video, we combine the process of smoothing processing to estimate the camera pose trajectory, obtain the predicted pose after smoothing processing, and combine it with the depth map prediction result to generate a finally stable video, thereby being able to improve the quality and visual experience of the stable video.
[0100] As Figure 3 shown, based on the same inventive concept, an embodiment of this application further provides a lightweight 3D video stabilization device based on a CNN-Transformer hybrid framework. The device includes: An input module 11, configured to input a sequence of jitter video frames to be processed into a pre-trained image reconstruction network to obtain the depth map prediction result and pose prediction result of the sequence of jitter video frames to be processed. The training samples of the pre-trained image reconstruction network are sample image pairs in a sequence of sample jitter video frames, and the sample image pairs include sample source images and corresponding sample target images; the pre-trained image reconstruction network is obtained through unsupervised training with the goal of reconstructing the sample target images; A determination module 12, configured to determine the predicted poses of the respective jitter video frames in the to-be-processed jitter video frame sequence by using the pose prediction results of the to-be-processed jitter video frame sequence; A smoothing processing module 13, configured to perform smoothing processing on the predicted poses of the respective jitter video frames in the to-be-processed jitter video frame sequence to obtain the smoothed predicted poses; A reconstruction module 14, configured to reconstruct the respective jitter video frames in the to-be-processed jitter video frame sequence according to the depth map prediction results of the to-be-processed jitter video frame sequence and the smoothed predicted poses to obtain a stabilized video frame sequence.
[0101] Optionally, the image reconstruction network includes: a depth estimation network and a pose estimation network; the apparatus further includes: a training module, configured to execute the training process of the image reconstruction network, and the training module includes: A sample image pair generation unit, configured to generate a plurality of sample image pairs according to the sample jitter video frame sequence; A first input unit, configured to, for each sample image pair in the plurality of sample image pairs, input the sample source image in the sample image pair into the depth estimation network to obtain the sample depth map prediction result of the sample source image; A second input unit, configured to input the sample source image and the sample target image in the sample image pair into the pose estimation network to obtain a sample pose prediction result, where the sample pose prediction result is the predicted relative pose between the sample source image and the sample target image; A sample target image reconstruction unit, configured to reconstruct the sample target image according to the sample depth map prediction result, the sample pose prediction result, and the sample source image to obtain a sample reconstructed image; A network parameter update unit, configured to update the parameters of the image reconstruction network based on the loss function value between the sample reconstructed image and the sample target image.
[0102] Optionally, the loss function value of the image reconstruction network includes: a graph reconstruction loss and an edge smoothing loss of the depth map, and the apparatus further includes: a loss function value determination module, configured to determine the loss function value of the image reconstruction network, and the loss function value determination module includes: A first threshold determination unit, configured to determine a first threshold based on the sample source image and the sample target image; A first weight value determination unit, configured to determine a first weight value corresponding to the graph reconstruction loss based on the magnitude relationship between the graph reconstruction loss and the first threshold; A loss function value determination unit, configured to determine a loss function value of the image reconstruction network according to a first weight value corresponding to the graph reconstruction loss, the graph reconstruction loss, an edge smoothing loss of the depth map, and a second weight value corresponding to the edge smoothing loss of the depth map.
[0103] Optionally, the first threshold determination unit includes: A sample image pair first determination unit, configured to use the previous frame of the sample target image as the sample source image, as a first sample image pair corresponding to the sample target image, and use the next frame of the sample target image as the sample source image, as a second sample image pair corresponding to the sample target image; A first photometric reconstruction loss calculation unit, configured to perform a photometric reconstruction calculation on the previous frame of the sample target image and the sample target image based on the first sample image pair, to obtain a first photometric reconstruction loss between the previous frame of the sample target image and the sample target image; A second photometric reconstruction loss calculation unit, configured to determine a second photometric reconstruction loss between the next frame of the sample target image and the sample target image based on the second sample image pair; A first threshold determination subunit, configured to determine the minimum value of the first photometric reconstruction loss and the second photometric reconstruction loss as the first threshold; Wherein, when the graph reconstruction loss is less than the first threshold, the first weight value is 1, and when the graph reconstruction loss is not less than the first threshold, the first weight value is 0.
[0104] Optionally, the apparatus further includes: a graph reconstruction loss determination module, configured to determine the graph reconstruction loss, and the graph reconstruction loss determination module includes: A photometric reconstruction loss determination unit, configured to determine a photometric reconstruction loss between the sample reconstruction image and the sample target image based on the sample reconstruction image and the sample target image; A graph reconstruction loss determination unit, configured to determine the graph reconstruction loss according to the photometric reconstruction loss between the sample reconstruction image and the sample target image.
[0105] Optionally, the graph reconstruction loss determination unit includes: A sample image pair second determination unit, configured to use the previous frame of the sample target image as the sample source image, as a first sample image pair corresponding to the sample target image, and use the next frame of the sample target image as the sample source image, as a second sample image pair corresponding to the sample target image; The sample reconstructed image includes: a first sample reconstructed image based on the first sample image pair, and a second sample reconstructed image based on the second sample image pair; A third photometric reconstruction loss determination unit, configured to determine, based on the first sample reconstructed image and the sample target image, that the photometric reconstruction loss between the first sample reconstructed image and the sample target image is a third photometric reconstruction loss; A fourth photometric reconstruction loss determination unit, configured to determine, based on the second sample reconstructed image and the sample target image, that the photometric reconstruction loss between the second sample reconstructed image and the sample target image is a fourth photometric reconstruction loss; A graphic reconstruction loss determination subunit, configured to determine the minimum value between the third photometric reconstruction loss and the fourth photometric reconstruction loss as the graphic reconstruction loss.
[0106] Optionally, the reconstruction module 14 includes: A reconstruction unit, configured to use each jitter video frame in the to-be-processed jitter video frame sequence as the current to-be-processed jitter video frame, and reconstruct the next frame of the current to-be-processed jitter video frame according to the depth map prediction result of the current to-be-processed jitter video frame, the smoothed predicted pose of the current to-be-processed jitter video frame, the smoothed predicted pose of the next frame of the current to-be-processed jitter video frame, and the current to-be-processed jitter video frame, to obtain the stabilized next frame of the current to-be-processed jitter video frame; A combination unit, configured to form a stabilized video frame sequence from the stabilized next frames of each jitter video frame in the to-be-processed jitter video frame sequence.
[0107] Optionally, the device further includes a photometric reconstruction loss calculation module, and the calculation module is configured to execute the calculation process of the photometric reconstruction loss. The calculation module includes: A structural similarity index determination unit, configured to determine the structural similarity index SSIM between a first image and a second image; An L1 loss determination unit, configured to determine the L1 loss between the first image and the second image; A photometric reconstruction loss calculation unit, configured to perform a weighted sum of the structural similarity index SSIM and the L1 loss according to a preset weight, to obtain the photometric reconstruction loss between the first image and the second image.
[0108] Based on the same inventive concept, another embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory. Wherein, the processor executes the computer program to implement the lightweight 3D video stabilization method based on the CNN-Transformer hybrid framework as described in any of the above embodiments.
[0109] For the device, since it is basically similar to the method embodiment, the description is relatively simple. For the related parts, please refer to the partial description of the method embodiment.
[0110] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same and similar parts among the embodiments, please refer to each other.
[0111] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0112] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0113] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0114] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 steps for implementing the functions specified in one block or multiple blocks.
[0115] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0116] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.
[0117] The above has introduced in detail a lightweight 3D video stabilization method, device and electronic device based on a CNN-Transformer hybrid framework provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A lightweight 3D video stabilization method based on a CNN-Transformer hybrid framework, characterized in that: The method comprises: Inputting the jittered video frame sequence to be processed into a pre-trained image reconstruction network to obtain a depth map prediction result and a pose prediction result of the jittered video frame sequence to be processed, wherein the training samples of the pre-trained image reconstruction network are sample image pairs in the sample jittered video frame sequence, and the sample image pairs include a sample source image and a corresponding sample target image; the pre-trained image reconstruction network is obtained through unsupervised training with the goal of reconstructing the sample target image; Determine the predicted pose of each jittery video frame in the jittery video frame sequence to be processed by using the pose prediction result of the jittery video frame sequence to be processed; Smoothing the predicted position and posture of each jittery video frame in the jittery video frame sequence to be processed to obtain a smoothed predicted position and posture; According to the depth map prediction result of the jittery video frame sequence to be processed and the predicted position and posture after the smoothing process, each jittery video frame in the jittery video frame sequence to be processed is reconstructed to obtain a stabilized video frame sequence.
2. The lightweight 3D video stabilization method based on the CNN-Transformer hybrid framework according to claim 1 is characterized in that: The image reconstruction network includes: a depth estimation network and a pose estimation network; the training process of the image reconstruction network is: Generating a plurality of sample image pairs according to the sample jitter video frame sequence; For each sample image pair in the plurality of sample image pairs, inputting a sample source image in the sample image pair into the depth estimation network to obtain a sample depth map prediction result of the sample source image; Inputting a sample source image and a sample target image in the sample image pair into the pose estimation network to obtain a sample pose prediction result, wherein the sample pose prediction result is a predicted relative pose between the sample source image and the sample target image; Reconstructing the sample target image according to the sample depth map prediction result, the sample pose prediction result and the sample source image to obtain a sample reconstructed image; Based on the loss function value between the sample reconstructed image and the sample target image, the parameters of the image reconstruction network are updated.
3. The lightweight 3D video stabilization method based on the CNN-Transformer hybrid framework according to claim 2 is characterized in that: The loss function value of the image reconstruction network includes: a graphics reconstruction loss and an edge smoothing loss of a depth map. The loss function value of the image reconstruction network is determined according to the following steps: Determining a first threshold based on the sample source image and the sample target image; Determining a first weight value corresponding to the graphics reconstruction loss based on a magnitude relationship between the graphics reconstruction loss and the first threshold; The loss function value of the image reconstruction network is determined according to a first weight value corresponding to the graphics reconstruction loss, the graphics reconstruction loss, the edge smoothing loss of the depth map, and a second weight value corresponding to the edge smoothing loss of the depth map.
4. The lightweight 3D video stabilization method based on the CNN-Transformer hybrid framework according to claim 3 is characterized in that: The determining of a first threshold based on the sample source image and the sample target image comprises: Taking a previous frame of the sample target image as the sample source image, as a first sample image pair corresponding to the sample target image, and taking a subsequent frame of the sample target image as the sample source image, as a second sample image pair corresponding to the sample target image; Based on the first sample image pair, performing photometric reconstruction calculation on a previous frame of the sample target image and the sample target image to obtain a first photometric reconstruction loss between the previous frame of the sample target image and the sample target image; Based on the second sample image pair, determining a second photometric reconstruction loss between a subsequent frame of the sample target image and the sample target image; Determine a minimum value between the first photometric reconstruction loss and the second photometric reconstruction loss as the first threshold; When the graphics reconstruction loss is less than the first threshold, the first weight value is 1, and when the graphics reconstruction loss is not less than the first threshold, the first weight value is 0.
5. The lightweight 3D video stabilization method based on the CNN-Transformer hybrid framework according to claim 3, characterized in that: The image reconstruction loss is determined according to the following steps: Based on the sample reconstructed image and the sample target image, determining a photometric reconstruction loss between the sample reconstructed image and the sample target image; The graphic reconstruction loss is determined according to the photometric reconstruction loss between the sample reconstructed image and the sample target image.
6. The lightweight 3D video stabilization method based on the CNN-Transformer hybrid framework according to claim 5, characterized in that: The determining the graphic reconstruction loss according to the photometric reconstruction loss between the sample reconstructed image and the sample target image comprises: Taking a previous frame of the sample target image as the sample source image, as a first sample image pair corresponding to the sample target image, and taking a subsequent frame of the sample target image as the sample source image, as a second sample image pair corresponding to the sample target image; The sample reconstructed image comprises: a first sample reconstructed image based on the first sample image pair, and a second sample reconstructed image based on the second sample image pair; Based on the first sample reconstructed image and the sample target image, determining a photometric reconstruction loss between the first sample reconstructed image and the sample target image as a third photometric reconstruction loss; Based on the second sample reconstructed image and the sample target image, determining a photometric reconstruction loss between the second sample reconstructed image and the sample target image as a fourth photometric reconstruction loss; The minimum value of the third photometric reconstruction loss and the fourth photometric reconstruction loss is determined as the graphic reconstruction loss.
7. The lightweight 3D video stabilization method based on the CNN-Transformer hybrid framework according to claim 1, characterized in that: The step of reconstructing each jittery video frame in the jittery video frame sequence to be processed according to the depth map prediction result of the jittery video frame sequence to be processed and the predicted position and posture after the smoothing process to obtain a stabilized video frame sequence includes: Taking each jittery video frame in the sequence of jittery video frames to be processed as a current jittery video frame to be processed, reconstructing the next frame of the current jittery video frame to be processed according to the depth map prediction result of the current jittery video frame to be processed, the predicted position and posture after smoothing of the current jittery video frame to be processed, the predicted position and posture after smoothing of the next frame of the current jittery video frame to be processed, and the current jittery video frame to be processed, so as to obtain a stabilized next frame of the current jittery video frame to be processed; The stabilized next frame of each shaking video frame in the shaking video frame sequence to be processed constitutes a stabilized video frame sequence.
8. The lightweight 3D video stabilization method based on the CNN-Transformer hybrid framework according to any one of claims 4 to 6, characterized in that: The process of photometric reconstruction calculation includes: For a first image and a second image, determining a structural similarity index SSIM between the first image and the second image; For the first image and the second image, determining an L1 loss between the first image and the second image; According to preset weights, the structural similarity index SSIM and the L1 loss are weightedly summed to obtain a photometric reconstruction loss between the first image and the second image.
9. A lightweight 3D video stabilization device based on a CNN-Transformer hybrid framework, characterized in that: The device comprises: An input module is used to input the jittered video frame sequence to be processed into a pre-trained image reconstruction network to obtain a depth map prediction result and a pose prediction result of the jittered video frame sequence to be processed, wherein the training samples of the pre-trained image reconstruction network are sample image pairs in the sample jittered video frame sequence, and the sample image pairs include a sample source image and a corresponding sample target image; the pre-trained image reconstruction network is obtained through unsupervised training with the goal of reconstructing the sample target image; A determination module, configured to determine a predicted pose of each jittery video frame in the jittery video frame sequence to be processed by using the pose prediction result of the jittery video frame sequence to be processed; A smoothing processing module, used for smoothing the predicted position and posture of each jittery video frame in the jittery video frame sequence to be processed to obtain a smoothed predicted position and posture; A reconstruction module is used to reconstruct each jittery video frame in the jittery video frame sequence to be processed according to the depth map prediction result of the jittery video frame sequence to be processed and the predicted posture after the smoothing process to obtain a stabilized video frame sequence.
10. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement a lightweight 3D video stabilization method based on a CNN-Transformer hybrid framework as described in any one of claims 1-8.