Video image stabilization method and device and electronic equipment
By segmenting videos into scene fragments and using deep learning and optimization models to calculate camera motion trajectories, the problem of stabilization in complex scenes and resource-constrained devices is solved, achieving efficient and accurate video stabilization results.
Patent Information
- Application Number
- CN202510382033.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-08-01
AI Technical Summary
Existing video stabilization technologies have limitations when dealing with complex scenes or low-texture areas. Traditional methods have high computational complexity, while deep learning methods have high resource requirements, making it difficult to achieve real-time stabilization on resource-constrained devices.
By segmenting the video into multiple scene segments, deep learning is used for shot transition detection and video scene segmentation. An optimized deep mesh flow model is combined to calculate the camera motion trajectory, and a smooth distortion field is generated through a camera path smoothing network to convert shaky video frames into stable video frames.
It significantly improves the visual stability of videos, enabling them to handle complex scenes and significant jitter, providing a smoother and more stable viewing experience, while reducing computational complexity and resource requirements.
Smart Images

Figure CN120416664A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of video processing, and more particularly, to a video stabilization method, apparatus, and electronic device. Background Art
[0002] Video has become an indispensable information dissemination medium in modern life. A high-quality video experience not only requires clear images and true colors, but also ensures stable images. Video stabilization technology is therefore particularly important, as it can effectively eliminate jitter during shooting and improve viewing comfort.
[0003] Current video stabilization technologies are mainly divided into traditional 2D / 3D methods and deep learning-based methods. Traditional methods estimate camera motion through pixel correspondence, but have limitations in dealing with complex scenes or low-texture areas; 3D methods have higher accuracy, but are computationally complex and require a large amount of resources. Deep learning methods predict camera jitter parameters by mining video frame information and perform well in dealing with complex scenes, but also face challenges such as high computational resource requirements and poor algorithm interpretability.
[0004] Traditional methods have poor effects in dealing with low-frequency jitter and large-amplitude jitter, and are prone to reducing the video field of view; deep learning methods are difficult to achieve real-time stabilization on resource-constrained devices due to high computational complexity. Summary of the Invention
[0005] In view of the problems existing in the above-mentioned prior art, embodiments of the present application provide a video stabilization method, apparatus, and electronic device, aiming to provide an efficient and accurate video stabilization solution for video shooting and application scenarios.
[0006] In a first aspect, embodiments of the present application provide a video stabilization method, including the following steps:
[0007] Divide the initial video into multiple scene segments;
[0008] Judge the video stability of the multiple scene segments;
[0009] If the video of the scene segment is unstable, calculate the motion trajectory of the camera according to the jitter of the video frames; and
[0010] Convert the jittery video frames into stable video frames according to the motion trajectory of the camera.
[0011] Further, the step of dividing the initial video into multiple scene segments includes:
[0012] Based on a deep learning-based video analysis model, perform shot transition detection and video scene division on the initial video to layer the initial video into multiple scene segments.
[0013] Further, the video stability judgment for the multiple scene segments includes:
[0014] Calculating the average motion vector between video frames of the scene segment; and
[0015] If the average motion vectors of multiple consecutive video frames exceed a preset threshold, it is determined that the scene segment is an unstable video.
[0016] Further, the calculation of the camera's motion trajectory according to the jitter of the video frames includes:
[0017] Precisely calculating the motion between video frames through an optimized deep grid flow model; and
[0018] Calculating the camera's motion trajectory according to the motion between the video frames.
[0019] Further, the conversion of the jittery video frames into stable video frames according to the camera's motion trajectory includes:
[0020] Obtaining a smooth warping field that matches the video frame size of the original video according to the camera's motion trajectory; and
[0021] Based on the smooth warping field, converting the jittery video frames into stable video frames.
[0022] Further, the obtaining of a smooth warping field that matches the video frame size of the original video according to the camera's motion trajectory includes:
[0023] Inputting the camera's motion trajectory into a camera path smoothing network to obtain a smooth warping field that matches the video frame size of the original video.
[0024] Further, after converting the jittery video frames into stable video frames according to the camera's motion trajectory, it further includes:
[0025] Stitching the stable video frames to generate a stable target video.
[0026] In a second aspect, an embodiment of the present application further provides a video stabilization device, including:
[0027] A video segmentation module for dividing an original video into multiple scene segments;
[0028] A stability judgment module for performing video stability judgment on the multiple scene segments;
[0029] A motion trajectory calculation module for calculating the camera's motion trajectory according to the jitter of the video frames if the video of the scene segment is unstable; and
[0030] A stable video conversion module, configured to convert jittery video frames into stable video frames according to the motion trajectory of the camera.
[0031] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor is configured to implement the video stabilization method according to the first aspect described above when executing the program.
[0032] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. The computer program is configured to implement the video stabilization method according to the first aspect described above.
[0033] The embodiments of the present application bring the following beneficial effects:
[0034] In the video stabilization method provided by the embodiments of the present application, first, an initial video is segmented into multiple scene segments for independent analysis of the stability of each segment. For the scene segments determined to be unstable, the method further calculates the motion trajectory of the camera in the video frames. Finally, based on the obtained motion trajectory, the jittery video frames are converted into stable video frames, thereby significantly improving the visual stability of the video. The video stabilization method provided by the embodiments of the present application ensures effective processing of each unstable segment through detailed scene division and accurate jitter analysis. Its advantage lies in being able to handle complex scenes and large jitters, providing a smoother and more stable viewing experience for the audience. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0036] Figure 1 It is a schematic flowchart of the video stabilization method provided by the embodiments of the present application;
[0037] Figure 2 It is a schematic diagram of the algorithm process of the video stabilization method provided by the embodiments of the present application;
[0038] Figure 3 It is a structural block diagram of the video stabilization device provided by the embodiments of the present application;
[0039] Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiments of the present application.
[0040] The realization of the purpose of this application, functional features and advantages will be further described in conjunction with embodiments with reference to the accompanying drawings. Detailed implementation manners
[0041] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0042] In the description of this application, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, unless otherwise specified, the meaning of "a plurality" is two or more. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0043] Refer to Figure 1 and Figure 2 , Figure 1 and Figure 2 are the flow framework diagram and algorithm process schematic diagram of the video stabilization method according to the embodiments of this application. As Figure 1 and Figure 2 shown, the video stabilization method of the embodiments of this application includes the following steps:
[0044] S101: Divide the initial video into multiple scene segments; [[ID=!28]]
[0045] S102: Determine the video stability of the multiple scene segments;
[0046] S103: If the video of the scene segment is unstable, calculate the motion trajectory of the camera according to the jitter of the video frames; and
[0047] S104: Convert the jittery video frames into stable video frames according to the motion trajectory of the camera.
[0048] In the video stabilization method provided by the embodiments of this application, first, the initial video is segmented into multiple scene segments for independent analysis of the stability of each segment. For the scene segments determined to be unstable, the method further calculates the motion trajectory of the camera in the video frames. Finally, based on the obtained motion trajectory, the jittery video frames are converted into stable video frames, thereby significantly improving the visual stability of the video.
[0049] The video stabilization method provided by the embodiments of this application ensures the effective processing of each unstable segment through detailed scene division and accurate jitter analysis. Its advantage lies in being able to handle complex scenes and large jitters, providing a smoother and more stable viewing experience for the audience.
[0050] Further, in some embodiments of this application, the dividing the initial video into multiple scene segments includes:
[0051] Performing shot transition detection and video scene division on the initial video based on a deep learning-based video analysis model to layer the initial video into multiple scene segments.
[0052] Specifically, the input initial video may be composed of videos of multiple scenes. By using a deep learning-based video analysis model, shot transition detection and video scene division are performed on the input video. This model identifies the positions of shot transitions by analyzing the visual feature differences between video frames, such as color histogram changes, edge structure changes, motion vector distributions, etc. For example, when there are significant changes in color distribution and inconsistent motion patterns between two adjacent frames, the model determines it as a shot transition. At the same time, according to the semantic coherence and visual similarity of the video content, the video is divided into multiple scene segments, providing a basis for subsequent processing of different segments.
[0053] The video stabilization method provided by the embodiments of this application introduces a preprocessing process of shot transition detection and video scene segmentation in the video stabilization process, which can accurately identify the shot switching points in the video and thus cut it into different video segments. Through shot transition detection, the transmission of motion estimation errors caused by shot switching can be effectively avoided, ensuring that the stabilization processing within each shot segment is based on its own independent and accurate motion model, thereby significantly improving the accuracy and stability of stabilization.
[0054] Further, in some embodiments of this application, the determining the video stability of the multiple scene segments includes:
[0055] Calculating the average motion vector between video frames of the scene segment; and
[0056] If the average motion vectors of multiple consecutive video frames exceed a preset threshold, then determine that the scene segment is an unstable video.
[0057] Specifically, in the video stabilization process, it is crucial to determine the stability of scene segments. First, calculate the average motion vector between video frames within each scene segment. This step is obtained by comparing the position changes of pixels or feature points in adjacent frames. Subsequently, based on these average motion vectors, if the vector values of consecutive multiple video frames exceed a preset threshold, it is determined that the scene segment is unstable. The preset threshold is set according to actual applications and stabilization requirements to distinguish normal motion from abnormal jitter.
[0058] By accurately calculating and analyzing the motion changes between video frames, it is possible to effectively identify scene segments with jitter. Its high efficiency and accuracy provide a solid foundation for subsequent stabilization processing, ensuring the improvement of the overall video stabilization effect.
[0059] Furthermore, in some embodiments of the present application, the calculating the motion trajectory of the camera according to the jitter of the video frames includes:
[0060] Precisely calculating the motion between video frames through an optimized deep grid flow model; and
[0061] Calculating the motion trajectory of the camera according to the motion between the video frames.
[0062] Specifically, for each input unstable video segment, an optimized deep grid flow model is used to precisely calculate the motion between video frames. Through learning a large amount of video data, this model can adaptively extract depth features and motion information in video frames, especially showing excellent performance in complex scenes and low-texture areas. During the calculation process, the model not only considers the local motion relationship between adjacent frames but also integrates global motion information through multi-scale feature fusion technology, thereby more comprehensively and accurately estimating the motion trajectory of the camera. At the same time, to reduce the computational amount and improve real-time performance, the model adopts a lightweight network structure design and an efficient feature extraction algorithm, and can quickly complete the motion estimation task without losing too much accuracy.
[0063] The video stabilization method provided by the embodiments of the present application adopts a deep neural network model that fuses multi-modal features. In addition to conventional image pixel information, it also integrates multi-modal data such as texture features, edge features, and depth information (if available) of the image. This way of fusing multi-modal features enriches the information source for motion estimation, enabling the model to more comprehensively and deeply understand the motion relationship in the video.
[0064] Furthermore, in some embodiments of the present application, the converting the jittery video frames into stable video frames according to the motion trajectory of the camera includes:
[0065] Obtaining a smooth warping field that matches the video frame size of the original video according to the motion trajectory of the camera; and
[0066] Convert the jittery video frames into stable video frames based on the smooth warping field.
[0067] First, generate a smooth warping field that matches the size of the video frames according to the camera motion trajectory. This warping field accurately describes the mapping relationship from the jittery state to the stable state. Subsequently, based on this warping field, transform each jittery video frame to adjust the pixel positions, thereby eliminating the jitter and obtaining stable video frames. This process ensures the visual stability of the final video frames and enhances the viewing experience.
[0068] Further, in some embodiments of the present application, obtaining a smooth warping field that matches the size of the video frames of the initial video according to the motion trajectory of the camera includes:
[0069] Input the motion trajectory of the camera into a camera path smoothing network to obtain a smooth warping field that matches the size of the video frames of the initial video.
[0070] Specifically, take the calculated camera motion information as the input and feed it into a specially designed camera path smoothing network. This network adopts an encoder-decoder architecture. The encoder part gradually extracts high-level features of the input motion information through a series of convolutional layers and pooling layers, and the decoder part restores these features to a smooth warping field that matches the size of the original video frames through deconvolutional layers and upsampling layers. In the network structure, a channel attention mechanism is innovatively introduced. This mechanism can automatically learn the importance of different channel features and assign higher weights to important channels, so that the network focuses more on key motion feature information, effectively improving the prediction accuracy. The channel attention mechanism in the camera path smoothing network has strong compensation ability for low-frequency camera jitter, can significantly reduce the low-frequency jitter amplitude, and improve the overall stability of the video after image stabilization. Whether it is a slowly moving surveillance video or a long video shot by hand, it can present a smooth and stable visual effect.
[0071] Further, in some embodiments of the present application, after converting the jittery video frames into stable video frames according to the motion trajectory of the camera, it further includes:
[0072] Stitch the stable video frames to generate a stable target video.
[0073] That is, recombine the processed scene segments of each scene after image stabilization to ensure the smoothness and coherence of the entire video. Through precise image stabilization processing and efficient video frame stitching, a target video with high visual stability can be generated, which not only enhances the viewing experience of the video but also provides a basis for subsequent video editing and dissemination.
[0074] Figure 3 is the structural block diagram of the video stabilization device 200 provided by the embodiments of the present application. As Figure 3 shown, the video stabilization device 200 of the embodiments of the present application includes: a video segmentation module 210, a stability determination module 220, a motion trajectory calculation module 230, and a stabilized video conversion module 240, where:
[0075] The video segmentation module 210 is configured to divide an initial video into multiple scene segments;
[0076] The stability determination module 220 is configured to perform video stability determination on the multiple scene segments;
[0077] The motion trajectory calculation module 230 is configured to, if the video of the scene segment is unstable, calculate the motion trajectory of the camera according to the jitter of the video frames; and
[0078] The stabilized video conversion module 240 is configured to convert the jittery video frames into stabilized video frames according to the motion trajectory of the camera.
[0079] In the video stabilization device provided by the embodiments of the present application, first, the initial video is segmented into multiple scene segments for independent analysis of the stability of each segment. For the scene segments determined to be unstable, the method further calculates the motion trajectory of the camera in the video frames. Finally, based on the obtained motion trajectory, the jittery video frames are converted into stabilized video frames, thereby significantly improving the visual stability of the video. The video stabilization device provided by the embodiments of the present application ensures effective processing of each unstable segment through detailed scene segmentation and accurate jitter analysis. Its advantage lies in being able to handle complex scenes and large jitters, providing a smoother and more stable viewing experience for the audience.
[0080] It should be noted that the specific implementation manner of the video stabilization device of the embodiments of the present application is similar to the specific implementation manner of the video stabilization method of the embodiments of the present application. For details, please refer to the description in the method part, and will not be elaborated here.
[0081] Figure 4 is the structural schematic diagram of the electronic device 300 of the embodiments of the present application.
[0082] As Figure 4As shown, the electronic device 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage section 302 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0083] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read from it can be installed into the storage section 308 as needed.
[0084] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, the above functions defined in the electronic device of the present application are executed.
[0085] It should be noted that the computer-readable medium shown in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. A computer-readable storage medium can be, for example, but not limited to, an electronic device, apparatus, or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0086] In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction-executing electronic device, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction-executing electronic device, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the foregoing.
[0087] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of a processing receiving device, method, and computer program product according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the foregoing module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based electronic device that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0088] The units or modules involved in the embodiments described in the present application may be implemented in software or in hardware. The described units or modules may also be provided in a processor, and when the processor executes the program, it implements the video stabilization method:
[0089] Divide the initial video into multiple scene segments;
[0090] Perform video stability judgment on the multiple scene segments;
[0091] If the video of the scene segment is unstable, calculate the motion trajectory of the camera according to the jitter of the video frame; and
[0092] Convert the jittery video frames into stable video frames according to the motion trajectory of the camera.
[0093] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the foregoing programs are executed by one or more processors to perform the video stabilization method described in the present application:
[0094] Divide the initial video into multiple scene segments;
[0095] Perform video stability judgment on the multiple scene segments;
[0096] If the video of the scene segment is unstable, calculate the motion trajectory of the camera according to the jitter of the video frames; and
[0097] Convert the jittery video frames into stable video frames according to the motion trajectory of the camera.
[0098] As another aspect, the present application also provides a computer program product, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer program product stores one or more programs, and when the foregoing programs are executed by one or more processors to perform the video stabilization method described in the present application:
[0099] Divide the initial video into multiple scene segments;
[0100] Perform video stability judgment on the multiple scene segments;
[0101] If the video of the scene segment is unstable, calculate the motion trajectory of the camera according to the jitter of the video frames; and
[0102] Convert the jittery video frames into stable video frames according to the motion trajectory of the camera.
[0103] The foregoing are only preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made using the specification and drawings of the present application under the application concept of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A video stabilization method, characterized in that It includes the following steps: Divide the initial video into multiple scene segments; Judge the video stability of the multiple scene segments; If the video of the scene segment is unstable, calculate the movement trajectory of the camera according to the jitter of the video frames; and Convert the jittery video frames into stable video frames according to the movement trajectory of the camera.
2. The video stabilization method according to claim 1, wherein The dividing the initial video into multiple scene segments includes: Based on a deep learning-based video analysis model, perform shot transition detection and video scene division on the initial video to layer the initial video into multiple scene segments.
3. The video stabilization method according to claim 1, wherein The judging the video stability of the multiple scene segments includes: Calculate the average motion vector between the video frames of the scene segment; and If the average motion vector of multiple consecutive video frames exceeds a preset threshold, judge that the scene segment is an unstable video.
4. The video stabilization method according to claim 1, wherein The calculating the movement trajectory of the camera according to the jitter of the video frames includes: Precisely calculate the motion between video frames through an optimized deep grid flow model; and Calculate the movement trajectory of the camera according to the motion between the video frames.
5. The video stabilization method according to claim 1, characterized in that, The converting the jittery video frames into stable video frames according to the movement trajectory of the camera includes: Obtain a smooth warping field that matches the video frame size of the initial video according to the movement trajectory of the camera; and Based on the smooth warping field, convert the jittery video frames into stable video frames.
6. The video stabilization method according to claim 5, characterized in that, The obtaining a smooth warping field that matches the video frame size of the initial video according to the movement trajectory of the camera includes: Input the movement trajectory of the camera into a camera path smoothing network to obtain a smooth warping field that matches the video frame size of the initial video.
7. The video stabilization method according to claim 1, wherein After the converting the jittery video frames into stable video frames according to the movement trajectory of the camera, it further includes: Stitch the stable video frames to generate a stable target video.
8. A video stabilization device, characterized in that, It includes: A video segmentation module for dividing the initial video into multiple scene segments; A stability judgment module for judging the video stability of the multiple scene segments; A movement trajectory calculation module for calculating the movement trajectory of the camera according to the jitter of the video frames if the video of the scene segment is unstable; and A stable video conversion module for converting the jittery video frames into stable video frames according to the movement trajectory of the camera.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor is used to implement the video image stabilization method according to any one of claims 1-7 when executing the program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is used to implement the video image stabilization method according to any one of claims 1-7.
Citation Information
Cited By
Dynamic scene adaptive image stabilization processing method and device, equipment and medium
CN121438167A