A video stabilization method, device, terminal and storage medium

CN115460341BActive Publication Date: 2026-09-04WUHAN TCL CORP RES CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110643649.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-09
Publication Date
2026-09-04
Estimated Expiration
2041-06-09

AI Technical Summary

Technical Problem

[0004]本发明要解决的技术问题在于,针对现有技术的上述缺陷,提供一种视频稳化方法、装置、终端及存储介质,旨在解决现有的稳化矩阵在对视频帧进行稳化后,容易出现稳化越界的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115460341B_ABST
    Figure CN115460341B_ABST
Patent Text Reader

Abstract

The application discloses a video stabilization method and device, a terminal and a storage medium. The method comprises the following steps: obtaining a to-be-processed video frame; determining a stabilized video frame of each video frame except a first frame in the to-be-processed video frame; wherein, for each video frame, an offset prediction value corresponding to the video frame is determined; the video frame is subjected to a stabilization treatment according to the offset prediction value, so that the stabilized video frame of the video frame is obtained; and the stabilized video corresponding to the to-be-processed video frame is determined according to the first frame in the to-be-processed video frame and the stabilized video frame of each video frame except the first frame in the to-be-processed video frame. Since the offset prediction value can be used to correct the offset of the original stabilized video frame of the video frame when the original stabilized video frame exceeds the boundary, the stabilized video frame obtained by the stabilization treatment of the video frame according to the offset prediction value of the video frame will not exceed the boundary, so that the problem that the existing stabilization matrix is prone to exceeding the boundary after the stabilization of the video frame is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing, and more particularly to a video stabilization method, apparatus, terminal, and storage medium. Background Technology

[0002] With the development of technology, various types of terminal devices are equipped with video recording functions. However, some terminal devices, due to the inability to control device stability, are more prone to video shakiness compared to other devices, reducing video quality. Video stabilization technology can effectively improve video quality and reduce the amplitude of shakiness between video frames. Video stabilization technology refers to converting the original recorded video into a stabilized video through special cropping and interpolation methods, so that the stabilized video playback content presents uniform motion. To ensure that the stabilized image content and the original video content have no visual distortion difference, the cropping method needs to be restricted to operations such as rotation, proportional scaling, and translation of a 2D plane in 3D space. Existing video stabilization technology usually calculates motion from the motion estimation and smooth motion estimation of the original video frames to obtain a stabilization matrix, and then uses this stabilization matrix to crop the original video frames to obtain the corresponding stabilized video frames. However, the stabilized video frames obtained based on this stabilization matrix are prone to the problem of stabilization exceeding the boundaries, that is, the stabilized video frames appear outside the content of the current frame.

[0003] Therefore, existing technologies still need improvement and development. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a video stabilization method, device, terminal and storage medium to address the above-mentioned defects of the prior art, and to solve the problem that existing stabilization matrices are prone to stabilization exceeding the limit after stabilizing video frames.

[0005] The technical solution adopted by this invention to solve the problem is as follows:

[0006] In a first aspect, embodiments of the present invention provide a video stabilization method, the method comprising:

[0007] Obtain the video frames to be processed;

[0008] Determine the stabilized video frame for each video frame that is not the first frame in the video frames to be processed; wherein, for each video frame, determine the offset prediction value corresponding to the video frame; and perform stabilization processing on the video frame according to the offset prediction value to obtain the stabilized video frame of the video frame.

[0009] The stabilized video corresponding to the video frame to be processed is determined based on the first video frame and the stabilized video frames of each non-first video frame in the video frame to be processed.

[0010] In one implementation, determining the offset prediction value corresponding to each video frame includes:

[0011] Obtain several reference video frames corresponding to the video frame; wherein, the several reference video frames are several consecutive video frames whose playback time is before the video frame and adjacent to the video frame;

[0012] Obtain the original stabilized video frame corresponding to the video frame and the original stabilized video frame corresponding to the plurality of reference video frames respectively;

[0013] The first image offset information is determined based on the video frame, the original stabilized video frame corresponding to the video frame, the plurality of reference video frames, and the original stabilized video frames corresponding to the plurality of reference video frames respectively.

[0014] The second image offset information is determined based on the video frame and the plurality of reference video frames;

[0015] The first image offset information and the second image offset information are input into the prediction model to obtain the offset prediction value.

[0016] In one embodiment, determining the first image offset information based on the video frame, the first stabilized video frame, the plurality of reference video frames, and the original stabilized video frames corresponding to the plurality of reference video frames respectively includes:

[0017] A first original stabilization matrix is ​​determined based on the video frame and the original stabilized video frame corresponding to the video frame. The first original stabilization matrix is ​​used to reflect the projective transformation relationship between the video frame and the first stabilized video frame.

[0018] Based on the plurality of reference video frames and the original stabilized video frames corresponding to the plurality of reference video frames, a second original stabilization matrix corresponding to the plurality of reference video frames is determined; wherein, for each video frame among the plurality of reference video frames, a second original stabilization matrix corresponding to the video frame is determined based on the video frame and the original stabilized video frame corresponding to the video frame, and the second original stabilization matrix is ​​used to reflect the projective transformation relationship between the video frame and the second stabilized video frame corresponding to the video frame;

[0019] The first image offset information is determined based on the first original stabilization matrix and the second original stabilization matrices corresponding to the plurality of reference video frames.

[0020] In one implementation, determining the second image offset information based on the video frame and the plurality of reference video frames includes:

[0021] Obtain the motion matrix between two adjacent video frames in the video frame and the plurality of reference video frames to obtain a plurality of motion matrices. Each of the plurality of motion matrices is used to reflect the projective transformation relationship between the two adjacent video frames corresponding to the motion matrix.

[0022] The second image offset information is determined based on the aforementioned motion matrices.

[0023] In one implementation, the offset prediction value includes: an angle offset prediction value; the step of inputting the first image offset information and the second image offset information into a prediction model to obtain the offset prediction value includes:

[0024] Obtain first angle offset information and second angle offset information, wherein the first angle offset information is the angle offset information in the first image offset information, and the second angle offset information is the angle offset information in the second image offset information;

[0025] The first angle offset information and the second angle offset information are input into the angle offset prediction model to obtain the angle offset prediction value, wherein the angle offset prediction model is the model used in the prediction model to predict the angle offset value between the video frame and the original stabilized video frame.

[0026] In one embodiment, the step of inputting the first angle offset information and the second angle offset information into the angle offset prediction model to obtain the angle offset prediction value includes:

[0027] Input the first angle offset information and the second angle offset information into the angle offset prediction model to obtain the initial angle offset prediction value;

[0028] The stabilization clipping range is determined based on the original stabilization matrix, and the angle offset range is determined based on the stabilization clipping range.

[0029] The initial angle offset prediction value is adjusted according to the angle offset range to obtain the angle offset prediction value.

[0030] In one embodiment, the offset prediction value includes: a horizontal displacement prediction value; the step of inputting the first image offset information and the second image offset information into a prediction model to obtain the offset prediction value includes:

[0031] Obtain first horizontal displacement information and second horizontal displacement information, wherein the first horizontal displacement information is horizontal displacement information generated by the first image offset information based on the angle offset prediction value, and the second horizontal displacement information is horizontal displacement information generated by the second image offset information based on the angle offset prediction value;

[0032] The first horizontal displacement information and the second horizontal displacement information are input into the horizontal displacement prediction model to obtain the horizontal displacement prediction value. The horizontal displacement prediction model is the model used in the prediction model to predict the horizontal displacement value between the video frame and the original stabilized video frame.

[0033] In one embodiment, the step of inputting the first horizontal displacement information and the second horizontal displacement information into a horizontal displacement prediction model to obtain the predicted horizontal displacement value includes:

[0034] Input the first horizontal displacement information and the second horizontal displacement information into the horizontal displacement prediction model to obtain the initial horizontal displacement prediction value;

[0035] The horizontal displacement range is determined based on the stabilized cutting range and the predicted angle offset value.

[0036] The initial horizontal displacement prediction value is adjusted according to the horizontal displacement range to obtain the horizontal displacement prediction value.

[0037] In one embodiment, the offset prediction value includes: a vertical displacement prediction value; the step of inputting the first image offset information and the second image offset information into the prediction model to obtain the offset prediction value includes:

[0038] Obtain first vertical displacement information and second vertical displacement information, wherein the first vertical displacement information is vertical displacement information generated by the first image offset information based on the angle offset prediction value, and the second vertical displacement information is vertical displacement information generated by the second image offset information based on the angle offset prediction value;

[0039] The first vertical displacement information and the second vertical displacement information are input into the vertical displacement prediction model to obtain the vertical displacement prediction value, wherein the vertical displacement prediction model is the model used in the prediction model to predict the vertical displacement value between the video frame and the original stabilized video frame.

[0040] In one embodiment, the step of inputting the first vertical displacement information and the second vertical displacement information into a vertical displacement prediction model to obtain the predicted vertical displacement value includes:

[0041] Input the first vertical displacement information and the second vertical displacement information into the vertical displacement prediction model to obtain the initial vertical displacement prediction value;

[0042] The vertical displacement range is determined based on the stabilized cutting range and the predicted angle offset value.

[0043] The initial vertical displacement prediction value is adjusted according to the vertical displacement range to obtain the vertical displacement prediction value.

[0044] In one embodiment, the step of stabilizing the video frame based on the offset prediction value to obtain a stabilized video frame includes:

[0045] A correction matrix is ​​generated based on the predicted offset values;

[0046] The first original stabilization matrix is ​​corrected according to the correction matrix to obtain the target stabilization matrix corresponding to the video frame;

[0047] The video frame is stabilized according to the target stabilization matrix to obtain the stabilized video frame.

[0048] In one implementation, the prediction model is a trained model, wherein the training process of the prediction model is as follows:

[0049] The third image offset information and the fourth image offset information in the training data are input into the original prediction model, and the original prediction model generates the offset prediction value corresponding to the third image offset information and the fourth image offset information. The training data includes multiple training information groups, each training information group includes the third image offset information, the fourth image offset information and the standard offset value, and the standard offset value is the offset value corresponding to the third image offset information and the fourth image offset information.

[0050] Based on the standard offset values ​​corresponding to the third and fourth image offset information and the offset prediction values ​​corresponding to the third and fourth image offset information, the model parameters of the original prediction model are adjusted, and the step of inputting the third and fourth image offset information from the training data into the original prediction model is continued until the preset training conditions are met to obtain the prediction model.

[0051] Secondly, embodiments of the present invention also provide a video stabilization device, wherein the device includes:

[0052] The acquisition module is used to acquire the video frames to be processed.

[0053] The stabilization module determines the stabilized video frame for each video frame other than the first frame in the video frame to be processed; wherein, for each video frame, the offset prediction value corresponding to the video frame is determined; and the video frame is stabilized according to the offset prediction value to obtain the stabilized video frame of the video frame.

[0054] The output module determines the stable video corresponding to the video frame to be processed based on the first video frame and the stabilized video frames of each video frame other than the first video frame in the video frame to be processed.

[0055] Thirdly, embodiments of the present invention also provide a terminal, wherein the terminal includes a memory and one or more processors; the memory stores one or more programs; the programs include instructions for performing the video stabilization method as described above; and the processor is used to execute the programs.

[0056] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a plurality of instructions, wherein the instructions are loaded and executed by a processor to implement the steps of any of the video stabilization methods described above.

[0057] The beneficial effects of this invention are as follows: This embodiment of the invention acquires a video frame to be processed; determines a stabilized video frame for each video frame other than the first frame in the video frame to be processed; for each video frame, determines the corresponding offset prediction value; performs stabilization processing on the video frame based on the offset prediction value to obtain a stabilized video frame; and determines the stabilized video corresponding to the video frame to be processed based on the first video frame and the stabilized video frames for each video frame other than the first frame in the video frame to be processed. Since the offset prediction value can be used to correct the offset when the original stabilized video frame corresponding to the video frame exceeds the stabilization limit, this invention performs stabilization processing on the video frame based on the offset prediction value, and the resulting stabilized video frame will no longer exceed the stabilization limit, thereby solving the problem that existing stabilization matrices easily cause stabilization limit exceedances after stabilizing video frames. Attached Figure Description

[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0059] Figure 1 This is a flowchart illustrating the video stabilization method provided in an embodiment of the present invention.

[0060] Figure 2 This is a schematic diagram illustrating an application scenario of video frame transformation provided in an embodiment of the present invention.

[0061] Figure 3 This is a reference schematic diagram of the stabilization cutting range provided in an embodiment of the present invention.

[0062] Figure 4This is a schematic diagram of the feature weighted matching model provided in an embodiment of the present invention.

[0063] Figure 5 This is a scatter plot comparing the video jitter pixel offset error generated after stabilization using a stabilization matrix generated based on existing technology and a stabilization matrix generated based on the present invention, as provided in this embodiment of the invention.

[0064] Figure 6 This is a comparison diagram of the border positions of the stabilization matrix generated based on the prior art and the video stabilization generated based on the present invention on the same video frame, provided by an embodiment of the present invention.

[0065] Figure 7 This is provided by the embodiments of the present invention. Figure 6 A comparison diagram of the positions of the borders of the stabilized video frames generated based on existing technologies and the video stabilization generated based on the present invention on adjacent video frames.

[0066] Figure 8 This is a schematic diagram of the video stabilization device provided in an embodiment of the present invention.

[0067] Figure 9 This is a schematic diagram of the terminal provided in the embodiment of the present invention.

[0068] Figure 10 This is a schematic diagram of the process for generating the correction matrix provided in an embodiment of the present invention.

[0069] Figure 11 This is a schematic diagram of the training and prediction model provided in this invention. Detailed Implementation

[0070] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0071] It should be noted that if the embodiments of the present invention involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.

[0072] With the development of technology, various types of terminal devices are equipped with video recording functions. However, some terminal devices, due to the inability to control device stability, are more prone to video shakiness compared to other devices, reducing video quality. Video stabilization technology can effectively improve video quality and reduce the amplitude of shakiness between video frames. Video stabilization technology refers to converting the original recorded video into a stabilized video through special cropping and interpolation methods, so that the stabilized video playback content presents uniform motion. To ensure that the stabilized image content and the original video content have no visual distortion difference, the cropping method needs to be restricted to operations such as rotation, proportional scaling, and translation of a 2D plane in 3D space. Existing video stabilization technology usually calculates motion from the motion estimation and smooth motion estimation of the original video frames to obtain a stabilization matrix, and then uses this stabilization matrix to crop the original video frames to obtain the corresponding stabilized video frames. However, the stabilized video frames obtained based on this stabilization matrix are prone to the problem of stabilization exceeding the boundaries, that is, the stabilized video frames appear outside the content of the current frame.

[0073] To address the aforementioned deficiencies in existing technologies, this invention provides a video stabilization method. This method involves: acquiring a video frame to be processed; determining a stabilized video frame for each non-first frame within the video frame to be processed; determining an offset prediction value for each video frame; performing stabilization processing on the video frame based on the offset prediction value to obtain a stabilized video frame; and determining the stabilized video corresponding to the video frame to be processed based on the first video frame and the stabilized video frames for each non-first frame within the video frame to be processed. Since the offset prediction value can be used to correct the offset when the original stabilized video frame corresponding to the video frame exceeds the stabilization limit, this invention stabilizes the video frame based on the offset prediction value, ensuring that the resulting stabilized video frame will not exceed the stabilization limit, thus solving the problem of existing stabilization matrices easily exceeding the stabilization limit after stabilizing video frames.

[0074] For example, such as Figure 2As shown, in this embodiment of the invention, video frames F1, F2, and F3 to be processed can be obtained. F1 is the first video frame, so F1 does not need to be processed. F2 and F3 are not the first video frames, so it is necessary to determine the offset prediction values ​​offset2 and offset3 corresponding to F2 and F3, respectively. Since offset2 can determine the offset between F2 and its corresponding original stabilized video frame S2, and offset3 can determine the offset between F3 and its corresponding original stabilized video frame S3, F2 is stabilized according to offset2 to obtain the stabilized video frame S2' corresponding to F2, and F3 is stabilized according to offset3 to obtain the stabilized video frame S3' corresponding to F3. Thus, neither S2' nor S3' will have a stabilization out-of-bounds situation. Finally, the stabilized video corresponding to the video frames F1, F2, and F3 to be processed is obtained according to F1, S2', and S3'.

[0075] Exemplary methods

[0076] Before describing the video stabilization method provided by this invention, we will first combine... Figure 2 An illustrative explanation of the concepts related to the stabilization matrix is ​​provided.

[0077] For ease of description, in the embodiments of this application, the video frames captured by the terminal device are represented by F. For example... Figure 2 As shown, the first video frame captured is denoted as F1, the second video frame as F2, the third video frame as F3, and so on.

[0078] There is a projective transformation relationship between two adjacent video frames, and this projective transformation relationship can be obtained through motion parameters. M can be represented by the value M. M can be calculated based on motion estimation. Specific implementation methods for motion estimation can be found in existing methods, which will not be described in this application.

[0079] For the t-th (t>1) video frame F captured i And the previous video frame F i-1 F i Image content and F i-1 Image content can be correlated with each other based on motion parameters M i The approximate representation of the two-dimensional planar projective transformation, i.e., F i ≈f(F i-1 M i ), where the function f represents the projective transformation.

[0080] Let (x) i-1 ,y i-1 ,1)∈F i-1 F represents i-1 The corresponding homogeneous coordinate point in (x)i ,y i ,1)∈F i F represents i The corresponding F in i-1 Any homogeneous coordinate point.

[0081] make:

[0082]

[0083] Then we have:

[0084]

[0085] Where F1 represents the first acquired raw video frame; S1 represents the first acquired stabilized video frame; and R1 represents the stabilization matrix. For example... Figure 2 As shown, F i Stabilized video frames S i It is based on a projective transformation matrix (R is used in this application for ease of description). i (to represent) for F i The result obtained by cropping, i.e., S i ≈f(F i ,R i Let the stabilization cropping ratio be r, and the original video height and width be img_height and img_width, respectively. Then we have:

[0086]

[0087] For example, such as Figure 3 As shown, the content of S1 comes from a sub-image region of F1. Its height and width are both r times that of F1. The horizontal offset of the top left corner origin is (1-r)*img_width / 2, and the vertical offset is (1-r)*img_height / 2.

[0088] like Figure 1 As shown, this embodiment provides a video stabilization method, which includes the following steps:

[0089] Step S100: Obtain the video frame to be processed.

[0090] Specifically, since video jitter usually occurs between multiple video frames, this embodiment first needs to acquire the video frames to be stabilized. By stabilizing the video frames to be stabilized, this embodiment can reduce jitter between adjacent video frames, resulting in a better viewing experience for the user.

[0091] Step S200: Determine the stabilized video frame for each video frame that is not the first frame in the video frames to be processed; wherein, for each video frame, determine the offset prediction value corresponding to the video frame; and perform stabilization processing on the video frame according to the offset prediction value to obtain the stabilized video frame of the video frame.

[0092] Specifically, since the offset prediction value can be used to correct the offset when the original stabilized video frame corresponding to the video frame exceeds the stabilization limit, this invention stabilizes the video frame according to the offset prediction value, and the resulting stabilized video frame will no longer exceed the stabilization limit. Furthermore, since users cannot perceive video jitter when watching the first video frame in the video frame to be processed, but only when watching the second frame and subsequent video frames, this embodiment only needs to stabilize each non-first video frame in the video frame to be processed according to its corresponding offset prediction value to obtain its corresponding stabilized video frame.

[0093] For example, assuming the video frames to be processed are video frames A, B, and C in sequence, this embodiment does not need to perform stabilization processing on video frame A. It only needs to obtain the offset prediction values ​​b and c corresponding to video frames B and C, respectively. Based on the offset prediction value b, video frame B is stabilized to obtain the stabilized video frame B' corresponding to video frame B. Then, based on the offset prediction value c, video frame C is stabilized to obtain the stabilized video frame C' corresponding to video frame C.

[0094] In one implementation, determining the offset prediction value corresponding to each video frame specifically includes:

[0095] Step S201: Obtain a plurality of reference video frames corresponding to the video frame; wherein, the plurality of reference video frames are a plurality of consecutive video frames whose playback time is before the video frame and adjacent to the video frame;

[0096] Step S202: Obtain the original stabilized video frame corresponding to the video frame and the original stabilized video frames corresponding to the plurality of reference video frames respectively;

[0097] Step S203: Determine the first image offset information based on the video frame, the original stabilized video frame corresponding to the video frame, the plurality of reference video frames, and the original stabilized video frames corresponding to the plurality of reference video frames respectively.

[0098] Step S204: Determine the second image offset information based on the video frame and the plurality of reference video frames;

[0099] Step S205: Input the first image offset information and the second image offset information into the prediction model to obtain the offset prediction value.

[0100] This embodiment uses a single video frame as an example to illustrate the video stabilization method of the present invention. To stabilize the video frame, this embodiment needs to obtain several reference video frames corresponding to the video frame. These reference video frames are consecutive video frames whose playback time is before and adjacent to the video frame. It is understood that the video jitter generated between only two video frames is very weak. When a user perceives video jitter, it is often the result of the cumulative video jitter between multiple video frames. Therefore, to provide a better viewing experience, this embodiment needs to refer to several consecutive video frames adjacent to the video frame when stabilizing a single video frame. Specifically, this embodiment needs to obtain the original stabilized video frame corresponding to the video frame and the original stabilized video frames corresponding to the several reference video frames. The original stabilized video frame refers to the stabilized video frame corresponding to the video frame obtained after stabilizing the video frame using the stabilization matrix obtained from motion estimation. Since the original stabilized video frame may experience stabilization boundary violations, this embodiment does not directly use the original stabilized video frame corresponding to the video frame. Instead, it first needs to determine whether the original stabilized video frame will experience stabilization boundary violations. Specifically, this embodiment needs to obtain first image offset information, which reflects the image offset information between the video frame and its corresponding original stabilized video frame, as well as the image offset information between the plurality of reference video frames and their respective corresponding original stabilized video frames. Second image offset information is also obtained, which reflects the image offset information between two adjacent video frames in the video frame and the plurality of reference video frames. Then, the first and second image offset information are input into a pre-trained prediction model. The prediction model can then predict the image offset information between the video frame and its corresponding original stabilized video frame based on the input first and second image offset information, thus obtaining the offset prediction value.

[0101] For example, such as Figure 2As shown, F3 is the video frame, S3 is the original stabilized video frame corresponding to F3, F2 and F1 are reference video frames of F3, and S2 and S1 are the original stabilized video frames of F2 and F1, respectively. First image offset information is determined based on F3, S3, F2, S2, F1, and S1. This first image offset information can reflect the image offset information between F3 and S3, between F2 and S2, and between F1 and S1. For example, the first image offset information can be determined based on the positional change information of the line segments formed by the same point pairs in F3 and S3, the positional change information of the line segments formed by the same point pairs in F2 and S2, and the positional change information of the line segments formed by the same point pairs in F1 and S1. The second image offset information is determined based on F1, F2, and F3. This second image offset information can determine the image offset information between F1 and F2, and between F2 and F3. For example, it can be determined based on the positional change information of the line segment formed by the same point pair in F1 and F2, and the positional change information of the line segment formed by the same point pair in F2 and F3. Finally, the first and second image offset information are input into the prediction model to obtain the offset prediction value corresponding to F3.

[0102] In one implementation, step S203 specifically includes:

[0103] Step A1: Determine a first original stabilization matrix based on the video frame and the original stabilized video frame corresponding to the video frame. The first original stabilization matrix is ​​used to reflect the projective transformation relationship between the video frame and the first stabilized video frame.

[0104] Step A2: Based on the plurality of reference video frames and the original stabilized video frames corresponding to the plurality of reference video frames, determine the second original stabilization matrix corresponding to the plurality of reference video frames respectively; wherein, for each video frame in the plurality of reference video frames, determine the second original stabilization matrix corresponding to the video frame based on the video frame and the original stabilized video frame corresponding to the video frame, and the second original stabilization matrix is ​​used to reflect the projective transformation relationship between the video frame and the second stabilized video frame corresponding to the video frame;

[0105] Step A3: Determine the first image offset information based on the first original stabilization matrix and the second original stabilization matrices corresponding to the plurality of reference video frames respectively.

[0106] Since the first image offset information includes the image offset information between the video frame and its corresponding original stabilized video frame, as well as the image offset information between the plurality of reference video frames and their respective corresponding original stabilized video frames, determining the first image offset information requires first determining the above two types of image offset information. Specifically, in order to obtain the image offset information between the video frame and its corresponding original stabilized video frame, this embodiment needs to determine a first original stabilization matrix based on the video frame and its corresponding original stabilized video frame. The first original stabilization matrix is ​​used to reflect the projective transformation relationship between the video frame and the first stabilized video frame. That is, after the video frame undergoes a projective transformation based on the first original stabilization matrix, its corresponding original stabilized video frame can be obtained. Therefore, the first original stabilization matrix can also reflect the image offset information between the video frame and its corresponding original stabilized video frame to a certain extent. Similarly, to obtain the image offset information between the plurality of reference video frames and their respective corresponding original stabilized video frames, this embodiment needs to determine the second original stabilization matrix corresponding to each of the plurality of reference video frames. Since the second original stabilization matrix can reflect the projective transformation relationship between each reference video frame and its corresponding original video frame, it can also reflect the image offset information between each reference video frame and its corresponding original video frame to a certain extent. Finally, based on the first original stabilization matrix and the second original stabilization matrices corresponding to the plurality of reference video frames, the image offset information between the video frame and its corresponding original stabilized video frame and the image offset information between the plurality of reference video frames and their respective original stabilized video frames can be obtained, that is, the first image offset information can be obtained.

[0107] For example, such as Figure 2 As shown, assuming the video frame is F3, and F1 and F2 are reference video frames of F3, the original stabilized video frame of F3 is S3, the original stabilized video frame of F2 is S2, and the original stabilized video frame of F1 is S1, then it is necessary to obtain the original stabilization matrix R3 reflecting the projective transformation relationship between F3 and S3, the original stabilization matrix R3 reflecting the projective transformation relationship between F2 and S3, and the original stabilization matrix R1 reflecting the projective transformation relationship between F1 and S1. Then, the first image offset information is determined based on R1, R2, and R3.

[0108] In one implementation, step S204 specifically includes:

[0109] Step B1: Obtain the motion matrix between two adjacent video frames in the video frame and the plurality of reference video frames to obtain a plurality of motion matrices. Each motion matrix in the plurality of motion matrices is used to reflect the projective transformation relationship between the two adjacent video frames corresponding to the motion matrix.

[0110] Step B2: Determine the second image offset information based on the aforementioned motion matrices.

[0111] Specifically, since the second image offset information reflects the image offset information between two adjacent video frames in the video frame and the plurality of reference video frames, this embodiment needs to obtain the motion matrix between two adjacent video frames in the video frame and the plurality of reference video frames to determine the second image offset information, thus obtaining a plurality of motion matrices. Because the motion matrix between two adjacent video frames can reflect the projective transformation relationship of the time between the two adjacent video frames—that is, the first video frame in two adjacent video frames can be obtained by performing a projective transformation based on the motion matrix to obtain the second video frame in the two adjacent video frames—this motion matrix can also reflect the image offset information between two adjacent video frames to a certain extent. Therefore, this embodiment can determine the image offset information between two adjacent video frames in the video frame and the plurality of reference video frames based on the plurality of motion matrices, thus obtaining the second image offset information.

[0112] For example, such as Figure 2 As shown, assuming the video frame is F3, and F1 and F2 are reference video frames of F3, it is necessary to obtain the motion matrix M3 reflecting the projective transformation relationship between F3 and F2, and the motion matrix M2 reflecting the projective transformation relationship between F2 and F1, and then determine the second image offset information based on M2 and M3.

[0113] In one implementation, the offset prediction value includes: an angle offset prediction value; step S205 specifically includes:

[0114] Step C1: Obtain first angle offset information and second angle offset information, wherein the first angle offset information is the angle offset information in the first image offset information, and the second angle offset information is the angle offset information in the second image offset information;

[0115] Step C2: Input the first angle offset information and the second angle offset information into the angle offset prediction model to obtain the angle offset prediction value, wherein the angle offset prediction model is the model used in the prediction model to predict the angle offset value between the video frame and the original stabilized video frame.

[0116] Specifically, the offset prediction value in this embodiment includes an angle offset prediction value, which can be used to correct the deflection angle when a stabilization boundary violation occurs between the video frame and its corresponding original stabilized video frame. To obtain the angle offset prediction value, this embodiment needs to extract first angle offset information from the first image offset information. The first angle offset information reflects the deflection angle between the video frame and its corresponding original stabilized video frame, as well as the deflection angle between the plurality of reference video frames and their respective corresponding original stabilized video frames. Second angle offset information is then extracted from the second image offset information. The second angle offset information reflects the deflection angle between the video frame and two adjacent video frames among the plurality of reference video frames. Therefore, the first angle offset information and the second angle offset information are input into the angle offset prediction model. The angle offset prediction model can then calculate the deflection angle between the video frame and its corresponding original stabilized video frame when a stabilization boundary violation occurs, based on the first angle offset information and the second angle offset information, and further calculate the angle offset prediction value used to correct this deflection angle.

[0117] For example, in practical applications, to ensure that the stabilized image content and the original video content have no visual distortion differences, when using a stabilization matrix to crop the original video frames, the cropping method needs to be restricted to operations such as rotation, proportional scaling, and translation of a 2D plane in 3D space. Figure 2 As shown, F1 is the original video frame, and S1 is the stabilized video frame corresponding to F1. It is clear that S1 has a certain rotation angle compared to F1. When this rotation angle is too large, a portion of S1 will fall outside the range of F1, resulting in stabilization out-of-bounds. Therefore, the magnitude of this rotation angle is crucial to whether stabilization out-of-bounds occurs. Therefore, assuming the current video frame is the i-th frame, the latent variables of the angle prediction model are... Input the first angle offset information and the second angle offset information into the angle offset prediction model. θ Model θ That is, output the predicted angle offset value θ corresponding to the i-th video frame. i Among them, Model θ The working principle is as follows:

[0118]

[0119] in, This is the first angle offset information. The second angle offset information is given, and L is the model prediction window length, j∈[1,L].

[0120] In one implementation, step C2 specifically includes:

[0121] Step C201: Input the first angle offset information and the second angle offset information into the angle offset prediction model to obtain the initial angle offset prediction value;

[0122] Step C202: Determine the stabilization clipping range based on the original stabilization matrix, and determine the angle offset range based on the stabilization clipping range;

[0123] Step C203: Adjust the initial angle offset prediction value according to the angle offset range to obtain the angle offset prediction value.

[0124] Specifically, to ensure the accuracy of the predicted angle offset, such as Figure 10 As shown, after inputting the first angle offset information and the second angle offset information into the angle offset prediction model in this embodiment, it is also necessary to adjust the value range of the initial angle offset prediction value output by the angle offset prediction model. During adjustment, a stabilization clipping range needs to be obtained by referring to the original stabilization matrix corresponding to the video frame. The stabilization clipping range is used to limit the position of the target stabilized video frame corresponding to the video frame. The target stabilized video frame is the stabilized video frame corresponding to the video frame that will not experience stabilization out-of-bounds errors. Therefore, if the position of the stabilized video frame exceeds the stabilization clipping range, a stabilization out-of-bounds error may occur. In this embodiment, the angle offset range can be determined first based on the stabilization clipping range. The angle offset range is used to limit the maximum deflection angle between the video frame and its corresponding target stabilized video frame. Then, the initial angle offset prediction value is adjusted based on the angle offset range so that the initial angle offset prediction value does not exceed the angle offset range. After adjustment, the angle offset prediction value is obtained. It is understood that if the initial angle offset prediction value is within the angle offset range, there is no need to adjust the initial angle offset prediction value; if the initial angle offset prediction value is outside the angle offset range, the value of the initial angle offset prediction value needs to be adjusted to the maximum value corresponding to the angle offset range.

[0125] For example, inputting the first angle offset information and the second angle offset information into the angle offset prediction model yields an initial angle offset prediction value of 10°. Figure 3 As shown, the stabilization clipping range is determined to be the rectangle corresponding to F1 based on the original stabilization matrix corresponding to the video frame. The angle offset range is determined to be [0°, θ°] based on the stabilization clipping range. If 10° is less than or equal to θ°, there is no need to adjust the initial angle offset prediction value, and 10° is used as the angle offset prediction value. If 10° is greater than θ°, the initial angle offset prediction value needs to be adjusted to θ°, and θ° is used as the angle offset prediction value.

[0126] In one implementation, the offset prediction value includes: a horizontal displacement prediction value; step S205 specifically includes:

[0127] Step D1: Obtain first horizontal displacement information and second horizontal displacement information, wherein the first horizontal displacement information is horizontal displacement information generated by the first image offset information based on the angle offset prediction value, and the second horizontal displacement information is horizontal displacement information generated by the second image offset information based on the angle offset prediction value;

[0128] Step D2: Input the first horizontal displacement information and the second horizontal displacement information into the horizontal displacement prediction model to obtain the horizontal displacement prediction value. The horizontal displacement prediction model is the model used in the prediction model to predict the horizontal displacement value between the video frame and the original stabilized video frame.

[0129] Specifically, the offset prediction value in this embodiment also includes a horizontal displacement prediction value. This horizontal displacement prediction value can be used to correct the horizontal displacement amount when a stabilization boundary violation occurs between the video frame and its corresponding original stabilized video frame. To obtain the horizontal displacement prediction value, this embodiment needs to extract first horizontal displacement information from the first image offset information based on the angle offset prediction value. The first horizontal displacement information reflects the horizontal displacement amount between the video frame and its corresponding original stabilized video frame, as well as the horizontal displacement amount between the plurality of reference video frames and their respective corresponding original stabilized video frames. And, based on the angle offset prediction value, second horizontal displacement information is extracted from the second image offset information. The second horizontal displacement information reflects the horizontal displacement amount between the video frame and two adjacent video frames among the plurality of reference video frames. Therefore, by inputting the first and second horizontal displacement information into the horizontal displacement prediction model, the horizontal displacement prediction model can calculate the horizontal displacement amount between the video frame and its corresponding original stabilized video frame when a stabilization boundary violation occurs, and then calculate the horizontal displacement prediction value used to correct this horizontal displacement amount.

[0130] For example, such as Figure 2 As shown, F1 is the original video frame, and S1 is the stabilized video frame corresponding to F1. It is clear that S1, compared to F1, not only undergoes a certain rotation angle but also a certain horizontal displacement. When this rotation angle is determined, if the horizontal displacement is too large, a portion of S1 will fall outside of F1, resulting in stabilization exceeding the limit. Therefore, when the predicted horizontal displacement value is determined, the magnitude of the horizontal displacement between F1 and S1 is crucial in determining whether stabilization exceeding the limit will occur. Therefore, assuming the current video frame is the i-th frame, the latent variables of the horizontal displacement prediction model are... Input the first horizontal displacement information and the second horizontal displacement information into the horizontal displacement prediction model. dx That is, output the predicted horizontal displacement value d corresponding to the i-th video frame. xi Among them, Model dx The working principle is as follows:

[0131]

[0132] in, This is the first horizontal displacement information. The second horizontal displacement information is given, and L is the model prediction window length, j∈[1,L].

[0133] In one implementation, the step of inputting the first horizontal displacement information and the second horizontal displacement information into a horizontal displacement prediction model to obtain the predicted horizontal displacement value specifically includes:

[0134] Step D201: Input the first horizontal displacement information and the second horizontal displacement information into the horizontal displacement prediction model to obtain the initial horizontal displacement prediction value;

[0135] Step D202: Determine the horizontal displacement range based on the stabilized cutting range and the predicted angle offset value;

[0136] Step D203: Adjust the initial horizontal displacement prediction value according to the horizontal displacement range to obtain the horizontal displacement prediction value.

[0137] Specifically, to ensure the accuracy of the horizontal displacement prediction value, this embodiment requires adjusting the initial horizontal displacement prediction value output by the horizontal displacement prediction model after inputting the first and second horizontal displacement information. The adjustment requires reference to the stabilization clipping range and the angle offset prediction value. The stabilization clipping range is used to limit the position of the target stabilized video frame corresponding to the video frame, and the angle offset prediction value can be used to correct the deflection angle when the video frame and its corresponding original stabilized video frame exceed the stabilization limit. Therefore, during adjustment, the horizontal displacement range can be determined first based on the stabilization clipping range and the angle offset prediction value. The horizontal displacement range is used to limit the maximum horizontal displacement between the video frame and its corresponding target stabilized video frame. Then, the initial horizontal displacement prediction value is adjusted based on the horizontal displacement range, and the adjusted value is the obtained horizontal displacement prediction value. It is understood that if the initial horizontal displacement prediction value is within the horizontal displacement range, there is no need to adjust the initial horizontal displacement prediction value; if the initial horizontal displacement prediction value is outside the horizontal displacement range, the initial horizontal displacement prediction value needs to be adjusted to the maximum value corresponding to the horizontal displacement range.

[0138] For example, inputting the first horizontal displacement information and the second horizontal displacement information into the horizontal displacement prediction model yields an initial horizontal displacement prediction value of 10. Figure 3 As shown, the stabilization clipping range is determined to be the rectangle corresponding to F1 based on the original stabilization matrix corresponding to the video frame. The horizontal displacement range is determined to be [0, X] based on the stabilization clipping range and the angular offset prediction value θ°. If 10 is less than or equal to X, there is no need to adjust the initial horizontal displacement prediction value, and 10 is used as the horizontal displacement prediction value. If 10 is greater than X, the initial horizontal displacement prediction value needs to be adjusted to X, and X is used as the horizontal displacement prediction value.

[0139] In one implementation, the offset prediction value includes: a vertical displacement prediction value; step S205 includes:

[0140] Step E1: Obtain first vertical displacement information and second vertical displacement information, wherein the first vertical displacement information is vertical displacement information generated by the first image offset information based on the angle offset prediction value, and the second vertical displacement information is vertical displacement information generated by the second image offset information based on the angle offset prediction value.

[0141] Step E2: Input the first vertical displacement information and the second vertical displacement information into the vertical displacement prediction model to obtain the vertical displacement prediction value, wherein the vertical displacement prediction model is the model used in the prediction model to predict the vertical displacement value between the video frame and the original stabilized video frame.

[0142] Specifically, the offset prediction value in this embodiment also includes a vertical displacement prediction value. This vertical displacement prediction value can be used to correct the vertical displacement amount when a stabilization boundary violation occurs between the video frame and its corresponding original stabilized video frame. To obtain the vertical displacement prediction value, this embodiment needs to extract first vertical displacement information from the first image offset information based on the angle offset prediction value. The first vertical displacement information reflects the vertical displacement amount between the video frame and its corresponding original stabilized video frame, as well as the vertical displacement amount between the plurality of reference video frames and their respective corresponding original stabilized video frames. And, based on the angle offset prediction value, second vertical displacement information is extracted from the second image offset information. The second vertical displacement information reflects the vertical displacement amount between the video frame and two adjacent video frames among the plurality of reference video frames. Therefore, by inputting the first vertical displacement information and the second vertical displacement information into the vertical displacement prediction model, the vertical displacement prediction model can calculate the vertical displacement amount between the video frame and its corresponding original stabilized video frame when a stabilization boundary violation occurs, and then calculate the vertical displacement prediction value used to correct this vertical displacement amount.

[0143] For example, such as Figure 2 As shown, F1 is the original video frame, and S1 is the stabilized video frame corresponding to F1. It is clear that S1, compared to F1, not only undergoes a certain rotation angle but also a certain displacement in the vertical direction. When this rotation angle is determined, if the vertical displacement is too large, a portion of S1 will fall outside of F1, resulting in stabilization exceeding the limit. Therefore, when the predicted angle offset is determined, the magnitude of the vertical displacement between F1 and S1 is crucial in determining whether stabilization exceeding the limit will occur. Therefore, assuming the current video frame is the i-th frame, the latent variables of the vertical displacement prediction model are... The first vertical displacement information and the second vertical displacement information are input into the vertical displacement prediction model. dy That is, output the vertical displacement prediction value corresponding to the i-th video frame as d. yi Among them, Model dy The working principle is as follows:

[0144]

[0145] in, This is the first vertical displacement information. The second vertical displacement information is given by L, where L is the model prediction window length and j∈[1,L].

[0146] In one implementation, step E2 specifically includes:

[0147] Step E201: Input the first vertical displacement information and the second vertical displacement information into the vertical displacement prediction model to obtain the initial vertical displacement prediction value;

[0148] Step E202: Determine the vertical displacement range based on the stabilized cutting range and the predicted angle offset value;

[0149] Step E203: Adjust the initial vertical displacement prediction value according to the vertical displacement range to obtain the vertical displacement prediction value.

[0150] Specifically, to ensure the accuracy of the vertical displacement prediction value, this embodiment requires adjusting the initial vertical displacement prediction value output by the vertical displacement prediction model after inputting the first and second vertical displacement information. The adjustment requires reference to the stabilization clipping range and the angle offset prediction value. The stabilization clipping range defines the position of the target stabilized video frame corresponding to the video frame, and the angle offset prediction value corrects the deflection angle when the video frame crosses the stabilization boundary with its corresponding original stabilized video frame. Therefore, during adjustment, the vertical displacement range can be determined first based on the stabilization clipping range and the angle offset prediction value. The vertical displacement range defines the maximum vertical displacement between the video frame and its corresponding target stabilized video frame. Then, the initial vertical displacement prediction value is adjusted based on the vertical displacement range, resulting in the final vertical displacement prediction value. It is understood that if the initial vertical displacement prediction value is within the vertical displacement range, there is no need to adjust the initial vertical displacement prediction value; if the initial vertical displacement prediction value is outside the vertical displacement range, the initial vertical displacement prediction value needs to be adjusted to the maximum value corresponding to the vertical displacement range.

[0151] For example, inputting the first vertical displacement information and the second vertical displacement information into the vertical displacement prediction model yields an initial vertical displacement prediction value of 8. Figure 3As shown, the stabilization clipping range is determined to be the rectangle corresponding to F1 based on the original stabilization matrix corresponding to the video frame. The vertical displacement range is determined to be [0, Y] based on the stabilization clipping range and the angular offset prediction value θ°. If 8 is less than or equal to Y, there is no need to adjust the initial vertical displacement prediction value, and 8 is used as the vertical displacement prediction value. If 8 is greater than Y, the initial vertical displacement prediction value needs to be adjusted to Y, and Y is used as the vertical displacement prediction value.

[0152] In one implementation, the step of stabilizing the video frame based on the offset prediction value to obtain a stabilized video frame specifically includes:

[0153] Step S206: Generate a correction matrix based on the predicted offset values;

[0154] Step S207: Correct the first original stabilization matrix according to the correction matrix to obtain the target stabilization matrix corresponding to the video frame;

[0155] Step S208: Stabilize the video frame according to the target stabilization matrix to obtain the stabilized video frame.

[0156] Specifically, since the offset prediction value can be used to correct the offset when the original stabilized video frame corresponding to the video frame exceeds the stabilization limit, this embodiment can generate a correction matrix based on the offset prediction value, and adjust the original stabilization matrix corresponding to the video frame, i.e., the first original stabilization matrix, based on the correction matrix to obtain the target stabilization matrix. Because this application corrects the offset between the video frame and its corresponding original stabilized video frame based on the offset prediction value, i.e., corrects the projective relationship between the video frame and its corresponding original stabilized video frame, the stabilized video frame obtained by stabilizing the video frame based on the obtained target stabilization matrix will not experience stabilization exceeding the limit.

[0157] For example, assuming that a correction matrix Ni is generated based on the offset prediction value, and the first original stabilization matrix is ​​R1, then the target stabilization matrix R1' = R1 * Ni.

[0158] In one implementation, the offset prediction value includes the angular offset prediction value, the horizontal displacement prediction value, and the vertical displacement prediction value, and the correction matrix is ​​generated based on the angular offset prediction value, the horizontal displacement prediction value, and the vertical displacement prediction value.

[0159] Specifically, since the predicted angle offset, the predicted horizontal displacement, and the predicted vertical displacement can be used to correct the rotation angle, horizontal displacement, and vertical displacement between the video frame and its corresponding stabilized video frame, respectively, in order to ensure that the target stabilized video frame corresponding to the video frame does not experience stabilization out of bounds, this embodiment needs to combine the three predicted values—the predicted angle offset, the predicted horizontal displacement, and the predicted vertical displacement—to obtain the correction matrix.

[0160] For example, the goal of this embodiment is to correct the first original stabilization matrix corresponding to the video frame, obtaining a target stabilization matrix with less distortion and better video stabilization effect. First, this embodiment needs to explain how to determine the video stabilization effect of the stabilization matrix:

[0161] In existing technology, two adjacent video frames F i-1 ,F i Motion matrix M i Smooth motion matrix K i and the stabilization matrix R i The derivation relationship between them can be obtained as follows:

[0162]

[0163]

[0164] Let P denote the set of sampling points in the image (P can be a matrix containing all points in the image). Clearly, as shown in the following formula, for the same set of sampling points, the smaller the difference in motion speed or acceleration between adjacent video frames, the better the stabilization effect.

[0165]

[0166] In this embodiment, the predicted values ​​of angular offset, horizontal displacement, and vertical displacement are used to generate a correction matrix. The first original stabilization matrix is ​​then corrected based on this correction matrix to obtain the target stabilization matrix. The correction matrix N is then... i The first original stabilization matrix R i and the target stabilization matrix The following relationship exists between them:

[0167]

[0168] This leads to the following formula:

[0169]

[0170] To achieve the best video stabilization effect, the goal of this embodiment is to determine the correction matrix N. iAnd this minimizes the following stabilization loss function:

[0171]

[0172] In one implementation, the correction matrix N i It is simplified to the following matrix form:

[0173]

[0174] in These are the predicted values ​​for angular offset, horizontal displacement, and vertical displacement, respectively. The correction matrix N is now determined. i Subsequently, based on the correction matrix N i The first original stabilization matrix is ​​corrected, and the target stabilization matrix is ​​obtained after the correction is completed.

[0175] In one implementation, the prediction model is a trained model, and the training process of the prediction model is as follows:

[0176] Step F1: Input the third image offset information and the fourth image offset information in the training data into the original prediction model, and generate the offset prediction value corresponding to the third image offset information and the fourth image offset information through the original prediction model. The training data includes multiple training information groups, each training information group includes the third image offset information, the fourth image offset information and the standard offset value, and the standard offset value is the offset value corresponding to the third image offset information and the fourth image offset information.

[0177] Step F2: Adjust the model parameters of the original prediction model according to the standard offset values ​​corresponding to the third image offset information and the fourth image offset information, and the offset prediction values ​​corresponding to the third image offset information and the fourth image offset information. Continue to execute the step of inputting the third image offset information and the fourth image offset information from the training data into the original prediction model until the preset training conditions are met to obtain the prediction model.

[0178] Specifically, the training data in this embodiment includes multiple sets of training information groups. Each set of training information groups includes third image offset information, fourth image offset information, and a standard offset value. The standard offset value is equivalent to the true label corresponding to the third and fourth image offset information. Then, the third and fourth image offset information are input into an untrained original prediction model to obtain the offset prediction value output by the original prediction model. This offset prediction value is compared with the standard offset value to obtain the difference between them. The model parameters of the original prediction model are then adjusted based on the standard offset value to reduce the difference between the offset prediction value output by the original prediction model and the standard offset value. This process is repeated until the difference between the offset prediction value output by the original prediction model and the standard offset value meets a preset training condition, thus obtaining the trained prediction model.

[0179] In one implementation, such as Figure 11 As shown, to improve training efficiency, this embodiment can also use multiple video samples to form a training package to train the original prediction model. Specifically, for a single video sample, a video frame idx is randomly selected, and prediction is performed from frame 1 to frame idx-1. Each prediction yields a set of model latent variables, namely the latent variables corresponding to the angle offset prediction model, the horizontal displacement prediction model, and the vertical displacement prediction model, respectively. This results in L sets of model latent variables and the angle offset prediction value corresponding to the last L-1 frames. Starting from frame idx, the prediction window size is determined by extracting the first angle offset information from the original stabilization matrix R corresponding to each frame from the past L frames, as well as the second angle offset information, second horizontal displacement information, and second vertical displacement information from the motion matrix M. All the first angle offset information is then merged into a single frame. All second-angle offset information is merged into All second-level displacement information is merged into All second vertical displacement information is merged into Then, the latent variables of the angle offset prediction model, horizontal displacement prediction model, and vertical displacement prediction model corresponding to multiple video samples are combined according to the type of latent variables to obtain the angle offset latent variable matrix {H}. θ}、Horizontal displacement latent variable matrix {H dx}, Vertical displacement latent variable matrix {H dy}. The multiple video samples are respectively... Combined into the first angle offset information parameter package The multiple video samples are respectively corresponding to Combined into a second angle offset information parameter package The multiple video samples are respectively corresponding to Combined into a second horizontal displacement parameter package The multiple video samples are respectively corresponding to Combined into a second vertical displacement parameter package The multiple video samples are respectively corresponding to Combined into angular offset prediction parameter package Θ pre Then {H θ The original angle offset prediction model is input to obtain the initial angle offset prediction value. The target original stabilization matrix R, generated from all the original stabilization matrices of each video sample, is used to determine the stabilization clipping range. The initial angle offset prediction value is adjusted according to the clipping range to obtain the angle offset prediction value Θ for the current frame. Θ is then... pre Combined with Θ, an updated angular offset prediction parameter package Θ is generated. merge According to Θ merge The first horizontal displacement information is extracted from the original stabilization matrix R of the target. First vertical displacement information Then and {H dx The input horizontal displacement prediction model yields initial horizontal displacement prediction values. A horizontal displacement range is generated using the target original stabilization matrix R and the current frame's angle offset prediction value Θ. The initial horizontal displacement prediction values ​​are then adjusted based on this range to obtain the current frame's horizontal displacement prediction value DX. Similarly, the input horizontal displacement prediction model is used to... and {H dy The input vertical displacement prediction model yields initial vertical displacement prediction values. A vertical displacement range is generated using the target original stabilization matrix R and the current frame's angle offset prediction value Θ. The initial vertical displacement prediction value is then adjusted based on this range to obtain the current frame's vertical displacement prediction value DY. A correction matrix N is generated by combining the current frame's angle offset prediction value Θ, the current frame's horizontal displacement prediction value DX, and the current frame's horizontal displacement prediction value DY. Standard angle offset prediction values, standard horizontal displacement prediction values, and standard vertical displacement prediction values ​​for the current frame are obtained. A standard correction matrix N' is generated based on these values. The correction matrix N and the standard correction matrix N' are substituted into the loss function, and the backpropagation gradient is calculated to update the model parameters of the angle offset prediction model, the horizontal displacement prediction model, and the vertical displacement prediction model, respectively.

[0180] The formula for the loss function is shown below:

[0181]

[0182] Where P is the set matrix consisting of all points of all video frames in all video samples; N i K is the correction matrix corresponding to the i-th video frame; i Let be the motion matrix corresponding to the i-th video frame.

[0183] In one implementation, in order to further improve the prediction performance of the prediction model, such as... Figure 4 As shown, this embodiment can use a feature-weighted matching model to establish a prediction model. The conceptual formula of the feature-weighted matching model is as follows:

[0184]

[0185] Where, x i Indicates the query condition characteristics, x j This indicates the feature to be searched. q i Indicates the query condition, k j Indicates a matching item, v j This indicates the content information returned by the retrieval. When q i With k j The higher the degree of matching, the higher the corresponding v j The greater the contribution to the final model output, the better. In this embodiment, the offset feature data extracted from the original motion matrix will be used as x. i The x represents the query condition feature, where the offset feature data extracted from the original stabilization matrix is ​​used as x. j This indicates that the features to be retrieved are being modeled. It is understandable that, for example... Figure 2 As shown, since one original motion matrix is ​​associated with only two original stabilization matrices, this embodiment will use L-1 matching models to form a network in the retrieval window, where L refers to the number of video frames to be processed. In one implementation, the feature-weighted matching model will also use the linear transformation result of the original motion matrix to predict the result.

[0186] To verify the technical effect of the present invention, the inventors used 50 1080p videos of movement, each longer than 3 minutes, for training, and used 30 videos for verification testing. The statistical results are shown in the scatter plot. Figure 5 As shown. The results indicate that the stabilization effect of this invention improves by an average of about 2 pixels per frame. Figure 6 As shown in Figure 7, Figure 6In the case of two adjacent video frames (frames 7 and 7), the target stabilization matrix is ​​determined using the method of this invention. The four borders of the cropped stabilized video frame are basically within the four borders of the original video frame. However, based on the original stabilization matrix determined by existing technology, the four borders of the cropped stabilized video frame overlap with or exceed the boundaries of the original video frame. Therefore, this invention can effectively reduce the problem of stabilization boundary overflow that easily occurs when performing video stabilization on the stabilization matrix obtained by motion inference from motion estimation and smooth motion estimation of the original video frame in existing technology.

[0187] Exemplary device

[0188] Based on the above embodiments, the present invention also provides a video stabilization device, such as... Figure 8 As shown, the device includes:

[0189] Acquisition module 01 is used to acquire video frames to be processed;

[0190] The stabilization module 02 determines the stabilized video frame for each video frame that is not the first frame in the video frame to be processed; wherein, for each video frame, the offset prediction value corresponding to the video frame is determined; and the video frame is stabilized according to the offset prediction value to obtain the stabilized video frame of the video frame.

[0191] Output module 03 determines the stable video corresponding to the video frame to be processed based on the first video frame and the stabilized video frames of each video frame other than the first video frame in the video frame to be processed.

[0192] Based on the above embodiments, the present invention also provides a terminal device, the principle block diagram of which can be as follows: Figure 9 As shown, the terminal device includes a processor, memory, network interface, and display screen connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a video stabilization method. The display screen can be a liquid crystal display (LCD) or an e-ink display.

[0193] Those skilled in the art will understand that Figure 9 The block diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal device to which the present invention is applied. The specific terminal device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0194] In one embodiment, a terminal device is provided, comprising a memory, a processor, and a video stabilization program stored in the memory and executable on the processor. When the processor executes the video stabilization program, it implements the following operation instructions:

[0195] Obtain the video frames to be processed;

[0196] Determine the stabilized video frame for each video frame that is not the first frame in the video frames to be processed; wherein, for each video frame, determine the offset prediction value corresponding to the video frame; and perform stabilization processing on the video frame according to the offset prediction value to obtain the stabilized video frame of the video frame.

[0197] The stabilized video corresponding to the video frame to be processed is determined based on the first video frame and the stabilized video frames of each non-first video frame in the video frame to be processed.

[0198] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0199] In summary, this invention provides a video stabilization method, apparatus, terminal, and storage medium. The method involves acquiring a video frame to be processed; determining a stabilized video frame for each non-first frame of the video frame to be processed; wherein, for each video frame, an offset prediction value is determined; the video frame is stabilized based on the offset prediction value to obtain a stabilized video frame; and the stabilized video corresponding to the video frame to be processed is determined based on the first video frame and the stabilized video frames for each non-first frame of the video frame to be processed. Since the offset prediction value can be used to correct the offset when the original stabilized video frame corresponding to the video frame exceeds the stabilization limit, this invention stabilizes the video frame based on the offset prediction value, and the resulting stabilized video frame will not exceed the stabilization limit, thus solving the problem that existing stabilization matrices easily cause stabilization limit exceedances after stabilizing video frames.

[0200] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A video stabilization method, characterized in that, The method includes: Obtain the video frames to be processed; The process involves determining a stabilized video frame for each video frame (excluding the first frame) in the video frame to be processed; wherein, for each video frame, determining the offset prediction value corresponding to the video frame includes: acquiring a plurality of reference video frames corresponding to the video frame; wherein the plurality of reference video frames are a plurality of consecutive video frames whose playback time is before the video frame and adjacent to the video frame; acquiring the original stabilized video frame corresponding to the video frame and the original stabilized video frames corresponding to the plurality of reference video frames respectively; determining first image offset information based on the video frame, the original stabilized video frame corresponding to the video frame, the plurality of reference video frames, and the original stabilized video frames corresponding to the plurality of reference video frames respectively; determining second image offset information based on the video frame and the plurality of reference video frames; and inputting the first image offset information and the second image offset information into a prediction model to obtain the offset prediction value. Based on the offset prediction value, the video frame is stabilized to obtain the stabilized video frame. The stabilized video corresponding to the video frame to be processed is determined based on the first video frame and the stabilized video frames of each non-first video frame in the video frame to be processed.

2. The video stabilization method according to claim 1, characterized in that, Based on the video frame, the corresponding original stabilized video frame, the plurality of reference video frames, and the original stabilized video frames corresponding to the plurality of reference video frames, the first image offset information is determined, including: A first original stabilization matrix is ​​determined based on the video frame and the original stabilized video frame corresponding to the video frame. The first original stabilization matrix is ​​used to reflect the projective transformation relationship between the video frame and the original stabilized video frame corresponding to the video frame. Based on the plurality of reference video frames and the original stabilized video frames corresponding to the plurality of reference video frames, a second original stabilization matrix corresponding to the plurality of reference video frames is determined; wherein, for each video frame in the plurality of reference video frames, a second original stabilization matrix corresponding to the video frame is determined based on the video frame and the original stabilized video frame corresponding to the video frame, and the second original stabilization matrix is ​​used to reflect the projective transformation relationship between the video frame and the original stabilized video frame corresponding to the video frame; The first image offset information is determined based on the first original stabilization matrix and the second original stabilization matrices corresponding to the plurality of reference video frames.

3. The video stabilization method according to claim 1, characterized in that, The step of determining the second image offset information based on the video frame and the plurality of reference video frames includes: Obtain the motion matrix between two adjacent video frames in the video frame and the plurality of reference video frames to obtain a plurality of motion matrices. Each of the plurality of motion matrices is used to reflect the projective transformation relationship between the two adjacent video frames corresponding to the motion matrix. The second image offset information is determined based on the aforementioned motion matrices.

4. The video stabilization method according to claim 3, characterized in that, The offset prediction value includes: an angle offset prediction value; the step of inputting the first image offset information and the second image offset information into the prediction model to obtain the offset prediction value includes: Obtain first angle offset information and second angle offset information, wherein the first angle offset information is the angle offset information in the first image offset information, and the second angle offset information is the angle offset information in the second image offset information; The first angle offset information and the second angle offset information are input into the angle offset prediction model to obtain the angle offset prediction value, wherein the angle offset prediction model is the model used in the prediction model to predict the angle offset value between the video frame and the original stabilized video frame.

5. The video stabilization method according to claim 4, characterized in that, The step of inputting the first angle offset information and the second angle offset information into the angle offset prediction model to obtain the angle offset prediction value includes: Input the first angle offset information and the second angle offset information into the angle offset prediction model to obtain the initial angle offset prediction value; The stabilization clipping range is obtained based on the original stabilization matrix corresponding to the video frame. The original stabilization matrix corresponding to the video frame is used to reflect the projective transformation relationship between the video frame and the original stabilized video frame corresponding to the video frame. The stabilization clipping range is used to limit the position of the target stabilized video frame corresponding to the video frame. The target stabilized video frame is the stabilized video frame corresponding to the video frame that will not have stabilization out-of-bounds. The angle offset range is determined based on the stabilization clipping range. The initial angle offset prediction value is adjusted according to the angle offset range to obtain the angle offset prediction value.

6. The video stabilization method according to claim 5, characterized in that, The offset prediction value includes: a horizontal displacement prediction value; the step of inputting the first image offset information and the second image offset information into the prediction model to obtain the offset prediction value includes: Obtain first horizontal displacement information and second horizontal displacement information, wherein the first horizontal displacement information is horizontal displacement information generated by the first image offset information based on the angle offset prediction value, and the second horizontal displacement information is horizontal displacement information generated by the second image offset information based on the angle offset prediction value; The first horizontal displacement information and the second horizontal displacement information are input into the horizontal displacement prediction model to obtain the horizontal displacement prediction value. The horizontal displacement prediction model is the model used in the prediction model to predict the horizontal displacement value between the video frame and the original stabilized video frame.

7. The video stabilization method according to claim 6, characterized in that, The step of inputting the first horizontal displacement information and the second horizontal displacement information into the horizontal displacement prediction model to obtain the predicted horizontal displacement value includes: Input the first horizontal displacement information and the second horizontal displacement information into the horizontal displacement prediction model to obtain the initial horizontal displacement prediction value; The horizontal displacement range is determined based on the stabilized cutting range and the predicted angle offset value. The initial horizontal displacement prediction value is adjusted according to the horizontal displacement range to obtain the horizontal displacement prediction value.

8. The video stabilization method according to claim 5, characterized in that, The offset prediction value includes: a vertical displacement prediction value; the step of inputting the first image offset information and the second image offset information into the prediction model to obtain the offset prediction value includes: Obtain first vertical displacement information and second vertical displacement information, wherein the first vertical displacement information is vertical displacement information generated by the first image offset information based on the angle offset prediction value, and the second vertical displacement information is vertical displacement information generated by the second image offset information based on the angle offset prediction value; The first vertical displacement information and the second vertical displacement information are input into the vertical displacement prediction model to obtain the vertical displacement prediction value, wherein the vertical displacement prediction model is the model used in the prediction model to predict the vertical displacement value between the video frame and the original stabilized video frame.

9. The video stabilization method according to claim 8, characterized in that, The step of inputting the first vertical displacement information and the second vertical displacement information into the vertical displacement prediction model to obtain the predicted vertical displacement value includes: Input the first vertical displacement information and the second vertical displacement information into the vertical displacement prediction model to obtain the initial vertical displacement prediction value; The vertical displacement range is determined based on the stabilized cutting range and the predicted angle offset value. The initial vertical displacement prediction value is adjusted according to the vertical displacement range to obtain the vertical displacement prediction value.

10. The video stabilization method according to claim 2, characterized in that, The step of stabilizing the video frame based on the offset prediction value to obtain a stabilized video frame includes: A correction matrix is ​​generated based on the predicted offset values; The first original stabilization matrix is ​​corrected according to the correction matrix to obtain the target stabilization matrix corresponding to the video frame; The video frame is stabilized according to the target stabilization matrix to obtain the stabilized video frame.

11. The video stabilization method according to claim 1, characterized in that, The prediction model is a trained model, and the training process of the prediction model is as follows: The third image offset information and the fourth image offset information in the training data are input into the original prediction model, and the original prediction model generates the offset prediction value corresponding to the third image offset information and the fourth image offset information. The training data includes multiple training information groups, each training information group includes the third image offset information, the fourth image offset information and the standard offset value, and the standard offset value is the offset value corresponding to the third image offset information and the fourth image offset information. Based on the standard offset values ​​corresponding to the third and fourth image offset information and the offset prediction values ​​corresponding to the third and fourth image offset information, the model parameters of the original prediction model are adjusted, and the step of inputting the third and fourth image offset information from the training data into the original prediction model is continued until the preset training conditions are met to obtain the prediction model.

12. A video stabilization device, characterized in that, The device includes: The acquisition module is used to acquire the video frames to be processed. A stabilization module determines a stabilized video frame for each video frame other than the first frame in the video frame to be processed. For each video frame, it determines the corresponding offset prediction value, including: acquiring several reference video frames corresponding to the video frame; wherein the several reference video frames are several consecutive video frames whose playback time is before the video frame and adjacent to the video frame; acquiring the original stabilized video frame corresponding to the video frame and the original stabilized video frames corresponding to the several reference video frames respectively; determining first image offset information based on the video frame, the original stabilized video frame corresponding to the video frame, the several reference video frames, and the original stabilized video frames corresponding to the several reference video frames respectively; determining second image offset information based on the video frame and the several reference video frames; inputting the first image offset information and the second image offset information into a prediction model to obtain the offset prediction value; and stabilizing the video frame based on the offset prediction value to obtain the stabilized video frame of the video frame. The output module determines the stable video corresponding to the video frame to be processed based on the first video frame and the stabilized video frames of each video frame other than the first video frame in the video frame to be processed.

13. A terminal, characterized in that, The terminal includes a memory and one or more processors; the memory stores one or more programs; the programs contain instructions for executing the video stabilization method as described in any one of claims 1-11; the processor is used to execute the programs.

14. A computer-readable storage medium storing a plurality of instructions, characterized in that, The processor loads and executes the instructions to implement the steps of the video stabilization method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Video stability augmentation method and device, computer equipment and storage medium

    CN110740247A