Video processing method, video processing device and related product
By acquiring and smoothing the transformation matrix of video frames, real-time anti-shake of the video picture is achieved, which solves the problem of long-term deep learning models in the prior art, and realizes real-time stable processing of the video picture without relying on complex calculations.
Patent Information
- Application Number
- CN202510250745.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The prior art is difficult to prevent the video image from being anti-shake in real time, and the deep learning model takes a long time to predict and process video frames, making it difficult to achieve real-time anti-shake.
By obtaining the original transformation matrix of the Nth frame in the video and the smoothing transformation matrix of the N-1th frame in the video, and smoothing processing is performed based on the weight to obtain the smoothing transformation matrix of the Nth frame. This matrix is used to transform and add coordinates to the target display elements in the video screen to achieve anti-shake on the video screen.
Real-time anti-shake of the video picture is achieved, time consumption of using deep learning models is avoided, and video frames can be processed in real time without relying on complex calculations.
Smart Images

Figure CN120111364A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of data processing technology, and in particular to a video processing method, a video processing device, and related products. Background Art
[0002] As the video shooting capabilities of electronic devices are enhanced, more and more users begin to use electronic devices to shoot videos. During the video shooting process, especially when users hold electronic devices to take mobile selfies, the electronic devices often shake or jitter to a certain extent, resulting in unstable and jittery video images.
[0003] Currently, the captured video frame sequence is often input into a deep learning model, and the video frames predicted by the deep learning model are used to replace the jittery video frames in the video frame sequence to achieve video image stabilization. However, a large number of data sets are required for deep learning models, which takes a long time to train the deep learning models. In addition, the deep learning models also need to spend a lot of time to perform complex calculations during the inference process in order to predict the output video frames, making it difficult to stabilize video images in real time. Summary of the invention
[0004] The embodiments of the present application provide a video processing method, a video processing device and related products, which can prevent video images from shaking in real time.
[0005] The present application provides a video processing method, including:
[0006] Obtaining an original transformation matrix of the Nth frame in the video and a smoothed transformation matrix of the N-1th frame; wherein N is greater than or equal to 2; the original transformation matrix of the Nth frame is a coordinate mapping relationship between a face position in a video screen to be displayed of the Nth frame and a face model; the face position in the video screen to be displayed is described by a video screen coordinate system, and the coordinates of the face model are described by a calibrated coordinate system, and the calibrated coordinate system is also used to describe the coordinates of candidate display elements associated with the face model;
[0007] Determine a first weight corresponding to the original transformation matrix of the Nth frame and a second weight corresponding to the smoothed transformation matrix of the N-1th frame;
[0008] Based on the first weight and the second weight, smoothing the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame to obtain the smoothed transformation matrix of the Nth frame;
[0009] Using the smoothed transformation matrix of the Nth frame, coordinate transformation is performed on the coordinates of the target display element in the to-be-displayed video screen of the Nth frame in the calibrated coordinate system to obtain the target coordinates of the target display element in the video screen coordinate system; the target display element belongs to the candidate display element;
[0010] The target display element is added to the video picture to be displayed according to the target coordinates to obtain the video picture displayed in the Nth frame.
[0011] Furthermore, the method further comprises:
[0012] Determine a coordinate offset matrix of the Nth frame based on the smoothed transformation matrix of the Nth frame and the original transformation matrix of the Nth frame;
[0013] Determine, based on the coordinate offset matrix and the first weight, a third weight corresponding to the original transformation matrix of the N+1th frame and a fourth weight corresponding to the smoothed transformation matrix of the Nth frame;
[0014] Based on the third weight and the fourth weight, the original transformation matrix of the N+1th frame and the smoothed transformation matrix of the Nth frame are smoothed to obtain the smoothed transformation matrix of the N+1th frame.
[0015] Further, the determining, based on the coordinate offset matrix and the first weight, a third weight corresponding to the original transformation matrix of the N+1th frame and a fourth weight corresponding to the smoothed transformation matrix of the Nth frame includes:
[0016] Obtaining the coordinate offset in the coordinate offset matrix of the Nth frame;
[0017] Within a preset weight range, based on the coordinate offset and the first weight, obtaining a third weight corresponding to the original transformation matrix of the N+1th frame;
[0018] Based on the third weight corresponding to the original transformation matrix of the N+1th frame, a fourth weight corresponding to the smoothed transformation matrix of the Nth frame is determined.
[0019] Further, when the Nth frame is the second frame, the N-1th frame is the first frame;
[0020] Obtaining the smoothed transformation matrix of the N-1th frame in the video includes:
[0021] Obtaining the original transformation matrix of the first frame in the video;
[0022] A smoothed transformation matrix of the first frame is obtained based on an original transformation matrix of the first frame and a preset weight corresponding to the original transformation matrix of the first frame.
[0023] Furthermore, obtaining the original transformation matrix of the Nth frame in the video includes:
[0024] Obtaining the first vertex coordinates of the face position in the video picture to be displayed of the Nth frame in the video picture coordinate system, and the second vertex coordinates of the face model in the calibration coordinate system;
[0025] Based on the mapping relationship between the first vertex coordinates and the second vertex coordinates, the original transformation matrix of the Nth frame is obtained.
[0026] Further, the determining of the first weight corresponding to the original transformation matrix of the Nth frame and the second weight corresponding to the smoothed transformation matrix of the N-1th frame includes:
[0027] Obtaining a first weight corresponding to the original transformation matrix of the Nth frame;
[0028] The second weight corresponding to the smoothed transformation matrix of the N-1th frame is set to be inversely proportional to the first weight.
[0029] Furthermore, the first weight is: a n , the second weight is: 1-a n ;
[0030] The step of performing smoothing processing on the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame based on the first weight and the second weight to obtain the smoothed transformation matrix of the Nth frame includes:
[0031] Use the formula: Get the smoothed transformation matrix of the Nth frame Among them, M n is the original transformation matrix of the Nth frame, is the smoothed transformation matrix of the N-1th frame.
[0032] Further, the target display element in the video picture to be displayed in the Nth frame includes: a face special effect map;
[0033] The step of using the smoothed transformation matrix of the Nth frame to transform the coordinates of the target display element in the to-be-displayed video screen of the Nth frame in the calibrated coordinate system to obtain the target coordinates of the target display element in the video screen coordinate system includes:
[0034] The smoothed transformation matrix of the Nth frame is multiplied by the vertex coordinates of the face special effect map corresponding to the calibrated coordinate system to obtain the target coordinates of the face special effect map in the video screen coordinate system.
[0035] The embodiment of the present application also provides a video processing device, including:
[0036] An acquisition unit, used to acquire an original transformation matrix of the Nth frame in the video and a smoothed transformation matrix of the N-1th frame; wherein N is greater than or equal to 2; the original transformation matrix of the Nth frame is a coordinate mapping relationship between a face position in a video screen to be displayed of the Nth frame and a face model; the face position in the video screen to be displayed is described by a video screen coordinate system, and the coordinates of the face model are described by a calibrated coordinate system, and the calibrated coordinate system is also used to describe the coordinates of a candidate display element associated with the face model;
[0037] A determination unit, configured to determine a first weight corresponding to the original transformation matrix of the Nth frame and a second weight corresponding to the smoothed transformation matrix of the N-1th frame;
[0038] a smoothing unit, configured to perform smoothing processing on the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame based on the first weight and the second weight, so as to obtain the smoothed transformation matrix of the Nth frame;
[0039] a transformation unit, configured to use the smoothed transformation matrix of the Nth frame to transform the coordinates of a target display element in the to-be-displayed video screen of the Nth frame in the calibrated coordinate system, to obtain target coordinates of the target display element in the video screen coordinate system; the target display element belongs to the candidate display element;
[0040] An execution unit is used to add the target display element to the video picture to be displayed according to the target coordinates to obtain the video picture displayed in the Nth frame.
[0041] The embodiment of the present application also provides a video processing device, including:
[0042] Processor, memory, and input and output interfaces;
[0043] The memory is a short-term storage memory or a persistent storage memory;
[0044] The processor is configured to communicate with the memory and execute instructions in the memory to perform the above method.
[0045] An embodiment of the present application also provides a computer-readable storage medium, comprising instructions, which, when executed on a computer, enable the computer to execute the method described above.
[0046] The embodiment of the present application also provides a computer program product including instructions or a computer program. When the computer program product is run on a computer, the computer executes the method described above.
[0047] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0048] In an embodiment of the present application, based on the first weight corresponding to the original transformation matrix of the Nth frame and the second weight corresponding to the smoothed transformation matrix of the N-1th frame, the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame are smoothed to obtain the smoothed transformation matrix of the Nth frame; the smoothed transformation matrix of the Nth frame is used to transform the coordinates of the target display element in the video picture to be displayed of the Nth frame in the calibrated coordinate system, and the target display element is added to the video picture to be displayed to achieve anti-shake for the video picture of the Nth frame; only the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame need to be used to transform the coordinates of the target display element in the video picture to be displayed of the Nth frame to achieve anti-shake for the video picture of the Nth frame, and the video picture can be anti-shake in real time without using a deep learning model. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0050] Figure 1 A schematic diagram of a communication architecture for video processing disclosed in an embodiment of the present application;
[0051] Figure 2 A flowchart of video processing disclosed in an embodiment of the present application;
[0052] Figure 3 A flowchart of obtaining a smoothed transformation matrix of the N+1th frame disclosed in an embodiment of the present application;
[0053] Figure 4 A flowchart of obtaining a smoothed transformation matrix of a first frame disclosed in an embodiment of the present application;
[0054] Figure 5 A flowchart of another video processing disclosed in an embodiment of the present application;
[0055] Figure 6 A schematic diagram of a video processing algorithm disclosed in an embodiment of the present application;
[0056] Figure 7 A schematic diagram of a video processing device disclosed in an embodiment of the present application;
[0057] Figure 8A schematic diagram of another video processing device disclosed in an embodiment of the present application. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0059] In the description of the embodiments of the present application, it should be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc. indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the embodiments of the present application and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on the embodiments of the present application.
[0060] In the description of the embodiments of the present application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific circumstances.
[0061] Existing video processing architectures such as Figure 1 As shown, it includes: a video processing device 101 and a video player 102, wherein the video player 102 can be an electronic device that can play videos, such as a mobile phone, a computer, etc. The video player 102 has a corresponding video file, which can be a pre-stored video file, or a video file that is recorded and played in real time (such as a video file when making a video call or live video broadcasting through a mobile phone). The video processing device 101 can be connected to one or more video players 102 by wire or wirelessly to stabilize the video image of the video file on the video player 102.
[0062] The video processing device often inputs the video frame sequence of the video file on the video player into the deep learning model, and replaces the jittery video frames in the video frame sequence with the video frames predicted by the deep learning model to achieve video picture anti-shake. However, a large number of data sets are required for the deep learning model, and it takes a long time to train the deep learning model. In addition, the deep learning model also needs to spend a lot of time to perform complex calculations during the inference process in order to predict the output video frames, which makes it difficult to anti-shake the video picture in real time. Therefore, an embodiment of the present application provides a video processing method that can anti-shake the video picture in real time, such as Figure 2 As shown, the specific steps include:
[0063] 201. Obtain an original transformation matrix of the Nth frame in the video and a smoothed transformation matrix of the N-1th frame.
[0064] In an embodiment of the present application, the video processing device can obtain the original transformation matrix of the Nth frame in the video, and the smoothed transformation matrix of the N-1th frame. Wherein, the video is a video file with multiple consecutive time-series frames, which can be a video file shot in real time, or a video file obtained from the network, and the specific details are not limited here. Wherein, N is greater than or equal to 2, that is, the Nth frame is the second video frame in the video and the video frame after the second video frame, and the N-1th frame is the previous frame of the Nth frame. If the Nth frame can be the third frame (current frame) to be played in the video, then the N-1th frame is the second frame.
[0065] The original transformation matrix of the Nth frame is the coordinate mapping relationship between the face position in the video screen to be displayed in the Nth frame and the face model, wherein the face model has basic components of the face, such as eyes, nose, mouth and head, etc.; the face position in the video screen to be displayed is described by the video screen coordinate system, and the coordinates of the face model are described by the calibration coordinate system, which is the coordinate system used as a reference in the video processing device; that is, the original transformation matrix of the Nth frame is used to transform the coordinates of the face position in the video screen to be displayed in the Nth frame and locate them in the calibration coordinate system. The calibration coordinate system is also used to describe the coordinates of the candidate display element associated with the face model, and the candidate display element is an element that may exist on the face in addition to the components contained in the face model, such as a hat, glasses, beard, etc.; it can be understood that the relative position between the candidate display element and the face model is determined, for example, if the candidate display element is a hat, then its position is on the key point of the head of the face model, and if the candidate display element is a beard, then its position is on the key point of the mouth of the face model, etc.
[0066] That is, before the video player displays the video picture of the Nth frame, the video processing device can obtain the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame, so as to perform adaptive smoothing on the original transformation matrix of the Nth frame to obtain the smoothed transformation matrix of the Nth frame, thereby realizing anti-shake of the video picture of the Nth frame; wherein, the smoothed transformation matrix of the N-1th frame is obtained by adaptive smoothing on the original transformation matrix of the N-1th frame.
[0067] The original transformation matrix of the Nth frame in the video can be obtained specifically as follows: obtaining the first vertex coordinates of the face position in the video screen to be displayed in the Nth frame in the video screen coordinate system, and the second vertex coordinates of the face model in the calibration coordinate system; based on the mapping relationship between the first vertex coordinates and the second vertex coordinates, the original transformation matrix of the Nth frame is obtained. For example, the face detection model can be used to detect the key points of the face position in the video screen to be displayed in the Nth frame, obtain the first vertex coordinates corresponding to the key points of the face position in the video screen coordinate system of the Nth frame, and the second vertex coordinates of the face key points of the face model in the calibration coordinate system; the first vertex coordinates are mapped point-to-point with the second vertex coordinates, and the original transformation matrix (mapping matrix) of the Nth frame is calculated using a matrix solution method.
[0068] 202. Determine a first weight corresponding to the original transformation matrix of the Nth frame and a second weight corresponding to the smoothed transformation matrix of the N-1th frame.
[0069] After obtaining the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame, the first weight corresponding to the original transformation matrix of the Nth frame and the second weight corresponding to the smoothed transformation matrix of the N-1th frame can be determined, that is, the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame can be assigned corresponding weights. It can be understood that the first weight and the second weight are used to smooth the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame. Generally, the sum of the first weight and the second weight is 1, wherein the first weight and the second weight can be set by oneself.
[0070] Preferably, the first weight corresponding to the original transformation matrix of the Nth frame can be obtained, and the second weight corresponding to the smoothed transformation matrix of the N-1th frame can be set to be inversely proportional to the first weight, that is, when the first weight corresponding to the original transformation matrix of the Nth frame is larger, the second weight corresponding to the smoothed transformation matrix of the N-1th frame is smaller, so that the original transformation matrix of the Nth frame is tended to be focused on during video processing, so as to accelerate the convergence to the original transformation matrix of the Nth frame during video processing and avoid delays in the video picture of the Nth frame; or, when the first weight corresponding to the original transformation matrix of the Nth frame is smaller, the second weight corresponding to the smoothed transformation matrix of the N-1th frame is larger, and the focus is tended to be on the smoothed transformation matrix of the N-1th frame, so as to eliminate the tiny jitter signal that occurs when the video continuously plays the video pictures of the N-1th frame and the Nth frame.
[0071] 203. Smoothing the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame based on the first weight and the second weight to obtain the smoothed transformation matrix of the Nth frame.
[0072] After determining the first weight corresponding to the original transformation matrix of the Nth frame and the second weight corresponding to the smoothed transformation matrix of the N-1th frame, the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame can be smoothed based on the first weight and the second weight to obtain the smoothed transformation matrix of the Nth frame. That is, the original transformation matrix of the Nth frame can be multiplied by the first weight, and the smoothed transformation matrix of the N-1th frame can be multiplied by the second weight to obtain the smoothed transformation matrix of the Nth frame.
[0073] 204 . Using the smoothed transformation matrix of the Nth frame, coordinate transformation is performed on the coordinates of the target display element in the to-be-displayed video screen of the Nth frame in the calibration coordinate system to obtain the target coordinates of the target display element in the video screen coordinate system.
[0074] After obtaining the smoothed transformation matrix of the Nth frame, the smoothed transformation matrix of the Nth frame can be used to transform the coordinates of the target display element in the video screen to be displayed in the Nth frame in the calibrated coordinate system to obtain the target coordinates of the target display element in the video screen coordinate system, and the target display element belongs to the candidate display element. That is, the coordinates of the target display element in the calibrated coordinate system are obtained, and the coordinates of the target display element in the calibrated coordinate system are transformed using the smoothed transformation matrix of the Nth frame to obtain the target coordinates of the target display element in the video screen coordinate system. If the target display element is a hat, the coordinates of the hat in the calibrated coordinate system can be transformed to obtain the target coordinates of the hat in the video screen coordinate system.
[0075] 205. Add the target display element to the video picture to be displayed according to the target coordinates to obtain the video picture displayed at the Nth frame.
[0076] After obtaining the target coordinates, the target display element can be added to the video screen to be displayed according to the target coordinates to obtain the video screen displayed in the Nth frame, that is, the target display element is rendered in the target coordinates of the video screen coordinate system. That is, the target coordinates corresponding to all target display elements in the video screen to be displayed in the Nth frame can be obtained, and the vertex coordinates of all target display elements in the video screen of the Nth frame are unified in the video screen coordinate system to avoid jitter in the video screen.
[0077] In an embodiment of the present application, based on the first weight corresponding to the original transformation matrix of the Nth frame and the second weight corresponding to the smoothed transformation matrix of the N-1th frame, the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame are smoothed to obtain the smoothed transformation matrix of the Nth frame; the smoothed transformation matrix of the Nth frame is used to transform the coordinates of the target display element in the video picture to be displayed of the Nth frame in the calibrated coordinate system, and the target display element is added to the video picture to be displayed to achieve anti-shake for the video picture of the Nth frame; only the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame need to be used to transform the coordinates of the target display element in the video picture to be displayed of the Nth frame to achieve anti-shake for the video picture of the Nth frame, and the video picture can be anti-shake in real time without using a deep learning model.
[0078] Furthermore, after obtaining the smoothed transformation matrix of the Nth frame, the coordinate offset matrix of the Nth frame can be determined, and the smoothed transformation matrix of the N+1th frame can be obtained based on the coordinate offset matrix of the Nth frame to achieve anti-shake for the video picture of the N+1th frame, and further anti-shake for the video picture in real time; Figure 3 As shown, the specific steps are as follows:
[0079] 301. Determine a coordinate offset matrix of the Nth frame based on the smoothed transformation matrix of the Nth frame and the original transformation matrix of the Nth frame.
[0080] After obtaining the smoothed transformation matrix of the Nth frame, the coordinate offset matrix of the Nth frame can be determined based on the smoothed transformation matrix of the Nth frame and the original transformation matrix of the Nth frame. Specifically, the coordinate offset matrix of the Nth frame can be obtained by subtracting the original transformation matrix of the Nth frame from the smoothed transformation matrix of the Nth frame.
[0081] The coordinate offset matrix represents the difference between the smoothed transformation matrix of the Nth frame after smoothing and the original transformation matrix of the Nth frame before smoothing; the smaller the coordinate offset matrix is, the smaller the offset of the display element in the Nth frame is, that is, the smaller the movement amplitude relative to the N-1th frame is; the larger the coordinate offset matrix is, the larger the offset of the display element in the Nth frame is, that is, the larger the movement amplitude relative to the N-1th frame is.
[0082] 302. Determine, based on the coordinate offset matrix and the first weight, a third weight corresponding to the original transformation matrix of the N+1th frame and a fourth weight corresponding to the smoothed transformation matrix of the Nth frame.
[0083] Next, based on the coordinate offset matrix and the first weight, the third weight corresponding to the original transformation matrix of the N+1th frame and the fourth weight corresponding to the smoothed transformation matrix of the Nth frame can be determined. That is, when obtaining the smoothed transformation matrix of the N+1th frame, the degree of attention paid to the original transformation matrix of the N+1th frame and the smoothed transformation matrix of the Nth frame is adaptively adjusted through the coordinate offset matrix and the first weight of the Nth frame, thereby improving the stability of the video picture.
[0084] Generally, if the coordinate offset matrix is larger, the third weight corresponding to the original transformation matrix of the N+1th frame is increased, and the fourth weight corresponding to the smoothed transformation matrix of the Nth frame is reduced, so as to increase the attention paid to the original transformation matrix of the N+1th frame, that is, when the motion amplitude of the video picture to be displayed is large, the convergence to the N+1th frame is accelerated to avoid the delay of the video picture of the N+1th frame; if the coordinate offset matrix is smaller, the third weight corresponding to the original transformation matrix of the N+1th frame is reduced, and the fourth weight corresponding to the smoothed transformation matrix of the Nth frame is increased, so as to increase the attention paid to the smoothed transformation matrix of the Nth frame, that is, when the motion amplitude of the video picture to be displayed is small, the video picture of the N+1th frame is made more continuous with the video picture of the Nth frame.
[0085] Specifically, the coordinate offset in the coordinate offset matrix of the Nth frame can be obtained; wherein, the absolute values of multiple elements in the coordinate offset matrix of the Nth frame can be summed to obtain the coordinate offset in the coordinate offset matrix of the Nth frame. Within the preset weight range, based on the coordinate offset and the first weight, the third weight corresponding to the original transformation matrix of the N+1th frame is obtained. Wherein, the preset weight range includes a weight upper limit and a weight lower limit, and the weight upper limit and the weight lower limit can be set by oneself. After obtaining the sum of the coordinate offset and the first weight, it can be determined whether it is within the preset weight range. If so, the sum is used as the third weight corresponding to the original transformation matrix of the N+1th frame. If the sum is greater than the weight upper limit, the weight upper limit is used as the third weight corresponding to the original transformation matrix of the N+1th frame. If the sum is less than the weight lower limit, the weight lower limit is used as the third weight corresponding to the original transformation matrix of the N+1th frame. Then, based on the third weight corresponding to the original transformation matrix of the N+1th frame, the fourth weight corresponding to the smoothed transformation matrix of the Nth frame is determined.
[0086] 303. Smoothing the original transformation matrix of the N+1th frame and the smoothed transformation matrix of the Nth frame based on the third weight and the fourth weight to obtain the smoothed transformation matrix of the N+1th frame.
[0087] After determining the third weight corresponding to the original transformation matrix of the N+1th frame and the fourth weight corresponding to the smoothed transformation matrix of the Nth frame, the original transformation matrix of the N+1th frame and the smoothed transformation matrix of the Nth frame can be smoothed based on the third weight and the fourth weight to obtain the smoothed transformation matrix of the N+1th frame, that is, the original transformation matrix of the N+1th frame and the smoothed transformation matrix of the Nth frame are respectively assigned the third weight and the fourth weight and then summed to obtain the smoothed transformation matrix of the N+1th frame.
[0088] The smoothed transformation matrix of the N+1th frame is used to transform the coordinates of the target display element in the video picture to be displayed in the N+1th frame in the calibration coordinate system to obtain the target coordinates of the target display element in the video picture coordinate system. The target display element is added to the video picture to be displayed according to the target coordinates to obtain the video picture displayed in the N+1th frame, thereby achieving anti-shake of the video picture of the N+1th frame.
[0089] Further, in the embodiment of the present application, it is necessary to use the smoothed transformation matrix of the N-1th frame to obtain the smoothed transformation matrix of the Nth frame. When the Nth frame is the second frame, it is necessary to obtain the smoothed transformation matrix of the first frame to provide a data basis for the subsequent Nth frame. The process of obtaining the smoothed transformation matrix of the first frame is as follows: Figure 4 As shown, the specific steps are as follows:
[0090] 401. Obtain the original transformation matrix of the first frame in the video.
[0091] It can be understood that obtaining the original transformation matrix of the first frame in the video is similar to obtaining the original transformation matrix of the Nth frame in the video in the above step 201, and the details are not repeated here.
[0092] 402. Obtain a smoothed transformation matrix of the first frame based on an original transformation matrix of the first frame and a preset weight of the original transformation matrix of the first frame.
[0093] After obtaining the original transformation matrix of the first frame in the video, the smoothed transformation matrix of the first frame can be obtained based on the original transformation matrix of the first frame and the preset weights corresponding to the original transformation matrix of the first frame; the preset weights corresponding to the original transformation matrix of the first frame can be set by oneself; that is, the original transformation matrix of the first frame is assigned weights to obtain the smoothed transformation matrix of the first frame.
[0094] It can be understood that after obtaining the smoothed transformation matrix of the first frame, the coordinate offset matrix of the first frame can be determined based on the smoothed transformation matrix of the first frame and the original transformation matrix of the first frame, and the weight corresponding to the original transformation matrix of the second frame and the weight corresponding to the smoothed transformation matrix of the first frame can be determined based on the coordinate offset matrix and the preset weights; the original transformation matrix of the second frame and the smoothed transformation matrix of the first frame are smoothed to obtain the smoothed transformation matrix of the second frame. The specific process is the same as Figure 3 It can be seen that after obtaining the smoothed transformation matrix of the first frame, the smoothed transformation matrix of the first frame can be used to participate in the determination process of the smoothed transformation matrix of the subsequent N-th frame, providing a data basis for the subsequent N-th frame.
[0095] The following will be combined Figure 5 as well as Figure 6 , the video processing process is described, the specific steps are as follows:
[0096] 501. Obtain an original transformation matrix of the Nth frame in the video and a smoothed transformation matrix of the N-1th frame.
[0097] It is understandable that step 501 is similar to step 201, and the details are not repeated here. The original transformation matrix can be a model matrix, and the original transformation matrix of the Nth frame is: n ∈R( 4×4 ), and are as follows:
[0098]
[0099] Among them, m 01 、m 02 、m 10 、m 12 、m 20 、m 22 is the rotation element, m 00 、m 11 、m 22 is the scaling element, m 03 、m 13 、m 23 To translate an element.
[0100] 502. Determine a first weight corresponding to the original transformation matrix of the Nth frame and a second weight corresponding to the smoothed transformation matrix of the N-1th frame.
[0101] Specifically, the original transformation matrix of the N-1th frame can be obtained, and the coordinate offset matrix of the N-1th frame can be determined based on the smoothed transformation matrix of the N-1th frame and the original transformation matrix of the N-1th frame. Based on the coordinate offset matrix of the N-1th frame and the weights corresponding to the original transformation matrix of the N-1th frame, the first weight corresponding to the original transformation matrix of the N-1th frame and the second weight corresponding to the smoothed transformation matrix of the N-1th frame are determined. The specific process is the same as Figure 3 It is understandable that by using the weights corresponding to the coordinate offset matrix of the N-1th frame and the original transformation matrix of the N-1th frame to determine the first weight corresponding to the original transformation matrix of the N-1th frame and the second weight corresponding to the smoothed transformation matrix of the N-1th frame, the smoothed transformation matrix obtained in the Nth frame can be more correlated with the smoothed transformation matrix of the N-1th frame, thereby further improving the stability of the video image of the Nth frame.
[0102] 503. Smoothing the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame based on the first weight and the second weight to obtain the smoothed transformation matrix of the Nth frame.
[0103] Specifically, the first weight corresponding to the original transformation matrix of the Nth frame is: n , the second weight corresponding to the smoothed transformation matrix of the N-1th frame is: 1-a n That is, the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame are used as indicators of the smoothed transformation matrix of the Nth frame, and the sum of the weights of the indicators is 1.
[0104] Use the formula: Get the smoothed transformation matrix of the Nth frame That is, the smoothed transformation matrix of the Nth frame after adaptive smoothing; where M n is the original transformation matrix of the Nth frame, is the smoothed transformation matrix of the N-1th frame.
[0105] 504. Multiply the smoothed transformation matrix of the Nth frame by the vertex coordinates corresponding to the face special effect map in the calibrated coordinate system to obtain the vertex coordinates corresponding to the face special effect map in the calibrated coordinate system, and obtain the target coordinates of the face special effect map in the video screen coordinate system.
[0106] Furthermore, when the user adds special effects during the live broadcast, the face special effects map in the displayed video screen will shake on the face. At this time, the smoothed transformation matrix of the Nth frame can be used to transform the vertex coordinates of the face special effects map in the calibrated coordinate system to obtain the target coordinates of the face special effects map in the video screen coordinate system.
[0107] Specifically, the smoothed transformation matrix of the Nth frame is multiplied by the vertex coordinates of the face special effect map in the calibration coordinate system to obtain the vertex coordinates of the face special effect map in the calibration coordinate system, and the target coordinates of the face special effect map in the video screen coordinate system are obtained. For example, the first vertex coordinates (homogeneous coordinates v local )(x,y,z,1), the transformation matrix after smoothing of the Nth frame By performing coordinate transformation, we can obtain the target coordinate v of any vertex of the face special effect map in the video screen coordinate system. world , the specific formula is:
[0108]
[0109] 505. Add the face special effect map to the video picture to be displayed according to the target coordinates to obtain the video picture displayed at the Nth frame.
[0110] After obtaining the target coordinates of the face special effects map in the video screen coordinate system, the face special effects map is added to the video screen to be displayed according to the target coordinates to obtain the video screen displayed by the Nth frame, that is, the face special effects map is rendered based on the target coordinates to obtain the special effects video screen displayed by the Nth frame, thereby achieving smooth optimization of the timing frames, avoiding the jitter of the face special effects map, and improving the stability of the special effects video screen.
[0111] Further, after obtaining the smoothed transformation matrix of the Nth frame in step 503, the smoothed transformation matrix of the Nth frame and the original transformation matrix of the Nth frame can be used to obtain the smoothed transformation matrix of the N+1th frame, as follows:
[0112] Based on the smoothed transformation matrix of the Nth frame and the original transformation matrix of the Nth frame, the coordinate offset matrix of the Nth frame is determined. The coordinate offset matrix is the offset matrix E before and after the smoothing of the Nth frame. n ∈R( 4×4 ),Right now
[0113] Get the coordinate offset in the coordinate offset matrix of the Nth frame. Specifically, you can find the coordinate offset in the offset matrix E n Select a specific element and use the following formula:
[0114] e n =|e 00 |+|e 01 |+|e 10 |+|e 11 |+|e 03 |+|e 13 |-b, get the coordinate offset e of the coordinate offset matrix n , where eij is the offset matrix E n The element in the i-th row and j-th column, b is a fixed compensation value.
[0115] Next, within the preset weight range, based on the coordinate offset and the first weight, the third weight corresponding to the original transformation matrix of the N+1th frame is obtained. Specifically, the third weight α corresponding to the original transformation matrix of the N+1th frame can be obtained by the following formula: n+1 :
[0116] α n+1 =min(max(α n +β×e n ,minimal),maximal)
[0117] Among them, β is the scaling factor of the coordinate offset, minimal and maximal are the upper and lower limits of the weight respectively. The fourth weight corresponding to the smoothed transformation matrix of the corresponding Nth frame is: 1-α n+1 .
[0118] It can be seen that if the coordinate offset matrix (i.e., the offset matrix E n ) is smaller, the corresponding third weight α of the original transformation matrix of the N+1th frame n+1 The smaller the value, the corresponding fourth weight (1-α n+1 ) is larger, so that the video picture is more continuous and the tiny jitter signal is eliminated; if the coordinate offset matrix (i.e., the offset matrix E n ) is larger, the corresponding third weight α of the original transformation matrix of the N+1th frame n+1 The larger the value, the corresponding fourth weight (1-α n+1 ) is smaller to accelerate the convergence to the N+1th frame and avoid delay in the video image of the N+1th frame.
[0119] Based on the third weight and the fourth weight, the original transformation matrix of the N+1th frame and the smoothed transformation matrix of the Nth frame are smoothed to obtain the smoothed transformation matrix of the N+1th frame. That is, the formula can be used: Get the smoothed transformation matrix of the N+1th frame The video image of the N+1th frame is obtained by using the smoothed transformation matrix of the N+1th frame.
[0120] The present application also provides a video processing device, such as Figure 7 As shown, including:
[0121] The acquisition unit 701 is used to acquire the original transformation matrix of the Nth frame in the video and the smoothed transformation matrix of the N-1th frame; wherein N is greater than or equal to 2; the original transformation matrix of the Nth frame is the coordinate mapping relationship between the face position in the video screen to be displayed of the Nth frame and the face model; the face position in the video screen to be displayed is described by the video screen coordinate system, and the coordinates of the face model are described by the calibration coordinate system, and the calibration coordinate system is also used to describe the coordinates of the candidate display elements associated with the face model;
[0122] A determining unit 702 is used to determine a first weight corresponding to the original transformation matrix of the Nth frame and a second weight corresponding to the smoothed transformation matrix of the N-1th frame;
[0123] A smoothing unit 703 is used to perform smoothing processing on the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame based on the first weight and the second weight to obtain the smoothed transformation matrix of the Nth frame;
[0124] The transformation unit 704 is used to use the smoothed transformation matrix of the Nth frame to perform coordinate transformation on the coordinates of the target display element in the to-be-displayed video screen of the Nth frame in the calibration coordinate system to obtain the target coordinates of the target display element in the video screen coordinate system; the target display element belongs to the candidate display element;
[0125] The execution unit 705 is configured to add the target display element to the video picture to be displayed according to the target coordinates to obtain the video picture displayed at the Nth frame.
[0126] The embodiment of the present application also provides a video processing device 800, such as Figure 8 As shown, the video processing device 800 of the embodiment of the present application may include one or more processors CPU (CPU, central processing units) 801 and a memory 802, and the memory 802 stores one or more application programs or data.
[0127] The memory 802 may be a volatile storage or a persistent storage. The program stored in the memory 802 may include one or more modules, each of which may include a series of instruction operations in the electronic device. Furthermore, the processor 801 may be configured to communicate with the memory 802 and execute a series of instruction operations in the memory 802 on the video processing device 800.
[0128] The video processing device 800 may also include one or more power supplies 805, one or more wired or wireless network interfaces 804, one or more input and output interfaces 803, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0129] The processor 801 can execute the operations performed by the aforementioned specific method embodiment, and the details will not be repeated here.
[0130] An embodiment of the present application also provides a computer-readable storage medium, comprising instructions, which, when executed on a computer, enable the computer to execute the method described above.
[0131] The embodiment of the present application also provides a computer program product including instructions or a computer program. When the computer program product is run on a computer, the computer executes the method described above.
[0132] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0133] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0134] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0135] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0136] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk and other media that can store program code.
[0137] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0138] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0139] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0140] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0141] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk and other media that can store program code.
Claims
1. A video processing method, characterized in that: include: Obtaining an original transformation matrix of the Nth frame in the video and a smoothed transformation matrix of the N-1th frame; wherein N is greater than or equal to 2; the original transformation matrix of the Nth frame is a coordinate mapping relationship between a face position in a video screen to be displayed of the Nth frame and a face model; the face position in the video screen to be displayed is described by a video screen coordinate system, and the coordinates of the face model are described by a calibrated coordinate system, and the calibrated coordinate system is also used to describe the coordinates of candidate display elements associated with the face model; Determine a first weight corresponding to the original transformation matrix of the Nth frame and a second weight corresponding to the smoothed transformation matrix of the N-1th frame; Based on the first weight and the second weight, smoothing the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame to obtain the smoothed transformation matrix of the Nth frame; Using the smoothed transformation matrix of the Nth frame, coordinate transformation is performed on the coordinates of the target display element in the to-be-displayed video screen of the Nth frame in the calibrated coordinate system to obtain the target coordinates of the target display element in the video screen coordinate system; the target display element belongs to the candidate display element; The target display element is added to the video picture to be displayed according to the target coordinates to obtain the video picture displayed in the Nth frame.
2. The video processing method according to claim 1, characterized in that: The method further comprises: Determine a coordinate offset matrix of the Nth frame based on the smoothed transformation matrix of the Nth frame and the original transformation matrix of the Nth frame; Determine, based on the coordinate offset matrix and the first weight, a third weight corresponding to the original transformation matrix of the N+1th frame and a fourth weight corresponding to the smoothed transformation matrix of the Nth frame; Based on the third weight and the fourth weight, the original transformation matrix of the N+1th frame and the smoothed transformation matrix of the Nth frame are smoothed to obtain the smoothed transformation matrix of the N+1th frame.
3. The video processing method according to claim 2, characterized in that: The determining, based on the coordinate offset matrix and the first weight, a third weight corresponding to the original transformation matrix of the N+1th frame and a fourth weight corresponding to the smoothed transformation matrix of the Nth frame includes: Obtaining the coordinate offset in the coordinate offset matrix of the Nth frame; Within a preset weight range, based on the coordinate offset and the first weight, obtaining a third weight corresponding to the original transformation matrix of the N+1th frame; Based on the third weight corresponding to the original transformation matrix of the N+1th frame, a fourth weight corresponding to the smoothed transformation matrix of the Nth frame is determined.
4. The video processing method according to claim 1, characterized in that: When the Nth frame is the second frame, the N-1th frame is the first frame; Obtaining the smoothed transformation matrix of the N-1th frame in the video includes: Obtaining the original transformation matrix of the first frame in the video; A smoothed transformation matrix of the first frame is obtained based on an original transformation matrix of the first frame and a preset weight corresponding to the original transformation matrix of the first frame.
5. The video processing method according to claim 1, characterized in that: Obtaining the original transformation matrix of the Nth frame in the video includes: Obtaining the first vertex coordinates of the face position in the video picture to be displayed of the Nth frame in the video picture coordinate system, and the second vertex coordinates of the face model in the calibration coordinate system; Based on the mapping relationship between the first vertex coordinates and the second vertex coordinates, the original transformation matrix of the Nth frame is obtained.
6. The video processing method according to claim 1, characterized in that: The determining of the first weight corresponding to the original transformation matrix of the Nth frame and the second weight corresponding to the smoothed transformation matrix of the N-1th frame includes: Obtaining a first weight corresponding to the original transformation matrix of the Nth frame; The second weight corresponding to the smoothed transformation matrix of the N-1th frame is set to be inversely proportional to the first weight.
7. The video processing method according to claim 1, characterized in that: The first weight is: a n , the second weight is: 1-a n ; The step of performing smoothing processing on the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame based on the first weight and the second weight to obtain the smoothed transformation matrix of the Nth frame includes: Use the formula: Get the smoothed transformation matrix of the Nth frame Among them, M n is the original transformation matrix of the Nth frame, is the smoothed transformation matrix of the N-1th frame.
8. The video processing method according to claim 1, characterized in that: The target display element in the video picture to be displayed in the Nth frame includes: a face special effect map; The step of using the smoothed transformation matrix of the Nth frame to transform the coordinates of the target display element in the to-be-displayed video screen of the Nth frame in the calibrated coordinate system to obtain the target coordinates of the target display element in the video screen coordinate system includes: The smoothed transformation matrix of the Nth frame is multiplied by the vertex coordinates of the face special effect map corresponding to the calibrated coordinate system to obtain the target coordinates of the face special effect map in the video screen coordinate system.
9. A video processing device, characterized in that: include: An acquisition unit, used to acquire an original transformation matrix of the Nth frame in the video and a smoothed transformation matrix of the N-1th frame; wherein N is greater than or equal to 2; the original transformation matrix of the Nth frame is a coordinate mapping relationship between a face position in a video screen to be displayed of the Nth frame and a face model; the face position in the video screen to be displayed is described by a video screen coordinate system, and the coordinates of the face model are described by a calibrated coordinate system, and the calibrated coordinate system is also used to describe the coordinates of a candidate display element associated with the face model; A determination unit, configured to determine a first weight corresponding to the original transformation matrix of the Nth frame and a second weight corresponding to the smoothed transformation matrix of the N-1th frame; a smoothing unit, configured to perform smoothing processing on the original transformation matrix of the Nth frame and the smoothed transformation matrix of the N-1th frame based on the first weight and the second weight, so as to obtain the smoothed transformation matrix of the Nth frame; a transformation unit, configured to use the smoothed transformation matrix of the Nth frame to transform the coordinates of a target display element in the to-be-displayed video screen of the Nth frame in the calibrated coordinate system, to obtain target coordinates of the target display element in the video screen coordinate system; the target display element belongs to the candidate display element; An execution unit is used to add the target display element to the video picture to be displayed according to the target coordinates to obtain the video picture displayed in the Nth frame.
10. A video processing device, characterized in that: include: Processor, memory, and input and output interfaces; The memory is a short-term storage memory or a persistent storage memory; The processor is configured to communicate with the memory and execute instructions in the memory to perform the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that: The method comprises instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 8.
12. A computer program product comprising instructions or a computer program, characterized in that When the computer program product is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Video anti-shake display method and device
CN109089015A
Key point positioning method and device, electronic equipment and storage medium
CN110807410A
Method and device for processing video
CN111314626A
Video processing method and device, electronic equipment and storage medium
CN113160244A
Video processing method and related equipment thereof
CN115546042A