Video processing method and apparatus thereof

By acquiring the motion vectors of the front and back frames in image alignment and weighting adjustment, the problem of low accuracy in the prior art is solved, and higher frame alignment accuracy and robustness are achieved.

CN115100236BActive Publication Date: 2025-06-27VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210828522.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2025-06-27
Estimated Expiration
2042-07-13

AI Technical Summary

Technical Problem

The existing optical flow method is difficult to meet the brightness changes and tiny movement conditions when the object moves during image alignment, resulting in low accuracy of image alignment.

Method used

By obtaining the first and last two frames of the target video, determining its motion vector, and adjusting the motion vector according to the weighted proportion, ensuring that the frame alignment remains robust and smooth in spatial and temporal dimensions.

Benefits of technology

It improves the accuracy and robustness of image alignment, combines spatial and temporal processing, reduces system power consumption and expands application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115100236B_ABST
    Figure CN115100236B_ABST
Patent Text Reader

Abstract

The present application discloses a video processing method and apparatus thereof, belonging to the technical field of video processing. The method includes obtaining a first frame and a second frame of a target video, wherein the second frame is the previous frame adjacent to the first frame; determining a first motion vector and a weighting ratio of the first frame according to the second frame; adjusting the first motion vector according to the weighting ratio to obtain a second motion vector of the first frame; and aligning the first frame and the second frame based on the second motion vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of video processing, and particularly relates to a video processing method and apparatus thereof. Background Art

[0002] Image alignment, also known as frame alignment or image registration, is a technique for warping and rotating an image to align it with another image. Image alignment is a key technique in many video processing scenarios, such as video denoising, motion detection, portrait segmentation, and so on.

[0003] Generally, the optical flow method is used for image alignment. The optical flow method is a method for inferring the motion information of an object by detecting the change in the intensity of image pixels over time. When applying the optical flow method for image alignment, it is necessary to satisfy the conditions that the brightness of the object does not change during movement and the movement is a small movement. However, in real-world scenarios, it is difficult to satisfy both of these conditions, resulting in low accuracy of using the optical flow method for image alignment. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a video processing method and apparatus thereof, which can improve the accuracy and robustness of image alignment.

[0005] In a first aspect, the embodiments of this application provide a video processing method, which includes: obtaining a first frame and a second frame of a target video, where the second frame is the previous frame adjacent to the first frame; determining a first motion vector of the first frame according to the second frame; adjusting the first motion vector according to the weighting ratio to obtain a second motion vector of the first frame; and aligning the first frame with the second frame based on the second motion vector.

[0006] In a second aspect, the embodiments of this application provide a video processing apparatus, which includes: a first obtaining module for obtaining a first frame and a second frame of a target video, where the second frame is the previous frame adjacent to the first frame; a first determining module for determining a first motion vector and a weighting ratio of the first frame according to the second frame; a first weighting module for adjusting the first motion vector according to the weighting ratio to obtain a second motion vector of the first frame; and a first aligning module for aligning the first frame with the second frame based on the second motion vector.

[0007] In a third aspect, the embodiments of this application provide an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, it implements the video processing method as described in the first aspect.

[0008] Fourthly, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the video processing method described in the first aspect is implemented.

[0009] Fifthly, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the video processing method described in the first aspect.

[0010] Sixthly, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the video processing method described in the first aspect.

[0011] In the embodiment of the present application, first, the motion vectors of two consecutive frames are determined according to the two consecutive frames, and the relationship between the two consecutive frames is determined from the spatial dimension to ensure the robustness of frame alignment in space; then, the motion vector is adjusted according to the weighting ratio of the previous frame to the next frame, so that the two consecutive frames are smooth in the time dimension and the robustness of frame alignment in time is improved. The combination of the time dimension and the spatial dimension can also improve the accuracy of frame alignment. Moreover, the technical solution of this embodiment has a simple process flow and a small data throughput, which is beneficial to reducing the power consumption pressure of the system and can increase the application scenarios. Description of the Drawings

[0012] Figure 1 is a flowchart of the video processing method provided by the embodiment of the present application;

[0013] Figure 2 is a schematic diagram of a sub-block in the video processing method provided by the embodiment of the present application;

[0014] Figure 3 is a schematic structural diagram of the video processing device provided by the embodiment of the present application;

[0015] Figure 4 is a schematic structural diagram of an electronic device provided by the embodiment of the present application;

[0016] Figure 5 is a schematic structural diagram of an electronic device provided by the embodiment of the present application. Detailed Embodiments

[0017] Next, the technical solutions in the embodiments of the present application will be clearly described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.

[0018] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.

[0019] The following will combine the accompanying drawings and, through specific embodiments and their application scenarios, elaborate in detail on the video processing method, video processing device, and electronic device provided by the embodiments of this application.

[0020] The embodiments of this application first provide a video processing method, which can be applied to electronic devices such as mobile phones, tablet computers, laptop computers, wearable electronic devices (such as smart watches), augmented reality (AR) / virtual reality (VR) devices, in-vehicle devices, etc. The embodiments of this application do not impose any restrictions on this.

[0021] Figure 1 A flowchart showing a video processing method provided by the embodiments of this application is shown. As Figure 1 shown, the video processing method includes the following steps:

[0022] Step 10: Obtain the first frame and the second frame of the target video, where the second frame is the previous frame adjacent to the first frame.

[0023] The image sensor can obtain images at a certain time period, such as obtaining one frame of image every 30 milliseconds, etc. Multiple frames of images can form a video. The first frame can refer to the currently obtained image or the image to be processed. The second frame refers to the image obtained before the first frame. For example, if the image sensor generates one frame of image every 30 milliseconds, then taking 30 milliseconds as a moment, the first frame can be the image at the 2nd moment, and the second frame can be the image at the 1st moment.

[0024] Step 20: Determine the first motion vector and the weighting ratio of the first frame according to the second frame.

[0025] The motion vector refers to the motion trajectory of an image or a region in the image relative to the reference frame, that is, the direction and size of the movement in the spatial perspective. In this embodiment, the first motion vector refers to the motion vector of the first frame relative to the second frame.

[0026] Exemplarily, the process of determining the first motion vector is as follows: First, both the first frame and the second frame are divided into multiple sub-blocks, such as M sub-blocks. M is a positive integer greater than 0. The first frame and the second frame have the same size. For example, for the first frame and the second frame with a size of W×H, where W is the width of the image and H is the height of the image, the image can be divided into m×n sub-blocks, and M = m×n. Then the size of each sub-block is W / m×H / n. For example, if the second frame and the first frame are respectively divided into 2×2 sub-blocks, as Figure 2 shown, the second frame 21 can be divided into sub-block 1, sub-block 2, sub-block 3, and sub-block 4, and the first frame 22 is also divided into four sub-blocks: sub-block 1, sub-block 2, sub-block 3, and sub-block 4.

[0027] Then, motion estimation is performed on each pair of sub-blocks of the first frame and the second frame respectively to determine the first local vector of the sub-block of the first frame relative to the corresponding sub-block in the second frame. The local vector refers to the estimated value of the motion vector of the sub-block, which can include the estimated value in the horizontal dimension and the estimated value in the vertical dimension. Exemplarily, the first frame and the second frame can be first grayscaled, and the value of each pixel point is converted into a gray value, thereby reducing the amount of calculation and improving the calculation efficiency. For each sub-block in the first frame and the second frame, the pixel values of each column in the sub-block, that is, the gray values, can be accumulated to obtain a one-dimensional array composed of the accumulated sums of the column pixel values. For example, if each sub-block includes 10 columns, the one-dimensional array composed of the accumulated sums of each column includes 10 elements. This one-dimensional array can be expressed as:

[0028]

[0029] For easy distinction, the sub-block 1 of the second frame 21 is denoted as pre1, and the sub-block 1 of the first frame 22 is denoted as cur1. In the above formula, Ver pre1 is the one-dimensional array corresponding to the first sub-block of the second frame, and Ver pre1 (i) represents the i-th element in this one-dimensional array; blockH is the width of each sub-block, that is, the total number of columns; blockW is the height of each sub-block, that is, the total number of rows; pre1(i,j) represents the pixel point in the first sub-block of the first frame, that is, the pixel point in the i-th row and the i-th column; there are a total of blockW elements in this one-dimensional array.

[0030] Moreover, the pixel values of each row of the first sub-block of the second frame are accumulated, and a one-dimensional array can also be obtained according to the accumulated sums of the row pixel values. For easy distinction, the one-dimensional array Ver composed of the above column pixel points is called the column array, and the one-dimensional array composed of the row pixel points is called the row array. The row array of the first sub-block of the second frame can be expressed as:

[0031]

[0032] Among them, Hor represents the column array. According to the above formulas (1) and (2), the row arrays and column arrays of other sub-blocks in the second frame, namely pre2, pre3, and pre4, as well as the row arrays and column arrays of each sub-block in the first frame, can also be determined. The horizontal and vertical dimension features of different regions in the first frame or the second frame can be counted through the row arrays and column arrays of each sub-block. After the above processing, m×n×2×2 one-dimensional arrays can be obtained, that is, 8 arrays of the first frame and 8 arrays of the second frame.

[0033] Optionally, the estimation values of the motion vectors in the horizontal and vertical dimensions between the second frame and the first frame are calculated by using the method of staggered subtraction. Staggering means that after moving the second frame or the first frame by a certain distance, it is subtracted from the other frame. The range of staggering is preset, such as [-5,5], [-9,9], etc., and this embodiment does not make special limitations on this. Exemplarily, different staggering ranges can be set for the horizontal and vertical dimensions respectively. After determining the staggering range, search for the staggering value with the smallest cumulative difference value between the second frame and the first frame within this range, and this staggering value can be determined as the first local vector of this sub-block. For example, when the staggering value is 0, for the first sub-block pre1 of the second frame and the first sub-block cur1 of the first frame, calculate the results of subtracting the first element of sub-block pre1 from the first element of sub-block cur1, the second element, the third element, etc. in turn, and then add the absolute values of the results of each pair of subtractions to obtain the cumulative difference value in the horizontal or vertical dimension when the staggering is 0 between sub-block pre1 and sub-block cur1. After calculating the difference values corresponding to each staggering position within the staggering range, determine the staggering value when the difference value is the smallest, and this value is the first local vector of sub-block pre1 and sub-block cur1. For example, for the first sub-block of the first frame, the calculation process in the horizontal dimension is expressed by the following formula:

[0034]

[0035] Among them, difVer block1 (k) represents the difference value between the first sub-block block1 of the first frame and the first sub-block of the second frame when the staggering is k; k takes values within the staggering range, such as k = -10, -9,..., 0, 1,..., 9, 10. Calculate the difference value difVer between the corresponding sub-blocks for each k value in turn within this range, and take the k value when the difference value difVer is the smallest as the estimated value of the horizontal dimension of this sub-block, denoted as X block1 . Similarly, the difference value difHor of the vertical dimension of the first sub-block block1 can be calculated, so as to obtain the estimated value Y block1 when the difference value difHor is the smallest, and the first local vector of the first sub-block block1 is: X block1 , Y block1 .

[0036] In the embodiment of the present application, by processing each sub-block in the first frame with the above formula (3), the first local vector of each sub-block can be obtained. Then, spatial filtering is performed on the local vectors of the first frame to improve the robustness and smoothness in the spatial dimension.

[0037] For the j-th sub-block in the first frame, the first local vector of the j-th sub-block is subjected to mean filtering according to the adjacent sub-blocks of the j-th sub-block to obtain the second local vector of the j-th sub-block. If counting starts from 0, then 0 ≤ j ≤ M. j can also start counting from 1. If the adjacent sub-blocks of the j-th sub-block are the (j + 1)-th sub-block and the (j - 1)-th sub-block, then the average value of the first local vectors of the (j - 1)-th, j-th, and (j + 1)-th sub-blocks is used as the second local vector of the j-th sub-block. Exemplarily, the 8 sub-blocks surrounding the j-th sub-block are used as the adjacent sub-blocks of the j-th sub-block, and the average value of the first local vectors of these 9 sub-blocks is calculated together with the j-th sub-block, and the obtained result is used as the second local vector of the j-th sub-block. This second local vector is the first motion vector of the first frame.

[0038] Exemplarily, the first motion vector of the first frame may include a vector map composed of the second local vectors of each sub-block. The length and width of this vector map are the number of sub-blocks in the horizontal direction and the number of sub-blocks in the vertical direction, that is, m and n. Figure 2 , when the image is divided into 2×2 sub-blocks, the vector map of the first frame can be expressed as MVmap1(x, y), where x = 1, 2; y = 1, 2. Specifically: MVmap1(1, 1) = {X block1 , Y block1}, MVmap1(1, 2) = {X block2 , Y block2}, MVmap1(2, 1) = {X block3 , Y block3}, MVmap1(2, 2) = {X block4 , Y block4}. Alternatively, the second local vector of each sub-block in the first frame can be expressed as a one-dimensional vector as the first motion vector. For example, the first element MVmap1(1) of the one-dimensional vector MVmap1 = {X block1 , Y block1}, the second element MVmap1(2) = {X block2 , Y block2}, and so on.

[0039] In this embodiment, first, the local vectors of the first frame are estimated according to the intra-frame information of the first frame, and then mean filtering is performed on the local vectors to enhance the smoothness of the local vectors, which can ensure the overall consistency of the local vectors.

[0040] Step 30: Adjust the first motion vector according to the weighting ratio to obtain the second motion vector of the first frame.

[0041] It can be understood that the first motion vector of the first frame can be determined according to the second frame, and the first motion vector of the second frame can also be determined according to the frames before the second frame. That is to say, the first motion vector of the subsequent frame can be determined according to every two adjacent frames in the video. Furthermore, the weighting ratio of the first frame can be determined according to the first motion vector of the second frame, and then the first motion vector of the first frame can be processed according to the weighting ratio. Similarly, the weighting ratio of the second frame can also be determined according to the frames before the second frame.

[0042] Exemplarily, the weighting ratio may include two coefficients, namely the first coefficient and the second coefficient. In terms of the horizontal dimension and the vertical dimension, the weighting ratio may include the first coefficient and the second coefficient in the horizontal dimension and the first coefficient and the second coefficient in the vertical dimension. The weighting ratio of the first frame can be determined according to the weighting ratio of the second frame. By incorporating the frames at historical moments into the motion vector result of the first frame through the weighting ratio, the robustness of the motion vector estimation result of the first frame in the time domain can be improved.

[0043] Specifically, first determine the first candidate coefficient of the first frame based on the first target parameter and the first coefficient in the weighting ratio of the second frame. If the second frame is the starting frame in the video, such as the 0th frame when counting starts from 0, then the weighting ratio of the 0th frame is 0, that is, both the first coefficient and the second coefficient are 0. By adding the first target parameter, weights can be introduced. Specifically, for the horizontal dimension or the vertical dimension, first update the first coefficient of the second frame based on the first target parameter, as shown in the following formula:

[0044] pre bx (i) = pre bx (i) + paramx1 (4)

[0045] where pre bx (i) is the first coefficient of the i-th sub-block in the horizontal dimension of the second frame, i is greater than or equal to 0 and less than or equal to m×n, and paramx1 represents the first target parameter. That is, the updated result is the sum of the first target parameter and the original first coefficient of the second frame. Then, normalize the updated first coefficient Pre bx to obtain the first candidate coefficient of the first frame. The normalization process can be expressed by the following formula:

[0046] tmp1(i) = pre bx (i) / (pre bx (i) + paramx2) (5)

[0047] Among them, tmp1(i) represents the first candidate coefficient of the horizontal dimension of the i-th sub-block after normalization, and paramx2 represents the second target parameter. The first candidate coefficient of the horizontal dimension of the first frame can be obtained through this formula. Similarly, by performing the same processing on the vertical dimension according to formula (4) and formula (5), the first candidate coefficient of the vertical dimension of the first frame can be obtained. Exemplarily, the first target parameter and the second target parameter of the vertical dimension can be different from or the same as those of the horizontal dimension, and this embodiment does not make special limitations on this.

[0048] Next, based on the second coefficient of the second frame and the first motion vector of the first frame, the second candidate coefficient of the first frame can be determined. Exemplarily, the first-order variance between the second coefficient of the second frame and the first motion vector of the first frame is determined, and the second candidate coefficient of the first frame is obtained based on this first-order variance. The formula is expressed as follows:

[0049] tmp2(i) = MVmap1(i) - pre xa (i) (6)

[0050] Among them, tmp2(i) is the second candidate coefficient of the first frame, and pre xa (i) is the second coefficient of the horizontal dimension of the second frame. If the second frame is the starting frame of the video, then pre xa (i) = 0. MVmap1(i) represents the i-th element in the first motion vector of the first frame, where i is greater than or equal to 0 and less than or equal to m×n. The first-order variance between each element in the first motion vector of the first frame and the second coefficient of the second frame is determined as the second candidate coefficient of the first frame.

[0051] Then, based on the first candidate coefficient and the second candidate coefficient of the first frame, the weighting ratio of the first frame is determined. Exemplarily, by performing a weighting process according to the first candidate coefficient and the second candidate coefficient of the first frame, the weighting ratio of the first frame can be obtained. Specifically, for the horizontal dimension, based on the first candidate coefficient and the second candidate coefficient of the horizontal dimension of the first frame, the first coefficient of the horizontal dimension of the first frame is determined. The formula is expressed as follows:

[0052] cur xb (i) = (1 - tmp1(i) * pre xb (i)) (7)

[0053] Among them, cur xb (i) is the first coefficient of the horizontal dimension of the i-th sub-block of the first frame. The first coefficient of the first frame is determined according to the first candidate coefficient of the first frame and the first coefficient of the second frame, that is, the updated first coefficient in formula (4). In this way, the first coefficient of the second frame is added to the first coefficient of the first frame in a certain proportion, enhancing the correlation on the time axis, thereby improving the smoothness between the first frame and the previous frame.

[0054] The second coefficient of the horizontal dimension of the first frame is calculated according to the following formula:

[0055] cur xa (i) = pre xa (i) + tmp1(i) * tmp2(i) (8)

[0056] Where cur xa (i) is the second coefficient of the horizontal dimension of the first frame. Incorporating the first motion vector of the first frame into the weighted ratio of the first frame according to a certain ratio and then continuing to perform weighted processing on the next frame can ensure the smoothness between the first frame and the next frame, thereby enhancing the robustness of the video in terms of time. Similarly, for the vertical dimension, by performing the same weighted processing based on the first candidate coefficient and the second candidate coefficient of the vertical dimension of the first frame, the first coefficient of the vertical dimension of the first frame can be obtained.

[0057] On the basis of determining the weighted ratio of the first frame, adjusting the determined first motion vector can obtain the second motion vector of the first frame. Exemplarily, based on the third target parameter and the above-mentioned weighted ratio of the first frame, the first component of the first frame can be determined, and based on the fourth target parameter and the first motion vector of the first frame, the second component can be determined. Then, by combining the first component and the second component, the second motion vector can be obtained. In this embodiment, the weighted ratio is used to perform weighted processing on the third target parameter, and the obtained result is fused with a certain ratio of the first motion vector to obtain the second motion vector. Adjusting the motion vector through a certain ratio of the target parameter can assist in the estimation of the motion vector and is beneficial to improving the accuracy of motion estimation. For example, the second motion vector of the first frame can be:

[0058] MVmap2(i) = pre xa (i) * paramx3 + MVmap1 x (i) * paramx4 (9)

[0059] MVmap1 x (i) represents the value of the horizontal dimension in the first motion vector of the first frame. According to this formula, the value of the horizontal dimension of the second motion vector MVmap2(i) of the first frame can be calculated. Where paramx3 represents the third target parameter of the horizontal dimension, and the result of weighting the third target parameter with the second coefficient of the first frame is the first component, that is, pre xa (i) * paramx3. paramx4 is the fourth target parameter of the horizontal dimension, and the result of weighting the horizontal value of the first motion vector with this fourth target parameter is the second component, that is, MVmap1 x(i) *paramx4. The calculation process in the vertical dimension is the same as that in the horizontal dimension and will not be elaborated here.

[0060] Continue to refer to Figure 1 , in step 40, align the first frame with the second frame based on the second motion vector.

[0061] Each point in the first frame can be moved according to the second motion vector to obtain an image with the same spatial perspective as the second frame. Each frame in the captured video can be aligned according to the above method, and then the aligned video is denoised to obtain a video with better image quality.

[0062] Exemplarily, after obtaining the second motion vector, the second motion vector can be further optimized in the spatial domain. Specifically, determine the average value of the second motion vectors of the first frame; then based on this average value, eliminate the target vector values in the second motion vector to obtain the third motion vector. For the average value in the horizontal dimension, through the following formula:

[0063]

[0064] where MVmap2 x (i) is the horizontal dimension value of the i-th element in the second motion vector, and m*n is the total number of sub-blocks. By calculating the variance between each value in the second motion vector and this average value, the value with the largest variance can be determined, that is, the target vector value. Taking the above horizontal dimension as an example, calculate the variance between the horizontal dimension average value and each horizontal dimension value in the second motion vector as follows:

[0065] x(i) = (MVmap2 x (i) - mean x ) 2 (11)

[0066] x(i) is the variance of the i-th element in the second motion vector. The horizontal dimension value of the element with the largest variance is the target vector value. After obtaining the variance of each element, use the horizontal dimension average value mean x value to replace this target vector value to complete the processing of the horizontal dimension. Then process the vertical dimension in the same way. Replace the target vector value with the average value, thereby eliminating the target vector value to obtain a smoother motion vector, that is, the third motion vector. Then align the first frame with the previous frame according to the third motion vector.

[0067] The video processing method provided in this embodiment has a simple process and can improve the energy efficiency ratio of software and hardware. And motion estimation is performed on video frames in terms of space and time, which can improve the accuracy of motion vectors and the smoothness in terms of time and space.

[0068] Further, for the video processing method provided in the embodiments of the present application, the execution subject may be a video processing device. Hereinafter, taking the execution of the video processing method in the video processing device as an example, the video processing device provided in the embodiments of the present application will be described.

[0069] As Figure 3 shown, the video processing device 30 provided in the embodiments of the present application may include a first acquisition module 31, a first determination module 32, a first weighting module 33, and a first alignment module 34. Specifically, the first acquisition module 31 is configured to acquire a first frame and a second frame of a target video, where the second frame is the previous frame adjacent to the first frame; the first determination module 32 is configured to determine a first motion vector and a weighting ratio of the first frame according to the second frame; the first weighting module 33 is configured to adjust the first motion vector according to the weighting ratio to obtain a second motion vector of the first frame; the first alignment module 34 is configured to align the first frame with the second frame based on the second motion vector.

[0070] The video processing device provided in this embodiment determines the motion vectors of two consecutive frames, determines the relationship between the two consecutive frames from the spatial dimension, and then adjusts the motion vector according to the weighting ratio of the previous frame to the next frame, so that the consecutive frames maintain smoothness in the time dimension. The combination of the time dimension and the spatial dimension can improve the accuracy and robustness of frame alignment. Moreover, the technical solution of this embodiment has a simple process and a small data throughput, which is beneficial to reducing the power consumption pressure of the system.

[0071] In the embodiments of the present application, the first determination module 32 includes: a first division module configured to divide each of the first frame and the second frame into M sub-blocks, where M is a positive integer; a second determination module configured to determine a first local vector of the sub-block in the first frame relative to the corresponding sub-block in the second frame; a first filtering module configured to perform mean filtering on the first local vector of the j-th sub-block in the first frame according to the adjacent sub-blocks of the j-th sub-block to obtain a second local vector of the j-th sub-block, where 0 ≤ j ≤ M; a third determination module configured to determine the first motion vector according to the second local vector of each sub-block in the first frame.

[0072] Exemplarily, the weighting ratio includes a first coefficient and a second coefficient. The first weighting module 33 specifically includes: a first coefficient determination module configured to determine a first candidate coefficient of the first frame based on a first target parameter and the first coefficient of the second frame; a second coefficient determination module configured to determine a second candidate coefficient of the first frame based on the second coefficient of the second frame and the first motion vector of the first frame; a second weighting module configured to determine the weighting ratio of the first frame based on the first candidate coefficient and the second candidate coefficient.

[0073] In an embodiment of the present application, the first coefficient determination module includes a first update module configured to update the first coefficient of the second frame based on a first target parameter; and a first normalization module configured to normalize the updated first coefficient based on a second target parameter to obtain a first candidate coefficient of the first frame.

[0074] In an embodiment of the present application, the second coefficient determination module is specifically configured to: determine a first-order variance between the second coefficient of the second frame and the first motion vector of the first frame to obtain a second candidate coefficient of the first frame.

[0075] In an embodiment of the present application, the first weighting module 33 specifically includes: a first component determination module configured to determine a first component based on a third target parameter and a weighting ratio of the first frame; a second component determination module configured to determine a second component based on a fourth target parameter and the first motion vector of the first frame; and a first obtaining module configured to obtain a second motion vector of the first frame based on the first component and the second component.

[0076] In an embodiment of the present application, the first alignment module includes: a first averaging module configured to determine an average value of the second motion vectors of the first frame; a first rejection module configured to reject a target vector in the second motion vectors of the first frame based on the average value to obtain a third motion vector of the first frame; and a second alignment module configured to align the first frame and the second frame based on the third motion vector.

[0077] The video processing device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0078] The video processing device in the embodiments of the present application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0079] The video processing device provided by the embodiments of the present application can implement Figure 1 each process implemented by the method embodiments in , and for the sake of brevity, details are not repeated here.

[0080] Optionally, as Figure 4 shown, the embodiments of the present application further provide an electronic device 400, including a processor 401 and a memory 402. A program or instruction that can run on the processor 401 is stored on the memory 402. When the program or instruction is executed by the processor 401, it implements each step of the above video processing method embodiment and can achieve the same technical effect. For the sake of brevity, details are not repeated here.

[0081] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0082] Figure 5 FIG. is a schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.

[0083] The electronic device 500 includes but is not limited to: a radio frequency unit 501, a network module 5102, an audio output unit 503, an input unit 504, a sensor 505, a display unit 506, a user input unit 507, an interface unit 508, a memory 509, and a processor 510, etc.

[0084] Those skilled in the art can understand that the electronic device 500 may further include a power supply (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 510 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 5 The structure of the electronic device shown in does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which are not elaborated here.

[0085] Among them, the processor 510 can be used to obtain the first frame and the second frame of the target video, where the second frame is the previous frame adjacent to the first frame; determine the first motion vector and the weighting ratio of the first frame according to the second frame; adjust the first motion vector according to the weighting ratio to obtain the second motion vector of the first frame; align the first frame and the second frame based on the second motion vector.

[0086] The input unit 504 may include a Graphics Processing Unit (GPU) 1041 and a microphone 5042. The graphics processor 5041 processes the image data of static pictures or videos obtained by an image capturing device (such as a camera) in a video capture mode or an image capture mode. The display unit 506 may include a display panel 5061, and the display panel 5061 may be configured in the form of, for example, a liquid crystal display, an organic light emitting diode, etc. The user input unit 507 includes at least one of a touch panel 5071 and other input devices 5072. The touch panel 5071 is also referred to as a touch screen. The touch panel 5071 may include two parts: a touch detection device and a touch controller. The other input devices 5072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0087] The memory 509 can be used to store software programs and various data. The memory 509 mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 509 can include a volatile memory or a non-volatile memory, or the memory 509 can include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 509 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.

[0088] The processor 510 may include one or more processing units; optionally, the processor 510 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor may not be integrated into the processor 510 either.

[0089] The embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned video processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0090] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.

[0091] The embodiment of the present application further provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above-mentioned video processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0092] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0093] The embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement each process of the above-mentioned video processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0094] It should be noted that, in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0096] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit and scope protected by the claims of the present application, can also make many forms, all of which fall within the protection scope of the present application.

Claims

1. A video processing method, characterized in that, Including: Obtain the first frame and the second frame of the target video, where the second frame is the previous frame adjacent to the first frame; Determine the first motion vector of the first frame according to the second frame; Update the first coefficient in the weighting ratio of the second frame based on the first target parameter; perform normalization processing on the updated first coefficient based on the second target parameter to obtain the first candidate coefficient of the first frame; determine the second candidate coefficient of the first frame based on the second coefficient in the weighting ratio of the second frame and the first motion vector of the first frame; determine the weighting ratio of the first frame based on the first candidate coefficient and the second candidate coefficient; Adjust the first motion vector according to the weighting ratio to obtain the second motion vector of the first frame; Align the first frame and the second frame based on the second motion vector.

2. The video processing method according to claim 1, wherein The determining the first motion vector of the first frame according to the second frame includes: Divide each of the first frame and the second frame into M sub-blocks, where M is a positive integer; Determine the first local vector of the sub-block in the first frame relative to the corresponding sub-block in the second frame; Perform mean filtering on the first local vector of the j-th sub-block in the first frame according to the adjacent sub-blocks of the j-th sub-block in the first frame to obtain the second local vector of the j-th sub-block, where 0 ≤ j ≤ M; Determine the first motion vector according to the second local vector of each sub-block in the first frame.

3. The video processing method according to claim 1, wherein The aligning the first frame and the second frame based on the second motion vector includes: Determine the average value of the second motion vector of the first frame; Eliminate the target vectors in the second motion vector based on the average value to obtain the third motion vector of the first frame; Align the first frame and the second frame based on the third motion vector.

4. A video processing device, characterized in that, Including: A first obtaining module, configured to obtain the first frame and the second frame of the target video, where the second frame is the previous frame adjacent to the first frame; A first determining module, configured to determine the first motion vector of the first frame according to the second frame, update the first coefficient in the weighting ratio of the second frame based on the first target parameter; perform normalization processing on the updated first coefficient based on the second target parameter to obtain the first candidate coefficient of the first frame; determine the second candidate coefficient of the first frame based on the second coefficient in the weighting ratio of the second frame and the first motion vector of the first frame; determine the weighting ratio of the first frame based on the first candidate coefficient and the second candidate coefficient; A first weighting module, configured to adjust the first motion vector according to the weighting ratio to obtain the second motion vector of the first frame; A first aligning module, configured to align the first frame and the second frame based on the second motion vector.

5. The video processing device according to claim 4, wherein The first determining module includes: A first dividing module, configured to divide each of the first frame and the second frame into M sub-blocks, where M is a positive integer; A second determining module, configured to determine the first local vector of the sub-block in the first frame relative to the corresponding sub-block in the second frame; The first filtering module is configured to perform mean filtering on the first local vector of the j-th sub-block in the first frame according to adjacent sub-blocks of the j-th sub-block in the first frame, to obtain a second local vector of the j-th sub-block, where 0 ≤ j ≤ M; The third determination module is configured to determine the first motion vector according to the second local vectors of each sub-block in the first frame.

6. The video processing device according to claim 4, wherein The first alignment module includes: The first averaging module is configured to determine an average value of the second motion vectors of the first frame; The first rejection module is configured to reject target vectors in the second motion vectors of the first frame based on the average value, to obtain a third motion vector of the first frame; The second alignment module is configured to align the first frame and the second frame based on the third motion vector.

Citation Information

Patent Citations

  • Motion Vector Prediction Through Scaling

    CN107205156A

  • Inter frame prediction method and device for video images and codec

    CN109587479A