Method for preprocessing video and related product
By preprocessing the video, variable-size chunking is used to perform motion estimation between the filtered frame and the reference frame, the problem that traditional methods cannot adapt to complex pixel motion is solved, and the time domain filtering and video encoding quality is improved.
Patent Information
- Application Number
- CN202411963300.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-30
AI Technical Summary
The traditional motion compensation time domain filtering method based on pyramid block matching algorithm cannot adapt to the complex pixel motion situation in video, resulting in incorrect motion estimation results, artifacts and block effects, reducing the time domain filtering effect.
By blocking the filtered frames, macroblocks with different sizes are obtained, motion estimation is performed based on the reference frames associated with the filtered frame, and motion estimation is performed between the filtered frame and the reference frame using variable-size chunking.
The time domain filtering effect and video encoding quality are improved, and larger motion units in the image can be better covered, and more fine-grained motion estimation is performed on pixel points.
Smart Images

Figure CN120075438A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of video coding technology. More specifically, the present disclosure relates to a method for preprocessing a video and related products. Background Art
[0002] The video encoding process reduces the redundant information in the original video file to obtain byte stream data that can be stored or transmitted at a lower bit rate. Since random noise is introduced during the video recording and generation process, the noise not only reduces the redundant information available in the video encoding process and reduces the video compression ratio, but also causes artifacts in the video image. Motion Compensated Temporal Filter (MCTF) is a preprocessing method for video encoding. MCTF uses frames near the filtered frame (i.e., frames that need to be filtered using MCTF) as reference frames. By adjusting the filtered frame, the redundant information between the filtered frame and the reference frame is richer and the changes are smoother.
[0003] The traditional MCTF algorithm performs motion estimation based on the Pyramidal Block Matching Algorithm, which assumes that the motion of all pixels in a filtered frame macroblock is consistent, and thus performs motion estimation between the filtered frame and the reference frame based on a fixed block size. However, motion estimation based on a fixed block size cannot adapt to the complex pixel motion in the video, and it is easy to cause erroneous motion estimation results, resulting in artifacts and block effects in the motion compensated image, thereby reducing the effect of temporal filtering.
[0004] In view of this, there is an urgent need to provide a video preprocessing solution to facilitate motion estimation between filtered frames and reference frames based on variable-sized blocks, improve the temporal filtering effect, and further improve the video encoding quality. Summary of the invention
[0005] In order to at least solve one or more technical problems described in the above background technology section, the present disclosure proposes the following technical solutions and multiple embodiments thereof.
[0006] In a first aspect, the present disclosure discloses a method for preprocessing a video, comprising: partitioning a filtered frame to obtain a first filtered frame macroblock having a first size; partitioning at least one first filtered frame macroblock to obtain a second filtered frame macroblock having a second size; determining a matching block that matches the second filtered frame macroblock based on a reference frame associated with the filtered frame; and updating the second filtered frame macroblock according to the matching block.
[0007] In a second aspect, the present disclosure discloses a computer-readable storage medium storing program instructions adapted to be loaded and executed by a processor to perform the method according to the first aspect.
[0008] In a third aspect, the present disclosure discloses an apparatus for preprocessing a video, including: a processor configured to execute program instructions; and a memory configured to store program instructions, which, when loaded and executed by the processor, cause the apparatus to perform the method according to the first aspect.
[0009] According to the method for preprocessing a video disclosed in the present disclosure, after the filtered frame is partitioned according to a first size to obtain first filtered-frame macroblocks, a part or all of the first filtered-frame macroblocks are further partitioned according to a second size to obtain second filtered-frame macroblocks; thus, there are relatively large first filtered-frame macroblocks and relatively small second filtered-frame macroblocks in the filtered frame. The relatively large filtered-frame macroblocks can better cover larger motion units in the image, and the relatively small filtered-frame macroblocks can perform motion estimation on pixel points in the image with finer granularity; in this way, motion estimation is performed between the filtered frame and the reference frame through variable-size partitioning, improving the temporal filtering and video coding quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of illustration and not limitation, and the same or corresponding reference numerals indicate the same or corresponding parts, wherein:
[0011] Figure 1 An exemplary schematic diagram of a video compression process in some embodiments of the present disclosure is shown.
[0012] Figure 2 An exemplary schematic diagram of performing MCTF processing on a filtered frame based on one reference frame in some embodiments of the present disclosure is shown.
[0013] Figure 3 An exemplary schematic diagram of the pyramid block matching algorithm in some embodiments of the present disclosure is shown.
[0014] Figure 4 An exemplary schematic diagram of performing block matching within a pyramid in some embodiments of the present disclosure is shown.
[0015] Figure 5 An exemplary flowchart of a method for preprocessing a video in some embodiments of the present disclosure is shown.
[0016] Figure 6Exemplary schematic diagrams showing the coding reference structure in some embodiments of the present disclosure.
[0017] Figure 7 Exemplary schematic diagrams showing the loop partitioning of the macroblocks of the first filtered frame in some embodiments of the present disclosure.
[0018] Figure 8 Exemplary schematic diagrams showing the partitioning of the first filtered frame in some embodiments of the present disclosure.
[0019] Figure 9 Block diagram showing the hardware configuration of a device in which embodiments of the present disclosure may be implemented. Detailed implementation manners
[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0021] It should be understood that the terms "including" and "comprising" as used in the specification and claims of the present invention indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0022] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification and claims of the present invention, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms. It should be further understood that the term " / and / " as used in the specification and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0023] As used in the specification and claims of the present invention, the term "if" may be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0024] Next, the detailed implementation manners of the present invention will be described in detail in conjunction with the accompanying drawings.
[0025] Figure 1 An exemplary schematic diagram showing the video compression process in some embodiments of the present disclosure is presented. The dynamic images presented in the video are essentially a series of static pictures; generally speaking, when 24 pictures are continuously played per second, the human eye cannot notice the lag between the pictures, thus perceiving the continuously played pictures as a dynamic image. In a video, the playback rate of the pictures is called the frame rate (Frames Per Second, FPS), and some other common frame rates include 25 FPS, 50 FPS, and 60 FPS. Each picture (or each frame) consists of many pixel points. Some common video resolutions are 1280*720 or 1920*1080, and the corresponding number of pixel points is 1280*720 and 1920*1080 respectively. For each pixel point, the color of the pixel point can be converted into a digital signal convenient for storage according to the adopted pixel coding format. Pixel coding formats include: Red-Green-Blue (RGB) format, Red-Green-Blue-Alpha (RGBA) format, Luminance-Chrominance (YUV) format; the YUV format can be further divided into sub-formats including YUV444, YUV422, and YUV420 according to the sampling method. As Figure 1 shown, the video source generates an original video file through shooting or production; the original video file usually has a large amount of pixel point data. In order to reduce the bitrate during the storage and transmission of the video file, an encoder is often used to convert the original video file into an encoded file with a smaller amount of data; after the encoded file is received at the remote end, the byte stream data in the encoded file is read through a decoder to obtain a playable video.
[0026] Generally speaking, the video encoding process reduces the redundant information in the original video file through the prediction encoding step, and then obtains the byte stream data that can be stored or transmitted at a lower bitrate through transformation, quantization, entropy encoding, and post-processing. Prediction encoding includes Intra-Coded and Inter-Coded. Intra-Coded encodes a single frame of the video without referring to other frames, thereby removing the redundant information in the encoded frame that can be predicted by some pixel points; the frame using Intra-Coded is called an I-frame. Inter-Coded encodes a frame of the video by referring to the previous frame and / or the subsequent frame, thereby removing the redundant information in the encoded frame that can be predicted by other frames; among them, the frame that only refers to the previous frame for encoding is called a P-frame, and the frame that refers to both the previous frame and the subsequent frame for encoding is called a B-frame.
[0027] Since random noise is introduced during the recording and generation of videos, the noise not only reduces the redundant information available in the prediction coding step and decreases the video compression ratio, but also causes artifacts to appear in the video frames. The preprocessing step reduces the noise doped in the video through smoothing between frames, thereby improving the video compression ratio and reducing artifacts in the video. Motion Compensation Temporal Filtering (MCTF) is a preprocessing method for video compression. Its main idea is to use the frames near the filtered frame as reference frames and adjust the filtered frame to make the redundant information between the filtered frame and the reference frames richer and the changes smoother.
[0028] Figure 2 FIG. shows an exemplary schematic diagram of MCTF processing of a filtered frame based on one reference frame in some embodiments of the present disclosure. As Figure 2 shown, the MCTF algorithm mainly includes four steps: filtered frame partitioning, motion estimation, motion compensation, and temporal filtering. First, the filtered frame is divided into multiple filtered frame macro-blocks (Macro-Block, MacB). Then, motion estimation is performed between the filtered frame and the reference frame, that is, for at least one filtered frame macro-block, a matching block (Match-Block, MatB) matching the filtered frame macro-block is searched in the reference frame using a matching block search algorithm, and a motion vector (Motion Vector, MV) is calculated between the filtered frame macro-block and the matching block. The motion vector represents the relative position between the filtered frame macro-block and the matching block. After motion estimation, the matching block in the reference frame can be moved according to the motion vector to obtain a motion compensated image, and each matching block in the motion compensated image can also be called a compensated block. Fusing the motion compensated image with the filtered frame image (for example, pixel-by-pixel weighted summation) gives the filtered image of MCTF processing of the filtered frame based on one reference frame.
[0029] It can be understood that MCTF processing of a filtered frame can also be based on multiple reference frames. Specifically, according to the Figure 2 processing flow shown, a motion compensated image is generated for the filtered frame based on each reference frame, thereby obtaining multiple motion compensated images, and then the multiple motion compensated images are fused with the filtered frame image (for example, element-by-element weighted summation) to obtain the filtered image of MCTF processing of the filtered frame based on multiple reference frames.
[0030] As can be understood from the foregoing, MCTF performs motion estimation through a block matching search algorithm to obtain the motion of blocks in a picture between a filtered frame and a reference frame, moves the blocks in the reference frame according to the motion vector to obtain a motion-compensated image, and the filtered image obtained by fusing the motion-compensated image and the filtered frame image plays a smoothing role between the filtered frame and the reference frame. The block matching search algorithm directly affects the estimation result of the motion between the filtered frame and the reference frame. Performing temporal filtering based on an accurate motion estimation result can achieve relatively good filtering effects, including but not limited to: reducing the noise in the video, making the redundant information between the filtered frame and the reference frame richer and the changes smoother, thereby improving the compression ratio of the video and reducing artifacts in the video. On the contrary, performing temporal filtering based on an incorrect motion estimation result will lead to poor filtering effects.
[0031] Traditional MCTF algorithms use a pyramid block matching algorithm to search for matching blocks for the macroblocks of the filtered frame; among them, the sizes of the macroblocks of the filtered frame include but are not limited to: 64 pixels * 64 pixels, 64 pixels * 32 pixels, 32 pixels * 32 pixels. Figure 3 An exemplary schematic diagram of the pyramid block matching algorithm in some embodiments of the present disclosure is shown. The pixels of the video frame can be encoded in the YUV format, and then the Y component images in the filtered frame and the reference frame are respectively taken for pyramid block matching. As Figure 3 shown, the original image of the filtered frame undergoes two downsamplings to form three levels of filtered frame pyramids of L 0 、L 1 、L 2 , and the three levels of images are denoted as f 0 、f 1 、f 2 ; the original image of the reference frame undergoes two downsamplings to form three levels of reference frame pyramids of L 0 、L 1 、L 2 , and the three levels of images are denoted as r 0 、r 1 、r 2 ; it can be understood that the bottom-level f 0 、r 0 are the original filtered frame image and the original reference frame image respectively. In some embodiments, the original image at the lower level of the pyramid can be downsampled according to a set scaling factor F to obtain the downsampled image at the upper level of the pyramid; for example, when the scaling factor is set to 2, the length and width of the downsampled image are 1 / 2 of the length and width of the original image respectively.
[0032] Continuing Figure 3 , after the pyramid is constructed, locate in f 1 、f 2 the ones corresponding to f 0The divided blocks MacB1 and MacB2 corresponding to the filtered frame macroblock MacB0 in; for example: project MacB0 onto f 1 and then scale the projection according to the scaling factor, that is, scale the length and width of the projection to 1 / F of the original to obtain MacB1; project MacB1 onto f 2 and then scale the projection according to the scaling factor, that is, scale the length and width of the projection to 1 / F of the original to obtain MacB2. Then, start performing block matching from the top layer of the pyramid; amplify the motion vectors obtained for each layer according to the scaling factor to obtain adjustment vectors; perform block matching for the next layer according to the adjustment vectors; until the block matching is completed at the bottom layer of the pyramid, find the block MatB0 that matches MacB0 and calculate the motion vector based on the positions of MacB0 and MatB0, and the pyramid block matching algorithm ends.
[0033] Figure 4 shows an exemplary schematic diagram of performing block matching within a pyramid layer in some embodiments of the present disclosure. In Figure 4 , f represents each layer image f 0 , f 1 , f 2 in the filtered frame pyramid, and r represents each layer image r 0 , r 1 , r 2 in the reference frame pyramid; hereinafter, f will be denoted as the filtered layer image and r will be denoted as the reference layer image. In some embodiments, according to the position of the filtered layer macroblock MacB in the filtered layer image and the adjustment vector, a search window is determined in the reference layer image, where the adjustment vector is obtained by amplifying the motion vector in the previous layer according to the scaling factor; for example, the motion vector in the previous layer is (1, 1) and the scaling factor is 2, and multiplying the motion vector in the previous layer by the scaling factor gives the adjustment vector as (2, 2); determining the search window in the reference layer image according to the position of the filtered layer macroblock and the adjustment vector can be: project the filtered layer macroblock into the reference layer image; the two components of the adjustment vector respectively correspond to the length dimension and the width dimension of the image, and move the projection along the length dimension and the width dimension in the reference layer image according to the adjustment vector to obtain an adjusted projection; with the adjusted projection as the center, move M 1 pixel points upward and downward on the reference layer image, and move M 2 pixel points left and right on the reference layer image, thereby obtaining the search window. Such a search window contains (2M 1 +1)(2M 2 +1) reference layer macroblocks to be compared, and it is necessary to calculate the matching degree between the filtered layer macroblock and each reference layer macroblock according to a preset matching criterion, and then use the reference layer macroblock with the highest matching degree in the search window as the matching block (MatB) corresponding to the filtered frame layer block.
[0034] In some embodiments, the mean absolute difference between the pixel values of the corresponding pixel points of the reference layer macroblock and the filtering layer macroblock is calculated based on the following formula (1), and the matching degree between the reference layer macroblock and the filtering layer macroblock is evaluated according to the mean absolute difference of the pixel values.
[0035]
[0036] Wherein, MAD represents the mean absolute difference (Mean Absolute Difference), M represents the length of the filtering layer macroblock, N represents the width of the filtering layer macroblock, and C ij represents the pixel value of the pixel point in the filtering layer macroblock, and R ij represents the pixel value of the pixel point in the reference layer macroblock.
[0037] In some embodiments, the mean squared error between the pixel values of the corresponding pixel points of the reference layer macroblock and the filtering layer macroblock is calculated based on formula (2), and the matching degree between the reference layer macroblock and the filtering layer macroblock is evaluated according to the mean squared error of the pixel values.
[0038]
[0039] Wherein, MSE represents the mean squared error (Mean Squared Error). It can be understood that the smaller the mean absolute difference and the mean squared error of the pixel values, the higher the matching degree between the reference layer macroblock and the filtering layer macroblock.
[0040] Continue Figure 4 After the corresponding matching block is found for the filtering layer macroblock, the motion vector can be calculated according to the positions of the filtering layer macroblock and the matching block. In some embodiments, the motion vector is calculated according to the coordinates of the upper left corner, upper right corner, lower left corner or lower right corner pixel points of the filtering layer macroblock and the matching block; for example, the difference between the coordinates of the upper left corner pixel points of the filtering layer macroblock and the matching block can be calculated and used as the motion vector; exemplarily, if the coordinates of the upper left corner pixel point of the filtering layer macroblock are (32, 32) and the coordinates of the upper left corner pixel point of the matching block are (30, 30), then the motion vector between the filtering layer macroblock and the matching block is (2, 2).
[0041] It can be understood that the pyramid block matching algorithm needs to perform downsampling and search for motion vectors layer by layer. The traditional pyramid block matching algorithm assumes that all pixel points in a filtered frame macroblock have the same motion, and thus performs motion estimation between the filtered frame and the reference frame based on a fixed block size. However, performing motion estimation based on a fixed block size cannot adapt to the complex pixel motion in the video: relatively large motion units will be divided into multiple blocks, and the motion vectors calculated for each of these blocks may not be consistent, resulting in the splitting of the boundaries of large motion units; in addition, the motion of pixel points within the same divided block is not the same. Some of the pixel points can move in the first motion mode, while some other pixel points can move in the second motion mode. Applying the same motion vector to all pixel points within a block will result in the incorrect estimation of the motion of some pixel points. Therefore, using a fixed block size to perform motion estimation between the filtered frame and the reference frame easily leads to incorrect motion estimation results, resulting in artifacts and blocking effects in the motion compensated image, and further reducing the effect of temporal filtering. In view of this, the present disclosure discloses a method for preprocessing a video, which performs motion estimation based on variable-sized blocks between the filtered frame and the reference frame, improves the temporal filtering effect, and further improves the video coding quality.
[0042] Figure 5 FIG. shows an exemplary flowchart of a method for preprocessing a video in some embodiments of the present disclosure. As Figure 5 shown, the method includes: Step 501, dividing a filtered frame into blocks to obtain first filtered frame macroblocks having a first size; Step 502, dividing at least one of the first filtered frame macroblocks into blocks to obtain second filtered frame macroblocks having a second size; Step 503, determining a matching block that matches the second filtered frame macroblock based on a reference frame associated with the filtered frame; Step 504, updating the second filtered frame macroblock according to the matching block.
[0043] For Step 501, it can be understood that the filtered frame can be divided into multiple first filtered frame macroblocks according to a preset first size. The size of the filtered frame macroblock is determined by the length and width. The length and width of the first size can be multiples of 4 pixel points. The first size includes but is not limited to: 64 pixels * 64 pixels, 64 pixels * 32 pixels, 32 pixels * 32 pixels.
[0044] For step 502, it can be understood that each first filtered frame macroblock can be divided into multiple second filtered frame macroblocks according to a preset second size, where the length of the second size is less than the length of the first size, or the width of the second size is less than the width of the first size, or the length and width of the second size are respectively less than the length and width of the first size. In some embodiments, the second size is set according to the first size; exemplarily, if the first size is 64 pixels * 64 pixels, the second size can be 64 pixels * 32 pixels, 32 pixels * 32 pixels, or 32 pixels * 16 pixels.
[0045] For step 503, it can be understood that the reference frame associated with the filtered frame refers to the nearby frame referred to when performing filtering processing on the filtered frame. Based on the matching block search algorithm, a matching block in the reference frame can be matched to the second filtered frame macroblock, where the matching block has the same size as the second filtered frame macroblock and approximate pixel values.
[0046] In some embodiments, updating the second filtered frame macroblock according to the matching block may include: performing weighted summation of the pixel values in the matching block and the pixel values in the second filtered frame macroblock pixel by pixel. The operation of updating the second filtered frame macroblock according to the matching block is equivalent to performing temporal filtering on the second filtered frame macroblock; it can be understood that when processing the entire filtered frame image, a corresponding matching block can also be searched for the first filtered frame macroblock in the filtered frame, and then temporal filtering is performed on the first filtered frame based on the matching block corresponding to the first filtered frame.
[0047] The method for preprocessing a video disclosed in this disclosure, after dividing the filtered frame into blocks according to the first size to obtain first filtered frame macroblocks, further divides a part or all of the first filtered frame macroblocks according to the second size to obtain second filtered frame macroblocks; thus, there are relatively large first filtered frame macroblocks and relatively small second filtered frame macroblocks in the filtered frame. The relatively large filtered frame macroblocks can better cover larger motion units in the image, and the relatively small filtered frame macroblocks can perform motion estimation on pixel points in the image with finer granularity; in this way, motion estimation is performed between the filtered frame and the reference frame through variable-size block division, improving the temporal filtering effect and video coding quality.
[0048] In some embodiments, the method for preprocessing a video disclosed in this disclosure further includes: traversing multiple frames of the video; for the current frame, determining whether the current frame is a filtered frame according to the coding reference structure, where the coding reference structure is used to determine the inter-frame reference relationship of video coding.
[0049] It can be understood that the predictive coding of B-frames utilizes future frames in the presentation order. To allow the coding order to be different from the presentation order, video frames are divided into Groups of Pictures (GOPs). A GOP includes multiple consecutively presented frames in the video, and the number of frames included in a GOP is referred to as the GOP size (GOPSize); in some embodiments, the GOPSize is set according to the frame rate; in other embodiments, the GOPSize can be set to 8, 12, or other multiples of 4. The GOP also specifies the coding reference structure for predictive coding, and the coding reference structure is used to determine the frames that need to be referred to for predictive coding of each frame in the GOP.
[0050] Figure 6 FIG. shows an exemplary schematic diagram of the coding reference structure in some embodiments of the present disclosure. In Figure 6 it, □ represents each frame, and the presentation order of the frames is from left to right; the coding order of the frames is shown by the numbers in □, where the frame corresponding to the number 0 can be the first frame of the video or the last frame of the previous GOP. Therefore Figure 6 the GOP shown includes the frames represented by the numbers 1-8. Further, in Figure 6 it, the arrows represent the reference relationships for predictive coding, that is, the frame at the start end of the arrow needs to refer to the frame at the end end of the arrow for predictive coding. Thus, 4 layers of coding hierarchies are formed within the GOP. Among them, the frames in layer 0 are first subjected to predictive coding. After the predictive coding of the frames in layer 0 is completed, the frames in layer 1 are subjected to predictive coding, and then the frames in layers 2 and 3 are sequentially subjected to predictive coding. Figure 6 The coding reference structure shown is only for ease of explanation. For those skilled in the art, different coding reference structures can be designed as needed, and the present disclosure places no restrictions thereon.
[0051] It can be understood that in the preprocessing step, it is possible to determine whether the current frame is a filtering frame according to the coding reference structure. In some embodiments, if the current frame is referred to by other frames in the coding reference structure, then the current frame is a filtering frame; otherwise, the current frame is not a filtering frame. In other embodiments, it is possible to determine whether the current frame is a filtering frame according to the number of times the current frame is referred to in the coding reference structure. If the number of times the current frame is referred to is greater than the set reference number threshold, then the current frame is considered a filtering frame; otherwise, the current frame is not a filtering frame; where the number of frames in the coding reference structure that refer to the current frame is the number of times the current frame is referred to.
[0052] It can be understood that the "reference frame associated with the filtering frame" described in step 503 above can also be determined according to the coding reference structure. In some embodiments, a method for preprocessing a video disclosed in the present disclosure further includes: for the current filtering frame, determining the reference frame associated with the current filtering frame according to the coding reference structure.
[0053] In some embodiments, partitioning at least one first filtered frame macroblock to obtain a second filtered frame macroblock having a second size includes: traversing a plurality of first filtered frame macroblocks; for a current first filtered frame macroblock, in response to the current first filtered frame macroblock satisfying a preset partitioning condition, partitioning the current first filtered frame macroblock to obtain a second filtered frame macroblock.
[0054] It can be understood that the video contains complex pixel motion situations; some pixel points can share the same motion state. Partitioning as many pixel points sharing the same motion state into the same block can reduce the number of blocks required for motion estimation, thereby reducing the computational amount of motion estimation and improving the running speed of the motion estimation algorithm; while the motion states of pixel points are not all the same, providing more partitions can obtain more accurate motion estimation results, but at the same time reduces the running speed of the motion estimation algorithm. Therefore, the operation of partitioning the filtered frame involves the problem of balancing the running speed of the motion estimation algorithm and the accuracy of the motion estimation result. In the embodiments of the present disclosure, if the current first filtered frame macroblock satisfies the preset partitioning condition, continue to partition the current first filtered frame macroblock to obtain smaller-sized partitions, so as to perform a more fine-grained and accurate motion estimation on the current filtered frame macroblock; if the current first filtered frame macroblock does not satisfy the preset partitioning condition, perform a faster motion estimation with the current size; thereby adaptively performing a partitioning operation on the filtered frame macroblock to accurately perform motion estimation on the filtered frame at a faster running speed.
[0055] In some embodiments, the method for preprocessing a video disclosed in the present disclosure further includes: traversing a plurality of first filtered frame macroblocks; for a current first filtered frame macroblock, in response to the current first filtered frame macroblock satisfying a preset partitioning condition, partitioning the current first filtered frame macroblock cyclically until the sub-blocks obtained by partitioning the current first filtered frame macroblock do not satisfy the preset partitioning condition.
[0056] Figure 7 Shows an exemplary schematic diagram of cyclically partitioning a first filtered frame macroblock in some embodiments of the present disclosure. As Figure 7As shown, the size of the current first filtered frame macroblock is 64*64. After determining that the current first filtered frame macroblock meets the preset partitioning condition, the first filtered frame macroblock is partitioned into two first sub-blocks of 64*32. The lower first sub-block does not meet the preset partitioning condition, so it is no longer partitioned; the upper first sub-block meets the preset partitioning condition, so it is further partitioned to obtain two second sub-blocks of 32*32. The left second sub-block does not meet the preset partitioning condition, so it is no longer partitioned; the right second sub-block meets the preset partitioning condition, so it is further partitioned to obtain two third sub-blocks of 32*16. The lower third sub-block does not meet the preset partitioning condition, so it is no longer partitioned; the upper third sub-block meets the preset partitioning condition, so it is further partitioned to obtain two fourth sub-blocks of 16*16, and neither of the two fourth sub-blocks meets the preset partitioning condition. Thus, the first filtered frame macroblock is partitioned into one first sub-block of 64*32, one second sub-block of 32*32, one third sub-block of 32*16, and two fourth sub-blocks of 16*16, and all sub-blocks do not meet the preset partitioning condition, so the partitioning operation on the first filtered frame macroblock is stopped.
[0057] It can be understood that there may be multiple motion units in the filtered frame. For example, the sky in the image may include motion units such as white clouds and a blue background, and the people in the image may include motion units such as heads, arms, and legs. Generally, the motion states of pixel points from the same motion unit are relatively close, while there are differences in the motion states of pixel points from different motion units. In the embodiments of the present disclosure, it is determined whether to perform a partitioning operation on the filtered frame macroblock according to whether the filtered frame macroblock includes different motion units: when the filtered frame macroblock includes different motion units, a partitioning operation is performed on the filtered frame macroblock; when the filtered frame macroblock has only a single motion unit, the partitioning operation on the filtered frame macroblock is no longer performed.
[0058] In some embodiments, the preset partitioning condition includes a preset texture condition and a preset size condition; the preset texture condition includes: the horizontal gradient value of the pixel values in the filtered frame macroblock is greater than a predetermined horizontal gradient threshold; the vertical gradient value of the pixel values in the filtered frame macroblock is greater than a predetermined vertical gradient threshold; the gradient value of the pixel values in the filtered frame macroblock is greater than a predetermined gradient threshold; the variance of the pixel values in the filtered frame macroblock is greater than a predetermined variance threshold; or the complexity level value of the filtered frame macroblock is greater than a predetermined complexity level threshold; the preset size condition includes: the length or width of the filtered frame macroblock to be partitioned is not less than 4.
[0059] It can be understood that the filtered frame macroblock includes a first filtered frame macroblock and sub-blocks obtained by dividing the first filtered frame macroblock. Whether there are different motion units in the current filtered frame macroblock can be determined according to the pixel values of the pixel points in the filtered frame macroblock: if the difference in pixel values between the pixel points in the filtered frame macroblock is relatively significant, there may be multiple motion units; if the difference in pixel values between the pixel points in the filtered frame macroblock is relatively weak, there may be only a single motion unit. Texture information such as the horizontal gradient value, vertical gradient value, gradient value (i.e., calculated by combining the gradient values in the horizontal and vertical directions), variance, and the complexity level value of the filtered frame macroblock can be used to measure the degree of difference in pixel values between the pixel points in the filtered frame macroblock.
[0060] Denote the coordinates of the pixel point at the lower left corner of the filtered frame macroblock as (0, 0). For a pixel point with coordinates (x, y), its pixel value is I(x, y), and its horizontal gradient value can be: G x (x,y) = |I(x + 1, y) - I(x, y)|, and its vertical gradient value can be:
[0061] G y (x,y) = |I(x, y + 1) - I(x, y)|, and its gradient value can be: In some embodiments, the horizontal gradient values are calculated for all pixel points in the filtered frame macroblock. When there is a pixel point whose horizontal gradient value is greater than a predetermined horizontal gradient threshold, it is considered that the preset texture condition is satisfied; in some embodiments, the vertical gradient values are calculated for all pixel points in the filtered frame macroblock. When there is a pixel point whose vertical gradient value is greater than a predetermined vertical gradient threshold, it is considered that the preset texture condition is satisfied; in some embodiments, the gradient values are calculated for all pixel points in the filtered frame macroblock. When there is a pixel point whose gradient value is greater than a predetermined gradient threshold, it is considered that the preset texture condition is satisfied.
[0062] In some embodiments, the total gradient energy of the filtered frame macroblock is calculated according to the following formula (3):
[0063]
[0064] where M is the number of rows of the filtered frame macroblock, and N is the number of columns of the filtered frame macroblock. In these embodiments, the complexity level value of the filtered frame macroblock is calculated according to the following formula (4):
[0065]
[0066] where E max and E min are respectively the maximum and minimum values of the total gradient energy in all filtered frame macroblocks, and BLCV is the complexity level value of the current filtered frame macroblock.
[0067] In some embodiments, the mean value of the pixel values of the pixels in the filtered frame macroblock is calculated according to the following formula (5):
[0068]
[0069] where μ represents the mean value. In these embodiments, the variance of the pixel values of the pixels in the filtered frame macroblock is calculated according to the following formula (6):
[0070]
[0071] where σ 2 represents the variance.
[0072] It can be understood that when the length or width of the filtered frame macroblock is less than 4, it is not necessary to continue to divide it into blocks, so it is considered that the preset size condition is not met and it is no longer divided into blocks.
[0073] In some embodiments, dividing at least one first filtered frame macroblock into blocks to obtain a second filtered frame macroblock with a second size includes: horizontally and evenly dividing the first filtered frame macroblock to obtain two second filtered frame macroblocks; vertically and evenly dividing the first filtered frame macroblock to obtain two second filtered frame macroblocks; or grid - evenly dividing the first filtered frame macroblock to obtain four second filtered frame macroblocks.
[0074] Figure 8 FIG. shows an exemplary schematic diagram of dividing the first filtered frame in some embodiments of the present disclosure. It can be understood that when both the length and width of the first filtered frame macroblock are even numbers, the first filtered frame macroblock can be horizontally divided into two second filtered frame macroblocks of the same size by a horizontal line, vertically divided into two second filtered frame macroblocks of the same size by a vertical line, or grid - divided into four second filtered frame macroblocks of the same size by a vertical line and a horizontal line. Figure 8 In FIG., the size of the first filtered frame macroblock is taken as 64*64 as an example to illustrate the division method. As Figure 8 shown, by horizontally dividing or vertically dividing the 64*64 first filtered frame macroblock, two 64*32 second filtered frame macroblocks can be obtained; while by grid - dividing the 64*64 first filtered frame macroblock, four 32*32 second filtered frame macroblocks can be obtained. It can be understood that for the sub - blocks obtained by dividing the first filtered frame macroblock, the above - mentioned division method can also be used to divide them to obtain new sub - blocks, which will not be elaborated here.
[0075] In some embodiments, after determining a matching block that matches a second filtered frame macroblock based on a reference frame associated with a filtered frame, the method includes: determining a motion vector based on the positions of the second filtered frame macroblock and the matching block, for determining the relative positions of the second filtered frame macroblock and the matching block.
[0076] It can be understood that based on a matching block search algorithm, a matching block that matches the second filtered frame macroblock can be searched in the reference frame, and a motion vector can be calculated based on the positions of the second filtered frame macroblock and the matching block. As described above in conjunction with Figure 3 , Figure 4 the method of searching for a matching block for a filtered frame macroblock and calculating a motion vector using a pyramid block matching algorithm has been explained, and will not be elaborated here. It should be noted that for those skilled in the art, different matching block search algorithms can be selected as needed, including but not limited to: the Three-Step Search algorithm, the Four-Step Search algorithm, and this disclosure does not make any restrictions thereon.
[0077] In some embodiments, updating the second filtered frame macroblock according to the matching block includes: setting a filtering weight according to the temporal distance between the filtered frame and the reference frame; performing weighted summation on the pixel points in the matching block and the second filtered frame macroblock according to the filtering weight.
[0078] It can be understood that the temporal distance between the filtered frame and the reference frame can be the temporal distance during playback between the two, and the temporal distance can measure the reference significance of the reference frame for filtering the filtered frame: the smaller the temporal distance, the greater the reference significance of the reference frame for the filtered frame, and at this time, a higher filtering weight can be set; the greater the temporal distance, the smaller the reference significance of the reference frame for the filtered frame, and at this time, a lower filtering weight can be set. In some embodiments, based on the following formula (7), weighted summation is performed on the pixel points in the matching block and the second filtered frame macroblock according to the filtering weight:
[0079]
[0080] where, I f represents the pixel value of the pixel point in the second filtered frame macroblock after weighted summation; I 0 represents the pixel value of the pixel point in the second filtered frame macroblock before weighted summation; w r represents the filtering weight; I r represents the pixel value of the pixel point in the matching block.
[0081] In some other embodiments, for a filtered frame with multiple reference frames, matching blocks can be obtained separately according to each reference frame, so as to obtain multiple matching blocks. In these embodiments, based on the following formula (8), the pixel points in the multiple matching blocks and the macroblock of the second filtered frame are weighted and summed according to the filtering weights:
[0082]
[0083] where N represents the number of reference frames; w r (i) represents the filtering weight corresponding to the i-th reference frame; I r (i) represents the pixel value of the pixel points in the compensation block corresponding to the i-th reference frame.
[0084] It can be understood that for the macroblocks (including the first filtered frame macroblocks and the sub-blocks obtained by dividing the first filtered frame macroblocks) in the filtered frame, each macroblock can search for a matching block in the reference frame and is processed by the method of performing temporal filtering on the above-mentioned second filtered frame macroblock, which will not be elaborated here.
[0085] Furthermore, the present disclosure also discloses a computer-readable storage medium, in which program instructions are stored, and the program instructions are adapted to be loaded and executed by a processor to perform the methods described in the foregoing embodiments of the present disclosure.
[0086] The present disclosure also discloses a device for preprocessing video, including: a processor configured to execute program instructions; and a memory configured to store program instructions, and when the program instructions are loaded and executed by the processor, the device is caused to execute the methods described in the foregoing embodiments of the present disclosure.
[0087] Figure 9 The block diagram showing the hardware configuration of the device 900 that can implement the embodiments of the present disclosure is as follows. As Figure 9 shown, the device 900 may include a processor 901 and a memory 902. The processor therein is configured to execute program instructions, and the memory is configured to store program instructions, and when the program instructions are loaded and executed by the processor, the device is caused to execute the method for preprocessing video described in any one of the above embodiments. In Figure 9 the device 900, only the constituent elements related to this embodiment are shown. Therefore, it is obvious to those of ordinary skill in the art that: the device 900 may further include common constituent elements different from those shown in Figure 9 The specific functions implemented by the memory 902 and the processor 901 of the device 900 provided in the embodiments of this specification can be explained in contrast to the foregoing embodiments in this specification and can achieve the technical effects of the foregoing embodiments, which will not be elaborated here.
[0088] Device 900 may correspond to a computing device having various processing functions. For example, device 900 may be implemented as various types of devices, such as a personal computer (PC), a server device, a mobile device, and the like.
[0089] Processor 901 may control the operation of device 900. For example, processor 901 may be implemented by a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), an artificial intelligence processor chip (IPU), etc. provided in device 900. However, the present invention is not limited thereto. In this embodiment, processor 901 may be implemented in any suitable manner. For example, processor 901 may take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and a form embedded microcontroller, and so on.
[0090] The memory 902 can be used to store various data and instructions processed in the device 900. For example, the memory 902 can store the processed data and the data to be processed in the device 900. The memory 902 can store the data that has been processed or is to be processed by the processor 901, such as filtered frame image data and reference frame image data. In addition, the memory 902 can store applications, drivers, etc. to be driven by the device 900. For example, the memory 902 can store various programs related to the method for preprocessing video to be executed by the processor 901. The memory 902 can be a DRAM, but the present invention is not limited thereto. The memory 902 can include at least one of volatile memory or non-volatile memory. The non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, phase change RAM (PRAM), magnetic RAM (MRAM), resistive RAM (RRAM), ferroelectric RAM (FRAM), etc. The volatile memory can include dynamic RAM (DRAM), static RAM (SRAM), synchronous DRAM (SDRAM), PRAM, MRAM, RRAM, ferroelectric RAM (FeRAM), etc. In some embodiments, the memory 902 can include at least one of a hard disk drive (HDD), a solid state drive (SSD), a high density flash (CF) card, a secure digital (SD) card, a micro secure digital (Micro-SD) card, a mini secure digital (Mini-SD) card, an extreme digital (xD) card, caches, or a memory stick.
[0091] In summary, the specific functions implemented by the computer-readable storage medium and the device for preprocessing video provided in the embodiments of this specification can be explained in contrast to the foregoing embodiments in this specification and can achieve the technical effects of the foregoing embodiments, which will not be elaborated herein.
[0092] It should be noted that for the sake of brevity, some methods and their embodiments of the present invention are expressed as a series of actions and their combinations. However, those skilled in the art can understand that the solution of the present invention is not limited by the order of the described actions. Therefore, based on the disclosure or teachings of the present invention, those skilled in the art can understand that some of the steps can be executed in other orders or simultaneously. Further, those skilled in the art can understand that the embodiments described in the present invention can be regarded as optional embodiments, that is, the actions or modules involved are not necessarily required for the implementation of a certain or some solutions of the present invention. Additionally, according to the differences in the solutions, the descriptions of some embodiments of the present invention also have different emphases. In view of this, those skilled in the art can understand that the parts not detailed in a certain embodiment of the present invention can also refer to the relevant descriptions of other embodiments.
Claims
1. A method for preprocessing a video, comprising: Dividing the filtered frame into blocks to obtain first filtered frame macroblocks having a first size; Dividing at least one first filtered frame macroblock into blocks to obtain a second filtered frame macroblock having a second size; determining a matching block that matches a macroblock of the second filtered frame based on a reference frame associated with the filtered frame; The second filtered frame macroblock is updated according to the matching block.
2. The method according to claim 1, further comprising: Iterate over multiple frames of the video; For a current frame, whether the current frame is a filtered frame is determined according to a coding reference structure, wherein the coding reference structure is used to determine an inter-frame reference relationship of video coding.
3. The method according to claim 1, wherein: Dividing at least one first filtered frame macroblock into blocks to obtain a second filtered frame macroblock having a second size, comprising: Traversing a plurality of first filtering frame macroblocks; For the current first filtered frame macroblock, in response to the current first filtered frame macroblock satisfying a preset division condition, the current first filtered frame macroblock is divided into blocks to obtain a second filtered frame macroblock.
4. The method according to claim 1, further comprising: Traversing a plurality of first filtering frame macroblocks; For the current first filtered frame macroblock, in response to the current first filtered frame macroblock satisfying a preset division condition, the current first filtered frame macroblock is cyclically divided into blocks until the sub-blocks obtained by dividing the current first filtered frame macroblock do not satisfy the preset division condition.
5. The method according to claim 3 or 4, wherein: The preset division conditions include preset texture conditions and preset size conditions; The preset texture conditions include: The horizontal gradient value of the pixel value in the macroblock of the filtered frame is greater than a predetermined horizontal gradient threshold; The vertical gradient value of the pixel value in the filtered frame macroblock is greater than a predetermined vertical gradient threshold; The gradient value of the pixel value in the filtered frame macroblock is greater than a predetermined gradient threshold; The variance of the pixel values in the filtered frame macroblock is greater than a predetermined variance threshold; or The complexity level value of the filtered frame macroblock is greater than a predetermined complexity level threshold; The preset size condition includes: the length or width of the filter frame macroblock to be divided is not less than 4.
6. The method according to claim 1, wherein: Dividing at least one first filtered frame macroblock into blocks to obtain a second filtered frame macroblock having a second size, comprising: Dividing the first filter frame macroblock horizontally and evenly to obtain two second filter frame macroblocks; Divide the first filter frame macroblock vertically and evenly to obtain two second filter frame macroblocks; or The first filter frame macroblock is gridded and evenly divided to obtain four second filter frame macroblocks.
7. The method according to claim 1, wherein: After determining a matching block matching the second filtered frame macroblock based on a reference frame associated with the filtered frame, comprising: According to the positions of the second filtered frame macroblock and the matching block, a motion vector is determined to determine the relative positions of the second filtered frame macroblock and the matching block.
8. The method according to claim 1, wherein: Updating the second filtered frame macroblock according to the matching block comprises: Setting a filtering weight according to a temporal distance between the filtering frame and the reference frame; According to the filtering weight, a weighted sum is performed on the pixel points in the matching block and the second filtering frame macroblock.
9. A computer-readable storage medium storing program instructions, wherein the program instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 8.
10. A device for preprocessing a video, comprising: a processor configured to execute program instructions; as well as A memory configured to store the program instructions, which, when loaded and executed by the processor, causes the apparatus to perform the method according to any one of claims 1 to 8.
Citation Information
Cited By
Image filtering processing method, system and device and storage medium
CN120355581A