Method for preprocessing video and related product
By blocking and boundary compensation processing of filtered frames in video encoding, the problem of boundary artifacts in traditional motion compensation technology is solved, and higher quality video encoding is achieved.
Patent Information
- Application Number
- CN202411982496.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-30
AI Technical Summary
In the video encoding process, traditional motion compensation technology will lead to the emergence of boundary artifacts and reduce the quality of video encoding.
By blocking the filtered frames, the boundary compensation area of adjacent matching blocks is determined, and the current matching block is compensated, thereby updating the filtered frame macroblock and reducing boundary artifacts.
The motion information between macroblocks of adjacent filtered frames is effectively integrated, which reduces the block effect and edge sawtooth caused by blocking processing, and significantly improves the quality of video encoding.
Smart Images

Figure CN120075439A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the field of video encoding technologies. More specifically, this disclosure relates to a method and related products for preprocessing videos. Background Art
[0002] The video encoding process obtains byte stream data that can be stored or transmitted at a lower bit rate by reducing redundant information in the original video file. Since random noise is introduced during the recording and generation of videos, the noise not only reduces the available redundant information during the video encoding process and decreases the video compression ratio, but also causes artifacts to appear in the video frames. Motion Compensated Temporal Filter (MCTF) is a preprocessing method for video encoding. MCTF uses frames near the filtered frame (i.e., the frame that needs to be filtered using MCTF) as reference frames, and adjusts the filtered frame to make the redundant information between the filtered frame and the reference frames richer and the changes smoother.
[0003] In the framework of the block motion compensation technology of MCTF, each image block performs motion estimation and motion compensation independently. Traditional motion compensation moves the matching block according to the motion vector to construct the motion-compensated image; during the process of moving the matching block, the relative positions between the matching blocks change. When two matching blocks are placed at adjacent positions in the motion-compensated image, it will cause the appearance of boundary artifacts. Performing temporal filtering on the filtered frame image based on the motion-compensated image containing boundary artifacts will reduce the redundant information amount and the smoothness between the filtered frame and the reference frame, thereby reducing the quality of video encoding. In view of this, there is an urgent need to provide a video preprocessing solution to suppress boundary artifacts in the motion-compensated image and improve the quality of video encoding. Summary of the Invention
[0004] To at least solve one or more of the technical problems described in the above Background Art section, the present disclosure proposes the following technical solutions and multiple embodiments thereof.
[0005] In a first aspect, the present disclosure discloses a method for preprocessing a video, including: partitioning a filtered frame to obtain a first filtered frame macroblock and a second filtered frame macroblock adjacent to the first filtered frame macroblock; determining, based on a reference frame associated with the filtered frame: a first matching block that matches the first filtered frame macroblock and a second matching block that matches the second filtered frame macroblock; compensating the first matching block based on the boundary compensation region of the second matching block to obtain a compensated block; and updating the first filtered frame macroblock according to the compensated block.
[0006] In a second aspect, the present disclosure discloses a computer-readable storage medium storing program instructions adapted to be loaded and executed by a processor to perform the method according to the first aspect.
[0007] In a third aspect, the present disclosure discloses an apparatus for preprocessing a video, comprising: a processor configured to execute program instructions; and a memory configured to store program instructions, which when loaded and executed by the processor, cause the apparatus to perform the method according to the first aspect.
[0008] According to the method for preprocessing a video disclosed in the present disclosure, in a motion-compensated image, a current matching block is compensated by a boundary compensation region of an adjacent matching block, thereby performing a filtering operation on a boundary region of adjacent matching blocks in the motion-compensated image and suppressing boundary artifacts. The strategy adopted in the present disclosure helps to achieve effective fusion of motion information between adjacent filtered frame macroblocks and visually seamless connection, effectively reducing the blocking effect and edge sawtooth phenomenon caused by block processing. Compared with traditional temporal filtering techniques, the present disclosure proposes a boundary-enhanced temporal filtering technique, making the boundary between blocks in the motion-compensated image smoother, significantly reducing the blocking effect, and thus improving the filtering effect. Further, this method can increase the redundant information between the processed filtered frame and the reference frame and the smoothness therebetween, thereby improving the quality of video coding. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0010] Figure 1 An exemplary diagram showing a video compression process in some embodiments of the present disclosure is shown.
[0011] Figure 2 An exemplary diagram showing MCTF processing of a filtered frame based on one reference frame in some embodiments of the present disclosure is shown.
[0012] Figure 3 An exemplary diagram showing boundary artifacts generated by motion compensation in some embodiments of the present disclosure is shown.
[0013] Figure 4 An exemplary flowchart of a method for preprocessing a video in some embodiments of the present disclosure is shown.
[0014] Figure 5An exemplary schematic diagram showing compensation for a first matching block based on a boundary compensation region of a second matching block in some embodiments of the present disclosure.
[0015] Figure 6 An exemplary schematic diagram showing preprocessing of a filtering frame macroblock according to multiple reference frames in some embodiments of the present disclosure.
[0016] Figure 7 An exemplary schematic diagram showing an encoding reference structure in some embodiments of the present disclosure.
[0017] Figure 8 An exemplary schematic diagram showing a matching block search algorithm in some embodiments of the present disclosure.
[0018] Figure 9 An exemplary schematic diagram showing a motion-compensated image in some embodiments of the present disclosure.
[0019] Figure 10 A block diagram showing the hardware configuration of a device in which embodiments of the present disclosure can be implemented. Detailed implementation manners
[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] It should be understood that the terms "including" and "comprising" used in the specification and claims of the present invention indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0022] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification and claims of the present invention, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms. It should further be understood that the term " / and / " used in the specification and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0023] As used in this specification and the claims, the term "if" can be construed, depending on the context, as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be construed, depending on the context, to mean "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]".
[0024] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0025] Figure 1 An exemplary schematic diagram showing the video compression process in some embodiments of this disclosure is presented. The dynamic images presented in a video are essentially a series of static pictures; generally speaking, when 24 pictures are continuously played per second, the human eye cannot notice the jerks between the pictures, thus perceiving the continuously played pictures as a dynamic image. In a video, the playing rate of pictures is called the frame rate (Frames Per Second, FPS), and some other common frame rates include 25 FPS, 50 FPS, and 60 FPS. Each picture (or each frame) consists of many pixel points. Some common video resolutions are 1280*720 or 1920*1080, and the corresponding number of pixel points is 1280*720, 1920*1080. For each pixel point, the color of the pixel point can be converted into a digital signal convenient for storage according to the adopted pixel coding format. Pixel coding formats include: Red-Green-Blue (RGB) format, Red-Green-Blue-Alpha (RGBA) format, Luminance-Chrominance (YUV) format; the YUV format can be further divided into sub-formats including YUV444, YUV422, and YUV420 according to the sampling method. As Figure 1 shown, the video source generates an original video file through shooting or production; the original video file usually has a large amount of pixel point data. In order to reduce the bitrate during the storage and transmission of the video file, an encoder is often required to convert the original video file into an encoded file with a smaller amount of data; after the encoded file is received at the remote end, the byte stream data in the encoded file is read through a decoder to obtain a playable video.
[0026] Generally speaking, the video encoding process reduces redundant information in the original video file through prediction encoding steps, and then obtains byte stream data that can be stored or transmitted at a lower bit rate through transformation, quantization, entropy encoding, and post-processing. Prediction encoding includes intra-coded and inter-coded. Intra-coded encodes a single frame of the video without referring to other frames, thereby removing redundant information in the encoded frame that can be predicted by some pixel points; the frame using intra-coded is called an I-frame. Inter-coded encodes a single frame of the video by referring to previous frames and / or subsequent frames, thereby removing redundant information in the encoded frame that can be predicted by other frames; among them, the frame that only refers to previous frames for encoding is called a P-frame, and the frame that refers to both previous frames and subsequent frames for encoding is called a B-frame.
[0027] Since random noise is introduced during the recording and generation of video, the noise not only reduces the available redundant information in the prediction encoding step and decreases the video compression ratio, but also causes artifacts to appear in the video image. The preprocessing step reduces the noise doped in the video through smoothing processing between frames, thereby improving the video compression ratio and reducing artifacts in the video. Motion-compensated temporal filtering is a preprocessing method for video compression, and its main idea is to use the frames near the filtered frame as reference frames, and adjust the filtered frame to make the redundant information between the filtered frame and the reference frames richer and the changes smoother.
[0028] Figure 2 An exemplary schematic diagram of performing MCTF processing on a filtered frame based on one reference frame in some embodiments of the present disclosure is shown. As Figure 2 shown, the MCTF algorithm mainly includes four steps: filtered frame block division, motion estimation, motion compensation, and temporal filtering. First, the filtered frame is divided into multiple filtered frame macro-blocks (Macro-Block, MacB). Then, motion estimation is performed between the filtered frame and the reference frame, that is, for at least one filtered frame macro-block, a matching block (Match-Block, MatB) that matches the filtered frame macro-block is searched in the reference frame, and a motion vector (MotionVector, MV) is calculated between the filtered frame macro-block and the matching block, and this motion vector represents the relative position between the filtered frame macro-block and the matching block; exemplarily, for the filtered frame macro-block MacBi, find the MatBi that matches MacBi in the reference frame, and the motion vector MV can be calculated by taking the difference between the coordinates corresponding to the top-left pixels of MacBi and MatBi i. After motion estimation, the matching blocks in the reference frame can be shifted according to the motion vectors to obtain a motion-compensated image, and each matching block in the motion-compensated image can also be called a compensated block. Fusing the motion-compensated image with the filtered frame image (e.g., weighted summation pixel by pixel) yields a filtered image obtained by performing MCTF processing on the filtered frame based on a single reference frame.
[0029] It can be understood that MCTF processing can also be performed on a filtered frame based on multiple reference frames. Specifically, according to the Figure 2 processing flow shown, a motion-compensated image is generated for the filtered frame based on each reference frame, thereby obtaining multiple motion-compensated images. Then, the multiple motion-compensated images are fused with the filtered frame image (e.g., weighted summation element by element) to obtain a filtered image obtained by performing MCTF processing on the filtered frame based on multiple reference frames.
[0030] As can be understood from the foregoing, MCTF obtains the motion of the blocks in the picture between the filtered frame and the reference frame through motion estimation, shifts the blocks in the reference frame according to the motion vectors to obtain a motion-compensated image, and the filtered image obtained by fusing the motion-compensated image and the filtered frame image plays a smoothing role between the filtered frame and the reference frame. However, in the framework of the block motion compensation technology in traditional MCTF, each image block independently performs motion estimation and motion compensation, which often produces significant boundary artifacts between the compensated blocks. Especially in the object edge region, due to the non-uniform motion of the pixels within the block, the compensation effect at the edge is often insufficient, affecting the final filtering effect.
[0031] Figure 3 An exemplary schematic diagram showing boundary artifacts generated by motion compensation in some embodiments of the present disclosure is shown. As Figure 3 shown, two adjacent macroblocks MacB1 and MacB2 in the filtered frame, where MacB1 is located to the left of MacB2. In motion estimation, the matching block search algorithm searches for matching blocks for the macroblocks in the filtered frame based on the similarity of the pixel data in the blocks. The matching block corresponding to MacB1 is MatB1, and the matching block corresponding to MacB2 is MatB2. The first motion vector MV 1 is calculated for MacB1 and MatB1, and the second motion vector MV 2 is calculated for MacB2 and MatB2. In accordance with the motion vectors MV 1 , MV 2When generating a motion-compensated image for a reference frame, the matching blocks MatB1 and MatB2 will be placed adjacent to each other on the left and right. The right end of MatB1 and the left end of MatB2 form the boundary region between the two matching blocks. Specifically, the boundary region includes a first boundary region on the MatB1 side and a second boundary region on the MatB2 side. It can be understood that when the pixel difference between the first boundary region and the second boundary region is large, boundary artifacts will appear in the boundary region between MatB1 and MatB2 due to non-conformity.
[0032] For the sake of convenience of explanation, Figure 3 taking two adjacent filter frame macroblocks on the left and right as an example to illustrate the boundary artifacts in the motion-compensated image. It can be understood that traditional motion compensation moves the matching blocks according to the motion vectors to construct the motion-compensated image. During the process of moving the matching blocks, the relative positions between the matching blocks will change. When two matching blocks are placed at adjacent positions (including adjacent on the left and right and adjacent up and down) in the motion-compensated image, boundary artifacts will appear. Performing temporal filtering on the filter frame image based on the motion-compensated image containing boundary artifacts will reduce the redundant information amount and smoothness between the filter frame and the reference frame, thereby reducing the quality of video coding. In view of this, the present disclosure discloses a method for preprocessing video to suppress boundary artifacts in the motion-compensated image and improve the quality of video coding.
[0033] Figure 4 shows an exemplary flowchart of a method for preprocessing video in some embodiments of the present disclosure, as Figure 4 shown. The method includes: Step 401, partitioning the filter frame to obtain a first filter frame macroblock and a second filter frame macroblock adjacent to the first filter frame macroblock; Step 402, determining, based on the reference frame associated with the filter frame, a first matching block matching the first filter frame macroblock and a second matching block matching the second filter frame macroblock; Step 403, compensating the first matching block based on the boundary compensation region of the second matching block to obtain a compensated block; Step 404, updating the first filter frame macroblock according to the compensated block.
[0034] For step 401, it can be understood that the filter frame can be divided into multiple filter frame macroblocks according to a preset macroblock size. The length of the filter frame macroblock can be a multiple of 4 pixel points, and the width of the macroblock can be a multiple of 4 pixel points. The macroblock size includes but is not limited to: 64 pixels * 64 pixels, 64 pixels * 32 pixels, 32 pixels * 32 pixels. The first filter frame macroblock being adjacent to the second filter frame macroblock includes: the first filter frame macroblock being adjacent to the second filter frame macroblock on the left and right or up and down.
[0035] For step 402, it can be understood that the reference frame associated with the filtered frame refers to the nearby frames referred to when filtering the filtered frame. Based on the matching block search algorithm, the first matching block in the reference frame can be matched to the macroblock of the first filtered frame, and the second matching block in the reference frame can also be matched to the macroblock of the second filtered frame.
[0036] Figure 5 FIG. shows an exemplary schematic diagram for compensating the first matching block based on the boundary compensation region of the second matching block in some embodiments of the present disclosure. As Figure 5 shown in sub - figure (a) therein, the first matching block matched to the macroblock (MacB1) of the first filtered frame is MatB1; and for the macroblock (MacB2) of the second filtered frame adjacent to the macroblock of the first filtered frame, its matching block is the second matching block MatB2. If MatB1 and MatB2 are placed adjacent to each other directly left - right, when the pixel difference between the first boundary region on the MatB1 side and the second boundary region on the MatB2 side is large, artifacts are likely to appear in the boundary region.
[0037] Continuing Figure 5 , the present disclosure compensates the first matching block based on the boundary compensation region of the second matching block, and suppresses the boundary artifacts in the motion - compensated image by reducing the pixel value difference between two adjacent matching blocks in the boundary region of the motion - compensated image. The boundary compensation region of the second matching block can be the region adjacent to the second boundary region in the reference frame, and the boundary compensation region is shown as a shaded region in sub - figures (b) and (c) of Figure 5 ; it can be understood that as long as it can be included within the first matching block, the boundary compensation region can be of any shape; in some embodiments, the boundary compensation region is a long - strip region with the same width as the second matching block and a length less than that of the second matching block; in other embodiments, the length of the boundary compensation region is 2 pixel points or 4 pixel points. Since they are in adjacent positions in the reference frame image, there are almost no boundary artifacts between the boundary compensation region and the second matching block. Compensating the pixel values of the boundary compensation region into the first matching block can obtain a compensated block with a boundary more consistent with that of the second matching block and a smoother boundary region. In some embodiments, compensating the pixel values of the boundary compensation region into the first matching block includes: performing a weighted sum of the pixel values pixel - by - pixel for the boundary compensation region and the corresponding compensated region in the first matching block, where the compensated region can be the region adjacent to the second matching block in the first matching block.
[0038] In some embodiments, updating the macroblock of the first filtered frame according to the compensated block may include: performing a weighted sum of the pixel values in the compensated block and the pixel values in the macroblock of the first filtered frame pixel - by - pixel.
[0039] For ease of explanation, the method disclosed in this disclosure for suppressing boundary artifacts of matching blocks corresponding to adjacent filtered frame macroblocks will be described below by taking two adjacent filtered frame macroblocks as an example; similarly, the method can also suppress boundary artifacts of matching blocks corresponding to two vertically adjacent filtered frame macroblocks, which will not be elaborated here. In addition, it should be noted that: for a filtered frame, the number of associated reference frames can be greater than 1; when processing a filtered frame with multiple reference frames, the compensation blocks for the filtered frame can be calculated based on multiple reference frames respectively, and the filtered frame macroblocks can be updated according to the multiple compensation blocks.
[0040] Figure 6 FIG. shows an exemplary schematic diagram of preprocessing a filtered frame macroblock according to multiple reference frames in some embodiments of this disclosure. As Figure 6 shown, the current filtered frame contains two adjacent filtered frame macroblocks MacB1 and MacB2; the first reference frame and the second reference frame are two reference frames associated with the current filtered frame. The matching block of MacB1 searched in the first reference frame is MatB11, and the matching block of MacB2 is MatB21; the matching block of MacB1 searched in the second reference frame is MatB12, and the matching block of MacB2 is MatB22. MatB11 is compensated based on the boundary compensation region of MatB21 to obtain the first compensation block; MatB12 is compensated based on the boundary compensation region of MatB22 to obtain the second compensation block. Then, the filtered frame macroblock MacB1 is updated according to the first compensation block and the second compensation block, for example, by pixel-by-pixel weighted summation of the pixel values in the first compensation block, the second compensation block, and MacB1.
[0041] The method disclosed in this disclosure for preprocessing video compensates the current matching block through the boundary compensation regions of adjacent matching blocks in the motion-compensated image, thereby performing a filtering operation on the boundary regions of adjacent matching blocks in the motion-compensated image and suppressing boundary artifacts. The strategy adopted in this disclosure helps to achieve effective fusion of motion information between adjacent filtered frame macroblocks and visually seamless connection, effectively reducing the blocking effect and edge jaggedness caused by block processing. Compared with traditional temporal filtering techniques, the boundary-enhanced temporal filtering technique proposed in the embodiments of this disclosure makes the boundaries between blocks in the motion-compensated image smoother, significantly reducing the blocking effect and thus improving the filtering effect. Further, this method can increase the redundancy information and smoothness between the processed filtered frame and the reference frame, thereby improving the quality of video coding.
[0042] In some embodiments, the method for preprocessing a video disclosed in this disclosure further includes: traversing multiple frames of the video; for the current frame, determining whether the current frame is a filtering frame according to the coding reference structure, where the coding reference structure is used to determine the inter-frame reference relationship for video coding.
[0043] It can be understood that predictive coding of B-frames utilizes future frames in the presentation order. To allow the coding order to be different from the presentation order, video frames are divided into Groups of Pictures (GOPs). A GOP includes multiple consecutively presented frames in the video, and the number of frames included in a GOP is called the GOP size (GOPSize); in some embodiments, the GOPSize is set according to the frame rate; in other embodiments, the GOPSize can be set to 8, 12, or other multiples of 4. The GOP also specifies the coding reference structure for predictive coding, and the coding reference structure is used to determine the frames that need to be referenced for predictive coding of each frame in the GOP.
[0044] Figure 7 An exemplary schematic diagram of the coding reference structure in some embodiments of this disclosure is shown. In Figure 7 it, the boxes represent each frame, and the presentation order of the frames is from left to right; the coding order of the frames is shown by the numbers in the boxes, where the frame corresponding to the number 0 can be the first frame of the video or the last frame in the previous GOP. Therefore Figure 7 the GOP shown includes the frames represented by the numbers 1 - 8. Further, in Figure 7 it, the arrows represent the reference relationships for predictive coding, that is, the frame at the start of the arrow needs to reference the frame at the end of the arrow for predictive coding. Thus, a 4-layer coding hierarchy is formed within the GOP. Among them, the frames in layer 0 are first subjected to predictive coding. After the predictive coding of the frames in layer 0 is completed, the frames in layer 1 are subjected to predictive coding, and then the frames in layers 2 and 3 are sequentially subjected to predictive coding. Figure 7 The coding reference structure shown is only for ease of illustration. For those skilled in the art, different coding reference structures can be designed as needed, and this disclosure places no restrictions on this.
[0045] It can be understood that in the preprocessing step, it is possible to determine whether the current frame is a filtering frame according to the coding reference structure. In some embodiments, if the current frame is referenced by other frames in the coding reference structure, then the current frame is a filtering frame; otherwise, the current frame is not a filtering frame. In other embodiments, it is possible to determine whether the current frame is a filtering frame according to the number of times the current frame is referenced in the coding reference structure. If the number of times the current frame is referenced is greater than the set reference number threshold, then the current frame is considered a filtering frame; otherwise, the current frame is not a filtering frame; where the number of frames in the coding reference structure that reference the current frame is the number of times the current frame is referenced.
[0046] It can be understood that the "reference frame associated with the filtered frame" described in the foregoing step 402 can also be determined according to the coding reference structure. In some embodiments, a method for preprocessing video disclosed in the present disclosure further includes: for a current filtered frame, determining a reference frame associated with the current filtered frame according to the coding reference structure.
[0047] In some embodiments, after determining, based on the reference frame associated with the filtered frame: a first matching block that matches a first filtered frame macroblock and a second matching block that matches a second filtered frame macroblock, it includes: determining a first motion vector according to the positions of the first filtered frame macroblock and the first matching block, for determining the relative positions of the first filtered frame macroblock and the first matching block; determining a second motion vector according to the positions of the second filtered frame macroblock and the second matching block, for determining the relative positions of the second filtered frame macroblock and the second matching block.
[0048] It can be understood that based on the matching block search algorithm, a first matching block that matches a first filtered frame macroblock and a second matching block that matches a second filtered frame macroblock can be searched in the reference frame. Based on the positions of the first filtered frame macroblock and the first matching block, the first motion vector can be calculated, and based on the positions of the second filtered frame macroblock and the second matching block, the second motion vector can be calculated.
[0049] Figure 8 An exemplary schematic diagram of the matching block search algorithm in some embodiments of the present disclosure is shown. In some embodiments, the matching block search algorithm includes: determining a search window in the reference frame according to the position of the filtered frame macroblock; determining a matching block according to the matching degree between the reference frame macroblock in the search window and the filtered frame macroblock. In these embodiments, taking the filtered frame macroblock as the center, move M 1 pixel points upward and downward respectively on the reference frame, and move M 2 pixel points leftward and rightward respectively on the reference frame, so as to obtain the search window. Such a search window contains (2M 1 +1)(2M 2 +1) reference frame macroblocks to be compared. It is necessary to calculate the matching degree between the filtered frame macroblock and each reference frame macroblock according to a preset matching criterion, and then use the reference frame macroblock with the highest matching degree in the search window as the matching block corresponding to the filtered frame macroblock.
[0050] In some embodiments, the mean absolute difference of the pixel values of the corresponding pixel points of the reference frame macroblock and the filtered frame macroblock is calculated based on formula (1), and the matching degree is evaluated according to the mean absolute difference of the pixel values.
[0051]
[0052] Among them, MAD represents the Mean Absolute Difference, M represents the length of the macroblock of the filtered frame, N represents the width of the macroblock of the filtered frame, and C ij represents the pixel value of the pixel point in the macroblock of the filtered frame, and R ij represents the pixel value of the pixel point in the macroblock of the reference frame.
[0053] In some embodiments, the mean squared error of the pixel values of the corresponding pixel points of the macroblock of the reference frame and the macroblock of the filtered frame is calculated based on formula (2), and the matching degree is evaluated according to the mean squared error of the pixel values.
[0054]
[0055] Among them, MSE represents the Mean Squared Error. It can be understood that the smaller the mean absolute difference and the mean squared error of the pixel values, the higher the matching degree between the macroblock of the reference frame and the macroblock of the filtered frame.
[0056] Continue Figure 8 After finding the corresponding matching block for the macroblock of the filtered frame, the motion vector can be calculated according to the positions of the macroblock of the filtered frame and the matching block. In some embodiments, the motion vector is calculated according to the coordinates of the upper left, upper right, lower left or lower right pixel points of the macroblock of the filtered frame and the matching block; for example, the difference between the coordinates of the upper left pixel points of the macroblock of the filtered frame and the matching block can be taken as the motion vector; exemplarily, if the coordinates of the upper left pixel point of the macroblock of the filtered frame are (32, 32) and the coordinates of the upper left pixel point of the matching block are (30, 30), then the motion vector between the macroblock of the filtered frame and the matching block is (2, 2).
[0057] It can be understood that Figure 8 the shown matching block search algorithm is a local search algorithm. For those skilled in the art, different matching block search algorithms can be selected as needed, including but not limited to: Pyramidal Block Matching Algorithm, Three-Step Block Matching Algorithm, and this disclosure does not make any restrictions on this.
[0058] In some embodiments, compensating the first matching block based on the boundary compensation region of the second matching block includes: performing weighted summation on the compensated pixel points in the boundary compensation region and the pixel points to be compensated in the first matching block.
[0059] Figure 9 shows an exemplary schematic diagram of a motion-compensated image in some embodiments of this disclosure. As Figure 9As shown, the current filtered frame includes macroblocks MacB1-9. After performing motion estimation and motion compensation on the current filtered frame, a motion-compensated image will be obtained. To suppress boundary artifacts in the entire motion-compensated image, starting from the first row of the current filtered frame, search for matching blocks for the macroblocks in the current filtered frame row by row, where the macroblocks in the filtered frame within each row are processed in order from left to right. After finding the current matching block for the macroblock in the current filtered frame, perform boundary compensation on the matching blocks corresponding to the adjacent macroblocks in the filtered frame above the current filtered frame macroblock and the matching blocks corresponding to the macroblocks in the filtered frame adjacent to the left of the current filtered frame macroblock. For example, after finding the matching block MatB2 for the filtered frame macroblock MacB2, compensate the right boundary of MatB1 based on the boundary compensation area adjacent to the left boundary of MatB2; after finding the matching block MatB4 for the filtered frame macroblock MacB4, compensate the lower boundary of MatB1 based on the boundary compensation area adjacent to the upper boundary of MatB4.
[0060] It can be understood that by performing boundary compensation in this way, for some matching blocks, only one boundary will be compensated once. For example, Figure 9 for MatB3 in Figure 9 , its lower boundary will be compensated once by MatB6, and its right boundary will not be compensated. Additionally, for some matching blocks, both their right and lower boundaries will be compensated once. For example,
[0061] for MatB1 in
[0062]
[0063] 0 n 1 0
[0065]
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] 0
[0064]
[0065] Represents the pixel value of the pixel before the secondary compensation point is compensated, I n1 Represents the pixel value of the compensated pixel point corresponding to the secondary compensation point in the boundary compensation area of the second matching block, w 1 Represents the weight corresponding to the second matching block, I n2 Represents the pixel value of the compensated pixel point corresponding to the secondary compensation point in the boundary compensation area of the third matching block, w 2 Represents the weight corresponding to the third matching block, I represents the pixel value of the pixel after the secondary compensation point is compensated.
[0066] In some embodiments, performing a weighted sum of the compensated pixel point and the compensated pixel point in the first matching block includes: obtaining the distance between the first motion vector and the second motion vector to obtain a vector distance; calculating a first weight according to the vector distance for controlling the compensation degree of the compensated pixel point; and performing a weighted sum of the compensated pixel point and the compensated pixel point according to the first weight.
[0067] Assuming that the first motion vector is (x1, y1) and the second motion vector is (x2, y2), the vector distance between the first motion vector and the second motion vector It can be understood that the vector distance can measure the distance between the first matching block and the second matching block in the reference frame: the larger the vector distance, the farther the distance between the first matching block and the second matching block in the reference frame, and at this time, using the compensated pixel point near the second matching block to compensate the first matching block is difficult to effectively suppress the boundary artifacts between the first matching block and the second matching block; the smaller the vector distance, the closer the distance between the first matching block and the second matching block in the reference frame, and at this time, using the compensated pixel point near the second matching block to compensate the first matching block can effectively suppress the boundary artifacts between the first matching block and the second matching block.
[0068] The compensated pixel point includes a primary compensation point and a secondary compensation point. In some embodiments, for the primary compensation point, compensation is performed based on the following formula (5):
[0069]
[0070] where, w mv Represents the first weight. In some other embodiments, for the secondary compensation point, compensation is performed based on the following formula (6):
[0071]
[0072] where, w mv1 Represents the first weight calculated for the first matching block and the second matching block, w mv2 Represents the first weight calculated for the first matching block and the third matching block.
[0073] In some embodiments, calculating the first weight according to the vector distance includes: in response to the vector distance being less than or equal to a set distance threshold, the first weight takes a value of 1; in response to the vector distance being greater than the set distance threshold, determining the maximum value of the distances between the motion vectors of adjacent filtered frame macroblocks for multiple filtered frame macroblocks of the filtered frame to obtain the maximum distance value; calculating the first weight according to the distance threshold, the maximum distance value, and the vector distance.
[0074] In these embodiments, the first weight is calculated according to the vector distance based on the following formula (7):
[0075]
[0076] where d th is the set distance threshold, and d max represents the maximum value of the distances between the motion vectors of adjacent filtered frame macroblocks. In some embodiments, d th can be set to 2, 3, or 4. In some other embodiments, d max is: the maximum value of the motion vector distances calculated for adjacent matching blocks among all the searched matching blocks up to the first matching block; exemplarily, referring to Figure 9 , assuming the first matching block corresponds to MatB3, d max is the maximum value of the motion vector distances calculated for all adjacent matching blocks in MatB1-3, that is, the maximum value of the motion vector distance between MatB1 and MatB2 and the motion vector distance between MatB2 and MatB3.
[0077] In some embodiments, performing a weighted sum of the compensated pixel points in the boundary compensation region and the compensated pixel points in the first matching block includes: traversing the compensated pixel points; for the current compensated pixel point, obtaining the distance from the current compensated pixel point to the boundary line between the first matching block and the second matching block to obtain the boundary distance; setting a second weight according to the boundary distance to control the compensation degree for the current compensated pixel point; performing a weighted sum of the pixel values of the current compensated pixel point and the compensated pixel point corresponding to the current compensated pixel point according to the second weight.
[0078] It can be understood that the boundary distance can measure the influence degree of the compensated pixel point on the boundary artifacts generated by the first matching block and the second matching block: the larger the boundary distance, the smaller the influence degree of the current compensated pixel point on generating boundary artifacts, so that a smaller second weight can be set for these compensated pixel points for a smaller degree of compensation; the smaller the boundary distance, the greater the influence degree of the current compensated pixel point on generating boundary artifacts, so that a larger second weight can be set for these compensated pixel points for a larger degree of compensation.
[0079] The compensated pixel points include primary compensation points and secondary compensation points. In some embodiments, for the primary compensation points, compensation is performed based on the following formula (8):
[0080]
[0081] where w d represents the second weight. In some other embodiments, for the secondary compensation points, compensation is performed based on the following formula (9):
[0082]
[0083] where w d1 represents the second weight set for the first matching block and the second matching block, and w d2 represents the second weight set for the first matching block and the third matching block.
[0084] In some embodiments, according to the first weight and the second weight, the pixel values of the current compensated pixel point and the compensated pixel point corresponding to the current compensated pixel point are weighted and summed. Among them, for the primary compensation points, compensation is performed based on the following formula (10):
[0085]
[0086] For the secondary compensation points, compensation is performed based on the following formula (11):
[0087]
[0088] In some embodiments, updating the first filtered frame macroblock according to the compensation block includes: setting a filtering weight according to the temporal distance between the filtered frame and the reference frame; according to the filtering weight, performing weighted summation on the pixel points in the compensation block and the first filtered frame macroblock.
[0089] It can be understood that the temporal distance between the filtered frame and the reference frame can be the temporal distance during their playback, and the temporal distance can measure the reference significance of the reference frame for filtering the filtered frame: the smaller the temporal distance, the greater the reference significance of the reference frame for the filtered frame, and at this time, a higher filtering weight can be set; the greater the temporal distance, the smaller the reference significance of the reference frame for the filtered frame, and at this time, a lower filtering weight can be set. In some embodiments, based on the following formula (12), weighted summation is performed on the pixel points in the compensation block and the first filtered frame macroblock according to the filtering weight:
[0090]
[0091] where I f represents the pixel value of the pixel point in the first filtered frame macroblock after weighted summation; I 0represents the pixel value of the pixel points in the first filtered frame macroblock before weighted summation; w r represents the filtering weight; I r represents the pixel value of the pixel points in the compensation block.
[0092] In some other embodiments, for a filtered frame with multiple reference frames, based on the following formula (13), the pixel points in multiple compensation blocks and the first filtered frame macroblock are weighted and summed according to the filtering weights:
[0093]
[0094] where N represents the number of reference frames; w r (i) represents the filtering weight corresponding to the i-th reference frame; I r (i) represents the pixel value of the pixel points in the compensation block corresponding to the i-th reference frame.
[0095] Furthermore, the present disclosure also discloses a computer-readable storage medium storing program instructions adapted to be loaded and executed by a processor to perform the methods described in the foregoing embodiments of the present disclosure.
[0096] The present disclosure also discloses a device for preprocessing video, including: a processor configured to execute program instructions; and a memory configured to store program instructions, which, when loaded and executed by the processor, cause the device to perform the methods described in the foregoing embodiments of the present disclosure.
[0097] Figure 10 The block diagram showing the hardware configuration of device 100 that can implement the embodiments of the present disclosure is as follows. As Figure 10 shown, device 100 may include a processor 101 and a memory 102. The processor therein is configured to execute program instructions, and the memory is configured to store program instructions, which, when loaded and executed by the processor, cause the device to perform the method for preprocessing video described in any one of the foregoing embodiments. In Figure 10 device 100, only the constituent elements related to this embodiment are shown. Therefore, it is obvious to those of ordinary skill in the art that: device 100 may further include common constituent elements different from those Figure 10 shown. The specific functions implemented by the memory 102 and the processor 101 of device 100 provided in the embodiments of this specification can be explained in contrast to the foregoing embodiments in this specification and can achieve the technical effects of the foregoing embodiments, which will not be elaborated here.
[0098] The device 100 may correspond to a computing device with various processing functions. For example, the device 100 may be implemented as various types of devices, such as a personal computer (PC), a server device, a mobile device, and the like.
[0099] The processor 101 may control the operation of the device 100. For example, the processor 101 may be implemented by a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), an artificial intelligence processor chip (IPU), etc. provided in the device 100. However, the present invention is not limited thereto. In the present embodiment, the processor 101 may be implemented in any suitable manner. For example, the processor 101 may take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and a form embedded in a microcontroller, and so on.
[0100] The memory 102 can be used to store various data and instructions processed in the device 100. For example, the memory 102 can store the processed data and the data to be processed in the device 100. The memory 102 can store the data that has been processed or is to be processed by the processor 101, such as filtered frame image data and reference frame image data. In addition, the memory 102 can store applications, drivers, etc. to be driven by the device 100. For example, the memory 102 can store various programs related to the method for preprocessing video to be executed by the processor 101, etc. The memory 102 can be a DRAM, but the present invention is not limited thereto. The memory 102 can include at least one of a volatile memory or a non-volatile memory. The non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, phase change RAM (PRAM), magnetic RAM (MRAM), resistive RAM (RRAM), ferroelectric RAM (FRAM), etc. The volatile memory can include dynamic RAM (DRAM), static RAM (SRAM), synchronous DRAM (SDRAM), PRAM, MRAM, RRAM, ferroelectric RAM (FeRAM), etc. In some embodiments, the memory 102 can include at least one of a hard disk drive (HDD), a solid state drive (SSD), a high density flash (CF), a secure digital (SD) card, a micro secure digital (Micro-SD) card, a mini secure digital (Mini-SD) card, an extreme digital (xD) card, caches, or a memory stick.
[0101] In summary, the specific functions implemented by the computer-readable storage medium and the device for preprocessing video provided in the embodiments of this specification can be explained in contrast to the foregoing embodiments in this specification and can achieve the technical effects of the foregoing embodiments, and will not be elaborated herein.
[0102] It should be noted that, for the purpose of simplicity, some methods and their embodiments of the present invention are expressed as a series of actions and their combinations. However, those skilled in the art can understand that the solution of the present invention is not limited by the order of the described actions. Therefore, based on the disclosure or teachings of the present invention, those skilled in the art can understand that some of the steps can be executed in other orders or simultaneously. Further, those skilled in the art can understand that the embodiments described in the present invention can be regarded as optional embodiments, that is, the actions or modules involved therein are not necessarily required for the implementation of a certain or some solutions of the present invention. In addition, according to the differences in the solutions, the present invention also focuses on the descriptions of some embodiments. In view of this, those skilled in the art can understand that the parts not detailed in a certain embodiment of the present invention can also refer to the relevant descriptions of other embodiments.
Claims
1. A method for preprocessing a video, comprising: Dividing the filtered frame into blocks to obtain a first filtered frame macroblock and a second filtered frame macroblock adjacent to the first filtered frame macroblock; Determining, based on a reference frame associated with the filtered frame: a first matching block that matches a first filtered frame macroblock and a second matching block that matches a second filtered frame macroblock; Compensating the first matching block based on the boundary compensation area of the second matching block to obtain a compensated block; as well as The first filtered frame macroblock is updated according to the compensation block.
2. The method according to claim 1, further comprising: Iterate over multiple frames of the video; For a current frame, whether the current frame is a filtered frame is determined according to a coding reference structure, wherein the coding reference structure is used to determine an inter-frame reference relationship of video coding.
3. The method according to claim 1, wherein: After determining, based on a reference frame associated with the filtered frame: a first matching block matching the first filtered frame macroblock and a second matching block matching the second filtered frame macroblock, comprising: Determining a first motion vector according to the positions of the first filtered frame macroblock and the first matching block, for determining the relative positions of the first filtered frame macroblock and the first matching block; and A second motion vector is determined according to the positions of the second filtered frame macroblock and the second matching block, and is used to determine the relative positions of the second filtered frame macroblock and the second matching block.
4. The method according to claim 3, wherein: Compensating the first matching block based on the boundary compensation area of the second matching block includes: A weighted sum is performed on the compensation pixel points in the boundary compensation area and the compensated pixel points in the first matching block.
5. The method according to claim 4, wherein: The weighted summing of the compensation pixel points in the boundary compensation area and the compensated pixel points in the first matching block includes: Obtaining a distance between the first motion vector and the second motion vector to obtain a vector distance; Calculating a first weight according to the vector distance, so as to control the degree of compensation for the compensated pixel point; The compensating pixel point and the compensated pixel point are weightedly summed according to the first weight.
6. The method according to claim 5, wherein: Calculating the first weight according to the vector distance includes: In response to the vector distance being less than or equal to a set distance threshold, the first weight is set to 1; In response to the vector distance being greater than a set distance threshold, Determining the maximum value of the distances between the motion vectors of adjacent filtered frame macroblocks for a plurality of filtered frame macroblocks of the filtered frame to obtain a maximum distance value; The first weight is calculated according to the distance threshold, the maximum distance and the vector distance.
7. The method according to claim 4, wherein: The weighted summing of the compensation pixel points in the boundary compensation area and the compensated pixel points in the first matching block includes: Traversing the compensated pixels; For the currently compensated pixel, acquiring the distance between the currently compensated pixel and the boundary line between the first matching block and the second matching block, To get the boundary distance; Setting a second weight according to the boundary distance to control the degree of compensation for the currently compensated pixel point; The pixel values of the current compensated pixel point and the compensating pixel point corresponding to the current compensated pixel point are weightedly summed according to the second weight.
8. The method according to claim 1, wherein: Updating the first filtered frame macroblock according to the compensation block comprises: Setting a filtering weight according to a temporal distance between the filtering frame and the reference frame; According to the filtering weight, a weighted sum is performed on the pixels in the compensation block and the first filtering frame macroblock.
9. A computer-readable storage medium storing program instructions, wherein the program instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 8.
10. A device for preprocessing a video, comprising: a processor configured to execute program instructions; as well as A memory configured to store the program instructions, which, when loaded and executed by the processor, causes the apparatus to perform the method according to any one of claims 1 to 8.